JPJPMorganChaseCiudad Autónoma de Buenos Aires, Argentinafull time
ActiveVerified 17m ago
Job description
Key responsibilities
Lead day-to-day operational ownership for network services, including complex troubleshooting and coordinated restoration.
Drive incident and problem management: run structured investigations, produce high-quality RCAs, and ensure corrective/preventive actions are delivered.
Participate in and support major incident management (communications support, technical lead support, mitigation execution).
Design and implement production-grade automation using Python, Shell, and Ansible (e.g., drift detection, change validation, automated diagnostics, safe rollout helpers).
Engineer/support across:
SD-WAN, SDA, and broader software-defined networking (SND)
Routing and Switching
Firewalls, Load Balancers, Proxies
Improve reliability via standardization, guardrails, repeatable runbooks, and continuous validation.
Build observability insights with developers/platform teams: actionable dashboards, high-signal alerting, and service health metrics.
Apply SRE practices: operational readiness, toil reduction, and reliability thinking; working understanding of NFRs and FMEA (or equivalent risk methods).
Required qualifications
Demonstrated experience in incident response and problem management, including ownership of RCAs through closure.
Strong hands-on networking skills across enterprise routing/switching and security/L4–L7 components.
Strong automation capability using Python, Shell, and Ansible in production operations.
SRE mindset with practical understanding of reliability concepts, NFRs, and risk analysis approaches (including FMEA familiarity).
Ability to work independently, prioritize effectively, and deliver with minimal oversight.
Preferred qualifications
Experience with Cisco ACI / Fabrics.
VMware NSX experience (nice-to-have).
Certifications such as CCNP (preferred), CCNA, and other vendor certifications.
Experience in a financial institution / regulated environment (advantage)