Senior Network Reliability Engineer

JPMorgan Chase & Co.Buenos Aires, Argentinafull time
ActiveVerified 3h ago

Job description

Detailed Job Description

Key responsibilities

  • Own reliability outcomes for mission-critical network services (availability, performance, recoverability) and drive measurable improvements.
  • Lead major incident response: establish response rigor, coordinate technical mitigation, coach responders, and ensure high-quality post-incident learning.
  • Own problem management end-to-end: author/approve RCAs, identify systemic root causes, and drive durable remediation and prevention.
  • Architect automation frameworks using Python, Shell, and Ansible (e.g., self-healing patterns where appropriate, automated remediation with guardrails, validation pipelines, standardized tooling).
  • Provide deep technical leadership across:
    • SD-WAN, SDA, and software-defined networking (SND)
    • Routing and Switching
    • Firewalls, Load Balancers, Proxies
    • Strong preference for Cisco ACI / Fabrics; VMware NSX as a plus
  • Partner with developers/platform teams to develop observability insights: define operational signals, reduce alert noise, improve actionability, and align monitoring to service health.
  • Embed SRE engineering practices into design and delivery:
    • Translate NFRs into controls, standards, and measurable targets
    • Apply FMEA-style risk analysis to designs/changes to reduce failure impact
    • Drive operational readiness reviews, resilience testing, and change safety mechanisms
  • Mentor engineers, set technical standards, influence cross-team roadmaps, and deliver outcomes with minimal supervision.

Required qualifications

  • Extensive experience operating and engineering large-scale networks with strong troubleshooting depth.
  • Proven leadership in major incidents, RCA, and delivery of long-term fixes that reduce repeat incidents.
  • Advanced automation track record using Python, Shell, and Ansible, with demonstrable toil reduction and reliability gains.
  • Strong applied understanding of SRE concepts, NFRs, and FMEA (or equivalent failure/risk analysis).
  • Strong ownership mindset, stakeholder management, and ability to drive delivery independently.

Preferred qualifications

  • Strong expertise with Cisco ACI / Fabrics.
  • VMware NSX experience (nice-to-have).
  • Certifications such as CCNP/CCIE (or equivalent), plus other vendor certifications.
  • Prior experience in financial institutions (advantage).

Similar jobs