DevOps / Site Reliability Engineer ID70127

Software Engineer (iOS Tech Lead) ID48363 in BostonLeón de los Aldama, Mexicofull timeSenior
Active

Job description

AgileEngine is an Inc. 5000 company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries. We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards.WHY JOIN USIf you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet you!ABOUT THE ROLEWe are looking for a DevOps / Site Reliability Engineer to maintain operational resilience and 24/7 stability for a multi-cloud enterprise security program, serving as Incident Commander on major and critical incidents while also owning IaC, CI/CD pipelines, and CSPM telemetry. You will drive major-incident calls, own post-incident remediation follow-through, draft stakeholder communications, and develop divisional incident-management playbooks alongside multi-cloud security guardrails using Terraform and Wiz. The role requires 5+ years of SRE experience with hands-on incident command in a 24/7 financial services environment.WHAT YOU WILL DO- Scale and maintain the ability to drive operational stability across multi-cloud environments (Azure, AWS, GCP).- Engineer unified security policies and configuration baselines using IaC (Terraform) to prevent misconfigurations.- Design, maintain, and optimize enterprise CI/CD pipelines to support continuous ASPM ingestion and deployment.- Act on continuous monitoring alerts, utilizing Cloud Security Posture Management (CSPM) tools like Wiz to secure workloads.- Serve as Incident Commander on major and critical incidents — running the bridge, directing technical workstreams, making time-critical decisions, and coordinating cross-functional responders under pressure.- Own the post-incident loop — track remediation items to closure, hold owning teams accountable to timelines, and drive systemic fixes and preventative actions across groups.- Draft and send clear, accurate, audience-appropriate incident notifications and status updates to technical teams, management, and stakeholders throughout the incident lifecycle.- Develop, maintain, and socialize divisional / group-level incident-management playbooks, runbooks, and escalation procedures that standardize response and reduce time-to-resolution.MUST HAVES- 5+ years of experience.- In-depth architectural expertise in multi-cloud defense, federated IAM, and zero-trust principles.- Strong practical experience with Kubernetes, Terraform, CI/CD orchestration, and Python/Go scripting.- Senior-level, hands-on incident-command experience driving major/critical incident calls to resolution in a 24x7 production environment.- Proven track record of remediation follow-up — coordinating with teams and holding owners accountable until issues are fully closed.- Demonstrated skill drafting and issuing incident notification communications to both technical and executive audiences.- Direct experience authoring divisional/group incident-management playbooks and escalation procedures.- Fully autonomous.- Drives the architecture of complex automated runbooks and mentors Middle-level SREs.- Extensive experience deploying and tuning APIs from modern CNAPP/CSPM platforms, ideally Wiz.- Prior experience building platforms subject to strict financial compliance standards (PCI-DSS, SOC2).- Upper-intermediate English level.NICE TO HAVES- PagerDuty — hands-on experience with on-call scheduling, alert routing, and incident orchestration.- ServiceNow — familiarity with incident, problem, and change management workflows and reporting.PERKS AND BENEFITS- Growth without limits: build your skills through mentorship, internal TechTalks, challenging projects, and a dedicated annual learning budget- Competitive compensation: get recognition that reflects your skills and impact, with regular performance and compensation reviews- Flexibility: work 100% remotely with flexible hours that support focus, autonomy, and a healthy work rhythm- Meaningful, modern projects: build impactful products using modern technologies alongside global teams and leading brands- Collaborative culture: join a supportive environment with zero micromanagement where ideas are welcomed and contributions are recognized- Well-being & support: access local well-being programs and people-focused support tailored to your locationAgileEngine is an Inc. 5000 company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries. We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards.WHY JOIN USIf you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet you!ABOUT THE ROLEWe are looking for a DevOps / Site Reliability Engineer to maintain operational resilience and 24/7 stability for a multi-cloud enterprise security program, serving as Incident Commander on major and critical incidents while also owning IaC, CI/CD pipelines, and CSPM telemetry. You will drive major-incident calls, own post-incident remediation follow-through, draft stakeholder communications, and develop divisional incident-management playbooks alongside multi-cloud security guardrails using Terraform and Wiz. The role requires 5+ years of SRE experience with hands-on incident command in a 24/7 financial services environment.WHAT YOU WILL DO- Scale and maintain the ability to drive operational stability across multi-cloud environments (Azure, AWS, GCP).- Engineer unified security policies and configuration baselines using IaC (Terraform) to prevent misconfigurations.- Design, maintain, and optimize enterprise CI/CD pipelines to support continuous ASPM ingestion and deployment.- Act on continuous monitoring alerts, utilizing Cloud Security Posture Management (CSPM) tools like Wiz to secure workloads.- Serve as Incident Commander on major and critical incidents — running the bridge, directing technical workstreams, making time-critical decisions, and coordinating cross-functional responders under pressure.- Own the post-incident loop — track remediation items to closure, hold owning teams accountable to timelines, and drive systemic fixes and preventative actions across groups.- Draft and send clear, accurate, audience-appropriate incident notifications and status updates to technical teams, management, and stakeholders throughout the incident lifecycle.- Develop, maintain, and socialize divisional / group-level incident-management playbooks, runbooks, and escalation procedures that standardize response and reduce time-to-resolution.MUST HAVES- 5+ years of experience.- In-depth architectural expertise in multi-cloud defense, federated IAM, and zero-trust principles.- Strong practical experience with Kubernetes, Terraform, CI/CD orchestration, and Python/Go scripting.- Senior-level, hands-on incident-command experience driving major/critical incident calls to resolution in a 24x7 production environment.- Proven track record of remediation follow-up — coordinating with teams and holding owners accountable until issues are fully closed.- Demonstrated skill drafting and issuing incident notification communications to both technical and executive audiences.- Direct experience authoring divisional/group incident-management playbooks and escalation procedures.- Fully autonomous.- Drives the architecture of complex automated runbooks and mentors Middle-level SREs.- Extensive experience deploying and tuning APIs from modern CNAPP/CSPM platforms, ideally Wiz.- Prior experience building platforms subject to strict financial compliance standards (PCI-DSS, SOC2).- Upper-intermediate English level.NICE TO HAVES- PagerDuty — hands-on experience with on-call scheduling, alert routing, and incident orchestration.- ServiceNow — familiarity with incident, problem, and change management workflows and reporting.PERKS AND BENEFITS- Growth without limits: build your skills through mentorship, internal TechTalks, challenging projects, and a dedicated annual learning budget- Competitive compensation: get recognition that reflects your skills and impact, with regular performance and compensation reviews- Flexibility: work 100% remotely with flexible hours that support focus, autonomy, and a healthy work rhythm- Meaningful, modern projects: build impactful products using modern technologies alongside global teams and leading brands- Collaborative culture: join a supportive environment with zero micromanagement where ideas are welcomed and contributions are recognized- Well-being & support: access local well-being programs and people-focused support tailored to your locationAgileEngine is an Inc. 5000 company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries. We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards.WHY JOIN USWHY JOIN USIf you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet you!ABOUT THE ROLEABOUT THE ROLEWe are looking for a DevOps / Site Reliability Engineer to maintain operational resilience and 24/7 stability for a multi-cloud enterprise security program, serving as Incident Commander on major and critical incidents while also owning IaC, CI/CD pipelines, and CSPM telemetry. You will drive major-incident calls, own post-incident remediation follow-through, draft stakeholder communications, and develop divisional incident-management playbooks alongside multi-cloud security guardrails using Terraform and Wiz. The role requires 5+ years of SRE experience with hands-on incident command in a 24/7 financial services environment.DevOps / Site Reliability EngineerWHAT YOU WILL DOWHAT YOU WILL DO- Scale and maintain the ability to drive operational stability across multi-cloud environments (Azure, AWS, GCP).- Engineer unified security policies and configuration baselines using IaC (Terraform) to prevent misconfigurations.- Design, maintain, and optimize enterprise CI/CD pipelines to support continuous ASPM ingestion and deployment.- Act on continuous monitoring alerts, utilizing Cloud Security Posture Management (CSPM) tools like Wiz to secure workloads.- Serve as Incident Commander on major and critical incidents — running the bridge, directing technical workstreams, making time-critical decisions, and coordinating cross-functional responders under pressure.- Own the post-incident loop — track remediation items to closure, hold owning teams accountable to timelines, and drive systemic fixes and preventative actions across groups.- Draft and send clear, accurate, audience-appropriate incident notifications and status updates to technical teams, management, and stakeholders throughout the incident lifecycle.- Develop, maintain, and socialize divisional / group-level incident-management playbooks, runbooks, and escalation procedures that standardize response and reduce time-to-resolution.MUST HAVESMUST HAVES- 5+ years of experience.5+ years of experience- In-depth architectural expertise in multi-cloud defense, federated IAM, and zero-trust principles.multi-cloud defense, federated IAM, and zero-trust principles- Strong practical experience with Kubernetes, Terraform, CI/CD orchestration, and Python/Go scripting.Kubernetes, Terraform, CI/CD orchestration, and Python/Go scripting- Senior-level, hands-on incident-command experience driving major/critical incident calls to resolution in a 24x7 production environment.Senior-level, hands-on incident-command experience- Proven track record of remediation follow-up — coordinating with teams and holding owners accountable until issues are fully closed.- Demonstrated skill drafting and issuing incident notification communications to both technical and executive audiences.- Direct experience authoring divisional/group incident-management playbooks and escalation procedures.- Fully autonomous.- Drives the architecture of complex automated runbooks and mentors Middle-level SREs.- Extensive experience deploying and tuning APIs from modern CNAPP/CSPM platforms, ideally Wiz.CNAPP/CSPM platforms, ideally Wiz- Prior experience building platforms subject to strict financial compliance standards (PCI-DSS, SOC2).financial compliance standards (PCI-DSS, SOC2)- Upper-intermediate English level.NICE TO HAVESNICE TO HAVES- PagerDuty — hands-on experience with on-call scheduling, alert routing, and incident orchestration.- ServiceNow — familiarity with incident, problem, and change management workflows and reporting.PERKS AND BENEFITSPERKS AND BENEFITS- Growth without limits: build your skills through mentorship, internal TechTalks, challenging projects, and a dedicated annual learning budgetGrowth without limits- Competitive compensation: get recognition that reflects your skills and impact, with regular performance and compensation reviewsCompetitive compensation- Flexibility: work 100% remotely with flexible hours that support focus, autonomy, and a healthy work rhythmFlexibility- Meaningful, modern projects: build impactful products using modern technologies alongside global teams and leading brandsMeaningful, modern projects- Collaborative culture: join a supportive environment with zero micromanagement where ideas are welcomed and contributions are recognizedCollaborative culture- Well-being & support: access local well-being programs and people-focused support tailored to your locationWell-being & supportI'm interested{{criteriaIndex}} {{field}} {{captialize(condition)}} ( ) {{cxPropHelpText}} {{cxPropField.consent_details.description}} {{record.Posting_Title}} {{record.Posting_Title}} {{record.Posting_Title}} {{candidate.Email}} {{topMessage}} Job Details {{topMessage}} previous next {{title}} {{score}}% Job Description {{unescape(sanitizeHTML(descriptionHTML))}} {{message}} Step {{curStepInMandatorySecPrompt}}/{{totalNumOfStepsInMandatorySecPrompt}} For privacy and security purposes, please go through the following points and provide consent. Accept Decline {{getI18n("zr.careers.publicpage.meta.joblisting")}} {{getI18n("zr.careers.publicpage.meta.joblisting")}} {{cxPropAssessment.name}} {{getI18n("zr.questionnaire.question.information.heading")}} {{getI18n("zr.cw.autofill.head")}} {{cxPropBodyMessage}} {{getI18n("zr.careers.autopopulate.success")}} {{getI18n("zr.cw.apply.sign",cxPropCompanyInfo.name)}} {{if(cxPropDetails.applyWithSeekButton,cxPropDetails.applyWithSeekButton.buttonLabel,getI18n("zr.quickapply.apply.toggle","Seek"))}} {{getI18n('zr.eeo.questionnaire.portal.maintitle')}} {{botName}} {{getI18n("zr.zia.sb.consent.msg1")}}{{criteriaIndex}}{{criteriaIndex}}{{field}} {{captialize(condition)}}{{field}}{{captialize(condition)}}( )(){{cxPropHelpText}}{{cxPropHelpText}}{{cxPropField.consent_details.description}}

{{cxPropField.consent_details.description}}

{{record.Posting_Title}}{{record.Posting_Title}}{{record.Posting_Title}}

{{record.Posting_Title}}

{{record.Posting_Title}}{{record.Posting_Title}}

{{record.Posting_Title}}

{{record.Posting_Title}}
  • {{record.Posting_Title}}

{{record.Posting_Title}}

{{record.Posting_Title}}{{candidate.Email}}{{candidate.Email}}{{candidate.Email}}

{{candidate.Email}}

{{topMessage}}Job DetailsJob DetailsJob DetailsJob Details{{topMessage}}previousnext{{title}} {{score}}% Job Description {{unescape(sanitizeHTML(descriptionHTML))}}{{title}} {{score}}% Job Description {{unescape(sanitizeHTML(descriptionHTML))}}{{title}} {{score}}%

{{title}}

{{title}}{{score}}%Job Description {{unescape(sanitizeHTML(descriptionHTML))}}

Job Description

{{message}}

Step {{curStepInMandatorySecPrompt}}/{{totalNumOfStepsInMandatorySecPrompt}}

For privacy and security purposes, please go through the following points and provide consent.Accept DeclineAccept DeclineAcceptDecline{{getI18n("zr.careers.publicpage.meta.joblisting")}}{{getI18n("zr.careers.publicpage.meta.joblisting")}}
  • {{getI18n("zr.careers.publicpage.meta.joblisting")}}
{{getI18n("zr.careers.publicpage.meta.joblisting")}}{{getI18n("zr.careers.publicpage.meta.joblisting")}}{{getI18n("zr.careers.publicpage.meta.joblisting")}}
  • {{getI18n("zr.careers.publicpage.meta.joblisting")}}
{{getI18n("zr.careers.publicpage.meta.joblisting")}}{{cxPropAssessment.name}} {{getI18n("zr.questionnaire.question.information.heading")}}{{cxPropAssessment.name}} {{getI18n("zr.questionnaire.question.information.heading")}}{{cxPropAssessment.name}}

{{getI18n("zr.questionnaire.question.information.heading")}}

{{getI18n("zr.cw.autofill.head")}} {{cxPropBodyMessage}} {{getI18n("zr.careers.autopopulate.success")}}

{{getI18n("zr.cw.autofill.head")}}

{{cxPropBodyMessage}}

{{getI18n("zr.careers.autopopulate.success")}}

{{getI18n("zr.cw.apply.sign",cxPropCompanyInfo.name)}}

{{getI18n('zr.eeo.questionnaire.portal.maintitle')}}{{botName}}{{botName}}{{getI18n("zr.zia.sb.consent.msg1")}}{{getI18n("zr.zia.sb.consent.msg1")}}{{getI18n("zr.zia.sb.consent.msg1")}}{{getI18n("zr.zia.sb.consent.msg1")}}{{getI18n("zr.zia.sb.consent.msg1")}}{{getI18n("zr.zia.sb.consent.msg1")}}{{getI18n("zr.zia.sb.consent.msg1")}}

Similar jobs