Site Reliability Engineer (SRE)

InfosysSantiagoMid Level
Active

Job description

Job Description – Site Reliability Engineer (SRE)Role PurposeThe Site Reliability Engineer (SRE) is responsible for ensuring the reliability, availability, and performance of digital services in production, balancing service stability with the ability to deliver change at speed.The role focuses on strengthening operational resilience through engineering, automation, and proactive reliability practices, working closely with application and platform teams.Scope of the RoleThe locally applied SRE role covers:

Job Description – Site Reliability Engineer (SRE)Job Description – Site Reliability Engineer (SRE)Job Description – Site Reliability Engineer (SRE)Job Description – Site Reliability Engineer (SRE)Role PurposeRole PurposeRole PurposeRole PurposeThe Site Reliability Engineer (SRE) is responsible for ensuring the reliability, availability, and performance of digital services in production, balancing service stability with the ability to deliver change at speed.The role focuses on strengthening operational resilience through engineering, automation, and proactive reliability practices, working closely with application and platform teams.The Site Reliability Engineer (SRE) is responsible for ensuring the reliability, availability, and performance of digital services in production, balancing service stability with the ability to deliver change at speed.The role focuses on strengthening operational resilience through engineering, automation, and proactive reliability practices, working closely with application and platform teams.The Site Reliability Engineer (SRE) is responsible for ensuring the reliability, availability, and performance of digital services in production, balancing service stability with the ability to deliver change at speed.The role focuses on strengthening operational resilience through engineering, automation, and proactive reliability practices, working closely with application and platform teams.Scope of the RoleScope of the RoleScope of the RoleScope of the RoleThe locally applied SRE role covers:The locally applied SRE role covers:The locally applied SRE role covers:
  • Production digital services (applications, platforms, data products)
Production digital services (applications, platforms, data products)Production digital services (applications, platforms, data products)Production digital services (applications, platforms, data products)Production digital services (applications, platforms, data products)
  • Associated infrastructure (cloud, CI/CD pipelines, integrations)
Associated infrastructure (cloud, CI/CD pipelines, integrations)Associated infrastructure (cloud, CI/CD pipelines, integrations)Associated infrastructure (cloud, CI/CD pipelines, integrations)Associated infrastructure (cloud, CI/CD pipelines, integrations)
  • Continuous operations (24/7 reliability mindset, not necessarily shift-based)
Continuous operations (24/7 reliability mindset, not necessarily shift-based)Continuous operations (24/7 reliability mindset, not necessarily shift-based)Continuous operations (24/7 reliability mindset, not necessarily shift-based)Continuous operations (24/7 reliability mindset, not necessarily shift-based)
  • Production changes (deployments, configurations)
Production changes (deployments, configurations)Production changes (deployments, configurations)Production changes (deployments, configurations)Production changes (deployments, configurations)
  • Incidents, problems, and service degradations
Incidents, problems, and service degradationsIncidents, problems, and service degradationsIncidents, problems, and service degradationsIncidents, problems, and service degradations
  • Continuous improvement of stability and operational efficiency
Continuous improvement of stability and operational efficiencyContinuous improvement of stability and operational efficiencyContinuous improvement of stability and operational efficiencyContinuous improvement of stability and operational efficiencyKey ResponsibilitiesKey ResponsibilitiesKey ResponsibilitiesKey ResponsibilitiesService Reliability & AvailabilityService Reliability & AvailabilityService Reliability & AvailabilityService Reliability & Availability
  • Define, implement, and maintain SLIs and SLOs(availability, latency, error rates)
Define, implement, and maintain SLIs and SLOs(availability, latency, error rates)Define, implement, and maintain SLIs and SLOs(availability, latency, error rates)Define, implement, and maintain SLIs and SLOs(availability, latency, error rates)Define, implement, and maintain SLIs and SLOs(availability, latency, error rates)
  • Continuously monitor service health and anticipate degradations
Continuously monitor service health and anticipate degradationsContinuously monitor service health and anticipate degradationsContinuously monitor service health and anticipate degradationsContinuously monitor service health and anticipate degradations
  • Ensure services operate within business‑agreed reliability thresholds
Ensure services operate within business‑agreed reliability thresholdsEnsure services operate within business‑agreed reliability thresholdsEnsure services operate within business‑agreed reliability thresholdsEnsure services operate within business‑agreed reliability thresholds
  • Manage reliability trade‑offs between speed and stability
Manage reliability trade‑offs between speed and stabilityManage reliability trade‑offs between speed and stabilityManage reliability trade‑offs between speed and stabilityManage reliability trade‑offs between speed and stabilityIncident & Problem ManagementIncident & Problem ManagementIncident & Problem ManagementIncident & Problem Management
  • Lead or coordinate response to relevant incidents (L2/L3)
Lead or coordinate response to relevant incidents (L2/L3)Lead or coordinate response to relevant incidents (L2/L3)Lead or coordinate response to relevant incidents (L2/L3)Lead or coordinate response to relevant incidents (L2/L3)
  • Ensure: Rapid and structured diagnosisSafe service restorationClear and effective communication
Ensure:Ensure:Ensure:Ensure:
  • Rapid and structured diagnosis
Rapid and structured diagnosisRapid and structured diagnosisRapid and structured diagnosisRapid and structured diagnosis
  • Safe service restoration
Safe service restorationSafe service restorationSafe service restorationSafe service restoration
  • Clear and effective communication
Clear and effective communicationClear and effective communicationClear and effective communicationClear and effective communication
  • Facilitate blameless postmortems
Facilitate blameless postmortemsFacilitate blameless postmortemsFacilitate blameless postmortemsFacilitate blameless postmortems
  • Convert recurring incidents into engineering improvement backlog
Convert recurring incidents into engineering improvement backlogConvert recurring incidents into engineering improvement backlogConvert recurring incidents into engineering improvement backlogConvert recurring incidents into engineering improvement backlog
  • Drive long‑term remediation rather than reactive firefighting
Drive long‑term remediation rather than reactive firefightingDrive long‑term remediation rather than reactive firefightingDrive long‑term remediation rather than reactive firefightingDrive long‑term remediation rather than reactive firefightingAutomation & Operational ExcellenceAutomation & Operational ExcellenceAutomation & Operational ExcellenceAutomation & Operational Excellence
  • Identify repetitive and manual operational tasks
Identify repetitive and manual operational tasksIdentify repetitive and manual operational tasksIdentify repetitive and manual operational tasksIdentify repetitive and manual operational tasks
  • Design and implement automation for: DeploymentsMonitoring and alertingHealth checksBasic recovery and self‑healing (where applicable)
Design and implement automation for:Design and implement automation for:Design and implement automation for:Design and implement automation for:
  • Deployments
DeploymentsDeploymentsDeploymentsDeployments
  • Monitoring and alerting
Monitoring and alertingMonitoring and alertingMonitoring and alertingMonitoring and alerting
  • Health checks
Health checksHealth checksHealth checksHealth checks
  • Basic recovery and self‑healing (where applicable)
Basic recovery and self‑healing (where applicable)Basic recovery and self‑healing (where applicable)Basic recovery and self‑healing (where applicable)Basic recovery and self‑healing (where applicable)
  • Reduce toil and increase system resilience through engineering solutions
Reduce toil and increase system resilience through engineering solutionsReduce toil and increase system resilience through engineering solutionsReduce toil and increase system resilience through engineering solutionsReduce toil and increase system resilience through engineering solutionsChange Governance & Production ReadinessChange Governance & Production ReadinessChange Governance & Production ReadinessChange Governance & Production Readiness
  • Support vendor and internal team change tracking
Support vendor and internal team change trackingSupport vendor and internal team change trackingSupport vendor and internal team change trackingSupport vendor and internal team change tracking
  • Ensure changes: Are traceableHave defined rollback strategiesMinimize operational risk
Ensure changes:Ensure changes:Ensure changes:Ensure changes:
  • Are traceable
Are traceableAre traceableAre traceableAre traceable
  • Have defined rollback strategies
Have defined rollback strategiesHave defined rollback strategiesHave defined rollback strategiesHave defined rollback strategies
  • Minimize operational risk
Minimize operational riskMinimize operational riskMinimize operational riskMinimize operational risk
  • Validate operational readiness before production
Validate operational readiness before productionValidate operational readiness before productionValidate operational readiness before productionValidate operational readiness before production
  • Participate early in solution and architecture design from a reliability perspective (early involvement)
Participate early in solution and architecture design from a reliability perspective (early involvement)Participate early in solution and architecture design from a reliability perspective (early involvement)Participate early in solution and architecture design from a reliability perspective (early involvement)Participate early in solution and architecture design from a reliability perspective (early involvement)Metrics, Observability & Continuous ImprovementMetrics, Observability & Continuous ImprovementMetrics, Observability & Continuous ImprovementMetrics, Observability & Continuous Improvement
  • Define and maintain near real‑time operational KPIs (“service pulse”)
Define and maintain near real‑time operational KPIs (“service pulse”)Define and maintain near real‑time operational KPIs (“service pulse”)Define and maintain near real‑time operational KPIs (“service pulse”)Define and maintain near real‑time operational KPIs (“service pulse”)
  • Ensure every deviation has: Clear ownershipDefined corrective actions
Ensure every deviation has:Ensure every deviation has:Ensure every deviation has:Ensure every deviation has:
  • Clear ownership
Clear ownershipClear ownershipClear ownershipClear ownership
  • Defined corrective actions
Defined corrective actionsDefined corrective actionsDefined corrective actionsDefined corrective actions
  • Prevent reactive operations by driving data‑driven decision making
Prevent reactive operations by driving data‑driven decision makingPrevent reactive operations by driving data‑driven decision makingPrevent reactive operations by driving data‑driven decision makingPrevent reactive operations by driving data‑driven decision making
  • Support identification, prioritization, and planning of technical debt remediation
Support identification, prioritization, and planning of technical debt remediationSupport identification, prioritization, and planning of technical debt remediationSupport identification, prioritization, and planning of technical debt remediationSupport identification, prioritization, and planning of technical debt remediationWhat This Role Is NotWhat This Role Is NotWhat This Role Is NotWhat This Role Is NotThe SRE role will not be:The SRE role will not be:The SRE role will not be:
  • A dedicated incident operator only
A dedicated incident operator onlyA dedicated incident operator onlyA dedicated incident operator onlyA dedicated incident operator only
  • An advanced Service Desk
An advanced Service DeskAn advanced Service DeskAn advanced Service DeskAn advanced Service Desk
  • The sole owner of service stability (reliability is shared)
The sole owner of service stability (reliability is shared)The sole owner of service stability (reliability is shared)The sole owner of service stability (reliability is shared)The sole owner of service stability (reliability is shared)
  • A gatekeeper blocking changes without technical justification
A gatekeeper blocking changes without technical justificationA gatekeeper blocking changes without technical justificationA gatekeeper blocking changes without technical justificationA gatekeeper blocking changes without technical justification
  • The owner of contractual MOPs
The owner of contractual MOPsThe owner of contractual MOPsThe owner of contractual MOPsThe owner of contractual MOPs
  • A commercial or account management role
A commercial or account management roleA commercial or account management roleA commercial or account management roleA commercial or account management role
  • The customer-side account or delivery lead
The customer-side account or delivery leadThe customer-side account or delivery leadThe customer-side account or delivery leadThe customer-side account or delivery leadExperience & Profile (Indicative)Experience & Profile (Indicative)Experience & Profile (Indicative)Experience & Profile (Indicative)
  • Proven experience as SRE, Production Engineer, or similar role
Proven experience as SRE, Production Engineer, or similar roleProven experience as SRE, Production Engineer, or similar roleProven experience as SRE, Production Engineer, or similar roleProven experience as SRE, Production Engineer, or similar role
  • Strong background in production systems and reliability engineering
Strong background in production systems and reliability engineeringStrong background in production systems and reliability engineeringStrong background in production systems and reliability engineeringStrong background in production systems and reliability engineering
  • Experience working with: Cloud platformsCI/CD pipelinesMonitoring and observability tools
Experience working with:Experience working with:Experience working with:Experience working with:
  • Cloud platforms
Cloud platformsCloud platformsCloud platformsCloud platforms
  • CI/CD pipelines
CI/CD pipelinesCI/CD pipelinesCI/CD pipelinesCI/CD pipelines
  • Monitoring and observability tools
Monitoring and observability toolsMonitoring and observability toolsMonitoring and observability toolsMonitoring and observability tools
  • Comfortable operating in product‑oriented or POD-based team models
Comfortable operating in product‑oriented or POD-based team modelsComfortable operating in product‑oriented or POD-based team modelsComfortable operating in product‑oriented or POD-based team modelsComfortable operating in product‑oriented or POD-based team models
  • Strong problem‑solving, communication, and collaboration skills
Strong problem‑solving, communication, and collaboration skillsStrong problem‑solving, communication, and collaboration skillsStrong problem‑solving, communication, and collaboration skillsStrong problem‑solving, communication, and collaboration skillsOperating Model AlignmentOperating Model AlignmentOperating Model AlignmentOperating Model Alignment
  • Works embedded or as an enabling function with PODs
Works embedded or as an enabling function with PODsWorks embedded or as an enabling function with PODsWorks embedded or as an enabling function with PODsWorks embedded or as an enabling function with PODs
  • Focused on enablement and reliability patterns, not centralized control
Focused on enablement and reliability patterns, not centralized controlFocused on enablement and reliability patterns, not centralized controlFocused on enablement and reliability patterns, not centralized controlFocused on enablement and reliability patterns, not centralized control
  • Promotes shared ownership of reliability
Promotes shared ownership of reliabilityPromotes shared ownership of reliabilityPromotes shared ownership of reliabilityPromotes shared ownership of reliability

About Us Infosys is a global leader in next-generation digital services and consulting. We enable clients in more than 50 countries to navigate their digital transformation. With over four decades of experience in managing the systems and workings of global enterprises, we expertly steer our clients through their digital journey. We do it by enabling the enterprise with an AI-powered core that helps prioritize the execution of change. We also empower the business with agile digital at scale to deliver unprecedented levels of performance and customer delight. Our always-on learning agenda drives their continuous improvement through building and transferring digital skills, expertise, and ideas from our innovation ecosystem. EEO Infosys provides equal employment opportunities to applicants and employees without regard to race; color; sex; gender identity; sexual orientation; religious practices and observances; national origin; pregnancy, childbirth, or related medical conditions; or disability.

About UsEEO

Similar jobs