Job description
Role Summary
Role SummaryRole SummaryRole SummaryRole SummaryRole SummaryThe Site Reliability Engineer (SRE) at iCareManager (iCM) is responsible for ensuring the reliability, scalability, performance, and availability of production systems. The role blends software engineering and operations, with a strong focus on automation, observability, incident management, and proactive reliability engineering.
The Site Reliability Engineer (SRE) at iCareManager (iCM) is responsible for ensuring the reliability, scalability, performance, and availability of production systems. The role blends software engineering and operations, with a strong focus on automation, observability, incident management, and proactive reliability engineering.The Site Reliability Engineer (SRE) at iCareManager (iCM) is responsible for ensuring the reliability, scalability, performance, and availability of production systems. The role blends software engineering and operations, with a strong focus on automation, observability, incident management, and proactive reliability engineering.The Site Reliability Engineer (SRE) at iCareManager (iCM) is responsible for ensuring the reliability, scalability, performance, and availability of production systems. The role blends software engineering and operations, with a strong focus on automation, observability, incident management, and proactive reliability engineering.The Site Reliability Engineer (SRE) at iCareManager (iCM) is responsible for ensuring the reliability, scalability, performance, and availability of production systems. The role blends software engineering and operations, with a strong focus on automation, observability, incident management, and proactive reliability engineering.SREs enable engineering teams to deliver features rapidly without compromising system stability, especially in regulated healthcare environments.
SREs enable engineering teams to deliver features rapidly without compromising system stability, especially in regulated healthcare environments.SREs enable engineering teams to deliver features rapidly without compromising system stability, especially in regulated healthcare environments.SREs enable engineering teams to deliver features rapidly without compromising system stability, especially in regulated healthcare environments.SREs enable engineering teams to deliver features rapidly without compromising system stability, especially in regulated healthcare environments.Key Responsibilities
Key ResponsibilitiesKey ResponsibilitiesKey ResponsibilitiesKey ResponsibilitiesKey Responsibilities1. Reliability Engineering
1. Reliability Engineering1. Reliability Engineering1. Reliability Engineering1. Reliability Engineering1. Reliability Engineering- Define, implement, and enforce Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets
Define, implement, and enforce Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets
Define, implement, and enforce Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgetsDefine, implement, and enforce Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgetsDefine, implement, and enforce Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgetsDefine, implement, and enforce Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets- Design and review resilience patterns (redundancy, failover, graceful degradation)
Design and review resilience patterns (redundancy, failover, graceful degradation)
Design and review resilience patterns (redundancy, failover, graceful degradation)Design and review resilience patterns (redundancy, failover, graceful degradation)Design and review resilience patterns (redundancy, failover, graceful degradation)Design and review resilience patterns (redundancy, failover, graceful degradation)- Perform capacity planning, load modeling, and scalability analysis
Perform capacity planning, load modeling, and scalability analysis
Perform capacity planning, load modeling, and scalability analysisPerform capacity planning, load modeling, and scalability analysisPerform capacity planning, load modeling, and scalability analysisPerform capacity planning, load modeling, and scalability analysis- Conduct chaos testing and failure injection to identify system weaknesses
Conduct chaos testing and failure injection to identify system weaknesses
Conduct chaos testing and failure injection to identify system weaknessesConduct chaos testing and failure injection to identify system weaknessesConduct chaos testing and failure injection to identify system weaknessesConduct chaos testing and failure injection to identify system weaknesses- Reduce Mean Time to Recovery (MTTR) through architectural improvements and tooling
Reduce Mean Time to Recovery (MTTR) through architectural improvements and tooling
Reduce Mean Time to Recovery (MTTR) through architectural improvements and toolingReduce Mean Time to Recovery (MTTR) through architectural improvements and toolingReduce Mean Time to Recovery (MTTR) through architectural improvements and toolingReduce Mean Time to Recovery (MTTR) through architectural improvements and tooling2. Observability & Monitoring
2. Observability & Monitoring2. Observability & Monitoring2. Observability & Monitoring2. Observability & Monitoring2. Observability & Monitoring- Instrument systems with metrics, logs, and distributed traces
Instrument systems with metrics, logs, and distributed traces
Instrument systems with metrics, logs, and distributed tracesInstrument systems with metrics, logs, and distributed tracesInstrument systems with metrics, logs, and distributed tracesInstrument systems with metrics, logs, and distributed traces- Build and maintain dashboards that reflect system health and performance
Build and maintain dashboards that reflect system health and performance
Build and maintain dashboards that reflect system health and performanceBuild and maintain dashboards that reflect system health and performanceBuild and maintain dashboards that reflect system health and performanceBuild and maintain dashboards that reflect system health and performance- Design alerting strategies that are actionable and minimize alert fatigue
Design alerting strategies that are actionable and minimize alert fatigue
Design alerting strategies that are actionable and minimize alert fatigueDesign alerting strategies that are actionable and minimize alert fatigueDesign alerting strategies that are actionable and minimize alert fatigueDesign alerting strategies that are actionable and minimize alert fatigue- Identify leading indicators of failure before customer impact
Identify leading indicators of failure before customer impact
Identify leading indicators of failure before customer impactIdentify leading indicators of failure before customer impactIdentify leading indicators of failure before customer impactIdentify leading indicators of failure before customer impact3. Incident Management & Postmortems
3. Incident Management & Postmortems3. Incident Management & Postmortems3. Incident Management & Postmortems3. Incident Management & Postmortems3. Incident Management & Postmortems- Participate in and lead production incident response
Participate in and lead production incident response
Participate in and lead production incident responseParticipate in and lead production incident responseParticipate in and lead production incident responseParticipate in and lead production incident response- Coordinate with engineering and infrastructure teams during incidents
Coordinate with engineering and infrastructure teams during incidents
Coordinate with engineering and infrastructure teams during incidentsCoordinate with engineering and infrastructure teams during incidentsCoordinate with engineering and infrastructure teams during incidentsCoordinate with engineering and infrastructure teams during incidents- Lead blameless postmortems and document root cause analysis
Lead blameless postmortems and document root cause analysis
Lead blameless postmortems and document root cause analysisLead blameless postmortems and document root cause analysisLead blameless postmortems and document root cause analysisLead blameless postmortems and document root cause analysis- Track and remediate reliability debt and systemic risks
Track and remediate reliability debt and systemic risks
Track and remediate reliability debt and systemic risksTrack and remediate reliability debt and systemic risksTrack and remediate reliability debt and systemic risksTrack and remediate reliability debt and systemic risks