Senior Site Reliability Engineer ID81653

Active

Job description

AgileEngine is an Inc. 5000 company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries. We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards.WHY JOIN USIf you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet you!ABOUT THE ROLEWe are looking for a Senior Site Reliability Engineer to support platform reliability, monitoring, and modernization across Kubernetes-based microservices environments with a strong observability focus. You will build and maintain Datadog solutions including dashboards, alerts, APM, metrics, logging, and tracing, integrate observability tooling into AWS and CI/CD pipelines, and automate monitoring and operational tasks using Python. The role blends software engineering (60–70%) with site reliability engineering (30–40%) and requires JST timezone overlap.WHAT YOU WILL DO- Support platform reliability, monitoring, and continuous improvement across internal systems.- Work in Kubernetes-based environments.- Build and maintain observability solutions, with a focus on Datadog.- Configure dashboards, alerts, APM, metrics, logging, and tracing.- Monitor containerized and microservices-based applications.- Integrate observability tools into AWS environments.- Integrate observability into CI/CD pipelines.- Automate monitoring and operational tasks using scripting (Python preferred).- Install and configure Datadog agents and integrations.- Manage API keys and secure configurations.- Manage user roles and access controls within observability platforms.- Lead maintenance efforts and platform improvements while driving reliability, scalability, and performance.MUST HAVES- Strong proficiency in Python, JavaScript (Node.js), or Java.- Hands-on experience with API integrations (designing, consuming, and integrating).- Strong experience working in Kubernetes environments (deployment, operations, monitoring).- Experience with Datadog (preferred) or similar tools (Prometheus, Grafana).- Ability to configure dashboards, alerts, and APM (tracing, metrics, logging).- Experience monitoring containerized/microservices architectures.- Hands-on experience with AWS.- Experience integrating observability tools into cloud environments.- Experience integrating observability into CI/CD pipelines.- Ability to automate monitoring and operational tasks using scripting (Python preferred).- Upper-intermediate English level.NICE TO HAVES- Experience owning and operating an internal engineering platform.- Demonstrated ownership of reliability, scalability, and performance.- Proven ability to proactively lead maintenance efforts and platform improvements (not just reactive support).- Familiarity with Golang.- Experience with additional observability tools such as New Relic, Dynatrace, Elastic, or Splunk Observability.PERKS AND BENEFITS- Growth without limits: build your skills through mentorship, internal TechTalks, challenging projects, and a dedicated annual learning budget- Competitive compensation: get recognition that reflects your skills and impact, with regular performance and compensation reviews- Flexibility: work 100% remotely with flexible hours that support focus, autonomy, and a healthy work rhythm- Meaningful, modern projects: build impactful products using modern technologies alongside global teams and leading brands- Collaborative culture: join a supportive environment with zero micromanagement where ideas are welcomed and contributions are recognized- Well-being & support: access local well-being programs and people-focused support tailored to your locationAgileEngine is an Inc. 5000 company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries. We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards.WHY JOIN USIf you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet you!ABOUT THE ROLEWe are looking for a Senior Site Reliability Engineer to support platform reliability, monitoring, and modernization across Kubernetes-based microservices environments with a strong observability focus. You will build and maintain Datadog solutions including dashboards, alerts, APM, metrics, logging, and tracing, integrate observability tooling into AWS and CI/CD pipelines, and automate monitoring and operational tasks using Python. The role blends software engineering (60–70%) with site reliability engineering (30–40%) and requires JST timezone overlap.WHAT YOU WILL DO- Support platform reliability, monitoring, and continuous improvement across internal systems.- Work in Kubernetes-based environments.- Build and maintain observability solutions, with a focus on Datadog.- Configure dashboards, alerts, APM, metrics, logging, and tracing.- Monitor containerized and microservices-based applications.- Integrate observability tools into AWS environments.- Integrate observability into CI/CD pipelines.- Automate monitoring and operational tasks using scripting (Python preferred).- Install and configure Datadog agents and integrations.- Manage API keys and secure configurations.- Manage user roles and access controls within observability platforms.- Lead maintenance efforts and platform improvements while driving reliability, scalability, and performance.MUST HAVES- Strong proficiency in Python, JavaScript (Node.js), or Java.- Hands-on experience with API integrations (designing, consuming, and integrating).- Strong experience working in Kubernetes environments (deployment, operations, monitoring).- Experience with Datadog (preferred) or similar tools (Prometheus, Grafana).- Ability to configure dashboards, alerts, and APM (tracing, metrics, logging).- Experience monitoring containerized/microservices architectures.- Hands-on experience with AWS.- Experience integrating observability tools into cloud environments.- Experience integrating observability into CI/CD pipelines.- Ability to automate monitoring and operational tasks using scripting (Python preferred).- Upper-intermediate English level.NICE TO HAVES- Experience owning and operating an internal engineering platform.- Demonstrated ownership of reliability, scalability, and performance.- Proven ability to proactively lead maintenance efforts and platform improvements (not just reactive support).- Familiarity with Golang.- Experience with additional observability tools such as New Relic, Dynatrace, Elastic, or Splunk Observability.PERKS AND BENEFITS- Growth without limits: build your skills through mentorship, internal TechTalks, challenging projects, and a dedicated annual learning budget- Competitive compensation: get recognition that reflects your skills and impact, with regular performance and compensation reviews- Flexibility: work 100% remotely with flexible hours that support focus, autonomy, and a healthy work rhythm- Meaningful, modern projects: build impactful products using modern technologies alongside global teams and leading brands- Collaborative culture: join a supportive environment with zero micromanagement where ideas are welcomed and contributions are recognized- Well-being & support: access local well-being programs and people-focused support tailored to your locationAgileEngine is an Inc. 5000 company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries. We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards.WHY JOIN USWHY JOIN USIf you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet you!ABOUT THE ROLEABOUT THE ROLEWe are looking for a Senior Site Reliability Engineer to support platform reliability, monitoring, and modernization across Kubernetes-based microservices environments with a strong observability focus. You will build and maintain Datadog solutions including dashboards, alerts, APM, metrics, logging, and tracing, integrate observability tooling into AWS and CI/CD pipelines, and automate monitoring and operational tasks using Python. The role blends software engineering (60–70%) with site reliability engineering (30–40%) and requires JST timezone overlap.Senior Site Reliability EngineerWHAT YOU WILL DOWHAT YOU WILL DO- Support platform reliability, monitoring, and continuous improvement across internal systems.- Work in Kubernetes-based environments.- Build and maintain observability solutions, with a focus on Datadog.- Configure dashboards, alerts, APM, metrics, logging, and tracing.- Monitor containerized and microservices-based applications.- Integrate observability tools into AWS environments.- Integrate observability into CI/CD pipelines.- Automate monitoring and operational tasks using scripting (Python preferred).- Install and configure Datadog agents and integrations.- Manage API keys and secure configurations.- Manage user roles and access controls within observability platforms.- Lead maintenance efforts and platform improvements while driving reliability, scalability, and performance.MUST HAVESMUST HAVES- Strong proficiency in Python, JavaScript (Node.js), or Java.Python, JavaScript (Node.js), or Java- Hands-on experience with API integrations (designing, consuming, and integrating).API integrations- Strong experience working in Kubernetes environments (deployment, operations, monitoring).Kubernetes- Experience with Datadog (preferred) or similar tools (Prometheus, Grafana).Datadog- Ability to configure dashboards, alerts, and APM (tracing, metrics, logging).dashboards, alerts, and APM- Experience monitoring containerized/microservices architectures.containerized/microservices architectures- Hands-on experience with AWS.AWS- Experience integrating observability tools into cloud environments.- Experience integrating observability into CI/CD pipelines.CI/CD pipelines- Ability to automate monitoring and operational tasks using scripting (Python preferred).Python- Upper-intermediate English level.NICE TO HAVESNICE TO HAVES- Experience owning and operating an internal engineering platform.- Demonstrated ownership of reliability, scalability, and performance.- Proven ability to proactively lead maintenance efforts and platform improvements (not just reactive support).- Familiarity with Golang.- Experience with additional observability tools such as New Relic, Dynatrace, Elastic, or Splunk Observability.PERKS AND BENEFITSPERKS AND BENEFITS- Growth without limits: build your skills through mentorship, internal TechTalks, challenging projects, and a dedicated annual learning budgetGrowth without limits- Competitive compensation: get recognition that reflects your skills and impact, with regular performance and compensation reviewsCompetitive compensation- Flexibility: work 100% remotely with flexible hours that support focus, autonomy, and a healthy work rhythmFlexibility- Meaningful, modern projects: build impactful products using modern technologies alongside global teams and leading brandsMeaningful, modern projects- Collaborative culture: join a supportive environment with zero micromanagement where ideas are welcomed and contributions are recognizedCollaborative culture- Well-being & support: access local well-being programs and people-focused support tailored to your locationWell-being & supportI'm interested{{criteriaIndex}} {{field}} {{captialize(condition)}} ( ) {{cxPropHelpText}} {{cxPropField.consent_details.description}} {{record.Posting_Title}} {{record.Posting_Title}} {{record.Posting_Title}} {{candidate.Email}} {{topMessage}} Job Details {{topMessage}} previous next {{title}} {{score}}% Job Description {{unescape(sanitizeHTML(descriptionHTML))}} {{message}} Step {{curStepInMandatorySecPrompt}}/{{totalNumOfStepsInMandatorySecPrompt}} For privacy and security purposes, please go through the following points and provide consent. Accept Decline {{getI18n("zr.careers.publicpage.meta.joblisting")}} {{getI18n("zr.careers.publicpage.meta.joblisting")}} {{cxPropAssessment.name}} {{getI18n("zr.questionnaire.question.information.heading")}} {{getI18n("zr.cw.autofill.head")}} {{cxPropBodyMessage}} {{getI18n("zr.careers.autopopulate.success")}} {{getI18n("zr.cw.apply.sign",cxPropCompanyInfo.name)}} {{if(cxPropDetails.applyWithSeekButton,cxPropDetails.applyWithSeekButton.buttonLabel,getI18n("zr.quickapply.apply.toggle","Seek"))}} {{getI18n('zr.eeo.questionnaire.portal.maintitle')}} {{botName}} {{getI18n("zr.zia.sb.consent.msg1")}}{{criteriaIndex}}{{criteriaIndex}}{{field}} {{captialize(condition)}}{{field}}{{captialize(condition)}}( )(){{cxPropHelpText}}{{cxPropHelpText}}{{cxPropField.consent_details.description}}

{{cxPropField.consent_details.description}}

{{record.Posting_Title}}{{record.Posting_Title}}{{record.Posting_Title}}

{{record.Posting_Title}}

{{record.Posting_Title}}{{record.Posting_Title}}

{{record.Posting_Title}}

{{record.Posting_Title}}
  • {{record.Posting_Title}}

{{record.Posting_Title}}

{{record.Posting_Title}}{{candidate.Email}}{{candidate.Email}}{{candidate.Email}}

{{candidate.Email}}

{{topMessage}}Job DetailsJob DetailsJob DetailsJob Details{{topMessage}}previousnext{{title}} {{score}}% Job Description {{unescape(sanitizeHTML(descriptionHTML))}}{{title}} {{score}}% Job Description {{unescape(sanitizeHTML(descriptionHTML))}}{{title}} {{score}}%

{{title}}

{{title}}{{score}}%Job Description {{unescape(sanitizeHTML(descriptionHTML))}}

Job Description

{{message}}

Step {{curStepInMandatorySecPrompt}}/{{totalNumOfStepsInMandatorySecPrompt}}

For privacy and security purposes, please go through the following points and provide consent.Accept DeclineAccept DeclineAcceptDecline{{getI18n("zr.careers.publicpage.meta.joblisting")}}{{getI18n("zr.careers.publicpage.meta.joblisting")}}
  • {{getI18n("zr.careers.publicpage.meta.joblisting")}}
{{getI18n("zr.careers.publicpage.meta.joblisting")}}{{getI18n("zr.careers.publicpage.meta.joblisting")}}{{getI18n("zr.careers.publicpage.meta.joblisting")}}
  • {{getI18n("zr.careers.publicpage.meta.joblisting")}}
{{getI18n("zr.careers.publicpage.meta.joblisting")}}{{cxPropAssessment.name}} {{getI18n("zr.questionnaire.question.information.heading")}}{{cxPropAssessment.name}} {{getI18n("zr.questionnaire.question.information.heading")}}{{cxPropAssessment.name}}

{{getI18n("zr.questionnaire.question.information.heading")}}

{{getI18n("zr.cw.autofill.head")}} {{cxPropBodyMessage}} {{getI18n("zr.careers.autopopulate.success")}}

{{getI18n("zr.cw.autofill.head")}}

{{cxPropBodyMessage}}

{{getI18n("zr.careers.autopopulate.success")}}

{{getI18n("zr.cw.apply.sign",cxPropCompanyInfo.name)}}

{{getI18n('zr.eeo.questionnaire.portal.maintitle')}}{{botName}}{{botName}}{{getI18n("zr.zia.sb.consent.msg1")}}{{getI18n("zr.zia.sb.consent.msg1")}}{{getI18n("zr.zia.sb.consent.msg1")}}{{getI18n("zr.zia.sb.consent.msg1")}}{{getI18n("zr.zia.sb.consent.msg1")}}{{getI18n("zr.zia.sb.consent.msg1")}}{{getI18n("zr.zia.sb.consent.msg1")}}

Similar jobs