Sr Platform/Infrastructure Engineer

Active

Job description

Summary

We are seeking a senior Sr Platform/Infrastructure Engineer to strengthen our platform team and drive cloud-native infrastructure initiatives. This role focuses on deploying and maintaining Kubernetes services, integrating monitoring and storage platforms, and troubleshooting distributed systems to ensure resilient, scalable operations.

You will work with Python-driven tooling, Prometheus-based monitoring, Ceph-backed storage, and public cloud environments (AWS and Azure) to modernize and operate our platform. This is an opportunity to shape platform reliability and performance in a hands-on engineering role.

Responsibilities

Design, deploy, and maintain production Kubernetes clusters and related services.

Build and maintain automation and tooling using Python to support platform operations.

Integrate and operate Prometheus for monitoring, alerting, and observability.

Deploy and manage Ceph storage solutions for distributed workloads.

Support platform modernization initiatives and migrate services to cloud-native patterns.

Troubleshoot and resolve issues in distributed systems across compute, storage, and network layers.

Collaborate with development, SRE, and operations teams to define platform requirements and SLAs.

Document platform designs, runbooks, and operational procedures.

Participate in on-call rotations and incident response to maintain platform availability.