Job description
We are hiring a Site Reliability Engineer to Build, support and improve Keystone, an ETL application/platform operating in the China region. Keystone supports data ingestion, transformation, and loading workflows across Kubernetes-based environments, including Airflow-based loader jobs that load data into Snowflake and Spark-on-EKS loader jobs that load data into a Datalake or Lakehouse environment.
The SRE will own reliability, availability, performance, data freshness, and operational readiness for production ETL pipelines. This role requires strong hands-on experience with Linux, Kubernetes, Spark, Airflow, Kafka or streaming systems, Snowflake, object storage, observability, and incident response.
The ideal candidate can troubleshoot distributed data platform issues from logs, metrics, timestamps, pipeline state, and infrastructure signals, and can drive permanent fixes through configuration improvements, automation, tuning, and engineering partnership.
This position is based in Shanghai, China. Key Responsibility
• Operate and support production ETL pipelines, including extractors, loaders, batch jobs, streaming ingestion, and data consolidation workflows.
• Support loader jobs running in an EKS Airflow cluster that load data into Snowflake.
• Support Spark jobs running on EKS that load data into the Datalake or Lakehouse.
• Own platform and pipeline reliability, including uptime, throughput, job success rate, SLA adherence, and data freshness.
• Triage and resolve production incidents such as failed jobs, delayed loads, OOMKilled pods, CrashLoopBackOff pods, stalled pipelines, driver/executor failures, credential issues, and network/connectivity problems.
• Perform root cause analysis for incidents and drive corrective actions through automation, configuration changes, capacity tuning, or code/platform improvements.
• Tune Kubernetes workloads, including CPU/memory requests and limits, autoscaling, pod scheduling, service accounts, secrets, config maps, and workload health.
• Tune Spark workloads, including executor count, cores, memory, memory overhead, shuffle partitions, dynamic allocation, adaptive query execution, spill reduction, and failed stage analysis.
• Monitor and troubleshoot Airflow DAGs, task retries, scheduling issues, dependency failures, executor/resource constraints, and SLA misses.
• Diagnose data latency across the full pipeline path, such as source event, Kafka or messaging layer, staging/object storage, ETL processing, Snowflake, and Datalake availability.
• Support Kafka or equivalent streaming systems, including consumer lag, offsets, partitions, topic health, producer delay, and SASL/TLS client configuration.
• Manage platform and job configuration through approved source-of-truth or GitOps-style processes instead of ad hoc production changes.
• Manage secrets, certificates, truststores, JDBC credentials, Snowflake credentials, object-storage credentials, and key rotations following least-privilege practices.
• Build and maintain dashboards, alerts, log queries, and runbooks for job health, freshness, backlog, failures, infrastructure usage, and incident response.
• Write scripts and automation using Python, Bash, or similar tools to reduce toil and improve recovery speed.
• Coordinate with application engineering, data engineering, platform, security, network, Snowflake, and Datalake teams.
• Participate in on-call or production support rotation for China business hours and critical incidents.