Machine Learning Infrastructure Engineers - Shopify

ShopifyRemote - AmericasRemoteMid Level
Activeverified Jun 13, 2026

Job description

About the role

Machine Learning Infrastructure Engineers build and operate the end-to-end platform that powers AI—from data ingestion and training to large-scale, low-latency inference. They design high-performance, GPU-accelerated systems on Kubernetes, craft self-serve developer experiences, and ship the paved roads that let ML teams move fast, safely, and at global scale. Some companies separate ML Infra, ML Platform and ML Ops- at Shopify- we call this ML Infrastructure. We have an agile workforce who can flex their experience and solve problems across these three domains.Responsibilities:Build and operate ML control planes, APIs, CLIs, SDKs, and self-serve golden pathsDesign and optimize multi-tenant GPU Kubernetes clusters, including autoscaling, scheduling, packing, and utilizationOwn model lifecycle: training orchestration/experiments, registries/versioning, CI/CD, canary/blue-green, and safe rollbackBuild real-time serving stacks (KServe/Seldon/TensorFlow Serving) and end-to-end pipelines for batch and streamingDesign feature platforms and engineer storage/data movement for datasets, features, and artifacts tuned for cost/performanceImplement observability and SLOs across pipelines, training, and inference; automate remediation and capacity planningPartner with ML, data, and product teams to unblock delivery and accelerate idea-to-impactQualifications:Proven platform/infrastructure engineering experience with a track record of shipping production systems and codeDeep Kubernetes/containerization expertise for ML workloads (operators, Helm, service mesh/gRPC) and multi-tenant clustersHands-on experience running GPU infrastructure at scale (NVIDIA ecosystem; scheduling/packing/optimization)Strong distributed systems and API/service design fundamentals; experience with high-scale inferenceProficiency with infrastructure-as-code and automation (Terraform, Helm, GitOps) on major clouds (GCP/AWS/Azure)Observability expertise (Prometheus/Grafana) and SLO-driven operations for ML systemsProficient in Python/Go/Java; experience building developer tooling and self-service platformsNice to Haves:Model serving and lifecycle tooling: KServe/Seldon/TensorFlow Serving, Kubeflow, MLflow/W&B, model registries, DVCFeature store experience (Feast/Tecton) with online/offline parity and SLAsData infrastructure familiarity (Kafka, Spark/Flink) and stateful stores (Redis/MySQL); CI/CD for online/batch inferenceModel performance optimization (batching, caching, quantization, distillation) and hardware-aware tuningExperience with experimentation/A/B testing platforms and online evaluation frameworks At Shopify, we pride ourselves on moving quickly—not just in shipping, but in our hiring process as well. If you're ready to apply, please be prepared to interview with us within the week. Our goal is to complete the entire interview loop within 30 days. You will be expected to complete a live pair programming session, come prepared with your own IDE.This role may require on-call workMachine Learning Infrastructure Engineers build and operate the end-to-end platform that powers AI—from data ingestion and training to large-scale, low-latency inference. They design high-performance, GPU-accelerated systems on Kubernetes, craft self-serve developer experiences, and ship the paved roads that let ML teams move fast, safely, and at global scale. Some companies separate ML Infra, ML Platform and ML Ops- at Shopify- we call this ML Infrastructure. We have an agile workforce who can flex their experience and solve problems across these three domains.Responsibilities:Build and operate ML control planes, APIs, CLIs, SDKs, and self-serve golden pathsDesign and optimize multi-tenant GPU Kubernetes clusters, including autoscaling, scheduling, packing, and utilizationOwn model lifecycle: training orchestration/experiments, registries/versioning, CI/CD, canary/blue-green, and safe rollbackBuild real-time serving stacks (KServe/Seldon/TensorFlow Serving) and end-to-end pipelines for batch and streamingDesign feature platforms and engineer storage/data movement for datasets, features, and artifacts tuned for cost/performanceImplement observability and SLOs across pipelines, training, and inference; automate remediation and capacity planningPartner with ML, data, and product teams to unblock delivery and accelerate idea-to-impactQualifications:Proven platform/infrastructure engineering experience with a track record of shipping production systems and codeDeep Kubernetes/containerization expertise for ML workloads (operators, Helm, service mesh/gRPC) and multi-tenant clustersHands-on experience running GPU infrastructure at scale (NVIDIA ecosystem; scheduling/packing/optimization)Strong distributed systems and API/service design fundamentals; experience with high-scale inferenceProficiency with infrastructure-as-code and automation (Terraform, Helm, GitOps) on major clouds (GCP/AWS/Azure)Observability expertise (Prometheus/Grafana) and SLO-driven operations for ML systemsProficient in Python/Go/Java; experience building developer tooling and self-service platformsNice to Haves:Model serving and lifecycle tooling: KServe/Seldon/TensorFlow Serving, Kubeflow, MLflow/W&B, model registries, DVCFeature store experience (Feast/Tecton) with online/offline parity and SLAsData infrastructure familiarity (Kafka, Spark/Flink) and stateful stores (Redis/MySQL); CI/CD for online/batch inferenceModel performance optimization (batching, caching, quantization, distillation) and hardware-aware tuningExperience with experimentation/A/B testing platforms and online evaluation frameworks At Shopify, we pride ourselves on moving quickly—not just in shipping, but in our hiring process as well. If you're ready to apply, please be prepared to interview with us within the week. Our goal is to complete the entire interview loop within 30 days. You will be expected to complete a live pair programming session, come prepared with your own IDE.This role may require on-call work

Machine Learning Infrastructure Engineers build and operate the end-to-end platform that powers AI—from data ingestion and training to large-scale, low-latency inference. They design high-performance, GPU-accelerated systems on Kubernetes, craft self-serve developer experiences, and ship the paved roads that let ML teams move fast, safely, and at global scale. Some companies separate ML Infra, ML Platform and ML Ops- at Shopify- we call this ML Infrastructure. We have an agile workforce who can flex their experience and solve problems across these three domains.

Machine Learning Infrastructure Engineers

Responsibilities:

Responsibilities
  • Build and operate ML control planes, APIs, CLIs, SDKs, and self-serve golden paths

Build and operate ML control planes, APIs, CLIs, SDKs, and self-serve golden paths

  • Design and optimize multi-tenant GPU Kubernetes clusters, including autoscaling, scheduling, packing, and utilization

Design and optimize multi-tenant GPU Kubernetes clusters, including autoscaling, scheduling, packing, and utilization

  • Own model lifecycle: training orchestration/experiments, registries/versioning, CI/CD, canary/blue-green, and safe rollback

Own model lifecycle: training orchestration/experiments, registries/versioning, CI/CD, canary/blue-green, and safe rollback

  • Build real-time serving stacks (KServe/Seldon/TensorFlow Serving) and end-to-end pipelines for batch and streaming

Build real-time serving stacks (KServe/Seldon/TensorFlow Serving) and end-to-end pipelines for batch and streaming

  • Design feature platforms and engineer storage/data movement for datasets, features, and artifacts tuned for cost/performance

Design feature platforms and engineer storage/data movement for datasets, features, and artifacts tuned for cost/performance

  • Implement observability and SLOs across pipelines, training, and inference; automate remediation and capacity planning

Implement observability and SLOs across pipelines, training, and inference; automate remediation and capacity planning

  • Partner with ML, data, and product teams to unblock delivery and accelerate idea-to-impact

Partner with ML, data, and product teams to unblock delivery and accelerate idea-to-impact

Qualifications:

Qualifications
  • Proven platform/infrastructure engineering experience with a track record of shipping production systems and code

Proven platform/infrastructure engineering experience with a track record of shipping production systems and code

  • Deep Kubernetes/containerization expertise for ML workloads (operators, Helm, service mesh/gRPC) and multi-tenant clusters

Deep Kubernetes/containerization expertise for ML workloads (operators, Helm, service mesh/gRPC) and multi-tenant clusters

  • Hands-on experience running GPU infrastructure at scale (NVIDIA ecosystem; scheduling/packing/optimization)

Hands-on experience running GPU infrastructure at scale (NVIDIA ecosystem; scheduling/packing/optimization)

  • Strong distributed systems and API/service design fundamentals; experience with high-scale inference

Strong distributed systems and API/service design fundamentals; experience with high-scale inference

  • Proficiency with infrastructure-as-code and automation (Terraform, Helm, GitOps) on major clouds (GCP/AWS/Azure)

Proficiency with infrastructure-as-code and automation (Terraform, Helm, GitOps) on major clouds (GCP/AWS/Azure)

  • Observability expertise (Prometheus/Grafana) and SLO-driven operations for ML systems

Observability expertise (Prometheus/Grafana) and SLO-driven operations for ML systems

  • Proficient in Python/Go/Java; experience building developer tooling and self-service platforms

Proficient in Python/Go/Java; experience building developer tooling and self-service platforms

Nice to Haves:

Nice to Haves:
  • Model serving and lifecycle tooling: KServe/Seldon/TensorFlow Serving, Kubeflow, MLflow/W&B, model registries, DVC

Model serving and lifecycle tooling: KServe/Seldon/TensorFlow Serving, Kubeflow, MLflow/W&B, model registries, DVC

  • Feature store experience (Feast/Tecton) with online/offline parity and SLAs

Feature store experience (Feast/Tecton) with online/offline parity and SLAs

  • Data infrastructure familiarity (Kafka, Spark/Flink) and stateful stores (Redis/MySQL); CI/CD for online/batch inference

Data infrastructure familiarity (Kafka, Spark/Flink) and stateful stores (Redis/MySQL); CI/CD for online/batch inference

  • Model performance optimization (batching, caching, quantization, distillation) and hardware-aware tuning

Model performance optimization (batching, caching, quantization, distillation) and hardware-aware tuning

  • Experience with experimentation/A/B testing platforms and online evaluation frameworks

Experience with experimentation/A/B testing platforms and online evaluation frameworks

At Shopify, we pride ourselves on moving quickly—not just in shipping, but in our hiring process as well. If you're ready to apply, please be prepared to interview with us within the week. Our goal is to complete the entire interview loop within 30 days. You will be expected to complete a live pair programming session, come prepared with your own IDE.

This role may require on-call work