Senior Staff Engineer - Machine Learning Inference - Shopify
Job description
About the role
Step into the engine room of Agentic Commerce! Imagine owning the bleeding edge of machine learning at Shopify, where your acceleration, optimization, and scaling of ML inference will shape the experience of millions of merchants, and influence how commerce AI is done worldwide. We’re seeking a Senior Staff Engineer to architect, optimize, and own the high-performance runtime that transforms innovative models into production breakthroughs. Your work will be the engine behind our real-time AI systems, driving game-changing cost and latency reductions, and enabling rapid launches of intelligent features that keep Shopify (and our merchants) years ahead. Join a remote-first team of world-class experts, experiment fearlessly, and see your code move the needle for some of the largest-scale ML workloads in commerce.ResponsibilitiesArchitect, optimize, and own Shopify’s production ML inference. Designing for high throughput, ultra-low latency, and global reliability.Leverage and extend technologies like CUDA, TensorRT, Triton, TVM, and custom GPU kernels to deliver state-of-the-art performance and efficiency at scale.Partner with ML, infrastructure, and product teams to seamlessly deploy, benchmark, and scale cutting-edge models powering our platform.Drive cost optimization and system efficiency, reducing cloud spend and carbon footprint by orders of magnitude without sacrificing model quality.Lead deep performance investigations, apply advanced techniques (pruning, quantization, distillation, batching), and implement robust solutions for serving models in production.Set technical strategy and culture for ML inference across Shopify, mentoring others and collaborating with global AI pioneers.QualificationsProven, hands-on expertise in building and optimizing large-scale ML inference systems, with measurable performance and cost wins.Deep experience in production model serving, runtime optimization, and acceleration. Especially leveraging GPUs (CUDA, TensorRT) and high-performance deep learning infrastructure.Strong software engineering skills (Python, C++, and/or other relevant languages) with a robust systems and distributed computing mindset.Demonstrated leadership in architecting or scaling reliable, real-time inference at scale, handling millions of queries per day.Track record of cross-functional impact: working closely with ML research/engineering, infra, and product teams to deliver production results.Advanced understanding of model compression, quantization, efficient deployment, and tradeoffs between speed, cost, and accuracy.Nice to HavesOpen source contributions to inference frameworks (TensorRT, TVM, Triton, DeepSpeed, ONNX, etc.) or technical talks/publications at leading AI conferences.Experience optimizing inference across a variety of hardware (NVIDIA, AMD, ARM, cloud TPUs).Familiarity with building or integrating robust monitoring, observability, and auto-scaling for inference platforms.Experience with modern MLOps pipelines and methodologies.Prior experience in e-commerce, large-scale product infra, or globally distributed inference workloads.At Shopify, we pride ourselves on moving quickly—not just in shipping, but in our hiring process as well. If you're ready to apply, please be prepared to interview with us within the week. Our goal is to complete the entire interview loop within 30 days. You will be expected to complete a live pair programming session, come prepared with your own IDE.Step into the engine room of Agentic Commerce! Imagine owning the bleeding edge of machine learning at Shopify, where your acceleration, optimization, and scaling of ML inference will shape the experience of millions of merchants, and influence how commerce AI is done worldwide. We’re seeking a Senior Staff Engineer to architect, optimize, and own the high-performance runtime that transforms innovative models into production breakthroughs. Your work will be the engine behind our real-time AI systems, driving game-changing cost and latency reductions, and enabling rapid launches of intelligent features that keep Shopify (and our merchants) years ahead. Join a remote-first team of world-class experts, experiment fearlessly, and see your code move the needle for some of the largest-scale ML workloads in commerce.ResponsibilitiesArchitect, optimize, and own Shopify’s production ML inference. Designing for high throughput, ultra-low latency, and global reliability.Leverage and extend technologies like CUDA, TensorRT, Triton, TVM, and custom GPU kernels to deliver state-of-the-art performance and efficiency at scale.Partner with ML, infrastructure, and product teams to seamlessly deploy, benchmark, and scale cutting-edge models powering our platform.Drive cost optimization and system efficiency, reducing cloud spend and carbon footprint by orders of magnitude without sacrificing model quality.Lead deep performance investigations, apply advanced techniques (pruning, quantization, distillation, batching), and implement robust solutions for serving models in production.Set technical strategy and culture for ML inference across Shopify, mentoring others and collaborating with global AI pioneers.QualificationsProven, hands-on expertise in building and optimizing large-scale ML inference systems, with measurable performance and cost wins.Deep experience in production model serving, runtime optimization, and acceleration. Especially leveraging GPUs (CUDA, TensorRT) and high-performance deep learning infrastructure.Strong software engineering skills (Python, C++, and/or other relevant languages) with a robust systems and distributed computing mindset.Demonstrated leadership in architecting or scaling reliable, real-time inference at scale, handling millions of queries per day.Track record of cross-functional impact: working closely with ML research/engineering, infra, and product teams to deliver production results.Advanced understanding of model compression, quantization, efficient deployment, and tradeoffs between speed, cost, and accuracy.Nice to HavesOpen source contributions to inference frameworks (TensorRT, TVM, Triton, DeepSpeed, ONNX, etc.) or technical talks/publications at leading AI conferences.Experience optimizing inference across a variety of hardware (NVIDIA, AMD, ARM, cloud TPUs).Familiarity with building or integrating robust monitoring, observability, and auto-scaling for inference platforms.Experience with modern MLOps pipelines and methodologies.Prior experience in e-commerce, large-scale product infra, or globally distributed inference workloads.At Shopify, we pride ourselves on moving quickly—not just in shipping, but in our hiring process as well. If you're ready to apply, please be prepared to interview with us within the week. Our goal is to complete the entire interview loop within 30 days. You will be expected to complete a live pair programming session, come prepared with your own IDE.Step into the engine room of Agentic Commerce! Imagine owning the bleeding edge of machine learning at Shopify, where your acceleration, optimization, and scaling of ML inference will shape the experience of millions of merchants, and influence how commerce AI is done worldwide. We’re seeking a Senior Staff Engineer to architect, optimize, and own the high-performance runtime that transforms innovative models into production breakthroughs. Your work will be the engine behind our real-time AI systems, driving game-changing cost and latency reductions, and enabling rapid launches of intelligent features that keep Shopify (and our merchants) years ahead. Join a remote-first team of world-class experts, experiment fearlessly, and see your code move the needle for some of the largest-scale ML workloads in commerce.
Responsibilities
Responsibilities- Architect, optimize, and own Shopify’s production ML inference. Designing for high throughput, ultra-low latency, and global reliability.
Architect, optimize, and own Shopify’s production ML inference. Designing for high throughput, ultra-low latency, and global reliability.
- Leverage and extend technologies like CUDA, TensorRT, Triton, TVM, and custom GPU kernels to deliver state-of-the-art performance and efficiency at scale.
Leverage and extend technologies like CUDA, TensorRT, Triton, TVM, and custom GPU kernels to deliver state-of-the-art performance and efficiency at scale.
- Partner with ML, infrastructure, and product teams to seamlessly deploy, benchmark, and scale cutting-edge models powering our platform.
Partner with ML, infrastructure, and product teams to seamlessly deploy, benchmark, and scale cutting-edge models powering our platform.
- Drive cost optimization and system efficiency, reducing cloud spend and carbon footprint by orders of magnitude without sacrificing model quality.
Drive cost optimization and system efficiency, reducing cloud spend and carbon footprint by orders of magnitude without sacrificing model quality.
- Lead deep performance investigations, apply advanced techniques (pruning, quantization, distillation, batching), and implement robust solutions for serving models in production.
Lead deep performance investigations, apply advanced techniques (pruning, quantization, distillation, batching), and implement robust solutions for serving models in production.
- Set technical strategy and culture for ML inference across Shopify, mentoring others and collaborating with global AI pioneers.
Set technical strategy and culture for ML inference across Shopify, mentoring others and collaborating with global AI pioneers.
Qualifications
Qualifications- Proven, hands-on expertise in building and optimizing large-scale ML inference systems, with measurable performance and cost wins.
Proven, hands-on expertise in building and optimizing large-scale ML inference systems, with measurable performance and cost wins.
- Deep experience in production model serving, runtime optimization, and acceleration. Especially leveraging GPUs (CUDA, TensorRT) and high-performance deep learning infrastructure.
Deep experience in production model serving, runtime optimization, and acceleration. Especially leveraging GPUs (CUDA, TensorRT) and high-performance deep learning infrastructure.
- Strong software engineering skills (Python, C++, and/or other relevant languages) with a robust systems and distributed computing mindset.
Strong software engineering skills (Python, C++, and/or other relevant languages) with a robust systems and distributed computing mindset.
- Demonstrated leadership in architecting or scaling reliable, real-time inference at scale, handling millions of queries per day.
Demonstrated leadership in architecting or scaling reliable, real-time inference at scale, handling millions of queries per day.
- Track record of cross-functional impact: working closely with ML research/engineering, infra, and product teams to deliver production results.
Track record of cross-functional impact: working closely with ML research/engineering, infra, and product teams to deliver production results.
- Advanced understanding of model compression, quantization, efficient deployment, and tradeoffs between speed, cost, and accuracy.
Advanced understanding of model compression, quantization, efficient deployment, and tradeoffs between speed, cost, and accuracy.
Nice to Haves
Nice to Haves- Open source contributions to inference frameworks (TensorRT, TVM, Triton, DeepSpeed, ONNX, etc.) or technical talks/publications at leading AI conferences.
Open source contributions to inference frameworks (TensorRT, TVM, Triton, DeepSpeed, ONNX, etc.) or technical talks/publications at leading AI conferences.
- Experience optimizing inference across a variety of hardware (NVIDIA, AMD, ARM, cloud TPUs).
Experience optimizing inference across a variety of hardware (NVIDIA, AMD, ARM, cloud TPUs).
- Familiarity with building or integrating robust monitoring, observability, and auto-scaling for inference platforms.
Familiarity with building or integrating robust monitoring, observability, and auto-scaling for inference platforms.
- Experience with modern MLOps pipelines and methodologies.
Experience with modern MLOps pipelines and methodologies.
- Prior experience in e-commerce, large-scale product infra, or globally distributed inference workloads.
Prior experience in e-commerce, large-scale product infra, or globally distributed inference workloads.
At Shopify, we pride ourselves on moving quickly—not just in shipping, but in our hiring process as well. If you're ready to apply, please be prepared to interview with us within the week. Our goal is to complete the entire interview loop within 30 days. You will be expected to complete a live pair programming session, come prepared with your own IDE.