Job description
About the role
About The RoleShopify is building SimGym, an AI-powered store testing system that simulates buyer behavior using browser automation and LLM/VLM agents. SimGym helps merchants understand how real shoppers might experience their storefronts, then turns those simulations into trustworthy, merchant-facing theme insights and recommendations.We are looking for a Staff Applied Machine Learning Engineer to set and drive the ML technical direction for SimGym. This is a senior technical leadership role for someone who can move between strategy and code: defining the roadmap, designing production ML systems, de-risking ambiguous research, mentoring senior engineers, and shipping model improvements that create real merchant impact.You will work across multimodal browsing agents, teacher-model distillation, learned buyer personas, SFT/post-training, evaluation systems, behavioral data pipelines, and production rollout infrastructure. The goal is not just to improve benchmarks, but to build AI agents that behave more like real buyers and produce recommendations merchants can trust.What You'll DoOwn the ML technical direction across SimGym's agent, persona, evaluation, and data flywheel workstreams.Develop, fine-tune, evaluate, and deploy LLM/VLM-based systems for browser automation, buyer simulation, and storefront analysis.Lead major model initiatives from research prototype to production readiness, such as VLM teacher distillation, learned buyer personas, or SFT model rollout.Design scalable ML architectures, data pipelines, and model-serving systems that support experimentation, evaluation, and production deployment.Improve human-vs-agent alignment and behavioral fidelity through rigorous offline and online evaluation.Build and stabilize golden datasets, trace replay systems, screenshot/A11y-tree evaluation flows, and reliable benchmark infrastructure.Use Shopify-scale behavioral and storefront data to improve buyer simulation, personalization, and merchant-facing recommendations.Make clear build/buy tradeoffs across models, infrastructure, evaluation tooling, and data systems.Communicate technical direction and tradeoffs clearly with engineering, product, data, research, and partner teams.Raise the technical bar for the team through hands-on implementation, design review, mentorship, and pragmatic execution.What You'll NeedStaff-level applied ML leadership experience, with a track record of setting strategy, writing code and designs, unblocking teams, and delivering production systems through ambiguity.Strong hands-on experience with LLM, VLM, or agent systems, including SFT/post-training, distillation, evaluation, model iteration, and serving tradeoffs.End-to-end experience training, evaluating, testing, deploying, and operating ML products at scale.Strong instincts for experimentation and evaluation, including designing meaningful metrics, diagnosing alignment failures, and avoiding misleading benchmark wins.Experience building production ML data pipelines and working with large-scale behavioral, event, or product data.Proficiency with Python, shell scripting, batch and streaming data pipelines, orchestration tools, vector databases, BigQuery/BigTable or equivalent systems.Experience with parallel or distributed ML environments, including GPU optimization, model serving, or large-scale training/inference workflows.Excellent communication with technical and non-technical audiences.Nice To HaveExperience with browser automation, AI agents, or simulated user behavior.Background in e-commerce, search, recommendations, personalization, or buyer behavior modeling.Experience with sequence or behavior representation learning, recommender embeddings, HSTU-style architectures, vLLM/GPU serving, Spark/GCS/Hugging Face-style data workflows.Prior experience translating research prototypes into reliable, production user-facing systems.Why This RoleThis is a high-leverage role at the center of Shopify's applied AI work for merchants. You will help shape SimGym's ML architecture and operating model while working on frontier LLM/VLM agents, learned buyer personas, evaluation science, and production systems with direct merchant impact.This role may require on-call work. Shopify's hiring process moves quickly, and candidates should be prepared for a technical interview loop that includes pair programming using their own IDE.About The RoleShopify is building SimGym, an AI-powered store testing system that simulates buyer behavior using browser automation and LLM/VLM agents. SimGym helps merchants understand how real shoppers might experience their storefronts, then turns those simulations into trustworthy, merchant-facing theme insights and recommendations.We are looking for a Staff Applied Machine Learning Engineer to set and drive the ML technical direction for SimGym. This is a senior technical leadership role for someone who can move between strategy and code: defining the roadmap, designing production ML systems, de-risking ambiguous research, mentoring senior engineers, and shipping model improvements that create real merchant impact.You will work across multimodal browsing agents, teacher-model distillation, learned buyer personas, SFT/post-training, evaluation systems, behavioral data pipelines, and production rollout infrastructure. The goal is not just to improve benchmarks, but to build AI agents that behave more like real buyers and produce recommendations merchants can trust.What You'll DoOwn the ML technical direction across SimGym's agent, persona, evaluation, and data flywheel workstreams.Develop, fine-tune, evaluate, and deploy LLM/VLM-based systems for browser automation, buyer simulation, and storefront analysis.Lead major model initiatives from research prototype to production readiness, such as VLM teacher distillation, learned buyer personas, or SFT model rollout.Design scalable ML architectures, data pipelines, and model-serving systems that support experimentation, evaluation, and production deployment.Improve human-vs-agent alignment and behavioral fidelity through rigorous offline and online evaluation.Build and stabilize golden datasets, trace replay systems, screenshot/A11y-tree evaluation flows, and reliable benchmark infrastructure.Use Shopify-scale behavioral and storefront data to improve buyer simulation, personalization, and merchant-facing recommendations.Make clear build/buy tradeoffs across models, infrastructure, evaluation tooling, and data systems.Communicate technical direction and tradeoffs clearly with engineering, product, data, research, and partner teams.Raise the technical bar for the team through hands-on implementation, design review, mentorship, and pragmatic execution.What You'll NeedStaff-level applied ML leadership experience, with a track record of setting strategy, writing code and designs, unblocking teams, and delivering production systems through ambiguity.Strong hands-on experience with LLM, VLM, or agent systems, including SFT/post-training, distillation, evaluation, model iteration, and serving tradeoffs.End-to-end experience training, evaluating, testing, deploying, and operating ML products at scale.Strong instincts for experimentation and evaluation, including designing meaningful metrics, diagnosing alignment failures, and avoiding misleading benchmark wins.Experience building production ML data pipelines and working with large-scale behavioral, event, or product data.Proficiency with Python, shell scripting, batch and streaming data pipelines, orchestration tools, vector databases, BigQuery/BigTable or equivalent systems.Experience with parallel or distributed ML environments, including GPU optimization, model serving, or large-scale training/inference workflows.Excellent communication with technical and non-technical audiences.Nice To HaveExperience with browser automation, AI agents, or simulated user behavior.Background in e-commerce, search, recommendations, personalization, or buyer behavior modeling.Experience with sequence or behavior representation learning, recommender embeddings, HSTU-style architectures, vLLM/GPU serving, Spark/GCS/Hugging Face-style data workflows.Prior experience translating research prototypes into reliable, production user-facing systems.Why This RoleThis is a high-leverage role at the center of Shopify's applied AI work for merchants. You will help shape SimGym's ML architecture and operating model while working on frontier LLM/VLM agents, learned buyer personas, evaluation science, and production systems with direct merchant impact.This role may require on-call work. Shopify's hiring process moves quickly, and candidates should be prepared for a technical interview loop that includes pair programming using their own IDE.About The Role
Shopify is building SimGym, an AI-powered store testing system that simulates buyer behavior using browser automation and LLM/VLM agents. SimGym helps merchants understand how real shoppers might experience their storefronts, then turns those simulations into trustworthy, merchant-facing theme insights and recommendations.
We are looking for a Staff Applied Machine Learning Engineer to set and drive the ML technical direction for SimGym. This is a senior technical leadership role for someone who can move between strategy and code: defining the roadmap, designing production ML systems, de-risking ambiguous research, mentoring senior engineers, and shipping model improvements that create real merchant impact.
You will work across multimodal browsing agents, teacher-model distillation, learned buyer personas, SFT/post-training, evaluation systems, behavioral data pipelines, and production rollout infrastructure. The goal is not just to improve benchmarks, but to build AI agents that behave more like real buyers and produce recommendations merchants can trust.
What You'll Do
- Own the ML technical direction across SimGym's agent, persona, evaluation, and data flywheel workstreams.
Own the ML technical direction across SimGym's agent, persona, evaluation, and data flywheel workstreams.
- Develop, fine-tune, evaluate, and deploy LLM/VLM-based systems for browser automation, buyer simulation, and storefront analysis.
Develop, fine-tune, evaluate, and deploy LLM/VLM-based systems for browser automation, buyer simulation, and storefront analysis.
- Lead major model initiatives from research prototype to production readiness, such as VLM teacher distillation, learned buyer personas, or SFT model rollout.
Lead major model initiatives from research prototype to production readiness, such as VLM teacher distillation, learned buyer personas, or SFT model rollout.
- Design scalable ML architectures, data pipelines, and model-serving systems that support experimentation, evaluation, and production deployment.
Design scalable ML architectures, data pipelines, and model-serving systems that support experimentation, evaluation, and production deployment.
- Improve human-vs-agent alignment and behavioral fidelity through rigorous offline and online evaluation.
Improve human-vs-agent alignment and behavioral fidelity through rigorous offline and online evaluation.
- Build and stabilize golden datasets, trace replay systems, screenshot/A11y-tree evaluation flows, and reliable benchmark infrastructure.
Build and stabilize golden datasets, trace replay systems, screenshot/A11y-tree evaluation flows, and reliable benchmark infrastructure.
- Use Shopify-scale behavioral and storefront data to improve buyer simulation, personalization, and merchant-facing recommendations.
Use Shopify-scale behavioral and storefront data to improve buyer simulation, personalization, and merchant-facing recommendations.
- Make clear build/buy tradeoffs across models, infrastructure, evaluation tooling, and data systems.
Make clear build/buy tradeoffs across models, infrastructure, evaluation tooling, and data systems.
- Communicate technical direction and tradeoffs clearly with engineering, product, data, research, and partner teams.
Communicate technical direction and tradeoffs clearly with engineering, product, data, research, and partner teams.
- Raise the technical bar for the team through hands-on implementation, design review, mentorship, and pragmatic execution.
Raise the technical bar for the team through hands-on implementation, design review, mentorship, and pragmatic execution.
What You'll Need
- Staff-level applied ML leadership experience, with a track record of setting strategy, writing code and designs, unblocking teams, and delivering production systems through ambiguity.
Staff-level applied ML leadership experience, with a track record of setting strategy, writing code and designs, unblocking teams, and delivering production systems through ambiguity.
- Strong hands-on experience with LLM, VLM, or agent systems, including SFT/post-training, distillation, evaluation, model iteration, and serving tradeoffs.
Strong hands-on experience with LLM, VLM, or agent systems, including SFT/post-training, distillation, evaluation, model iteration, and serving tradeoffs.
- End-to-end experience training, evaluating, testing, deploying, and operating ML products at scale.
End-to-end experience training, evaluating, testing, deploying, and operating ML products at scale.
- Strong instincts for experimentation and evaluation, including designing meaningful metrics, diagnosing alignment failures, and avoiding misleading benchmark wins.
Strong instincts for experimentation and evaluation, including designing meaningful metrics, diagnosing alignment failures, and avoiding misleading benchmark wins.
- Experience building production ML data pipelines and working with large-scale behavioral, event, or product data.
Experience building production ML data pipelines and working with large-scale behavioral, event, or product data.
- Proficiency with Python, shell scripting, batch and streaming data pipelines, orchestration tools, vector databases, BigQuery/BigTable or equivalent systems.
Proficiency with Python, shell scripting, batch and streaming data pipelines, orchestration tools, vector databases, BigQuery/BigTable or equivalent systems.
- Experience with parallel or distributed ML environments, including GPU optimization, model serving, or large-scale training/inference workflows.
Experience with parallel or distributed ML environments, including GPU optimization, model serving, or large-scale training/inference workflows.
- Excellent communication with technical and non-technical audiences.
Excellent communication with technical and non-technical audiences.
Nice To Have
- Experience with browser automation, AI agents, or simulated user behavior.
Experience with browser automation, AI agents, or simulated user behavior.
- Background in e-commerce, search, recommendations, personalization, or buyer behavior modeling.
Background in e-commerce, search, recommendations, personalization, or buyer behavior modeling.
- Experience with sequence or behavior representation learning, recommender embeddings, HSTU-style architectures, vLLM/GPU serving, Spark/GCS/Hugging Face-style data workflows.
Experience with sequence or behavior representation learning, recommender embeddings, HSTU-style architectures, vLLM/GPU serving, Spark/GCS/Hugging Face-style data workflows.
- Prior experience translating research prototypes into reliable, production user-facing systems.
Prior experience translating research prototypes into reliable, production user-facing systems.
Why This Role
This is a high-leverage role at the center of Shopify's applied AI work for merchants. You will help shape SimGym's ML architecture and operating model while working on frontier LLM/VLM agents, learned buyer personas, evaluation science, and production systems with direct merchant impact.
This role may require on-call work. Shopify's hiring process moves quickly, and candidates should be prepared for a technical interview loop that includes pair programming using their own IDE.