Senior Applied Machine Learning Engineer, SimGym - Shopify

ShopifyRemote - AmericasRemoteSenior
Active

Job description

About the role

About The RoleShopify is building SimGym, an AI-powered store testing system that simulates buyer behavior using browser automation and LLM/VLM agents. SimGym helps merchants understand how real shoppers might experience their storefronts, then turns those simulations into trustworthy, merchant-facing theme insights and recommendations.We are looking for a Senior Applied Machine Learning Engineer to lead high-value ML work within SimGym. This is a hands-on technical role for someone who can execute through ambiguity, write production-quality code, make strong technical decisions, and help turn model experiments into reliable systems with real merchant impact.You will work on problems across multimodal browsing agents, teacher-model distillation, learned buyer personas, SFT/post-training, evaluation systems, behavioral data pipelines, and production rollout infrastructure. The goal is not just to improve benchmarks, but to help build AI agents that behave more like real buyers and produce recommendations merchants can trust.What You'll DoLead high-impact ML projects within SimGym's agent, persona, evaluation, or data flywheel workstreams.Develop, fine-tune, evaluate, and deploy LLM/VLM-based systems for browser automation, buyer simulation, and storefront analysis.Move model improvements from experiment toward production readiness, such as VLM teacher distillation, learned buyer personas, or SFT model rollout.Design and build scalable ML components, data pipelines, and model-serving paths that are easy to understand, operate, and maintain.Improve human-vs-agent alignment and behavioral fidelity through strong offline and online evaluation practices.Contribute to golden datasets, trace replay systems, screenshot/A11y-tree evaluation flows, and reliable benchmark infrastructure.Use Shopify-scale behavioral and storefront data to improve buyer simulation, personalization, and merchant-facing recommendations.Diagnose complex model, data, and system failures with attention to the important details.Make practical technical tradeoffs across models, infrastructure, evaluation tooling, and data systems.Communicate clearly with engineering, product, data, research, and partner teams.Help raise the technical bar for the team through strong execution, thoughtful design, code quality, and mentorship by example.What You'll NeedStrong applied ML engineering experience, with a track record of delivering production systems through ambiguous technical problems.Hands-on experience with LLM, VLM, or agent systems, including SFT/post-training, distillation, evaluation, model iteration, or serving tradeoffs.End-to-end experience training, evaluating, testing, deploying, and operating ML products at meaningful scale.Strong experimentation and evaluation instincts, including designing useful metrics, diagnosing alignment failures, and avoiding misleading benchmark wins.Experience building production ML data pipelines and working with large-scale behavioral, event, or product data.Proficiency with Python, shell scripting, batch and streaming data pipelines, orchestration tools, vector databases, BigQuery/BigTable or equivalent systems.Experience working with parallel or distributed ML environments, including GPU optimization, model serving, or large-scale training/inference workflows.Strong software engineering fundamentals: code quality, operational awareness, maintainability, and speed of execution.Clear communication with both technical and non-technical audiences.Nice To HaveExperience with browser automation, AI agents, or simulated user behavior.Background in e-commerce, search, recommendations, personalization, or buyer behavior modeling.Experience with sequence or behavior representation learning, recommender embeddings, HSTU-style architectures, vLLM/GPU serving, Spark/GCS/Hugging Face-style data workflows.Prior experience translating research prototypes into reliable, production user-facing systems.Why This RoleThis is a high-leverage role in Shopify's applied AI work for merchants. You will help build the ML systems behind SimGym while working on frontier LLM/VLM agents, learned personas, evaluation science, and production systems with direct merchant impact.This role may require on-call work. Shopify's hiring process moves quickly, and candidates should be prepared for a technical interview loop that includes pair programming using their own IDE.About The RoleShopify is building SimGym, an AI-powered store testing system that simulates buyer behavior using browser automation and LLM/VLM agents. SimGym helps merchants understand how real shoppers might experience their storefronts, then turns those simulations into trustworthy, merchant-facing theme insights and recommendations.We are looking for a Senior Applied Machine Learning Engineer to lead high-value ML work within SimGym. This is a hands-on technical role for someone who can execute through ambiguity, write production-quality code, make strong technical decisions, and help turn model experiments into reliable systems with real merchant impact.You will work on problems across multimodal browsing agents, teacher-model distillation, learned buyer personas, SFT/post-training, evaluation systems, behavioral data pipelines, and production rollout infrastructure. The goal is not just to improve benchmarks, but to help build AI agents that behave more like real buyers and produce recommendations merchants can trust.What You'll DoLead high-impact ML projects within SimGym's agent, persona, evaluation, or data flywheel workstreams.Develop, fine-tune, evaluate, and deploy LLM/VLM-based systems for browser automation, buyer simulation, and storefront analysis.Move model improvements from experiment toward production readiness, such as VLM teacher distillation, learned buyer personas, or SFT model rollout.Design and build scalable ML components, data pipelines, and model-serving paths that are easy to understand, operate, and maintain.Improve human-vs-agent alignment and behavioral fidelity through strong offline and online evaluation practices.Contribute to golden datasets, trace replay systems, screenshot/A11y-tree evaluation flows, and reliable benchmark infrastructure.Use Shopify-scale behavioral and storefront data to improve buyer simulation, personalization, and merchant-facing recommendations.Diagnose complex model, data, and system failures with attention to the important details.Make practical technical tradeoffs across models, infrastructure, evaluation tooling, and data systems.Communicate clearly with engineering, product, data, research, and partner teams.Help raise the technical bar for the team through strong execution, thoughtful design, code quality, and mentorship by example.What You'll NeedStrong applied ML engineering experience, with a track record of delivering production systems through ambiguous technical problems.Hands-on experience with LLM, VLM, or agent systems, including SFT/post-training, distillation, evaluation, model iteration, or serving tradeoffs.End-to-end experience training, evaluating, testing, deploying, and operating ML products at meaningful scale.Strong experimentation and evaluation instincts, including designing useful metrics, diagnosing alignment failures, and avoiding misleading benchmark wins.Experience building production ML data pipelines and working with large-scale behavioral, event, or product data.Proficiency with Python, shell scripting, batch and streaming data pipelines, orchestration tools, vector databases, BigQuery/BigTable or equivalent systems.Experience working with parallel or distributed ML environments, including GPU optimization, model serving, or large-scale training/inference workflows.Strong software engineering fundamentals: code quality, operational awareness, maintainability, and speed of execution.Clear communication with both technical and non-technical audiences.Nice To HaveExperience with browser automation, AI agents, or simulated user behavior.Background in e-commerce, search, recommendations, personalization, or buyer behavior modeling.Experience with sequence or behavior representation learning, recommender embeddings, HSTU-style architectures, vLLM/GPU serving, Spark/GCS/Hugging Face-style data workflows.Prior experience translating research prototypes into reliable, production user-facing systems.Why This RoleThis is a high-leverage role in Shopify's applied AI work for merchants. You will help build the ML systems behind SimGym while working on frontier LLM/VLM agents, learned personas, evaluation science, and production systems with direct merchant impact.This role may require on-call work. Shopify's hiring process moves quickly, and candidates should be prepared for a technical interview loop that includes pair programming using their own IDE.

About The Role

Shopify is building SimGym, an AI-powered store testing system that simulates buyer behavior using browser automation and LLM/VLM agents. SimGym helps merchants understand how real shoppers might experience their storefronts, then turns those simulations into trustworthy, merchant-facing theme insights and recommendations.

We are looking for a Senior Applied Machine Learning Engineer to lead high-value ML work within SimGym. This is a hands-on technical role for someone who can execute through ambiguity, write production-quality code, make strong technical decisions, and help turn model experiments into reliable systems with real merchant impact.

You will work on problems across multimodal browsing agents, teacher-model distillation, learned buyer personas, SFT/post-training, evaluation systems, behavioral data pipelines, and production rollout infrastructure. The goal is not just to improve benchmarks, but to help build AI agents that behave more like real buyers and produce recommendations merchants can trust.

What You'll Do

  • Lead high-impact ML projects within SimGym's agent, persona, evaluation, or data flywheel workstreams.

Lead high-impact ML projects within SimGym's agent, persona, evaluation, or data flywheel workstreams.

  • Develop, fine-tune, evaluate, and deploy LLM/VLM-based systems for browser automation, buyer simulation, and storefront analysis.

Develop, fine-tune, evaluate, and deploy LLM/VLM-based systems for browser automation, buyer simulation, and storefront analysis.

  • Move model improvements from experiment toward production readiness, such as VLM teacher distillation, learned buyer personas, or SFT model rollout.

Move model improvements from experiment toward production readiness, such as VLM teacher distillation, learned buyer personas, or SFT model rollout.

  • Design and build scalable ML components, data pipelines, and model-serving paths that are easy to understand, operate, and maintain.

Design and build scalable ML components, data pipelines, and model-serving paths that are easy to understand, operate, and maintain.

  • Improve human-vs-agent alignment and behavioral fidelity through strong offline and online evaluation practices.

Improve human-vs-agent alignment and behavioral fidelity through strong offline and online evaluation practices.

  • Contribute to golden datasets, trace replay systems, screenshot/A11y-tree evaluation flows, and reliable benchmark infrastructure.

Contribute to golden datasets, trace replay systems, screenshot/A11y-tree evaluation flows, and reliable benchmark infrastructure.

  • Use Shopify-scale behavioral and storefront data to improve buyer simulation, personalization, and merchant-facing recommendations.

Use Shopify-scale behavioral and storefront data to improve buyer simulation, personalization, and merchant-facing recommendations.

  • Diagnose complex model, data, and system failures with attention to the important details.

Diagnose complex model, data, and system failures with attention to the important details.

  • Make practical technical tradeoffs across models, infrastructure, evaluation tooling, and data systems.

Make practical technical tradeoffs across models, infrastructure, evaluation tooling, and data systems.

  • Communicate clearly with engineering, product, data, research, and partner teams.

Communicate clearly with engineering, product, data, research, and partner teams.

  • Help raise the technical bar for the team through strong execution, thoughtful design, code quality, and mentorship by example.

Help raise the technical bar for the team through strong execution, thoughtful design, code quality, and mentorship by example.

What You'll Need

  • Strong applied ML engineering experience, with a track record of delivering production systems through ambiguous technical problems.

Strong applied ML engineering experience, with a track record of delivering production systems through ambiguous technical problems.

  • Hands-on experience with LLM, VLM, or agent systems, including SFT/post-training, distillation, evaluation, model iteration, or serving tradeoffs.

Hands-on experience with LLM, VLM, or agent systems, including SFT/post-training, distillation, evaluation, model iteration, or serving tradeoffs.

  • End-to-end experience training, evaluating, testing, deploying, and operating ML products at meaningful scale.

End-to-end experience training, evaluating, testing, deploying, and operating ML products at meaningful scale.

  • Strong experimentation and evaluation instincts, including designing useful metrics, diagnosing alignment failures, and avoiding misleading benchmark wins.

Strong experimentation and evaluation instincts, including designing useful metrics, diagnosing alignment failures, and avoiding misleading benchmark wins.

  • Experience building production ML data pipelines and working with large-scale behavioral, event, or product data.

Experience building production ML data pipelines and working with large-scale behavioral, event, or product data.

  • Proficiency with Python, shell scripting, batch and streaming data pipelines, orchestration tools, vector databases, BigQuery/BigTable or equivalent systems.

Proficiency with Python, shell scripting, batch and streaming data pipelines, orchestration tools, vector databases, BigQuery/BigTable or equivalent systems.

  • Experience working with parallel or distributed ML environments, including GPU optimization, model serving, or large-scale training/inference workflows.

Experience working with parallel or distributed ML environments, including GPU optimization, model serving, or large-scale training/inference workflows.

  • Strong software engineering fundamentals: code quality, operational awareness, maintainability, and speed of execution.

Strong software engineering fundamentals: code quality, operational awareness, maintainability, and speed of execution.

  • Clear communication with both technical and non-technical audiences.

Clear communication with both technical and non-technical audiences.

Nice To Have

  • Experience with browser automation, AI agents, or simulated user behavior.

Experience with browser automation, AI agents, or simulated user behavior.

  • Background in e-commerce, search, recommendations, personalization, or buyer behavior modeling.

Background in e-commerce, search, recommendations, personalization, or buyer behavior modeling.

  • Experience with sequence or behavior representation learning, recommender embeddings, HSTU-style architectures, vLLM/GPU serving, Spark/GCS/Hugging Face-style data workflows.

Experience with sequence or behavior representation learning, recommender embeddings, HSTU-style architectures, vLLM/GPU serving, Spark/GCS/Hugging Face-style data workflows.

  • Prior experience translating research prototypes into reliable, production user-facing systems.

Prior experience translating research prototypes into reliable, production user-facing systems.

Why This Role

This is a high-leverage role in Shopify's applied AI work for merchants. You will help build the ML systems behind SimGym while working on frontier LLM/VLM agents, learned personas, evaluation science, and production systems with direct merchant impact.

This role may require on-call work. Shopify's hiring process moves quickly, and candidates should be prepared for a technical interview loop that includes pair programming using their own IDE.