Job description
Apple's AIML Evaluation team builds the systems and methodologies that measure and improve the quality of foundation models and agentic experiences. We are looking for a senior, hands-on Machine Learning Engineering Manager to lead a small team working at the intersection of model evaluation, agent optimization, and data generation. In this role, you will help define how evaluation closes the loop with model and product development, turning observed quality gaps into targeted improvements to prompts, agent harnesses, datasets, and models.
You will combine technical depth with people leadership. You should be comfortable moving from research papers and experimental results to production-quality ML pipelines, while mentoring engineers and aligning teams around a clear technical direction. Your work will span Apple Foundation Models and product teams, with the goal of creating repeatable evaluation and refinement loops that improve the quality of Apple intelligence experiences. As a Senior Machine Learning Engineering Manager in AIML Evaluation, you will lead the technical strategy and execution for agent evaluation and automatic optimization. You will own systems that evaluate foundation models and agents, diagnose failure modes, and use those signals to drive automated prompt, context, tool, rubric, and agent-harness improvements. You will also help establish the interfaces between evaluation and post-training so that high-value failures can be converted into targeted data, environments, reward signals, and measurable model improvements.
This is a hands-on leadership role. You will prototype new approaches, participate in architecture and code reviews, design experiments, and help your team translate emerging research into scalable evaluation and optimization pipelines. You will partner closely with Apple Foundation Models, product engineering teams, and other AIML groups to build an evaluation flywheel that connects real product behavior with model and agent refinement. You will also work across the organization to advance synthetic data generation for both evaluation and post-training, with strong attention to data quality, representativeness, privacy, and reproducibility.
You will combine technical depth with people leadership. You should be comfortable moving from research papers and experimental results to production-quality ML pipelines, while mentoring engineers and aligning teams around a clear technical direction. Your work will span Apple Foundation Models and product teams, with the goal of creating repeatable evaluation and refinement loops that improve the quality of Apple intelligence experiences. As a Senior Machine Learning Engineering Manager in AIML Evaluation, you will lead the technical strategy and execution for agent evaluation and automatic optimization. You will own systems that evaluate foundation models and agents, diagnose failure modes, and use those signals to drive automated prompt, context, tool, rubric, and agent-harness improvements. You will also help establish the interfaces between evaluation and post-training so that high-value failures can be converted into targeted data, environments, reward signals, and measurable model improvements.
This is a hands-on leadership role. You will prototype new approaches, participate in architecture and code reviews, design experiments, and help your team translate emerging research into scalable evaluation and optimization pipelines. You will partner closely with Apple Foundation Models, product engineering teams, and other AIML groups to build an evaluation flywheel that connects real product behavior with model and agent refinement. You will also work across the organization to advance synthetic data generation for both evaluation and post-training, with strong attention to data quality, representativeness, privacy, and reproducibility.