Principal Research Engineer

MicrosoftRedmond, WA,US, map[@type:Country name:US]full timePrincipal
Active

Job description

Lead model-training programs spanning pretraining from scratch when appropriate, continued pretraining of existing base models, supervised fine-tuning, and post-training. Improve model quality through changes to model architecture, training data, learning objectives, optimizers, hyperparameters, and end-to-end training recipes. Design and implement controlled experiments, ablation studies, and evaluation methods for accuracy, reasoning, robustness, generalization, safety, and domain performance. Diagnose and resolve training instability, convergence failures, numerical issues, data-quality defects, overfitting, catastrophic forgetting, and capability regressions. Own technical and design decisions and develop reusable training, evaluation, experimentation, and model-release practices that support multiple research initiatives. Partner with machine learning systems and infrastructure specialists to scale successful approaches across GPU clusters, improve training efficiency, and ensure reliable and reproducible execution. Provide technical direction across research and engineering teams, mentor engineers, communicate trade-offs, and translate ambiguous goals into measurable milestones. Bachelor's Degree in Computer Science, Machine Learning, Applied Mathematics or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, Python, C, C++, C#, or Java Master's Degree or Doctorate in Computer Science, Machine Learning, Applied Mathematics, or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, Python, C, C++, C#, or Java OR Bachelor's Degree in a related technical field AND 12+ years technical engineering experience with coding in languages including, but not limited to, Python, C, C++, C#, or Java OR equivalent experience. Experience training language or foundation models, including developing models through pretraining or continuing the training of existing base models. Experience improving model quality through changes to training data, model architecture, learning objectives, optimization methods, hyperparameters, or post-training techniques. Experience developing and debugging machine learning systems using Python and a modern machine learning framework such as PyTorch or JAX. Experience designing model evaluations or ablation studies used to assess training changes and guide technical decisions. Experience with one or more model-adaptation methods, such as continued pretraining, instruction tuning, parameter-efficient fine-tuning, preference optimization, reinforcement-learning-based post-training, or model distillation. Experience diagnosing training instability, convergence, numerical, data-quality, overfitting, forgetting, or model-regression issues. Experience with distributed model training, mixed precision, checkpointing, experiment tracking, GPU performance, or multi-node training. Experience improving training data through filtering, deduplication, mixture design, sampling, tokenization, synthetic data, provenance, or contamination controls. Experience contributing to significant model-training efforts, research publications, patents, open-source machine learning projects, or research-to-product transfers. Experience providing technical direction across ambiguous, cross-functional work, including mentoring engineers and driving execution across team boundaries. Experience applying responsible model-development practices involving privacy, safety evaluation, data governance, documentation, or reproducibility.

Similar jobs

USA, CA, Pleasanton / USA, GA, Atlanta / USA, WA, Seattlefull timePrincipal
verified Jun 11, 2026posted 3 months ago