Compute Orchestration & Scheduling

Microsoftfull timeSenior
Activeverified Jun 12, 2026

Job description

Develop and tune the pretraining scalable software for Nvidia Grace-Blackwell (GB), Vera-Rubin (VR) and AMD MIxxx architectures Scale GB, VR and AMD MIxxx GPU clusters beyond thousands of GPUs Gather data and insights to develop the GPU & CPU compute roadmap suitable for large-scale AI labs Actively contribute to the development of AI models that are powering our innovative products Find a path to get things done despite roadblocks to get your work into the hands of users quickly and iteratively Enjoy working in a fast-paced, design-driven, product development cycle Embody our Culture and Values Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, Python, Go or JavaScript Master's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR Bachelor's Degree in Computer Science or related technical field AND 12+ years technical engineering experience with coding in languages including, but not limited to, C, C++, Python, Go or JavaScript OR equivalent experience. Expertise in one of Ray, Kubernetes, Kueue, Volcano, or any other AI-focused systems for scheduling / scaling / fault-tolerance.