Principal Technical Program Manager

MicrosoftRedmond, WA,US, map[@type:Country name:US]full timePrincipal
Active

Job description

Turn record-breaking AI infrastructure acceleration into the global operating standard.This role sits at the intersection of AI infrastructure engineering, data-center deployment, network architecture, capacity delivery, operational excellence, and organizational transformation.The Principal Technical Program Manager will take proven AI Network and GPU capacity acceleration practices—from design validation and deployment readiness through regional network readiness, infrastructure optimization, and deployment execution—and convert them into repeatable, measurable, and globally adopted capabilities.This is not a traditional program-management role.It is a technical execution and scaling leadership role designed for someone who can move comfortably between architecture conversations, critical-path deployment problems, executive decisions, engineering teams, regional operations, and global process transformation.The successful candidate will help provide solutions to make extraordinary deployment performance repeatable everywhere. What You Will Do:Lead Global AI Infrastructure Deployment Acceleration:Own the development, training workforce and scaling of a global AI network and infrastructure acceleration operating model spanning Commercial Cloud, new AI regions, new data centers, colo expansions, and large-scale GPU deployments. Translate successful deployment practices into repeatable architectures, operating mechanisms, technical decision frameworks, runbooks and deployment patterns that can be reused globally.Identify where AI infrastructure delivery remains dependent on manual coordination, repeated escalation, organizational handoffs or individual expertise and systematically eliminate those dependencies. Provide Deep Technical Program Leadership:Develop strong end-to-end understanding of the AI infrastructure deployment lifecycle and its critical dependencies across areas including: GPU infrastructure and large-scale AI clusters Data-center network architecture High-performance Ethernet and InfiniBand environments Front-end and back-end network infrastructure Regional and WAN connectivity RNG and IDF readiness Data-center fit-out and infrastructure readiness Capacity planning and deployment sequencing Facility/network dependencies Infrastructure validation and acceptance High-speed optical technologies and evolving network architectures Cloud and distributed infrastructure systems Bachelor's Degree AND 6+ years experience in Technical Program Management, AI Data center Infrastructure engineering,Cloud or data center platform delivery OR equivalent experience. 3+ years of experience managing cross-functional and/or cross-team projects. These requirements include, but are not limited to the following specialized security screenings: Strong technical judgment with the ability to engage credibly with engineering partners. 6+ years' experience in AI data center infrastructure, GPU platforms, or hyperscale networking. Backend fabrics using InfiniBand and/or advanced technologies, including optics Large-scale GPU clusters (training and/or inference) High-density rack architectures, optics, and cabling systems Design standardization Automation and tooling Parallel bring-up and validation models Exposure to AI platform bring-up, capacity activation, or production readiness reviews.

Similar jobs