Principal Software Engineer(M365 Storage Fabric team)

MicrosoftBeijing, Beijing,CN, map[@type:Country name:CN]full time
Active

Job description

Lead the design and development of large-scale distributed systems that manage resource allocation, workload execution, and service protection. Define and drive technical strategy for platform capabilities that improve reliability, efficiency, scalability, and operational excellence. Build intelligent, signal-driven automation using telemetry, health indicators, and real-time platform insights. Develop solutions that balance customer experience, infrastructure utilization, operational cost, and service performance. Drive innovations that proactively identify, mitigate, and prevent service disruptions. Influence architecture, design, and engineering best practices across multiple teams. Mentor engineers and contribute to a culture of technical excellence and continuous improvement. Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python Master's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR Bachelor's Degree in Computer Science or related technical field AND 12+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience. Deep understanding of distributed systems, concurrency, reliability engineering, scalability, and platform architecture. Experience developing and operating highly available backend services at scale. Demonstrated ability to lead technically complex initiatives across multiple teams and organizations. Solid problem-solving skills involving system performance, resiliency, resource management, and operational excellence. Demonstrated ability to effectively leverage AI-assisted engineering tools and autonomous coding agents to improve software development productivity, quality, and operational effectiveness. Ability to critically evaluate, validate, and refine AI-generated code, designs, tests, diagnostics, and recommendations while maintaining full engineering ownership and accountability. Experience building infrastructure platforms, resource management systems, scheduling systems, load balancing systems, storage platforms, or large-scale backend services. Solid understanding of workload management, traffic engineering, fault tolerance, admission control, and capacity planning. Experience using telemetry, monitoring, and service health signals to drive automated operational decisions. Experience supporting services with rapidly changing demand patterns and large-scale customer workloads. Hands-on experience with cloud-native architectures and distributed platforms such as Azure or similar cloud environments. Microservices and service-oriented architectures Event-driven systems Large-scale telemetry systems Containerized and cloud-native environments Experience building or supporting AI-powered services and high-throughput systems. Familiarity with AI workload characteristics, including bursty traffic, latency sensitivity, resource contention, and dynamic scaling requirements. Experience integrating AI-assisted engineering workflows into software development, testing, debugging, code review, and operational processes. Proven ability to drive technical alignment and influence engineering decisions across organizational boundaries. Solid communication skills with the ability to explain complex technical concepts to diverse audiences.