Job description
Design and own the architecture for provisioning and operator services that bring third-party graphics processing unit (GPU) servers online and keep them healthy, evaluating design options and documenting trade-offs for team and cross-team review. Write extensible, secure, well-tested code for bare-metal automation — spanning firmware, imaging, and network configuration — and mentor other engineers on coding patterns, code review, and quality bar. Serve as a designated responsible individual (DRI) on a rotational on-call basis, resolving live-site incidents within service level agreement (SLA) targets and driving root-cause fixes to prevent recurrence. Apply artificial intelligence (AI) tools and practices across the software development lifecycle (SDLC), while taking responsibility for the quality and safety of AI-generated designs, code, and other assets. Build the automation, telemetry, and monitoring needed to safely deploy and operate GPU fleets at scale, targeting zero-touch provisioning and rapid detection of failures. Bachelor's Degree in Computer Science or related technical field AND 4+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python Master's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR Bachelor's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience. 5+ years experience building or operating large-scale distributed systems or control planes (for example, orchestration, scheduling, or state-management services). 5+ years experience with Linux systems programming and networking fundamentals (switches, network interface cards).