Turkey
verified Jun 17, 2026posted 4 months ago
This job appears expired.
Last verified Jun 18, 2026.View similar active jobs →
You are a hands-on engineer who builds the software and processes that keep a large fleet of GPU servers healthy and productive. You write systems and tooling for managing 1000s of servers including provisioning, health monitoring, error detection, and recovery — and when something breaks that automation can’t fix, you drive resolution with partners.
San Francisco, CA (we are open to remote in the US for Senior and Staff levels)