The client is seeking a Head of GPU Cloud to lead the engineering organization responsible for transforming large-scale GPU capacity into high-performance cloud services for AI workloads. This role will shape platform architecture, production inference, developer experiences, and engineering strategy, with technical decisions expected to influence customer adoption, infrastructure economics, and long-term growth. The position will involve close collaboration with infrastructure, networking, product, operations, and commercial leadership, working in a highly technical environment at the intersection of software engineering, distributed systems, and AI infrastructure.
Key responsibilities include building and scaling the engineering organization behind a nationwide GPU cloud platform, with end-to-end ownership across inference services, orchestration, APIs, platform reliability, and technical execution. The role will define architecture for efficient workload flow from customer request to accelerator, including routing, placement, model lifecycle management, caching, and memory-aware scheduling. It will also lead production model serving and optimization to improve throughput, latency, availability, accelerator utilization, and cost, including across relevant serving technologies such as vLLM, TensorRT LLM, and TGI, where applicable.
The head of GPU cloud will own Kubernetes-based GPU infrastructure strategy covering workload scheduling, elasticity, observability, deployment automation, multi-tenant isolation, and operational resilience. Additional responsibilities include developing hosted model and private inference capabilities, usage measurement and customer controls, and developer-facing services, as well as designing orchestration policies that align workloads to accelerators based on memory, performance objectives, capacity, and cluster topology. Required qualifications include 12+ years of progressive software engineering experience with substantial leadership responsibility across complex production systems and 5+ years leading engineering teams for business-critical infrastructure or platform services, along with demonstrated production experience in large-scale model inference, strong model-serving optimization expertise, advanced Kubernetes/containerized infrastructure knowledge, practical accelerator infrastructure knowledge, and the ability to balance performance, reliability, security, customer experience, and infrastructure economics while recruiting and developing senior technical talent.