Senior GPU Infrastructure Engineer
Location: San Francisco
Salary: $180K - $250K
About Hyperbolic Hyperbolic Labs is pioneering AI infrastructure with its open-access GPU cloud, aggregating computing resources across the globe to offer an innovative GPU marketplace and AI inference service at up to 75% cost savings compared to traditional cloud providers. The mission is to democratise AI by breaking down barriers to computing power, making it affordable and accessible to developers and researchers everywhere.
Founded by co-founders with PhDs in AI, Math, and Computer Science, Hyperbolic raised a Series A and is preparing for significant growth. The team is approximately 30 people with an engineering-driven culture emphasising ownership, execution, and building at the cutting edge of GPU infrastructure. About the Role This is a foundational infrastructure role: building the multi-tenancy provisioning and virtualisation layer that transforms raw GPUs from diverse global suppliers into a programmable, orchestrated pool serving thousands of AI developers and researchers.
This is hands-on, deep infrastructure work - bare-metal provisioning, GPU scheduling, storage systems, and networking - not a wrapper around AWS. Key Responsibilities - Build and scale the GPU Cloud Marketplace by designing multi-tenancy provisioning and virtualisation solutions - Manage bare-metal provisioning and lifecycle management: IPMI/Redfish, BMC-based remote management, PXE boot, automated OS deployment - Design and implement GPU scheduling and orchestration: GPU type awareness, memory management, topology considerations, placement strategies for multi-GPU jobs, fragmentation minimisation - Build and maintain infrastructure automation using Terraform or Pulumi, CI/CD pipelines, secrets management, configuration management, and observability stacks - Design storage and data infrastructure for AI/ML workloads: object storage, high-IOPS block storage, distributed file systems for training data and checkpoints - Implement API design and cloud-init for automated provisioning and configuration - Work with hardware vendors and vendor engineering teams to troubleshoot issues and optimise integrations Requirements - Deep understanding of bare-metal provisioning and lifecycle management, including IPMI/Redfish, BMC-based remote management, PXE boot, and automated OS deployment workflows - Deep understanding of GPU scheduling and orchestration: GPU type awareness, memory management, topology considerations, placement strategies for multi-GPU jobs, fragmentation minimisation - Strong infrastructure and DevOps engineering skills: Terraform or Pulumi, CI/CD for infrastructure, secrets management, configuration management, and observability stack implementation - Experience with storage and data infrastructure for AI/ML workloads: object storage, high-IOPS block storage, and distributed file systems - Solid understanding of GPU architecture, CUDA, and GPU compute optimisation - Proven experience building and scaling cloud infrastructure or distributed systems in production environments - Excellent communication skills across technical and non-technical stakeholders Bonus Skills - Familiarity with high-performance networking: InfiniBand and RoCE (RDMA over Converged Ethernet) - Experience with distributed storage systems: Ceph, Weka, or VAST Data - Experience with Kubernetes GPU operators, Slurm, or Ray for distributed training - Background at GPU cloud or AI infrastructure companies