← All roles

Principal Engineer, Cluster Orchestration

CoreWeave ·Bellevue, WA / Sunnyvale, CA ·Onsite 30d ago
KubernetesGoHPC

About the role

As a Principal Engineer in AI Infrastructure, you will lead the design and evolution of the cluster orchestration systems that make this possible. This includes Slurm, Kubernetes, SUNK, and the control planes that support AI training, inference, and model onboarding at scale. You will define long-term architecture, solve hard scaling problems, and set technical direction across teams. Your work will directly affect how quickly customers can run models, how efficiently we use GPUs, and how the AI

Requirements

15+ years of experience building and operating large-scale distributed systems. Deep, practical knowledge of Kubernetes and Slurm internals. Experience running GPU-heavy platforms for AI training, inference, or HPC workloads. Strong background in Go and cloud-native systems development. Proven ability to set technical direction across teams without direct authority. Comfortable making high-impact technical decisions in complex systems. Bachelor’s or Master’s degree in a relevant field, or equi

About the company

CoreWeave

CoreWeave is a cloud provider specializing in an AI-native platform built to power complex AI workloads. They offer GPU and CPU compute, storage, and infrastructure control solutions, positioning themselves as the essential cloud for AI innovation and significantly reducing total cost of ownership.

View company page →
$206k–303k
Onsite Bellevue, WA / Sunnyvale, CABellevue, USSunnyvale, US
Posted 30d ago
Apply for this role
Opens coreweave.com ↗
About the company
CoreWeave
https://coreweave.com

CoreWeave is a cloud provider specializing in an AI-native platform built to power complex AI workloads. They offer GPU and CPU compute, storage, and infrastructure control solutions, positioning themselves as the essential cloud for AI innovation and significantly reducing total cost of ownership.

View company page →