← All roles

Principal Engineer, Cluster Orchestration

CoreWeave ·Bellevue, WA / Sunnyvale, CA ·Onsite 75d ago
KubernetesGoHPCDistributed Systems

About the role

CoreWeave is seeking a Principal Engineer to lead the design and evolution of cluster orchestration systems for AI infrastructure. This role involves defining long-term architecture for Kubernetes, Slurm, SUNK, and related systems, solving scaling problems, and setting technical direction. The engineer will work on scheduling, quota enforcement, multi-tenant GPU isolation, and reliability at scale. The role impacts how efficiently GPUs are used and how quickly customers run models.

Requirements

15+ years of experience in large-scale distributed systems. Deep knowledge of Kubernetes and Slurm internals. Experience with GPU-heavy platforms for AI training/inference. Strong background in Go and cloud-native development. Proven ability to set technical direction across teams. Bachelor's or Master's degree in relevant field or equivalent experience. Preferred: experience with Kueue, Kubeflow, Argo, Ray, Istio, Knative; ML platform engineering; scheduling strategies; open-source contributors

About the company

CoreWeave

CoreWeave is a cloud provider specializing in an AI-native platform built to power complex AI workloads. They offer GPU and CPU compute, storage, and infrastructure control solutions, positioning themselves as the essential cloud for AI innovation and significantly reducing total cost of ownership.

View company page →
$206k–303k
Onsite Bellevue, WA / Sunnyvale, CABellevue, USSunnyvale, USBellevue, WASunnyvale, CA
Posted 75d ago
Apply for this role
Opens coreweave.com ↗
About the company
CoreWeave
https://coreweave.com

CoreWeave is a cloud provider specializing in an AI-native platform built to power complex AI workloads. They offer GPU and CPU compute, storage, and infrastructure control solutions, positioning themselves as the essential cloud for AI innovation and significantly reducing total cost of ownership.

View company page →