← All roles

Senior Software Engineer - AI Infrastructure Performance Insights & Observability

CoreWeave ·Sunnyvale, CA / Bellevue, WA ·Onsite 31d ago
PrometheusGrafanaKubernetesPyTorchPythonGoC++OTELCUDAvLLM

About the role

This role involves building performance insights and observability systems for AI infrastructure, focusing on GPU, interconnect, and distributed training telemetry. You will own detection and diagnosis of performance signals across data centers, turning raw telemetry into real-time insight for engineers and executives. Responsibilities include designing time-series infrastructure, building data lakes, and optimizing query performance.

Requirements

5+ years building distributed systems, observability platforms, or performance engineering tooling. Strong coding in Python or Go. Hands-on Kubernetes at production scale, CI/CD, observability stacks (Prometheus, Grafana, OpenTelemetry). Working knowledge of time-series databases and PromQL/MetricsQL. Familiarity with data lake architectures (Iceberg, Parquet, Avro). Comfortable with GPU utilization, interconnect health, distributed training/inference. Strong communicator.

About the company

CoreWeave

CoreWeave is a cloud provider specializing in an AI-native platform built to power complex AI workloads. They offer GPU and CPU compute, storage, and infrastructure control solutions, positioning themselves as the essential cloud for AI innovation and significantly reducing total cost of ownership.

View company page →
$182k–242k
Onsite Sunnyvale, CA / Bellevue, WASunnyvale, USBellevue, US
Posted 31d ago
Apply for this role
Opens coreweave.com ↗
About the company
CoreWeave
https://coreweave.com

CoreWeave is a cloud provider specializing in an AI-native platform built to power complex AI workloads. They offer GPU and CPU compute, storage, and infrastructure control solutions, positioning themselves as the essential cloud for AI innovation and significantly reducing total cost of ownership.

View company page →