← All roles

Software Engineer, Observability

xAI ·Palo Alto, CA ·Onsite 55d ago
GoRustScalaPrometheusGrafanaOTELClickHouseApache KafkaRedisKubernetesDistributed Systems

About the role

The Observability team builds and operates core infrastructure for monitoring, debugging, and optimizing system performance at massive scale, handling billions of time series and petabytes of logs. This role involves designing scalable observability infrastructure, building high-performance telemetry pipelines, developing APIs and UIs, and enforcing best practices across the company.

Requirements

Production-level proficiency in Go, Rust, Scala, or similar language; deep understanding of distributed systems and telemetry architecture; experience building and operating infrastructure at scale; familiarity with observability stacks like Prometheus, Grafana, OpenTelemetry, VictoriaMetrics, or ClickHouse; experience with Kafka, Redis, or large-scale time series databases; experience operating observability pipelines in Kubernetes or similar orchestration environments.

About the company

xAI

This company provides frontier AI models for developers, enabling them to build applications powered by reasoning, code, voice, image, and video generation. Their unified API, trained on a massive supercluster, offers a powerful and versatile platform for integrating advanced AI capabilities into various products and solutions.

View company page →
$180k–440k
Onsite Palo Alto, CAPalo Alto, United States
Posted 55d ago
Apply for this role
Opens job-boards.greenhouse.io ↗
About the company
xAI
https://x.ai

This company provides frontier AI models for developers, enabling them to build applications powered by reasoning, code, voice, image, and video generation. Their unified API, trained on a massive supercluster, offers a powerful and versatile platform for integrating advanced AI capabilities into various products and solutions.

View company page →