← All roles

Member of Technical Staff - Infrastructure

Gimlet Labs ·San Francisco, CA ·Onsite 11d ago
KubernetesTerraformAnsibleHelmPythonGoCUDALinuxPrometheusOTELGrafana

About the role

Gimlet is building the first multi-silicon neocloud designed for fast, efficient inference. As AI workloads become more complex and new hardware architectures emerge, simply deploying more GPUs isn't enough. The challenge is making increasingly diverse compute work together. Gimlet's platform intelligently partitions and routes workloads across heterogeneous hardware, enabling step-function improvements in performance and efficiency. Customers deploy through production-grade APIs without needing

Requirements

Experience in infrastructure, cluster engineering, platform engineering, SRE, HPC, or distributed systems. Deep Linux systems experience, including debugging performance, networking, storage, processes, and kernel-level issues. Experience operating Kubernetes, Slurm, Nomad, or similar orchestration and scheduling systems. Strong automation skills using tools such as Terraform, Ansible, Helm, Python, Go, or equivalent. Experience with GPU or accelerator infrastructure, including drivers, firmware

About the company

Gimlet Labs

Gimlet Labs is an applied research lab focused on developing the next generation of computing systems for AI workloads. They offer serverless inference for AI agents through their Gimlet Cloud platform and provide autonomous kernel generation for optimized AI performance with kforge. Their value proposition lies in efficiently and scalably serving AI, accelerating both training and inference without manual code changes.

View company page →
$150k–350k
Onsite San Francisco, CASan Francisco, US
Posted 11d ago
Apply for this role
Opens jobs.ashbyhq.com ↗
About the company
Gimlet Labs
https://www.gimletlabs.ai

Gimlet Labs is an applied research lab focused on developing the next generation of computing systems for AI workloads. They offer serverless inference for AI agents through their Gimlet Cloud platform and provide autonomous kernel generation for optimized AI performance with kforge. Their value proposition lies in efficiently and scalably serving AI, accelerating both training and inference without manual code changes.

View company page →