anthropic

Staff + Sr. Software Engineer, Cloud Inference Launch Engineering

Apply Now

At a Glance

Location
San Francisco, California, United States
Posted
2026-07-14T11:49:09-04:00

Key Requirements

Required Skills

AWSAzureGCPKubernetesPythonRust

Domain Knowledge

  • Automation
  • Banking
  • Education
  • Engineering
  • Logistics

Requirements

Have significant software engineering experience, with a strong background in high-performance, large-scale distributed systems serving millions of users

Have a track record of building automation or test infrastructure that measurably improved release velocity or reliability

Have experience building or operating services on at least one major cloud platform (AWS, GCP, or Azure), with exposure to Kubernetes, Infrastructure as Code, or container orchestration

Thrive in cross-functional collaboration with both internal teams and external partners

Are a fast learner who can quickly ramp up on new technologies, hardware platforms, and provider ecosystems

LLM inference optimization, batching, and caching strategies

Responsibilities

The Cloud Inference team scales and optimizes Claude to serve the massive audiences of developers and enterprise companies across AWS, GCP, Azure, and future cloud service providers (CSPs).

We own the end-to-end product of Claude on each cloud platform, from API integration and intelligent request routing to inference execution, capacity management, and day-to-day operations.

Within Cloud Inference, the model & inference launch team owns the validation pipeline for our inference server and load balancer on these platforms.

We're responsible for every inference change — model launches, performance improvements, safeguard integrations — landing on cloud platforms with correctness, performance, and reliability intact.

This is high-leverage infrastructure work: validation has to be fast and cheap enough to run on the same accelerators that serve customers, trustworthy enough to replace manual checks, and consistent enough that a change working on Anthropic first-party means it works everywhere.

This directly determines how fast frontier models and features ship to every cloud platform, and how quickly performance wins reach production — reclaiming capacity at a time when compute is our scarcest resource.