anthropic
Staff + Sr. Software Engineer, Cloud Inference Launch Engineering
At a Glance
- Location
- San Francisco, California, United States
- Posted
- 2026-07-14T11:49:09-04:00
Key Requirements
Required Skills
Domain Knowledge
- Automation
- Banking
- Education
- Engineering
- Logistics
Requirements
Have significant software engineering experience, with a strong background in high-performance, large-scale distributed systems serving millions of users
Have a track record of building automation or test infrastructure that measurably improved release velocity or reliability
Have experience building or operating services on at least one major cloud platform (AWS, GCP, or Azure), with exposure to Kubernetes, Infrastructure as Code, or container orchestration
Thrive in cross-functional collaboration with both internal teams and external partners
Are a fast learner who can quickly ramp up on new technologies, hardware platforms, and provider ecosystems
LLM inference optimization, batching, and caching strategies
Responsibilities
The Cloud Inference team scales and optimizes Claude to serve the massive audiences of developers and enterprise companies across AWS, GCP, Azure, and future cloud service providers (CSPs).
We own the end-to-end product of Claude on each cloud platform, from API integration and intelligent request routing to inference execution, capacity management, and day-to-day operations.
Within Cloud Inference, the model & inference launch team owns the validation pipeline for our inference server and load balancer on these platforms.
We're responsible for every inference change — model launches, performance improvements, safeguard integrations — landing on cloud platforms with correctness, performance, and reliability intact.
This is high-leverage infrastructure work: validation has to be fast and cheap enough to run on the same accelerators that serve customers, trustworthy enough to replace manual checks, and consistent enough that a change working on Anthropic first-party means it works everywhere.
This directly determines how fast frontier models and features ship to every cloud platform, and how quickly performance wins reach production — reclaiming capacity at a time when compute is our scarcest resource.