doordashusa
Engineering Manager, Reliability Platform
At a Glance
- Location
- San Francisco, CA; Sunnyvale, CA; New York City, NY
- Experience
- 5+ years
- Posted
- 2026-06-12T17:39:27-04:00
Key Requirements
Required Skills
Domain Knowledge
- Automation
- Engineering
- Healthcare
- Insurance
- Logistics
- Medical
- Regulatory
Benefits & Perks
at’s why we offer a comprehensive benefits package to all regular employees, which
Requirements
You have 5+ years of experience in an infrastructure, platform, or backend engineering role, showing you can deliver and maintain complex systems through a team or as an individual contributor.
You think in terms of products/platforms/customers while designing systems that other engineers depend on every day.
Cloud/Infra Fundamentals:
You’re comfortable broadly across the infra discussing topics related to AWS primitives, security best practices, containerization, and Infrastructure as Code.
You understand concepts like SLOs, error budgets, and incident response though this is a platform development team, not an SRE/oncall team.
You embrace the use of AI tools to be a more productive Engineering Manager, and instill the same mindset in your team.
Responsibilities
As a Software Engineer on the Reliability Platform team, you’ll help design, build, and operate services and infrastructure that deliver on the team’s broad mandate described above.
This team has a unique opportunity for breadth, often in collaboration with expert peers across the Infrastructure and Product teams.
Depending on need and interest, you may be working on mission-critical back-end services or pipelines, complex orchestration workflows, self-service UI, or AI Agent continuous improvements.
We have fully embraced the use of AI tools in everything we do, and believe in the incredible potential this provides while remaining pragmatic enough to ensure the critical infrastructure we maintain cannot be compromised.
Our goal is to deliver innovative next generation capabilities, as well as make data in our custody available to others pursuing the same.
A few examples of efforts the team has owned in recent years:
Team
The Reliability Platform role is a key pillar of DoorDash’s Production Lifecycle team, alongside Observability and Deploy Platform. This group’s mandate is to enable users and agents to reason about the health of our services, facilitate change control safety, and provide the means to rapidly address any unexpected state.
Ownership is fundamental in DoorDash culture, and all teams own what they build. We are not here to operate services on others’ behalf, but to provide tools that enable their success and ensure a consistently high level of quality for everything we do. We approach challenges with the pragmatic perspective of an SRE, and deliver solutions with the mindset of a SWE who detests toil and repetitive tasks. We use software and agents to “keep the lights on” and focus our energy on innovation that will level up the entire organization.
This mission falls into three main categories.
Service Health
– Providing SLO frameworks, analytics tools, and AI Agent enablement to extract high quality insights from our telemetry to pinpoint faults, or highlight deficiencies
Change Orchestration