anthropic

Engineering Manager, Safeguards Review Tooling

Apply Now

At a Glance

Location
San Francisco, California, United States
Experience
4+ years
Posted
2026-06-10T17:27:09-04:00

Key Requirements

Domain Knowledge

  • Automation
  • Banking
  • Education
  • Engineering
  • Legal
  • Logistics
  • Regulatory

Requirements

Experience managing software engineering teams, including hiring, coaching, and developing engineers

A technical background in full-stack or platform engineering, with the ability to engage deeply in architecture and design discussions

Experience shipping internal tools or platforms with demanding operational users, and a track record of improving their workflows measurably

Experience working cross-functionally with non-engineering partners such as operations, policy, or legal teams

Care about the societal impacts of AI and want your work to make powerful systems safer

Experience building trust and safety, integrity, fraud, or abuse-prevention tooling, or other systems supporting human review at scale

Responsibilities

The Safeguards team is responsible for ensuring our models and products are developed and deployed safely.

We're looking for an Engineering Manager to lead our Review Tooling team, which builds the systems that humans — and increasingly Claude — use to investigate potential harms and take enforcement actions across Anthropic's first-party products and third-party cloud platforms.

This is a foundational role: you'll own the tools our safety investigators rely on to understand what's happening on our platforms and act on it, as well as the platform underneath those tools.

That platform includes analytics capabilities, privacy-preserving primitives that keep review workflows compatible with our data retention commitments, and a sandbox environment where new review interfaces and workflows can be built and iterated quickly.

As model capabilities and usage grow, you'll also drive how we scale review through automation — building systems where Claude meaningfully extends what human reviewers can do, while keeping people in the loop where their judgment matters most.

You'll partner closely with policy, operations, data science, and legal teams to ensure our enforcement systems are effective, accurate, and trustworthy.