datadog
Distinguished Architect, AI
At a Glance
- Location
- New York, New York, USA; San Francisco, California, USA
- Experience
- 10+ years
- Posted
- 2026-05-29T13:54:09-04:00
Key Requirements
Domain Knowledge
- Education
- Engineering
Requirements
10+ years of experience with at-scale distributed systems architecture, high-performance computing, or large-scale infrastructure.
Deep familiarity with the AI/LLM ecosystem, accelerator hardware (GPUs/TPUs), and modern orchestration frameworks.
: Able to think long term and creatively about a wide variety of technical and business challenges
: A true close-to-the-metal technologist who maintains deep hands-on credibility and can white-board architectural solutions seamlessly with senior engineers.
Strong understanding of best practices and real world challenges AI/LLM Ops and LLM Observability (LLMO).
: Excellent customer-facing presentation skills for large audiences, comfortable discussing complex technical details as well as with briefing executives or non-technical personas.
Compensation & Benefits
Best-in-breed onboarding
Generous global benefits
Intra-departmental mentor and buddy program for in-house networking
New hire stock equity (RSUs) and employee stock purchase plan (ESPP)
Continuous professional development, product training, and career pathing
An inclusive company culture, able to join our Community Guilds and Inclusion Talks
Responsibilities
: Demonstrate thought leadership in the AI/LLM space.
Influence key decision makers and stakeholders by connecting technical capabilities to organizational and business impact.
: Strategically partner with highly technical Founders, Heads of Infrastructure, and Research Lead peers.
Guide them on best practices and emerging industry trends in the AI/LLM space.
Lead high-level technical and architectural conversations around AI adoption.
: Lead deep-dive architecture reviews and design engagements with customer teams and their leaders to share industry trends, best practices, and demonstrate how Datadog can support high-throughput hyper scale AI workloads.