datadog

Product Manager II - Model Lab

Apply Now

At a Glance

Location
New York, United States
Experience
4+ years
Posted
2026-03-10T10:10:24-04:00

Key Requirements

Required Skills

PyTorchTensorFlow

Domain Knowledge

  • Engineering

Requirements

You have 4+ years of Product Management experience, with at least one 0→1 product or major feature launch

You have experience with developer tools, ML infrastructure, data platforms, or AI/LLM systems

You understand (or can quickly ramp on) modern ML training workflows — distributed training, experiment tracking, hyperparameter tuning, artifact storage, evaluation pipelines

You have familiarity with frameworks such as PyTorch, TensorFlow, JAX, or distributed training frameworks

You can translate highly technical concepts into compelling product narratives

You thrive in cross-functional environments and can align engineering, design, and GTM around a clear strategy

Compensation & Benefits

New hire stock equity (RSUs) and employee stock purchase plan (ESPP)

Continuous professional development, product training, and career pathing

Intradepartmental mentor and buddy program

An inclusive company culture, ability to join our Community Guilds (Datadog employee resource groups)

Access to Inclusion Talks, our Internal panel discussions

Global mental health benefits for employees and dependents age 6+

Responsibilities

Define the vision and strategy for Model Lab, establishing Datadog’s position in experiment tracking and model training observability

Lead 0→1 product discovery with AI research teams, ML platform engineers, and infrastructure leaders to deeply understand experiment tracking workflows and pain points

Design a system that unifies training metrics, hyperparameters, artifacts, dataset lineage, and model evaluation into a coherent and scalable experience

Identify differentiation opportunities vs.

competitive alternatives and homegrown internal tooling

Partner closely with engineering and design to ship foundational capabilities such as experiment lineage, artifact versioning, distributed training visibility, and reproducibility workflows