reddit

Senior Machine Learning Engineer, ML Efficiency

Apply Now

At a Glance

Location
United States
Work Regime
remote
Posted
2026-07-24T17:02:41-04:00

Key Requirements

Required Skills

PyTorch

Benefits & Perks

Health Insurance

edical, dental, and vision insurance, 401(k) program with employer match, ge

Requirements

Deep ML systems experience close to real production models and workloads, not just generic infra exposure.

Good customer and platform instincts: can solve concrete bottlenecks while keeping maintainability, adoption, and future reuse in mind.

Experience with GPU training or serving migrations.

Experience with PyTorch, distributed training frameworks, or kernel/runtime optimization.

Experience in organizations where platform and applied modeling responsibilities are split across multiple teams.

Experience with model compression or deployment optimizations such as quantization, pruning, distillation, or checkpoint optimization.

Compensation & Benefits

Comprehensive Healthcare Benefits and Income Replacement Programs

401k with Employer Match

Global Benefit programs that fit your lifestyle, from workspace to professional development to caregiving support

Family Planning Support

Gender-Affirming Care

Mental Health & Coaching Benefits

Responsibilities

Reddit is building a dedicated Ads ML Efficiency function to make model training and inference materially faster, cheaper, safer, and more scalable.

This person will be a key senior engineer on that team, owning meaningful efficiency work across training systems, inference and serving paths, launch-readiness tooling, and reusable optimization capabilities for Ads ML.

This role sits at the intersection of ML modeling, systems optimization, and engineering leverage.

The engineer will partner closely with ranking teams, serving owners, and ML Platform to identify important bottlenecks, land measurable efficiency wins, and help build the mechanisms that make those wins repeatable.

Independently own high-value optimization initiatives across training, inference, or launch-readiness for important Ads ML workloads.

Diagnose bottlenecks in real production systems using profiling, benchmarking, and observability rather than intuition-first debugging.