xai

Member of Technical Staff - Multimodal Understanding

Apply Now

At a Glance

Location
Palo Alto, California, United States
Compensation
s. COMPENSATION AND BENEFITS: $180,000 - $440,000 USD Base salary is just one p
Posted
2026-04-17T17:05:55-04:00

Key Requirements

Required Skills

KubernetesPyTorchPythonRust

Requirements

Expert-level proficiency in Python (core language), with strong experience in at least one of: JAX / PyTorch / XLA.

Proven track record building or optimizing large-scale distributed ML systems (training/inference optimization, GPU utilization, multi-GPU/TPU setups, hardware co-design).

Deep experience designing and running data pipelines at scale: curation, filtering, generation, quality studies, especially for noisy/real-world multimodal data.

Strong fundamentals in evaluation design, benchmarks, reward modeling, or RL techniques (particularly for interactive/agentic behaviors).

Willingness to own end-to-end initiatives and do whatever it takes to deliver breakthrough user experiences.

Experience leading major improvements in model capabilities through better data, modeling, algorithms, or scaling.

Compensation & Benefits

$180,000 - $440,000 USD

Base salary is just one part of our total rewards package at xAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.

xAI is an equal opportunity employer. For details on data processing, view our

Recruitment Privacy Notice

.

Responsibilities

Advance understanding and generation across modalities—image, video, audio, and text—spanning the full stack: data curation/acquisition, tokenizer training, large-scale pre-training, post-training/alignment, infrastructure/scaling, evaluation, tooling/demos, and end-to-end product experiences.

Collaborate cross-functionally with pre-training, post-training, reasoning, data, applied, and product teams to deliver frontier capabilities in multimodal reasoning, world modeling, tool use, agentic behaviors, and interactive human-AI collaboration.

Contribute to building models that can see, hear, reason about, and interact with the world in real time at unprecedented levels.

Develop high-throughput pipelines for data acquisition, preprocessing, filtering, generation, decoding, loading, crawling, visualization, and management (images, videos, audio + text).

Advance multimodal capabilities including spatial-temporal compression, cross-modal alignment, world modeling, reasoning, emergent abilities, audio/image/video understanding & generation, real-time video processing, and noisy data handling.

Drive data quality and studies: curation (human/synthetic), filtering techniques, analysis, and scalable pipelines to support trillion-parameter models.