lightningai

Senior Network Engineer - InfiniBand / UFM

Apply Now

At a Glance

Location
New York, New York, United States; San Francisco, California, United States; Seattle, Washington, United States
Experience
7+ years
Posted
2026-07-20T13:33:46-04:00

Key Requirements

Required Skills

LinuxPythonTerraform

Domain Knowledge

  • Automation
  • SaaS

Benefits & Perks

Health Insurance

Private medical and dental insurance (U.K.) Retirement and financial wellnes

Requirements

7+ years of data center networking experience.

3+ years supporting NVIDIA InfiniBand environments.

Hands-on experience with NVIDIA UFM Enterprise.

Experience deploying and operating Quantum and Quantum-2 InfiniBand switches.

InfiniBand Architecture

Partition Keys (PKeys)

Responsibilities

to design, deploy, automate, and operate next-generation AI Factory networking infrastructure supporting large-scale GPU clusters.

This role is responsible for building and maintaining high-performance NVIDIA Quantum InfiniBand fabrics that power AI training and inference environments utilizing NVIDIA UFM, NCCL, RoCEv2, and modern data center technologies.

The ideal candidate has deep expertise in InfiniBand networking, NVIDIA Unified Fabric Manager (UFM), large-scale GPU deployments, Linux networking, automation, and troubleshooting distributed AI workloads.

Design, deploy, and maintain large-scale NVIDIA InfiniBand fabrics supporting AI/ML GPU clusters.

Deploy and administer NVIDIA Unified Fabric Manager (UFM) Enterprise for monitoring, provisioning, telemetry, and fabric health.

Configure and optimize NVIDIA Quantum and Quantum-2 InfiniBand switches.