lightningai
Senior Network Engineer - InfiniBand / UFM
At a Glance
- Location
- New York, New York, United States; San Francisco, California, United States; Seattle, Washington, United States
- Experience
- 7+ years
- Posted
- 2026-07-20T13:33:46-04:00
Key Requirements
Required Skills
Domain Knowledge
- Automation
- SaaS
Benefits & Perks
Private medical and dental insurance (U.K.) Retirement and financial wellnes
Requirements
7+ years of data center networking experience.
3+ years supporting NVIDIA InfiniBand environments.
Hands-on experience with NVIDIA UFM Enterprise.
Experience deploying and operating Quantum and Quantum-2 InfiniBand switches.
InfiniBand Architecture
Partition Keys (PKeys)
Responsibilities
to design, deploy, automate, and operate next-generation AI Factory networking infrastructure supporting large-scale GPU clusters.
This role is responsible for building and maintaining high-performance NVIDIA Quantum InfiniBand fabrics that power AI training and inference environments utilizing NVIDIA UFM, NCCL, RoCEv2, and modern data center technologies.
The ideal candidate has deep expertise in InfiniBand networking, NVIDIA Unified Fabric Manager (UFM), large-scale GPU deployments, Linux networking, automation, and troubleshooting distributed AI workloads.
Design, deploy, and maintain large-scale NVIDIA InfiniBand fabrics supporting AI/ML GPU clusters.
Deploy and administer NVIDIA Unified Fabric Manager (UFM) Enterprise for monitoring, provisioning, telemetry, and fabric health.
Configure and optimize NVIDIA Quantum and Quantum-2 InfiniBand switches.