Lightning AI logo
Monthly
€12,298 - €15,192
Posted August 28, 2026 · 1 day agoLast seen August 27, 2026Est. expiry October 2, 2026

Network Engineer

Senior Network Engineer - InfiniBand / UFM
How this salary compares
Salary Context: Network Engineer

Hover or tap a row for full statistics (EUR / month on this chart).

Salary analysis

Compared with the selected benchmark ("Company in New York, United States"), this listing's salary midpoint is about 8% lower. The offer still falls within the benchmark range (€11,936–€22,426). The listed pay band (€14,167–€17,500) is tighter than the benchmark, which suggests lower salary variability. This benchmark is based on 1 comparable listings.

Monthly salary comparison for Network Engineer
MarketLower bound (25th percentile)MedianUpper bound (75th percentile)
Market Average: Network Engineer€4,192/per month€7,301/per month€10,557/per month
Company in New York, United States€11,936/per month€17,181/per month€22,426/per month
About the role

We are seeking an experienced Senior Network Engineer (InfiniBand / UFM) to design, deploy, automate, and operate next-generation AI Factory networking infrastructure supporting large-scale GPU clusters. This role is responsible for building and maintaining high-performance NVIDIA Quantum InfiniBand fabrics that power AI training and inference environments utilizing NVIDIA UFM, NCCL, RoCEv2, and modern data center technologies. The ideal candidate has deep expertise in InfiniBand networking, NVIDIA Unified Fabric Manager (UFM), large-scale GPU deployments, Linux networking, automation, and troubleshooting distributed AI workloads. What You’ll Do: Design, deploy, and maintain large-scale NVIDIA InfiniBand fabrics supporting AI/ML GPU clusters. Deploy and administer NVIDIA Unified Fabric Manager (UFM) Enterprise for monitoring, provisioning, telemetry, and fabric health. Configure and optimize NVIDIA Quantum and Quantum-2 InfiniBand switches. Troubleshoot fabric performance issues impacting NCCL, MPI, GPUDirect RDMA, and AI training jobs. Implement and validate fat-tree, Dragonfly+, Clos, and spine-leaf network architectures. Perform firmware lifecycle management for InfiniBand switches, adapters (HCAs), and UFM infrastructure. Optimize congestion control, adaptive routing, QoS, and traffic engineering for high-performance GPU communication. Work closely with AI platform, GPU infrastructure, storage, and systems engineering teams to deploy scalable AI Factory environments. Automate network provisioning using Python, Ansible, Git, REST APIs, and Infrastructure-as-Code methodologies. Monitor network health using UFM telemetry, Prometheus, Grafana, and other observability platforms. Support high availability, maintenance windows, incident response, root cause analysis, and capacity planning. Participate in architecture reviews and define networking standards for AI infrastructure. Required Qualifications: 7+ years of data center networking experience. 3+ years supporting NVIDIA InfiniBand environments. Hands-on experience with NVIDIA UFM Enterprise. Experience deploying and operating Quantum and Quantum-2 InfiniBand switches. Strong understanding of: InfiniBand Architecture, Subnet Manager (SM), Adaptive Routing, Congestion Control, Partition Keys (PKeys), LIDs, Queue Pairs (QP), Virtual Lanes (VL), Service Levels (SL). Strong Linux administration experience (Ubuntu). Experience with automation using Python and Ansible. Deep understanding of Layer 2 and Layer 3 networking. Experience with BGP, EVPN, VXLAN, and modern spine-leaf architectures. Experience with packet captures and troubleshooting using tcpdump, Wireshark, and ibdiagnet tools. Excellent troubleshooting and communication skills. What You Bring: 10+ years in large-scale data center networking; Deep expertise in spine-leaf architectures and L3 fabrics; Strong experience with BGP, EVPN, VXLAN; HPC/GPU environments experience; Designing networks for hyperscalers, neoclouds, or high-scale SaaS; Strong automation background (Python, Ansible, Terraform); Network observability tooling; Proven ability to design systems scalable to thousands of nodes; Strong documentation and architectural communication skills. Nice to Have: Netris and Terraform experience; multi-region backbone design; bare-metal provisioning; startup-scale infra exposure. Why This Role Matters: The network is the foundation of distributed AI training and will shape how Voltage Park scales; compensation includes base salary plus discretionary bonus, equity, and comprehensive benefits. The anticipated annual base salary range is: $170,000 - $210,000 USD. Benefits and Perks include health coverage, RSUs, retirement matching, unlimited PTO, winter break, parental leave, professional development, wellness stipends, sabbatical, flexible work, in-office meals, and more. We are an inclusive employer.

Job Details

Responsibilities

  • Design, deploy, and maintain large-scale NVIDIA InfiniBand fabrics
  • Deploy and administer NVIDIA Unified Fabric Manager (UFM) Enterprise
  • Configure and optimize NVIDIA Quantum and Quantum-2 InfiniBand switches
  • Troubleshoot fabric performance issues (NCCL, MPI, GPUDirect RDMA)
  • Implement fat-tree, Dragonfly+, Clos, and spine-leaf architectures
  • Firmware lifecycle management for InfiniBand switches and adapters
  • Optimize congestion control, adaptive routing, QoS, traffic engineering
  • Collaborate with AI platform, storage, and systems teams
  • Automate network provisioning (Python, Ansible, Git, REST APIs)
  • Monitor network health using UFM telemetry, Prometheus, Grafana
  • Support high availability, incident response, capacity planning
  • Participate in architecture reviews and networking standards

Requirements

  • 7+ years of data center networking experience
  • 3+ years supporting NVIDIA InfiniBand environments
  • Hands-on experience with NVIDIA UFM Enterprise
  • Experience deploying and operating Quantum and Quantum-2 InfiniBand switches
  • InfiniBand Architecture
  • Subnet Manager (SM)
  • Adaptive Routing
  • Congestion Control
  • Partition Keys (PKeys)
  • LIDs
  • Queue Pairs (QP)
  • Virtual Lanes (VL)
  • Service Levels (SL)
  • Ubuntu Linux
  • Python
  • Ansible
  • BGP
  • EVPN
  • VXLAN
  • tcpdump
  • Wireshark
  • ibdiagnet

Skills & Technologies

InfiniBandNVIDIA UFMNVIDIA Quantum switchesNCCLRoCEv2Linux networkingAnsiblePythonGitREST APIsTerraformPrometheusGrafanaIP networkingBGPEVPNVXLANtcpdumpWiresharkibdiagnet

Recruitment Process

  1. 1
    Application
  2. 2
    Screening
  3. 3
    Interview
  4. 4
    Offer
  5. 5
    Onboarding
Seen 3 days agoPartial Schema
Lightning AI logo
Lightning AI · 5 open roles
Top locations: London, United Kingdom · 1 · New York, United States · 1 · San Francisco, United States · 1+2 other locations
View company
Current open roles at Lightning AI on JobCrawls
LocationActive listings
London, United Kingdom1
New York, United States1
San Francisco, United States1
Seattle, United States1
Remote - Global1
Current role mix at Lightning AI on JobCrawls
Role typeActive listings
Research Engineer1
Current role-level mix at Lightning AI on JobCrawls
Role levelActive listings
Mid-Level1

Help us improve JobCrawls — sign in to sync saved jobs across devices, or send feedback anytime.