Nebius B.V. logo
Nebius B.V.
Company insightsEmployer Score
Remote - Europe
Last seen September 24, 2026
Posted September 24, 2026 · 2 days ago
Est. expiry October 29, 2026

Senior HPC Engineer

HPC Engineer
Remote - Europe
HybridEuropeFull-time, Senior
At a glance

Senior HPC Engineer at Nebius; GPU compute focus, hybrid EU remote.

Original job description

The role We’re looking for a Senior HPC Cluster Engineer to join our team and play a key role in the development of our cutting-edge hyperscaler platform. The GPU & InfiniBand team is responsible for enhancing and optimizing the core components of our Cloud platform, with a specific focus on GPU computing, InfiniBand networks, and the KVM/QEMU stack. You’ll work closely with hardware virtualization and device emulation technologies, ensuring high performance and security in multi-GPU, HPC environments. The role involves analyzing, troubleshooting, and improving infrastructure to support new hardware, fine-tuning system performance, and automating fault detection and resolution in a complex system. In this position, you will be responsible for: Tuning the performance of GPU clusters and InfiniBand networks to ensure optimal operation in HPC and GPU-based environments. Analyzing and troubleshooting the root cause of issues related to GPUs and InfiniBand networks, and proposing corrective actions. Integrating new hardware into the existing infrastructure, including support for new GPU hardware through software stacks like Kubernetes, QEMU, and KVM. Enhancing automation systems for proactive monitoring, detecting, and resolving issues in GPU and InfiniBand environments. Configuring and managing GPU devices and InfiniBand fabrics, ensuring efficient and reliable operation. We expect you to have: 5+ years of professional experience in system-level software development (focused on performance optimization, low-level programming). 3+ years of hands-on experience with Linux systems (administration, troubleshooting, and performance tuning). In-depth understanding of server architecture, including PCIe devices, NICs, Linux OS/Kernel, and high-performance computing (HPC) systems. Strong proficiency in one or more performance-oriented programming languages (C/C++, Go, Python). It would be a plus if you have: Experience with GPU end-to-end testing in a cluster environment using InfiniBand networking. Proven track record of analyzing and optimizing the performance of HPC workloads (e.g., simulations, data analysis, AI/ML workloads). Familiarity with RDMA, RoCE, and InfiniBand protocols for high-performance communication. Background in Software-Defined Networking (SDN) and experience with HPC cluster networking. Understanding of QEMU/KVM virtualization and managing virtualized environments. Experience with deep learning frameworks such as PyTorch and TensorFlow, and their integration with HPC systems. Familiarity with collective communication libraries like MPI and NCCL for distributed computing. We conduct coding interviews as part of the process. #LI-LH2

The text above is the employer's original job description, extracted as written. Other details on this page, like salary, responsibilities, and requirements, are interpreted from that text by our system, not values the employer explicitly confirmed, so treat them as our best interpretation rather than verified facts.

Responsibilities
  • Tuning the performance of GPU clusters and InfiniBand networks to ensure optimal operation in HPC and GPU-based environments
  • Analyzing and troubleshooting the root cause of issues related to GPUs and InfiniBand networks, and proposing corrective actions
  • Integrating new hardware into the existing infrastructure, including support for new GPU hardware through software stacks like Kubernetes, QEMU, and KVM
  • Enhancing automation systems for proactive monitoring, detecting, and resolving issues in GPU and InfiniBand environments
  • Configuring and managing GPU devices and InfiniBand fabrics, ensuring efficient and reliable operation
Requirements
  • 5+ years of professional experience in system-level software development (focused on performance optimization, low-level programming)
  • 3+ years of hands-on experience with Linux systems (administration, troubleshooting, and performance tuning)
  • In-depth understanding of server architecture, including PCIe devices, NICs, Linux OS/Kernel, and high-performance computing (HPC) systems
  • Strong proficiency in one or more performance-oriented programming languages (C/C++, Go, Python)
Recruitment Process
  1. Screening
  2. Coding interview
  3. HR interview
  4. Offer
Nebius B.V. logo
Nebius B.V. · 110 open roles
Top locations: Remote - Global · 396 · Remote - Europe · 212 · Amsterdam, Netherlands · 84+65 other locations
View company
Current open roles at Nebius B.V. on JobCrawls
LocationActive listings
Remote - Global396
Remote - Europe212
Amsterdam, Netherlands84
Remote - United States30
London, United Kingdom27
Berlin, Germany21
Prague, Czech Republic19
Lappeenranta, Finland12
Mäntsälä, Finland12
Remote - Finland11
Helsinki, Finland9
Amsterdam8
United Kingdom8
London5
Canada4
Israel4
Canada · Remote - United States3
Germany · Remote - Europe3
Béthune, Pas-de-Calais3
Singapore3
Remote · Singapore3
Philadelphia, Pennsylvania2
Abu Dhabi2
New York City, United States2
Tel Aviv, Israel2
France · Paris2
Austin, United States2
Oklahoma2
Dubai2
France, Paris2
New Jersey2
Israel, Israel2
Remote - EU2
Abu Dhabi · Dubai2
Canada, Remote - United States1
East London, United Kingdom1
Singapore, Singapore1
Minnesota, United States1
Prague, Czechia1
UK1
Berlin1
New Jersey, United States1
Abu Dhabi, Dubai1
Prague1
New Jersey, US1
Oklahoma, United States1
Finland1
California, United States1
Abu Dhabi, United Arab Emirates1
Paris, France1
Czechia1
Netherlands, Netherlands1
Alabama, US1
Béthune, France1
Finland, Finland1
Mäntsälä1
Netherlands1
Poland1
Austin, Texas1
San Francisco Bay Area, United States1
Kansas City, United States1
Béthune, Pas-de-Calais, France1
Minnesota1
Philadelphia, United States1
Dallas, United States1
London, UK1
Belgrade, Serbia1
Paris1
Current role mix at Nebius B.V. on JobCrawls
Role typeActive listings
Backend Engineer293
Software Engineer87
Account Executive57
Sales Representative29
Product Manager22
Technical Program Manager8
Technical Project Manager7
System Engineer7
ML Engineer6
Site Reliability Engineer5
Data Center Technician3
Technical Product Manager3
Data Center Operations Technician3
Senior Site Reliability Engineer3
IT Technician2
Senior ML Engineer2
Data Scientist2
Project Manager2
Director2
Senior HPC Cluster Engineer2
Senior Hypervisor Engineer2
Hypervisor Engineer2
Machine Learning Engineer2
ML Solutions Architect2
Delivery Manager2
Open Positions at Nebius2
Applied AI Researcher2
Senior Technical Product Manager2
Product Designer2
Technical Due Diligence Manager2
Backend engineers, Frontend engineers, Site reliability engineers2
Cloud Solutions Architect2
Security Engineer1
IT Support Manager1
Customer Support Specialist1
Application Security Engineer1
MEP Engineer1
Datacenter IT Technician1
Partner Solutions Architect1
Solutions Architect Lead1
Solutions Architecture Leader1
Network Planning Project Manager1
Data Center IT Manager1
Technical Account Manager1
Support Engineer1
Operations Specialist1
Backend Developer1
Data Center Operations Manager1
HPC Engineer1
Senior Software Engineer1
Pricing Director1
Accountant1
Group Product Manager1
GTM Recruiting Manager1
Technical Support Engineer1
AI/ML Specialist Solutions Architect1
Instructional Designer1
AI Engineer1
Vulnerability Operations Center Lead1
IT Risk and Control Manager1
VP of Strategic Sales1
Human Resources Specialist1
Educational Content Author1
Offensive Security Lead1
Field Technical Lead1
Data Center Facilities Manager1
Backend Engineers1
Forward Deployment Engineer1
Head of Channel Marketing1
Site Selection & Colocation Manager1
Network Engineer1
Microsoft 365 Engineer1
Applied AI Solutions Engineer1
Electrical Engineer1
IT Infrastructure Engineer1
Structured Cabling Design Engineer1
VP of Developer Relations & Community1
Transportation Security Manager1
Regulatory Counsel1
Generalist1
Development Manager1
Data Center Electrical Lead1
Product Growth Analytics Lead1
Infrastructure Security Engineer1
Infrastructure Site Reliability Engineer1
Partner GTM Planning and Analytics1
Senior Network Engineer1
Application Integration Developer1
Specialist Solutions Architect1
Key Customers Solutions Architect1
Vendor Security & Standards Manager1
Financial Reporting Lead1
Communications Manager1
Physical Security Systems Technician1
Senior Applied ML Engineer1
Mechanical Engineer1
Software Developer1
Mechanical Data Center Operations Technician1
Data Center Logistics Specialist1
Detection Engineer1
Current role-level mix at Nebius B.V. on JobCrawls
Role levelActive listings
Mid-Level514
Senior102
Manager35
Director2
Executive2
Core values
Fast movingBold thinkingConstant growthMeaningful impactTrust and real ownership

Never miss a new HPC Engineer job in Remote - Europe

Weekly or daily digest. Unsubscribe anytime.

Similar jobs

From JobCrawls search: same role title and primary location as this listing (this job excluded). Up to 8 results.