Nebius B.V. logo
Nebius B.V.
Company insightsEmployer Score
Remote - Europe
Last seen September 22, 2026
Posted September 22, 2026 · 0 days ago
Est. expiry October 27, 2026

Senior HPC Engineer

Remote - Europe
RemoteEuropeFull-time, Senior
At a glance

Senior HPC Engineer at Nebius focusing on GPU compute and InfiniBand in a European remote/hybrid setup.

Original job description

The role We’re looking for a Senior HPC Cluster Engineer to join our team and play a key role in the development of our hyperscaler platform. The GPU & InfiniBand team is responsible for enhancing and optimizing the core components of our Cloud platform, with a specific focus on GPU computing, InfiniBand networks, and the KVM/QEMU stack. You’ll work closely with hardware virtualization and device emulation technologies, ensuring high performance and security in multi-GPU, HPC environments. The role involves analyzing, troubleshooting, and improving infrastructure to support new hardware, fine-tuning system performance, and automating fault detection and resolution in a complex system. In this position, you will be responsible for: Tuning the performance of GPU clusters and InfiniBand networks to ensure optimal operation in HPC and GPU-based environments. Analyzing and troubleshooting the root cause of issues related to GPUs and InfiniBand networks, and proposing corrective actions. Integrating new hardware into the existing infrastructure, including support for new GPU hardware through software stacks like Kubernetes, QEMU, and KVM. Enhancing automation systems for proactive monitoring, detecting, and resolving issues in GPU and InfiniBand environments. Configuring and managing GPU devices and InfiniBand fabrics, ensuring efficient and reliable operation. We expect you to have: 5+ years of professional experience in system-level software development (focused on performance optimization, low-level programming). 3+ years of hands-on experience with Linux systems (administration, troubleshooting, and performance tuning). In-depth understanding of server architecture, including PCIe devices, NICs, Linux OS/Kernel, and high-performance computing (HPC) systems. Strong proficiency in one or more performance-oriented programming languages (C/C++, Go, Python). It would be a plus if you have: Experience with GPU end-to-end testing in a cluster environment using InfiniBand networking. Proven track record of analyzing and optimizing the performance of HPC workloads (e.g., simulations, data analysis, AI/ML workloads). Familiarity with RDMA, RoCE, and InfiniBand protocols for high-performance communication. Background in Software-Defined Networking (SDN) and experience with HPC cluster networking. Understanding of QEMU/KVM virtualization and managing virtualized environments. Experience with deep learning frameworks such as PyTorch and TensorFlow, and their integration with HPC systems. Familiarity with collective communication libraries like MPI and NCCL for distributed computing. We conduct coding interviews as part of the process.

The text above is the employer's original job description, extracted as written. Other details on this page, like salary, responsibilities, and requirements, are interpreted from that text by our system, not values the employer explicitly confirmed, so treat them as our best interpretation rather than verified facts.

Responsibilities
  • Tune the performance of GPU clusters and InfiniBand networks to ensure optimal operation in HPC and GPU-based environments
  • Analyze and troubleshoot the root cause of issues related to GPUs and InfiniBand networks, and propose corrective actions
  • Integrate new hardware into the existing infrastructure, including support for new GPU hardware through software stacks like Kubernetes, QEMU, and KVM
  • Enhance automation systems for proactive monitoring, detecting, and resolving issues in GPU and InfiniBand environments
  • Configure and manage GPU devices and InfiniBand fabrics, ensuring efficient and reliable operation
Requirements
  • 5+ years of professional experience in system-level software development (performance optimization, low-level programming)
  • 3+ years of hands-on experience with Linux systems (administration, troubleshooting, and performance tuning)
  • In-depth understanding of server architecture, including PCIe devices, NICs, Linux OS/Kernel, and HPC systems
  • Strong proficiency in C/C++, Go, Python
Nebius B.V. logo
Nebius B.V. · 114 open roles
Top locations: Remote - Global · 393 · Remote - Europe · 213 · Amsterdam, Netherlands · 86+69 other locations
View company
Current open roles at Nebius B.V. on JobCrawls
LocationActive listings
Remote - Global393
Remote - Europe213
Amsterdam, Netherlands86
Remote - United States30
London, United Kingdom30
Berlin, Germany21
Prague, Czech Republic19
Lappeenranta, Finland12
Mäntsälä, Finland12
Helsinki, Finland10
United Kingdom9
Amsterdam8
Remote - Finland7
Israel6
London5
Canada4
Canada · Remote - United States3
Germany · Remote - Europe3
Béthune, Pas-de-Calais3
Remote - United Kingdom3
Tel Aviv, Israel3
Singapore3
Remote · Singapore3
Remote - France2
Prague, Czechia2
Philadelphia, Pennsylvania2
Remote - Sweden2
Remote - Germany2
Abu Dhabi2
New York City, United States2
France · Paris2
Austin, United States2
Oklahoma2
Dubai2
Netherlands2
Remote - Netherlands2
France, Paris2
New Jersey2
Abu Dhabi · Dubai2
Canada, Remote - United States1
East London, United Kingdom1
Singapore, Singapore1
Minnesota, United States1
UK1
Berlin1
New Jersey, United States1
Abu Dhabi, Dubai1
Prague1
New Jersey, US1
Oklahoma, United States1
Finland1
California, United States1
Abu Dhabi, United Arab Emirates1
Paris, France1
Czechia1
Alabama, US1
Béthune, France1
Mäntsälä1
Poland1
Austin, Texas1
San Francisco Bay Area, United States1
Kansas City, United States1
Béthune, Pas-de-Calais, France1
Minnesota1
Philadelphia, United States1
Dallas, United States1
Remote - Spain1
London, UK1
Belgrade, Serbia1
Israel, Israel1
Remote - EU1
Paris1
Current role mix at Nebius B.V. on JobCrawls
Role typeActive listings
Backend Engineer288
Software Engineer84
Account Executive57
Sales Representative29
Product Manager22
Technical Program Manager7
System Engineer7
ML Engineer6
Technical Project Manager5
Technical Product Manager5
Site Reliability Engineer4
Senior Site Reliability Engineer4
Data Center Technician3
Data Center Operations Technician3
IT Technician2
Senior ML Engineer2
Data Scientist2
Project Manager2
Senior HPC Cluster Engineer2
Senior Technical Project Manager2
Senior Hypervisor Engineer2
Hypervisor Engineer2
ML Solutions Architect2
Delivery Manager2
Open Positions at Nebius2
Applied AI Researcher2
Product Designer2
Backend engineers, Frontend engineers, Site reliability engineers2
Security Engineer1
IT Support Manager1
Application Security Engineer1
MEP Engineer1
Datacenter IT Technician1
Partner Solutions Architect1
Customer Engineer1
Internal Control Business Partner1
Senior Technical Program Manager1
Solutions Architect Lead1
Solutions Architecture Leader1
Network Planning Project Manager1
Data Center IT Manager1
Technical Account Manager1
Security Solutions Engineer1
Support Engineer1
HPC Cluster Engineer1
ML Infrastructure Engineer1
Operations Specialist1
Backend Developer1
Data Center Operations Manager1
Mechanical Design Engineer1
Pricing Director1
Accountant1
Group Product Manager1
GTM Recruiting Manager1
Technical Support Engineer1
Machine Learning Engineer1
Mechanical Data Center Technician1
Instructional Designer1
AI Engineer1
Vulnerability Operations Center Lead1
IT Risk and Control Manager1
VP of Strategic Sales1
Human Resources Specialist1
Educational Content Author1
Offensive Security Lead1
Field Technical Lead1
Data Center Facilities Manager1
Detection Engineering & Response1
Forward Deployment Engineer1
Head of Channel Marketing1
Microsoft 365 Engineer1
Manager1
Applied AI Solutions Engineer1
Electrical Engineer1
Cloud Solution Architect1
IT Infrastructure Engineer1
Structured Cabling Design Engineer1
Principal, EMEA GTM - Physical AI1
VP of Developer Relations & Community1
Transportation Security Manager1
Regulatory Counsel1
Generalist1
Development Manager1
Data Center Electrical Lead1
Product Growth Analytics Lead1
Infrastructure Security Engineer1
Partner GTM Planning and Analytics1
Technical Due Diligence Manager1
GTM Lead1
Senior Network Engineer1
Application Integration Developer1
Specialist Solutions Architect1
Senior Software Developer1
Vendor Security & Standards Manager1
Financial Reporting Lead1
Communications Manager1
Physical Security Systems Technician1
Mechanical Engineer1
Data Center Logistics Specialist1
Applied ML Engineer1
Current role-level mix at Nebius B.V. on JobCrawls
Role levelActive listings
Mid-Level507
Senior97
Manager36
Executive5
Director1
Core values
Fast movingBold thinkingConstant growthMeaningful impactTrust and ownership

Never miss a new Senior HPC Engineer job in Remote - Europe

Weekly or daily digest. Unsubscribe anytime.

Similar jobs

From JobCrawls search: same role title and primary location as this listing (this job excluded). Up to 8 results.