
Senior HPC Engineer
Senior HPC Engineer at Nebius; GPU compute focus, hybrid EU remote.
The role We’re looking for a Senior HPC Cluster Engineer to join our team and play a key role in the development of our cutting-edge hyperscaler platform. The GPU & InfiniBand team is responsible for enhancing and optimizing the core components of our Cloud platform, with a specific focus on GPU computing, InfiniBand networks, and the KVM/QEMU stack. You’ll work closely with hardware virtualization and device emulation technologies, ensuring high performance and security in multi-GPU, HPC environments. The role involves analyzing, troubleshooting, and improving infrastructure to support new hardware, fine-tuning system performance, and automating fault detection and resolution in a complex system. In this position, you will be responsible for: Tuning the performance of GPU clusters and InfiniBand networks to ensure optimal operation in HPC and GPU-based environments. Analyzing and troubleshooting the root cause of issues related to GPUs and InfiniBand networks, and proposing corrective actions. Integrating new hardware into the existing infrastructure, including support for new GPU hardware through software stacks like Kubernetes, QEMU, and KVM. Enhancing automation systems for proactive monitoring, detecting, and resolving issues in GPU and InfiniBand environments. Configuring and managing GPU devices and InfiniBand fabrics, ensuring efficient and reliable operation. We expect you to have: 5+ years of professional experience in system-level software development (focused on performance optimization, low-level programming). 3+ years of hands-on experience with Linux systems (administration, troubleshooting, and performance tuning). In-depth understanding of server architecture, including PCIe devices, NICs, Linux OS/Kernel, and high-performance computing (HPC) systems. Strong proficiency in one or more performance-oriented programming languages (C/C++, Go, Python). It would be a plus if you have: Experience with GPU end-to-end testing in a cluster environment using InfiniBand networking. Proven track record of analyzing and optimizing the performance of HPC workloads (e.g., simulations, data analysis, AI/ML workloads). Familiarity with RDMA, RoCE, and InfiniBand protocols for high-performance communication. Background in Software-Defined Networking (SDN) and experience with HPC cluster networking. Understanding of QEMU/KVM virtualization and managing virtualized environments. Experience with deep learning frameworks such as PyTorch and TensorFlow, and their integration with HPC systems. Familiarity with collective communication libraries like MPI and NCCL for distributed computing. We conduct coding interviews as part of the process. #LI-LH2
The text above is the employer's original job description, extracted as written. Other details on this page, like salary, responsibilities, and requirements, are interpreted from that text by our system, not values the employer explicitly confirmed, so treat them as our best interpretation rather than verified facts.
- Tuning the performance of GPU clusters and InfiniBand networks to ensure optimal operation in HPC and GPU-based environments
- Analyzing and troubleshooting the root cause of issues related to GPUs and InfiniBand networks, and proposing corrective actions
- Integrating new hardware into the existing infrastructure, including support for new GPU hardware through software stacks like Kubernetes, QEMU, and KVM
- Enhancing automation systems for proactive monitoring, detecting, and resolving issues in GPU and InfiniBand environments
- Configuring and managing GPU devices and InfiniBand fabrics, ensuring efficient and reliable operation
- 5+ years of professional experience in system-level software development (focused on performance optimization, low-level programming)
- 3+ years of hands-on experience with Linux systems (administration, troubleshooting, and performance tuning)
- In-depth understanding of server architecture, including PCIe devices, NICs, Linux OS/Kernel, and high-performance computing (HPC) systems
- Strong proficiency in one or more performance-oriented programming languages (C/C++, Go, Python)
- Screening
- Coding interview
- HR interview
- Offer
| Location | Active listings |
|---|---|
| Remote - Global | 396 |
| Remote - Europe | 212 |
| Amsterdam, Netherlands | 84 |
| Remote - United States | 30 |
| London, United Kingdom | 27 |
| Berlin, Germany | 21 |
| Prague, Czech Republic | 19 |
| Lappeenranta, Finland | 12 |
| Mäntsälä, Finland | 12 |
| Remote - Finland | 11 |
| Helsinki, Finland | 9 |
| Amsterdam | 8 |
| United Kingdom | 8 |
| London | 5 |
| Canada | 4 |
| Israel | 4 |
| Canada · Remote - United States | 3 |
| Germany · Remote - Europe | 3 |
| Béthune, Pas-de-Calais | 3 |
| Singapore | 3 |
| Remote · Singapore | 3 |
| Philadelphia, Pennsylvania | 2 |
| Abu Dhabi | 2 |
| New York City, United States | 2 |
| Tel Aviv, Israel | 2 |
| France · Paris | 2 |
| Austin, United States | 2 |
| Oklahoma | 2 |
| Dubai | 2 |
| France, Paris | 2 |
| New Jersey | 2 |
| Israel, Israel | 2 |
| Remote - EU | 2 |
| Abu Dhabi · Dubai | 2 |
| Canada, Remote - United States | 1 |
| East London, United Kingdom | 1 |
| Singapore, Singapore | 1 |
| Minnesota, United States | 1 |
| Prague, Czechia | 1 |
| UK | 1 |
| Berlin | 1 |
| New Jersey, United States | 1 |
| Abu Dhabi, Dubai | 1 |
| Prague | 1 |
| New Jersey, US | 1 |
| Oklahoma, United States | 1 |
| Finland | 1 |
| California, United States | 1 |
| Abu Dhabi, United Arab Emirates | 1 |
| Paris, France | 1 |
| Czechia | 1 |
| Netherlands, Netherlands | 1 |
| Alabama, US | 1 |
| Béthune, France | 1 |
| Finland, Finland | 1 |
| Mäntsälä | 1 |
| Netherlands | 1 |
| Poland | 1 |
| Austin, Texas | 1 |
| San Francisco Bay Area, United States | 1 |
| Kansas City, United States | 1 |
| Béthune, Pas-de-Calais, France | 1 |
| Minnesota | 1 |
| Philadelphia, United States | 1 |
| Dallas, United States | 1 |
| London, UK | 1 |
| Belgrade, Serbia | 1 |
| Paris | 1 |
| Role type | Active listings |
|---|---|
| Backend Engineer | 293 |
| Software Engineer | 87 |
| Account Executive | 57 |
| Sales Representative | 29 |
| Product Manager | 22 |
| Technical Program Manager | 8 |
| Technical Project Manager | 7 |
| System Engineer | 7 |
| ML Engineer | 6 |
| Site Reliability Engineer | 5 |
| Data Center Technician | 3 |
| Technical Product Manager | 3 |
| Data Center Operations Technician | 3 |
| Senior Site Reliability Engineer | 3 |
| IT Technician | 2 |
| Senior ML Engineer | 2 |
| Data Scientist | 2 |
| Project Manager | 2 |
| Director | 2 |
| Senior HPC Cluster Engineer | 2 |
| Senior Hypervisor Engineer | 2 |
| Hypervisor Engineer | 2 |
| Machine Learning Engineer | 2 |
| ML Solutions Architect | 2 |
| Delivery Manager | 2 |
| Open Positions at Nebius | 2 |
| Applied AI Researcher | 2 |
| Senior Technical Product Manager | 2 |
| Product Designer | 2 |
| Technical Due Diligence Manager | 2 |
| Backend engineers, Frontend engineers, Site reliability engineers | 2 |
| Cloud Solutions Architect | 2 |
| Security Engineer | 1 |
| IT Support Manager | 1 |
| Customer Support Specialist | 1 |
| Application Security Engineer | 1 |
| MEP Engineer | 1 |
| Datacenter IT Technician | 1 |
| Partner Solutions Architect | 1 |
| Solutions Architect Lead | 1 |
| Solutions Architecture Leader | 1 |
| Network Planning Project Manager | 1 |
| Data Center IT Manager | 1 |
| Technical Account Manager | 1 |
| Support Engineer | 1 |
| Operations Specialist | 1 |
| Backend Developer | 1 |
| Data Center Operations Manager | 1 |
| HPC Engineer | 1 |
| Senior Software Engineer | 1 |
| Pricing Director | 1 |
| Accountant | 1 |
| Group Product Manager | 1 |
| GTM Recruiting Manager | 1 |
| Technical Support Engineer | 1 |
| AI/ML Specialist Solutions Architect | 1 |
| Instructional Designer | 1 |
| AI Engineer | 1 |
| Vulnerability Operations Center Lead | 1 |
| IT Risk and Control Manager | 1 |
| VP of Strategic Sales | 1 |
| Human Resources Specialist | 1 |
| Educational Content Author | 1 |
| Offensive Security Lead | 1 |
| Field Technical Lead | 1 |
| Data Center Facilities Manager | 1 |
| Backend Engineers | 1 |
| Forward Deployment Engineer | 1 |
| Head of Channel Marketing | 1 |
| Site Selection & Colocation Manager | 1 |
| Network Engineer | 1 |
| Microsoft 365 Engineer | 1 |
| Applied AI Solutions Engineer | 1 |
| Electrical Engineer | 1 |
| IT Infrastructure Engineer | 1 |
| Structured Cabling Design Engineer | 1 |
| VP of Developer Relations & Community | 1 |
| Transportation Security Manager | 1 |
| Regulatory Counsel | 1 |
| Generalist | 1 |
| Development Manager | 1 |
| Data Center Electrical Lead | 1 |
| Product Growth Analytics Lead | 1 |
| Infrastructure Security Engineer | 1 |
| Infrastructure Site Reliability Engineer | 1 |
| Partner GTM Planning and Analytics | 1 |
| Senior Network Engineer | 1 |
| Application Integration Developer | 1 |
| Specialist Solutions Architect | 1 |
| Key Customers Solutions Architect | 1 |
| Vendor Security & Standards Manager | 1 |
| Financial Reporting Lead | 1 |
| Communications Manager | 1 |
| Physical Security Systems Technician | 1 |
| Senior Applied ML Engineer | 1 |
| Mechanical Engineer | 1 |
| Software Developer | 1 |
| Mechanical Data Center Operations Technician | 1 |
| Data Center Logistics Specialist | 1 |
| Detection Engineer | 1 |
| Role level | Active listings |
|---|---|
| Mid-Level | 514 |
| Senior | 102 |
| Manager | 35 |
| Director | 2 |
| Executive | 2 |
Never miss a new HPC Engineer job in Remote - Europe
Weekly or daily digest. Unsubscribe anytime.
Similar jobs
From JobCrawls search: same role title and primary location as this listing (this job excluded). Up to 8 results.
