
Senior HPC Engineer, GPU Compute
Senior HPC Engineer at Nebius; work on GPU Compute in a hybrid Europe-wide setup.
The role We’re looking for a Senior HPC Cluster Engineer to join our team and play a key role in the development of our cutting-edge hyperscaler platform. The GPU & InfiniBand team is responsible for enhancing and optimizing the core components of our Cloud platform, with a specific focus on GPU computing, InfiniBand networks, and the KVM/QEMU stack. You’ll work closely with hardware virtualization and device emulation technologies, ensuring high performance and security in multi-GPU, HPC environments. The role involves analyzing, troubleshooting, and improving infrastructure to support new hardware, fine-tuning system performance, and automating fault detection and resolution in a complex system. In this position, you will be responsible for: Tuning the performance of GPU clusters and InfiniBand networks to ensure optimal operation in HPC and GPU-based environments. Analyzing and troubleshooting the root cause of issues related to GPUs and InfiniBand networks, and proposing corrective actions. Integrating new hardware into the existing infrastructure, including support for new GPU hardware through software stacks like Kubernetes, QEMU, and KVM. Enhancing automation systems for proactive monitoring, detecting, and resolving issues in GPU and InfiniBand environments. Configuring and managing GPU devices and InfiniBand fabrics, ensuring efficient and reliable operation. We expect you to have: 5+ years of professional experience in system-level software development (focused on performance optimization, low-level programming). 3+ years of hands-on experience with Linux systems (administration, troubleshooting, and performance tuning). In-depth understanding of server architecture, including PCIe devices, NICs, Linux OS/Kernel, and high-performance computing (HPC) systems. Strong proficiency in one or more performance-oriented programming languages (C/C++, Go, Python). It would be a plus if you have: Experience with GPU end-to-end testing in a cluster environment using InfiniBand networking. Proven track record of analyzing and optimizing the performance of HPC workloads (e.g., simulations, data analysis, AI/ML workloads). Familiarity with RDMA, RoCE, and InfiniBand protocols for high-performance communication. Background in Software-Defined Networking (SDN) and experience with HPC cluster networking. Understanding of QEMU/KVM virtualization and managing virtualized environments. Experience with deep learning frameworks such as PyTorch and TensorFlow, and their integration with HPC systems. Familiarity with collective communication libraries like MPI and NCCL for distributed computing. We conduct coding interviews as part of the process. #LI-LH2
The text above is the employer's original job description, extracted as written. Other details on this page, like salary, responsibilities, and requirements, are interpreted from that text by our system, not values the employer explicitly confirmed, so treat them as our best interpretation rather than verified facts.
Estimated from 75 comparable listings
- Screening
- Technical interview
- Coding interview
- Offer
| Location | Active listings |
|---|---|
| Remote - Global | 391 |
| Remote - Europe | 208 |
| Amsterdam, Netherlands | 82 |
| Remote - United States | 30 |
| London, United Kingdom | 27 |
| Berlin, Germany | 21 |
| Prague, Czech Republic | 20 |
| Remote - Finland | 17 |
| Lappeenranta, Finland | 12 |
| Mäntsälä, Finland | 12 |
| United Kingdom | 12 |
| Helsinki, Finland | 11 |
| Israel | 9 |
| Amsterdam | 6 |
| Canada | 4 |
| London | 3 |
| Remote | 3 |
| Singapore | 3 |
| Abu Dhabi | 2 |
| France, Paris | 2 |
| Remote - EU | 2 |
| Dubai | 2 |
| Tel Aviv, Israel | 2 |
| New York City, United States | 2 |
| Remote - North America | 2 |
| Remote - United Kingdom | 2 |
| Austin, United States | 2 |
| Netherlands | 2 |
| Finland | 1 |
| Oklahoma | 1 |
| Remote - Netherlands | 1 |
| Singapore, Singapore | 1 |
| San Francisco Bay Area, United States | 1 |
| Prague | 1 |
| Béthune, France | 1 |
| Remote - Sweden | 1 |
| Czechia | 1 |
| Canada · Remote - United States | 1 |
| Mäntsälä | 1 |
| Poland | 1 |
| Abu Dhabi, United Arab Emirates | 1 |
| Belgrade, Serbia | 1 |
| Abu Dhabi · Dubai | 1 |
| Béthune, Pas-de-Calais, France | 1 |
| UK | 1 |
| Alabama, US | 1 |
| France · Paris | 1 |
| Abu Dhabi, Dubai | 1 |
| Canada, Remote - United States | 1 |
| California, United States | 1 |
| Israel, Israel | 1 |
| East London, United Kingdom | 1 |
| Paris, France | 1 |
| London, UK | 1 |
| Oklahoma, United States | 1 |
| Philadelphia, Pennsylvania | 1 |
| Kansas City, United States | 1 |
| Remote - France | 1 |
| New Jersey, US | 1 |
| Remote · Singapore | 1 |
| New Jersey | 1 |
| New Jersey, United States | 1 |
| Austin, Texas | 1 |
| Remote - Germany | 1 |
| Minnesota, United States | 1 |
| Béthune, Pas-de-Calais | 1 |
| Philadelphia, United States | 1 |
| Dallas, United States | 1 |
| Paris | 1 |
| Berlin | 1 |
| Germany · Remote - Europe | 1 |
| Role type | Active listings |
|---|---|
| Backend Engineer | 298 |
| Software Engineer | 88 |
| Account Executive | 55 |
| Sales Representative | 29 |
| Product Manager | 20 |
| Technical Program Manager | 6 |
| Technical Project Manager | 6 |
| System Engineer | 6 |
| Site Reliability Engineer | 5 |
| ML Engineer | 5 |
| Technical Product Manager | 5 |
| Data Center Operations Technician | 3 |
| Senior Site Reliability Engineer | 3 |
| Data Center Technician | 3 |
| Data Scientist | 2 |
| Delivery Manager | 2 |
| Hypervisor Engineer | 2 |
| Product Designer | 2 |
| Senior Technical Program Manager | 2 |
| Project Manager | 2 |
| Applied AI Researcher | 2 |
| Open Positions at Nebius | 2 |
| Solutions Architect | 2 |
| Senior Hypervisor Engineer | 2 |
| IT Technician | 2 |
| Backend engineers, Frontend engineers, Site reliability engineers | 2 |
| Senior HPC Cluster Engineer | 2 |
| Microsoft 365 Engineer | 1 |
| Data Center Facilities Manager | 1 |
| Vulnerability Operations Center Lead | 1 |
| Accountant | 1 |
| Senior System Engineer | 1 |
| Mechanical Data Center Technician | 1 |
| Senior Technical Project Manager | 1 |
| Pricing Lead | 1 |
| Project Development Manager | 1 |
| Transportation Security Manager | 1 |
| Forward Deployment Engineer | 1 |
| Cloud Solution Architect | 1 |
| Group Product Manager | 1 |
| Applied ML Engineer | 1 |
| Solutions Architecture Leader | 1 |
| Generalist | 1 |
| HPC Cluster Engineer | 1 |
| Machine Learning Engineer | 1 |
| Physical Security Systems Technician | 1 |
| Manager, ML Solutions Architecture | 1 |
| Specialist Solutions Architect | 1 |
| Educational Content Author | 1 |
| IT Risk and Control Manager | 1 |
| Security Solutions Engineer | 1 |
| Field Technical Lead | 1 |
| Operations Specialist | 1 |
| Site Selection & Colocation Manager | 1 |
| IT Infrastructure Engineer | 1 |
| Offensive Security Lead | 1 |
| GTM Lead | 1 |
| Network Planning Project Manager | 1 |
| Applied AI Solutions Engineer | 1 |
| Software Developer | 1 |
| Partner Solutions Architect | 1 |
| Human Resources Specialist | 1 |
| Technical Account Manager | 1 |
| Senior Backend Developer | 1 |
| Electrical Engineer | 1 |
| Product Growth Analytics Lead | 1 |
| VP of Strategic Sales | 1 |
| Mechanical Engineer | 1 |
| Data Center IT Manager | 1 |
| Technical Support Engineer | 1 |
| Vendor Security & Standards Manager | 1 |
| Infrastructure Security Engineer | 1 |
| Tax Reporting Manager | 1 |
| Data Center Electrical Lead | 1 |
| AI and ISV Partner Business Development Manager | 1 |
| GTM Recruiting Manager | 1 |
| Financial Reporting Lead | 1 |
| Data Center Logistics Specialist | 1 |
| Mechanical Design Engineer | 1 |
| Solutions Architect Lead | 1 |
| MEP Engineer | 1 |
| Data Center IT Technician | 1 |
| Communications Manager | 1 |
| Instructional Designer | 1 |
| Senior Capacity Operations Manager | 1 |
| VP of Developer Relations & Community | 1 |
| Data Center Operations Manager | 1 |
| Detection Engineer | 1 |
| Structured Cabling Design Engineer | 1 |
| Data Engineer | 1 |
| Technical Due Diligence Manager | 1 |
| Support Engineer | 1 |
| Senior ML Solutions Architect | 1 |
| IT Support Manager | 1 |
| Backend Engineers | 1 |
| ML Infrastructure Engineer | 1 |
| Network Engineer | 1 |
| Partner GTM Planning and Analytics | 1 |
| Head of Channel Marketing | 1 |
| Senior Network Engineer | 1 |
| Role level | Active listings |
|---|---|
| Mid-Level | 515 |
| Senior | 95 |
| Manager | 40 |
| Executive | 3 |
| Director | 1 |
Never miss a new HPC Cluster Engineer job in Remote - Europe
Free weekly or daily digest. Unsubscribe anytime.