
Senior HPC Engineer
Senior HPC Engineer at Nebius focusing on GPU compute and InfiniBand in a European remote/hybrid setup.
The role We’re looking for a Senior HPC Cluster Engineer to join our team and play a key role in the development of our hyperscaler platform. The GPU & InfiniBand team is responsible for enhancing and optimizing the core components of our Cloud platform, with a specific focus on GPU computing, InfiniBand networks, and the KVM/QEMU stack. You’ll work closely with hardware virtualization and device emulation technologies, ensuring high performance and security in multi-GPU, HPC environments. The role involves analyzing, troubleshooting, and improving infrastructure to support new hardware, fine-tuning system performance, and automating fault detection and resolution in a complex system. In this position, you will be responsible for: Tuning the performance of GPU clusters and InfiniBand networks to ensure optimal operation in HPC and GPU-based environments. Analyzing and troubleshooting the root cause of issues related to GPUs and InfiniBand networks, and proposing corrective actions. Integrating new hardware into the existing infrastructure, including support for new GPU hardware through software stacks like Kubernetes, QEMU, and KVM. Enhancing automation systems for proactive monitoring, detecting, and resolving issues in GPU and InfiniBand environments. Configuring and managing GPU devices and InfiniBand fabrics, ensuring efficient and reliable operation. We expect you to have: 5+ years of professional experience in system-level software development (focused on performance optimization, low-level programming). 3+ years of hands-on experience with Linux systems (administration, troubleshooting, and performance tuning). In-depth understanding of server architecture, including PCIe devices, NICs, Linux OS/Kernel, and high-performance computing (HPC) systems. Strong proficiency in one or more performance-oriented programming languages (C/C++, Go, Python). It would be a plus if you have: Experience with GPU end-to-end testing in a cluster environment using InfiniBand networking. Proven track record of analyzing and optimizing the performance of HPC workloads (e.g., simulations, data analysis, AI/ML workloads). Familiarity with RDMA, RoCE, and InfiniBand protocols for high-performance communication. Background in Software-Defined Networking (SDN) and experience with HPC cluster networking. Understanding of QEMU/KVM virtualization and managing virtualized environments. Experience with deep learning frameworks such as PyTorch and TensorFlow, and their integration with HPC systems. Familiarity with collective communication libraries like MPI and NCCL for distributed computing. We conduct coding interviews as part of the process.
The text above is the employer's original job description, extracted as written. Other details on this page, like salary, responsibilities, and requirements, are interpreted from that text by our system, not values the employer explicitly confirmed, so treat them as our best interpretation rather than verified facts.
- Tune the performance of GPU clusters and InfiniBand networks to ensure optimal operation in HPC and GPU-based environments
- Analyze and troubleshoot the root cause of issues related to GPUs and InfiniBand networks, and propose corrective actions
- Integrate new hardware into the existing infrastructure, including support for new GPU hardware through software stacks like Kubernetes, QEMU, and KVM
- Enhance automation systems for proactive monitoring, detecting, and resolving issues in GPU and InfiniBand environments
- Configure and manage GPU devices and InfiniBand fabrics, ensuring efficient and reliable operation
- 5+ years of professional experience in system-level software development (performance optimization, low-level programming)
- 3+ years of hands-on experience with Linux systems (administration, troubleshooting, and performance tuning)
- In-depth understanding of server architecture, including PCIe devices, NICs, Linux OS/Kernel, and HPC systems
- Strong proficiency in C/C++, Go, Python
| Location | Active listings |
|---|---|
| Remote - Global | 393 |
| Remote - Europe | 213 |
| Amsterdam, Netherlands | 86 |
| Remote - United States | 30 |
| London, United Kingdom | 30 |
| Berlin, Germany | 21 |
| Prague, Czech Republic | 19 |
| Lappeenranta, Finland | 12 |
| Mäntsälä, Finland | 12 |
| Helsinki, Finland | 10 |
| United Kingdom | 9 |
| Amsterdam | 8 |
| Remote - Finland | 7 |
| Israel | 6 |
| London | 5 |
| Canada | 4 |
| Canada · Remote - United States | 3 |
| Germany · Remote - Europe | 3 |
| Béthune, Pas-de-Calais | 3 |
| Remote - United Kingdom | 3 |
| Tel Aviv, Israel | 3 |
| Singapore | 3 |
| Remote · Singapore | 3 |
| Remote - France | 2 |
| Prague, Czechia | 2 |
| Philadelphia, Pennsylvania | 2 |
| Remote - Sweden | 2 |
| Remote - Germany | 2 |
| Abu Dhabi | 2 |
| New York City, United States | 2 |
| France · Paris | 2 |
| Austin, United States | 2 |
| Oklahoma | 2 |
| Dubai | 2 |
| Netherlands | 2 |
| Remote - Netherlands | 2 |
| France, Paris | 2 |
| New Jersey | 2 |
| Abu Dhabi · Dubai | 2 |
| Canada, Remote - United States | 1 |
| East London, United Kingdom | 1 |
| Singapore, Singapore | 1 |
| Minnesota, United States | 1 |
| UK | 1 |
| Berlin | 1 |
| New Jersey, United States | 1 |
| Abu Dhabi, Dubai | 1 |
| Prague | 1 |
| New Jersey, US | 1 |
| Oklahoma, United States | 1 |
| Finland | 1 |
| California, United States | 1 |
| Abu Dhabi, United Arab Emirates | 1 |
| Paris, France | 1 |
| Czechia | 1 |
| Alabama, US | 1 |
| Béthune, France | 1 |
| Mäntsälä | 1 |
| Poland | 1 |
| Austin, Texas | 1 |
| San Francisco Bay Area, United States | 1 |
| Kansas City, United States | 1 |
| Béthune, Pas-de-Calais, France | 1 |
| Minnesota | 1 |
| Philadelphia, United States | 1 |
| Dallas, United States | 1 |
| Remote - Spain | 1 |
| London, UK | 1 |
| Belgrade, Serbia | 1 |
| Israel, Israel | 1 |
| Remote - EU | 1 |
| Paris | 1 |
| Role type | Active listings |
|---|---|
| Backend Engineer | 288 |
| Software Engineer | 84 |
| Account Executive | 57 |
| Sales Representative | 29 |
| Product Manager | 22 |
| Technical Program Manager | 7 |
| System Engineer | 7 |
| ML Engineer | 6 |
| Technical Project Manager | 5 |
| Technical Product Manager | 5 |
| Site Reliability Engineer | 4 |
| Senior Site Reliability Engineer | 4 |
| Data Center Technician | 3 |
| Data Center Operations Technician | 3 |
| IT Technician | 2 |
| Senior ML Engineer | 2 |
| Data Scientist | 2 |
| Project Manager | 2 |
| Senior HPC Cluster Engineer | 2 |
| Senior Technical Project Manager | 2 |
| Senior Hypervisor Engineer | 2 |
| Hypervisor Engineer | 2 |
| ML Solutions Architect | 2 |
| Delivery Manager | 2 |
| Open Positions at Nebius | 2 |
| Applied AI Researcher | 2 |
| Product Designer | 2 |
| Backend engineers, Frontend engineers, Site reliability engineers | 2 |
| Security Engineer | 1 |
| IT Support Manager | 1 |
| Application Security Engineer | 1 |
| MEP Engineer | 1 |
| Datacenter IT Technician | 1 |
| Partner Solutions Architect | 1 |
| Customer Engineer | 1 |
| Internal Control Business Partner | 1 |
| Senior Technical Program Manager | 1 |
| Solutions Architect Lead | 1 |
| Solutions Architecture Leader | 1 |
| Network Planning Project Manager | 1 |
| Data Center IT Manager | 1 |
| Technical Account Manager | 1 |
| Security Solutions Engineer | 1 |
| Support Engineer | 1 |
| HPC Cluster Engineer | 1 |
| ML Infrastructure Engineer | 1 |
| Operations Specialist | 1 |
| Backend Developer | 1 |
| Data Center Operations Manager | 1 |
| Mechanical Design Engineer | 1 |
| Pricing Director | 1 |
| Accountant | 1 |
| Group Product Manager | 1 |
| GTM Recruiting Manager | 1 |
| Technical Support Engineer | 1 |
| Machine Learning Engineer | 1 |
| Mechanical Data Center Technician | 1 |
| Instructional Designer | 1 |
| AI Engineer | 1 |
| Vulnerability Operations Center Lead | 1 |
| IT Risk and Control Manager | 1 |
| VP of Strategic Sales | 1 |
| Human Resources Specialist | 1 |
| Educational Content Author | 1 |
| Offensive Security Lead | 1 |
| Field Technical Lead | 1 |
| Data Center Facilities Manager | 1 |
| Detection Engineering & Response | 1 |
| Forward Deployment Engineer | 1 |
| Head of Channel Marketing | 1 |
| Microsoft 365 Engineer | 1 |
| Manager | 1 |
| Applied AI Solutions Engineer | 1 |
| Electrical Engineer | 1 |
| Cloud Solution Architect | 1 |
| IT Infrastructure Engineer | 1 |
| Structured Cabling Design Engineer | 1 |
| Principal, EMEA GTM - Physical AI | 1 |
| VP of Developer Relations & Community | 1 |
| Transportation Security Manager | 1 |
| Regulatory Counsel | 1 |
| Generalist | 1 |
| Development Manager | 1 |
| Data Center Electrical Lead | 1 |
| Product Growth Analytics Lead | 1 |
| Infrastructure Security Engineer | 1 |
| Partner GTM Planning and Analytics | 1 |
| Technical Due Diligence Manager | 1 |
| GTM Lead | 1 |
| Senior Network Engineer | 1 |
| Application Integration Developer | 1 |
| Specialist Solutions Architect | 1 |
| Senior Software Developer | 1 |
| Vendor Security & Standards Manager | 1 |
| Financial Reporting Lead | 1 |
| Communications Manager | 1 |
| Physical Security Systems Technician | 1 |
| Mechanical Engineer | 1 |
| Data Center Logistics Specialist | 1 |
| Applied ML Engineer | 1 |
| Role level | Active listings |
|---|---|
| Mid-Level | 507 |
| Senior | 97 |
| Manager | 36 |
| Executive | 5 |
| Director | 1 |
Never miss a new Senior HPC Engineer job in Remote - Europe
Weekly or daily digest. Unsubscribe anytime.
Similar jobs
From JobCrawls search: same role title and primary location as this listing (this job excluded). Up to 8 results.