
Senior HPC Engineer
Senior HPC Engineer at Nebius; GPU compute, InfiniBand, Linux HPC; 5+ yrs exp; hybrid Europe role.
Wedre looking for a Senior HPC Cluster Engineer to join our team and play a key role in the development of our cutting-edge hyperscaler platform. The GPU & InfiniBand team is responsible for enhancing and optimizing the core components of our Cloud platform, with a specific focus on GPU computing, InfiniBand networks, and the KVM/QEMU stack. Youll work closely with hardware virtualization and device emulation technologies, ensuring high performance and security in multi-GPU, HPC environments. The role involves analyzing, troubleshooting, and improving infrastructure to support new hardware, fine-tuning system performance, and automating fault detection and resolution in a complex system. In this position, you will be responsible for: Tuning GPU clusters and InfiniBand networks to ensure optimal operation in HPC and GPU-based environments; Analyzing root causes of issues related to GPUs and InfiniBand networks and proposing corrective actions; Integrating new hardware into existing infrastructure, including support for new GPU hardware through software stacks like Kubernetes, QEMU, and KVM; Enhancing automation systems for proactive monitoring, detecting, and resolving issues in GPU and InfiniBand environments; Configuring and managing GPU devices and InfiniBand fabrics, ensuring efficient operation. We expect you to have: 5+ years of professional experience in system-level software development (focused on performance optimization, low-level programming); 3+ years of hands-on experience with Linux systems (administration, troubleshooting, and performance tuning); In-depth understanding of server architecture, including PCIe devices, NICs, Linux OS/Kernel, and HPC systems; Strong proficiency in one or more performance-oriented programming languages (C/C++, Go, Python). It would be a plus if you have: Experience with GPU end-to-end testing in a cluster environment using InfiniBand networking; Proven track record of analyzing and optimizing HPC workloads; Familiarity with RDMA, RoCE, and InfiniBand protocols; Background in Software-Defined Networking (SDN) and experience with HPC cluster networking; Understanding of QEMU/KVM virtualization and managing virtualized environments; Experience with deep learning frameworks such as PyTorch and TensorFlow, and their integration with HPC systems; Familiarity with collective communication libraries like MPI and NCCL for distributed computing. We conduct coding interviews as part of the process. Benefits & Perks: Competitive compensation; Career growth and learning opportunities; Flexibility and ownership; Collaborative and innovative culture; Opportunity to work on impactful AI projects; International environment and talented teams.
The text above is the employer's original job description, extracted as written. Other details on this page, like salary, responsibilities, and requirements, are interpreted from that text by our system, not values the employer explicitly confirmed, so treat them as our best interpretation rather than verified facts.
- Tune GPU clusters and InfiniBand networks to ensure optimal operation in HPC and GPU-based environments
- Analyze and troubleshoot root causes of GPU/InfiniBand issues and propose corrective actions
- Integrate new hardware into existing infrastructure, including support for new GPU hardware through software stacks like Kubernetes, QEMU, and KVM
- Enhance automation systems for proactive monitoring, detecting, and resolving issues in GPU/InfiniBand environments
- Configure and manage GPU devices and InfiniBand fabrics for efficient operation
- 5+ years of professional experience in system-level software development focused on performance optimization and low-level programming
- 3+ years of hands-on experience with Linux systems (administration, troubleshooting, and performance tuning)
- Deep understanding of server architecture including PCIe devices, NICs, Linux OS/Kernel, and HPC systems
- Strong proficiency in one or more performance-oriented programming languages (C/C++, Go, Python)
| Location | Active listings |
|---|---|
| Remote - Global | 395 |
| Remote - Europe | 211 |
| Amsterdam, Netherlands | 85 |
| Remote - United States | 30 |
| London, United Kingdom | 25 |
| Berlin, Germany | 20 |
| Prague, Czech Republic | 19 |
| Lappeenranta, Finland | 12 |
| Mäntsälä, Finland | 12 |
| Helsinki, Finland | 10 |
| United Kingdom | 10 |
| Amsterdam | 8 |
| Remote - Finland | 6 |
| Israel | 6 |
| London | 5 |
| Canada | 4 |
| Canada · Remote - United States | 3 |
| Germany · Remote - Europe | 3 |
| Béthune, Pas-de-Calais | 3 |
| Singapore | 3 |
| Remote · Singapore | 3 |
| Philadelphia, Pennsylvania | 2 |
| Abu Dhabi | 2 |
| New York City, United States | 2 |
| Remote - United Kingdom | 2 |
| Tel Aviv, Israel | 2 |
| France · Paris | 2 |
| Austin, United States | 2 |
| Oklahoma | 2 |
| Dubai | 2 |
| Netherlands | 2 |
| France, Paris | 2 |
| New Jersey | 2 |
| Israel, Israel | 2 |
| Abu Dhabi · Dubai | 2 |
| Canada, Remote - United States | 1 |
| East London, United Kingdom | 1 |
| Remote - France | 1 |
| Singapore, Singapore | 1 |
| Minnesota, United States | 1 |
| Remote - Israel | 1 |
| UK | 1 |
| Remote - Sweden | 1 |
| Berlin | 1 |
| New Jersey, United States | 1 |
| Remote - Germany | 1 |
| Abu Dhabi, Dubai | 1 |
| Prague | 1 |
| New Jersey, US | 1 |
| Finland | 1 |
| Oklahoma, United States | 1 |
| California, United States | 1 |
| Abu Dhabi, United Arab Emirates | 1 |
| Paris, France | 1 |
| Czechia | 1 |
| Alabama, US | 1 |
| Béthune, France | 1 |
| Mäntsälä | 1 |
| Remote - Netherlands | 1 |
| Poland | 1 |
| Austin, Texas | 1 |
| San Francisco Bay Area, United States | 1 |
| Kansas City, United States | 1 |
| Béthune, Pas-de-Calais, France | 1 |
| Minnesota | 1 |
| Philadelphia, United States | 1 |
| Dallas, United States | 1 |
| London, UK | 1 |
| Belgrade, Serbia | 1 |
| Remote - EU | 1 |
| Paris | 1 |
| Role type | Active listings |
|---|---|
| Backend Engineer | 287 |
| Software Engineer | 84 |
| Account Executive | 57 |
| Sales Representative | 29 |
| Product Manager | 21 |
| Technical Program Manager | 7 |
| System Engineer | 7 |
| Technical Project Manager | 6 |
| Site Reliability Engineer | 6 |
| ML Engineer | 6 |
| Technical Product Manager | 5 |
| Data Center Technician | 3 |
| Data Center Operations Technician | 3 |
| Senior Site Reliability Engineer | 3 |
| IT Technician | 2 |
| Data Scientist | 2 |
| Senior ML Engineer | 2 |
| Project Manager | 2 |
| Senior HPC Cluster Engineer | 2 |
| Senior Hypervisor Engineer | 2 |
| Hypervisor Engineer | 2 |
| ML Solutions Architect | 2 |
| Delivery Manager | 2 |
| Open Positions at Nebius | 2 |
| Applied AI Researcher | 2 |
| Development Manager | 2 |
| Product Designer | 2 |
| Technical Due Diligence Manager | 2 |
| Backend engineers, Frontend engineers, Site reliability engineers | 2 |
| Security Engineer | 1 |
| IT Support Manager | 1 |
| MEP Engineer | 1 |
| Datacenter IT Technician | 1 |
| Partner Solutions Architect | 1 |
| Senior Technical Program Manager | 1 |
| Solutions Architect Lead | 1 |
| Solutions Architecture Leader | 1 |
| Network Planning Project Manager | 1 |
| Data Center IT Manager | 1 |
| Technical Account Manager | 1 |
| Security Solutions Engineer | 1 |
| ML Infrastructure Engineer | 1 |
| Operations Specialist | 1 |
| Backend Developer | 1 |
| Data Center Operations Manager | 1 |
| Data Center IT Technician | 1 |
| HPC Engineer | 1 |
| Pricing Director | 1 |
| Accountant | 1 |
| Senior Technical Project Manager | 1 |
| Group Product Manager | 1 |
| GTM Recruiting Manager | 1 |
| Technical Support Engineer | 1 |
| Data Engineer | 1 |
| Detection Engineering & Response Lead | 1 |
| Instructional Designer | 1 |
| AI Engineer | 1 |
| Vulnerability Operations Center Lead | 1 |
| IT Risk and Control Manager | 1 |
| VP of Strategic Sales | 1 |
| Human Resources Specialist | 1 |
| Educational Content Author | 1 |
| Director, GSI Partnerships | 1 |
| Field Technical Lead | 1 |
| Data Center Facilities Manager | 1 |
| Backend Engineers | 1 |
| Forward Deployment Engineer | 1 |
| Head of Channel Marketing | 1 |
| Site Selection & Colocation Manager | 1 |
| Network Engineer | 1 |
| Microsoft 365 Engineer | 1 |
| Applied AI Solutions Engineer | 1 |
| Electrical Engineer | 1 |
| Cloud Solution Architect | 1 |
| Offensive Security | 1 |
| IT Infrastructure Engineer | 1 |
| Structured Cabling Design Engineer | 1 |
| VP of Developer Relations & Community | 1 |
| Senior Support Engineer | 1 |
| Transportation Security Manager | 1 |
| Regulatory Counsel | 1 |
| Generalist | 1 |
| Data Center Electrical Lead | 1 |
| Product Growth Analytics Lead | 1 |
| Infrastructure Security Engineer | 1 |
| Partner GTM Planning and Analytics | 1 |
| GTM Lead | 1 |
| Senior Network Engineer | 1 |
| Application Integration Developer | 1 |
| Specialist Solutions Architect | 1 |
| Vendor Security & Standards Manager | 1 |
| Financial Reporting Lead | 1 |
| Communications Manager | 1 |
| Physical Security Systems Technician | 1 |
| Senior Applied ML Engineer | 1 |
| Mechanical Engineer | 1 |
| Hardware Engineer | 1 |
| Software Developer | 1 |
| Mechanical Data Center Operations Technician | 1 |
| Data Center Logistics Specialist | 1 |
| Role level | Active listings |
|---|---|
| Mid-Level | 506 |
| Senior | 95 |
| Manager | 37 |
| Executive | 7 |
| Director | 1 |
Nebius is a Nasdaq-listed company building a full-stack AI cloud platform for developers and enterprises, with a global footprint and 1,500+ engineers across hardware, software, and AI R&D.
Never miss a new HPC Engineer job in Remote - Europe
Weekly or daily digest. Unsubscribe anytime.
Similar jobs
From JobCrawls search: same role title and primary location as this listing (this job excluded). Up to 8 results.
