
ML Infrastructure Engineer
ML Infra Engineer at Nebius; benchmark GPU platforms for ML workloads; Amsterdam + remote Europe/US.
The role: We are seeking a highly skilled ML/AI Engineer to join our team to lead and support benchmarking of GPU platforms for machine learning and AI workloads. You will play a critical role in evaluating the performance of GPU-based hardware for various deep learning and AI frameworks, enabling data-driven decisions for platform optimisation and next-generation hardware development. Your responsibilities will include: Work closely with hardware, development teams to profile and analyse GPU performance at the system and kernel level. Evaluate and compare GPU performance across different platforms, architectures, and software stacks (e.g.,CUDA, ROCm). Debug and optimise ML workloads to run efficiently on GPU hardware, identifying and resolving performance bottlenecks. Perform acceptance testing for new GPU clusters, ensuring hardware and software meet performance, stability, and compatibility requirements for AI workloads. Perform experiments across diverse GPU system configurations to assess the impact of varying interconnect strategies and system-level optimisations on performance and scalability. Develop tools and dashboards to visualise performance metrics, bottlenecks, and trends. Contribute to internal tooling, frameworks, and best practices. We expect you to have: A profound understanding of theoretical foundations of machine learning Deep understanding of performance aspects of large neural networks training and inference (data/tensor/context/expert parallelism, offloading, custom kernels, hardware features, attention optimisations, dynamic batching etc.). Deep experience with modern deep learning frameworks (PyTorch, JAX, Megatron-LM, Tensort-LLM). Good understanding of the GPU stack: CUDA,NCCL, drivers, and relevant libraries. Familiarity with containerized environments (e.g., Docker, Kubernetes). Strong communication and ability to work independently. Ways to stand out from the crowd: Familiarity with modern LLM inference frameworks (vLLM, SGLang, TensorRT). Experience in Python and performance profiling tools (e.g., Nsight, nvprof, perf). Familiarity with cloud ML platforms like AWS, GCP, Azure ML. Contributions to open-source ML benchmarking tools.
The text above is the employer's original job description, extracted as written. Other details on this page, like salary, responsibilities, and requirements, are interpreted from that text by our system, not values the employer explicitly confirmed, so treat them as our best interpretation rather than verified facts.
- Profile GPU performance at system and kernel level
- Evaluate GPU performance across platforms
- Debug and optimise ML workloads
- Perform acceptance testing for GPU clusters
- Experiment with system configurations
- Develop tools and dashboards
- Contribute to internal tooling
- Strong knowledge of ML theory
- Deep understanding of training/inference of large neural networks
- Experience with PyTorch, JAX, Megatron-LM, Tensor-LLM
- GPU stack knowledge CUDA, NCCL, drivers
- Containerization experience Docker, Kubernetes
| Location | Active listings |
|---|---|
| Remote - Global | 398 |
| Remote - Europe | 210 |
| Amsterdam, Netherlands | 88 |
| Remote - United States | 30 |
| London, United Kingdom | 29 |
| Berlin, Germany | 23 |
| Prague, Czech Republic | 22 |
| United Kingdom | 13 |
| Lappeenranta, Finland | 12 |
| Mäntsälä, Finland | 12 |
| Helsinki, Finland | 10 |
| Israel | 10 |
| Amsterdam | 8 |
| London | 5 |
| Tel Aviv, Israel | 5 |
| Canada | 4 |
| Remote - Finland | 4 |
| Canada · Remote - United States | 3 |
| Germany · Remote - Europe | 3 |
| Béthune, Pas-de-Calais | 3 |
| Singapore | 3 |
| Remote · Singapore | 3 |
| Philadelphia, Pennsylvania | 2 |
| Abu Dhabi | 2 |
| New York City, United States | 2 |
| France · Paris | 2 |
| Austin, United States | 2 |
| Oklahoma | 2 |
| Dubai | 2 |
| Netherlands | 2 |
| France, Paris | 2 |
| New Jersey | 2 |
| Remote - EU | 2 |
| Abu Dhabi · Dubai | 2 |
| Canada, Remote - United States | 1 |
| East London, United Kingdom | 1 |
| Singapore, Singapore | 1 |
| Minnesota, United States | 1 |
| Prague, Czechia | 1 |
| UK | 1 |
| Berlin | 1 |
| New Jersey, United States | 1 |
| Abu Dhabi, Dubai | 1 |
| Prague | 1 |
| New Jersey, US | 1 |
| Finland | 1 |
| Oklahoma, United States | 1 |
| California, United States | 1 |
| Abu Dhabi, United Arab Emirates | 1 |
| Paris, France | 1 |
| Czechia | 1 |
| Alabama, US | 1 |
| Béthune, France | 1 |
| Mäntsälä | 1 |
| Poland | 1 |
| Austin, Texas | 1 |
| San Francisco Bay Area, United States | 1 |
| Kansas City, United States | 1 |
| Béthune, Pas-de-Calais, France | 1 |
| Minnesota | 1 |
| Philadelphia, United States | 1 |
| Dallas, United States | 1 |
| London, UK | 1 |
| Belgrade, Serbia | 1 |
| Paris | 1 |
| Role type | Active listings |
|---|---|
| Backend Engineer | 290 |
| Software Engineer | 86 |
| Account Executive | 57 |
| Sales Representative | 29 |
| Product Manager | 21 |
| Technical Program Manager | 8 |
| Technical Project Manager | 7 |
| ML Engineer | 7 |
| System Engineer | 7 |
| Data Center Technician | 4 |
| Site Reliability Engineer | 4 |
| Technical Product Manager | 4 |
| Senior Site Reliability Engineer | 4 |
| IT Technician | 3 |
| Senior Hypervisor Engineer | 3 |
| Data Center Operations Technician | 3 |
| Data Scientist | 2 |
| Project Manager | 2 |
| Senior HPC Cluster Engineer | 2 |
| Solutions Architect | 2 |
| ML Solutions Architect | 2 |
| Delivery Manager | 2 |
| Open Positions at Nebius | 2 |
| Applied AI Researcher | 2 |
| Product Designer | 2 |
| Backend engineers, Frontend engineers, Site reliability engineers | 2 |
| Security Engineer | 1 |
| Customer Support Specialist | 1 |
| DC IT Support Manager | 1 |
| Partner Solutions Architect | 1 |
| Customer Engineer | 1 |
| Internal Control Business Partner | 1 |
| Solutions Architect Lead | 1 |
| Solutions Architecture Leader | 1 |
| Network Planning Project Manager | 1 |
| Data Center IT Manager | 1 |
| Technical Account Manager | 1 |
| Security Solutions Engineer | 1 |
| Support Engineer | 1 |
| HPC Cluster Engineer | 1 |
| ML Infrastructure Engineer | 1 |
| Operations Specialist | 1 |
| Backend Developer | 1 |
| Data Center Operations Manager | 1 |
| Data Center IT Technician | 1 |
| Mechanical Design Engineer | 1 |
| Pricing Director | 1 |
| Accountant | 1 |
| Group Product Manager | 1 |
| Hypervisor Engineer | 1 |
| Technical Support Engineer | 1 |
| Data Engineer | 1 |
| Instructional Designer | 1 |
| AI Engineer | 1 |
| Vulnerability Operations Center Lead | 1 |
| IT Risk and Control Manager | 1 |
| VP of Strategic Sales | 1 |
| Human Resources Specialist | 1 |
| Educational Content Author | 1 |
| Detection Engineer & Response Lead | 1 |
| Field Technical Lead | 1 |
| Data Center Facilities Manager | 1 |
| Backend Engineers | 1 |
| Head of Channel Marketing | 1 |
| Site Selection & Colocation Manager | 1 |
| Microsoft 365 Engineer | 1 |
| Director, Global Systems Integrator Partnerships | 1 |
| Applied AI Solutions Engineer | 1 |
| Site Selection Manager | 1 |
| Senior Technical Product Manager | 1 |
| Electrical Engineer | 1 |
| Cloud Solution Architect | 1 |
| IT Infrastructure Engineer | 1 |
| Structured Cabling Design Engineer | 1 |
| AI/ML Specialist | 1 |
| VP of Developer Relations & Community | 1 |
| Transportation Security Manager | 1 |
| Regulatory Counsel | 1 |
| Generalist | 1 |
| Development Manager | 1 |
| Data Center Electrical Lead | 1 |
| Product Growth Analytics Lead | 1 |
| Infrastructure Security Engineer | 1 |
| Partner GTM Planning and Analytics | 1 |
| Technical Due Diligence Manager | 1 |
| GTM Lead | 1 |
| Senior Network Engineer | 1 |
| Application Integration Developer | 1 |
| Specialist Solutions Architect | 1 |
| Backend Software Engineer | 1 |
| Vendor Security & Standards Manager | 1 |
| Financial Reporting Lead | 1 |
| Communications Manager | 1 |
| Physical Security Systems Technician | 1 |
| Mechanical Engineer | 1 |
| Hardware Engineer | 1 |
| Software Developer | 1 |
| Presentation & Customer-facing Assets Designer | 1 |
| Data Center Logistics Specialist | 1 |
| Applied ML Engineer | 1 |
| Role level | Active listings |
|---|---|
| Mid-Level | 505 |
| Senior | 103 |
| Manager | 35 |
| Executive | 4 |
| Director | 1 |
| Junior | 1 |
Never miss a new ML Infrastructure Engineer job in Remote - Europe
Free weekly or daily digest. Unsubscribe anytime.
Similar jobs
From JobCrawls search: same role title and primary location as this listing (this job excluded). Up to 8 results.