
ML Systems Engineer - Nebius B.V. - Palo Alto, United States
Tap this card for salary charts and full compensation details.
Expand to unlock full salary context
See benchmark placement, pay-band comparison graph, and localized salary narrative.
Job Description
Nebius Token Factory is building an AI training and model post-training capability for frontier model improvement. This role owns the infrastructure that makes large-scale training and RL experiments possible, reliable, reproducible, and efficient. The work sits at the intersection of distributed systems, GPU performance, model training frameworks, RL pipelines, and production engineering. A Senior ML Systems Engineer owns substantial training or RL infrastructure components end to end. They are deeply hands-on, can debug difficult distributed training failures independently, and can deliver measurable improvements in experiment throughput, stability, and GPU utilization. Your responsibilities: - Build and maintain distributed training infrastructure for SFT, continued pretraining, preference optimization, and RL workloads. - Integrate and extend frameworks such as Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, Ray, verl, slime, AReaL, OpenRLHF, or equivalent internal systems. - Implement and debug parallelism strategies including tensor, pipeline, sequence/context, expert, and data parallelism. - Build reliable rollout, reward model serving, replay/data buffer, checkpointing, evaluation, and experiment orchestration components for RL training. - Profile and improve GPU utilization, communication efficiency, memory usage, and training throughput. - Diagnose failures across NCCL, CUDA, PyTorch, Ray, schedulers, storage, networking, and checkpointing layers. - Create reproducible training runs, launch scripts, dashboards, runbooks, and operational tooling for research users. - Partner with research scientists to turn algorithmic training recipes into scalable, debuggable systems. - Write clear design docs, incident reports, benchmark reports, and operating guides. Must-haves: - Strong Python and PyTorch engineering skills. - Hands-on experience with distributed model training, large-scale ML systems, or GPU cluster workloads. - Practical understanding of transformer training bottlenecks, memory pressure, gradient/optimizer state, communication overhead, and checkpointing. - Experience debugging production or research training jobs across multiple GPUs or nodes. - Ability to reason quantitatively about throughput, utilization, memory, reliability, cost, and research velocity. - Strong communication skills and ability to collaborate with researchers, ML engineers, platform engineers, and leadership. Nice-to-haves: - Experience with Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, Ray, Slurm, Kubernetes, or large internal training platforms. - Experience with RL infrastructure frameworks such as verl, slime, AReaL, OpenRLHF, TRL, or custom PPO/GRPO/RLHF systems. - Familiarity with NCCL, CUDA, Triton, Nsight, InfiniBand, RDMA, RoCE, H100/H200/B200 clusters, or storage/network bottlenecks. - Experience supporting SFT, DPO, PPO, GRPO, RLAIF, reward model serving, rollout generation, or agent training workloads. - Open-source contributions to distributed training, RL infrastructure, PyTorch, Ray, Megatron, DeepSpeed, or related systems.
Company Information
| Location | Active listings |
|---|---|
| Remote - Global | 559 |
| Remote - Europe | 57 |
| Remote - Finland | 25 |
| Remote - United States | 20 |
| Amsterdam, Netherlands | 19 |
| Berlin, Germany | 13 |
| Mäntsälä, Finland | 11 |
| Helsinki, Finland | 11 |
| London, United Kingdom | 7 |
| Amsterdam | 5 |
| Canada | 4 |
| Israel | 4 |
| Remote | 3 |
| Singapore | 3 |
| Abu Dhabi | 2 |
| France, Paris | 2 |
| London | 2 |
| Dubai | 2 |
| Oklahoma, United States | 1 |
| California, United States | 1 |
| Finland | 1 |
| Dallas, United States | 1 |
| Singapore, Singapore | 1 |
| Abu Dhabi, Dubai | 1 |
| Czechia | 1 |
| Alabama, US | 1 |
| Berlin | 1 |
| Remote - France | 1 |
| Abu Dhabi, United Arab Emirates | 1 |
| Austin, Texas | 1 |
| Remote - Benelux | 1 |
| Remote - Middle East | 1 |
| Paris, France | 1 |
| San Francisco Bay Area, United States | 1 |
| Philadelphia, United States | 1 |
| Tel Aviv, Israel | 1 |
| Kansas City, United States | 1 |
| Béthune, Pas-de-Calais, France | 1 |
| United Kingdom | 1 |
| New Jersey, United States | 1 |
| Paris | 1 |
| Remote - Singapore | 1 |
| Austin, United States | 1 |
| Canada, Remote - United States | 1 |
| East London, United Kingdom | 1 |
| Remote - Asia | 1 |
| Remote - North America | 1 |
| London, UK | 1 |
| New Jersey, US | 1 |
| Minnesota, United States | 1 |
| New York City, United States | 1 |
| Netherlands | 1 |
| Béthune, France | 1 |
| Remote - DACH | 1 |
| Prague, Czech Republic | 1 |
| UK | 1 |
| Prague | 1 |
| Role type | Active listings |
|---|---|
| Backend Engineer | 484 |
| Software Engineer | 77 |
| Account Executive | 76 |
| Sales Representative | 4 |
| Product Manager | 3 |
| Open Positions at Nebius | 2 |
| Backend engineers, Frontend engineers, Site reliability engineers | 2 |
| Data Center Operations Technician | 2 |
| Data Center Technician | 2 |
| Data Center IT Manager | 1 |
| Data Engineer | 1 |
| Operations Specialist | 1 |
| Data Scientist | 1 |
| Data Center IT Technician | 1 |
| Generalist | 1 |
| Data Center Logistics Specialist | 1 |
| System Engineer | 1 |
| Accountant | 1 |
| Backend Engineers | 1 |
| Human Resources Specialist | 1 |
| VP of Developer Relations & Community | 1 |
| Head of Channel Marketing | 1 |
| Role level | Active listings |
|---|---|
| Mid-Level | 561 |
Nebius B.V. appears in 788 indexed job postings in JobCrawls' Finland dataset since October 2023. In that historical index, the strongest location signals for this employer are Remote - Global, Remote - Europe, and Remote - Finland.
Data shown is based on historical job postings from our database.
Job Details
Responsibilities
- Build and maintain distributed training infrastructure for SFT, continued pretraining, preference optimization, and RL workloads
- Integrate and extend frameworks such as Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, Ray, verl, slime, AReaL, OpenRLHF, or equivalent internal systems
- Implement and debug parallelism strategies including tensor, pipeline, sequence/context, expert, and data parallelism
- Build reliable rollout, reward model serving, replay/data buffer, checkpointing, evaluation, and experiment orchestration components for RL training
- Profile and improve GPU utilization, communication efficiency, memory usage, and training throughput
- Diagnose failures across NCCL, CUDA, PyTorch, Ray, schedulers, storage, networking, and checkpointing layers
- Create reproducible training runs, launch scripts, dashboards, runbooks, and operational tooling for research users
- Partner with research scientists to turn algorithmic training recipes into scalable, debuggable systems
- Write clear design docs, incident reports, benchmark reports, and operating guides
Requirements
- Strong Python and PyTorch engineering skills
- Hands-on experience with distributed model training, large-scale ML systems, or GPU cluster workloads
- Practical understanding of transformer training bottlenecks, memory pressure, gradient/optimizer state, communication overhead, and checkpointing
- Experience debugging production or research training jobs across multiple GPUs or nodes
- Ability to reason quantitatively about throughput, utilization, memory, reliability, cost, and research velocity
- Strong communication skills and ability to collaborate with researchers, ML engineers, platform engineers, and leadership
Skills & Technologies
