
Senior Machine Learning Engineer - Nebius B.V. - Palo Alto, United States
Machine Learning Engineer
Tap this card for salary charts and full compensation details.
Expand to unlock full salary context
See benchmark placement, pay-band comparison graph, and localized salary narrative.
Job Description
Nebius Token Factory is building an AI training and model post-training capability for frontier model improvement. This role owns the infrastructure that makes large-scale training and RL experiments possible, reliable, reproducible, and efficient. The work sits at the intersection of distributed systems, GPU performance, model training frameworks, RL pipelines, and production engineering. A Senior MLE owns substantial model and endpoint optimization projects end to end. They are deeply hands-on, can debug difficult serving problems independently, and can deliver measurable improvements without needing heavy supervision. Your responsibilities: - Own optimization work for specific model families, customer endpoints, or serving backends. - Run engine comparisons and recommend practical serving configurations for specific workloads. - Debug model quality or performance regressions during production rollouts. - Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, and cost per token. - Deploy, configure, benchmark, and extend inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, or similar systems. - Build and productionize model-compression workflows, including quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery. - Implement or integrate speculative decoding, draft-model approaches, KV-cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving. - Build reproducible benchmark harnesses for TTFT, TPOT, tokens per second per GPU, p95/p99 latency, GPU memory, reliability, and cost per token. - Partner with GPU kernel engineers and platform engineers to diagnose bottlenecks across model code, kernels, runtime, scheduler, gateway, and cluster layers. - Write clear design docs, performance reports, rollout plans, and customer-facing technical explanations. Must-haves: - Strong Python and PyTorch engineering skills. - Hands-on experience deploying or optimizing LLM, VLM, or high-throughput transformer inference systems. - Practical knowledge of at least one modern inference stack such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, or equivalent internal systems. - Strong understanding of transformer inference bottlenecks, including KV cache, attention, memory bandwidth, batching, parallelism, and long-context serving. - Ability to reason quantitatively about latency, throughput, quality, utilization, and cost tradeoffs. - Strong communication skills and ability to collaborate with research, kernel, infrastructure, product, and customer teams. Nice-to-haves: - Experience with quantization-aware training, post-training quantization, FP8, INT8, INT4, NVFP4, MXFP4, AWQ, GPTQ, SmoothQuant, or related techniques. - Experience with distillation, speculative decoding, EAGLE, Medusa, multi-token prediction, or other inference acceleration methods. - Experience with agentic workloads, including tool calling, structured outputs, streaming APIs, high concurrency, and multi-step orchestration. - CUDA or Triton familiarity, even if the role is not primarily a kernel-engineering role. - Open-source contributions to vLLM, SGLang, TensorRT-LLM, FlashInfer, LMCache, PyTorch, Triton, Ray, KServe, or related projects.
Company Information
| Location | Active listings |
|---|---|
| Remote - Global | 559 |
| Remote - Europe | 57 |
| Remote - Finland | 25 |
| Remote - United States | 20 |
| Amsterdam, Netherlands | 19 |
| Berlin, Germany | 13 |
| Helsinki, Finland | 11 |
| Mäntsälä, Finland | 11 |
| London, United Kingdom | 7 |
| Amsterdam | 5 |
| Israel | 4 |
| Canada | 4 |
| Singapore | 3 |
| Remote | 3 |
| London | 2 |
| Dubai | 2 |
| Abu Dhabi | 2 |
| France, Paris | 2 |
| Tel Aviv, Israel | 1 |
| New Jersey, United States | 1 |
| Remote - DACH | 1 |
| London, UK | 1 |
| San Francisco Bay Area, United States | 1 |
| Kansas City, United States | 1 |
| Remote - France | 1 |
| Remote - Benelux | 1 |
| Prague, Czech Republic | 1 |
| Béthune, France | 1 |
| Finland | 1 |
| East London, United Kingdom | 1 |
| Minnesota, United States | 1 |
| Oklahoma, United States | 1 |
| Abu Dhabi, Dubai | 1 |
| Alabama, US | 1 |
| Prague | 1 |
| Remote - Asia | 1 |
| Netherlands | 1 |
| California, United States | 1 |
| Béthune, Pas-de-Calais, France | 1 |
| United Kingdom | 1 |
| Remote - Singapore | 1 |
| Berlin | 1 |
| Canada, Remote - United States | 1 |
| Paris, France | 1 |
| Austin, United States | 1 |
| Remote - Middle East | 1 |
| Paris | 1 |
| Dallas, United States | 1 |
| New Jersey, US | 1 |
| New York City, United States | 1 |
| Czechia | 1 |
| UK | 1 |
| Philadelphia, United States | 1 |
| Singapore, Singapore | 1 |
| Remote - North America | 1 |
| Abu Dhabi, United Arab Emirates | 1 |
| Austin, Texas | 1 |
| Role type | Active listings |
|---|---|
| Backend Engineer | 484 |
| Software Engineer | 77 |
| Account Executive | 76 |
| Sales Representative | 4 |
| Product Manager | 3 |
| Data Center Operations Technician | 2 |
| Backend engineers, Frontend engineers, Site reliability engineers | 2 |
| Open Positions at Nebius | 2 |
| Data Center Technician | 2 |
| Backend Engineers | 1 |
| Data Scientist | 1 |
| VP of Developer Relations & Community | 1 |
| Generalist | 1 |
| Human Resources Specialist | 1 |
| Accountant | 1 |
| Head of Channel Marketing | 1 |
| Operations Specialist | 1 |
| Data Engineer | 1 |
| Data Center Logistics Specialist | 1 |
| Data Center IT Technician | 1 |
| Data Center IT Manager | 1 |
| System Engineer | 1 |
| Role level | Active listings |
|---|---|
| Mid-Level | 561 |
Nebius B.V. appears in 788 indexed job postings in JobCrawls' Finland dataset since October 2023. In that historical index, the strongest location signals for this employer are Remote - Global, Remote - Europe, and Remote - Finland.
Data shown is based on historical job postings from our database.
Job Details
Responsibilities
- Own optimization work for specific model families, customer endpoints, or serving backends
- Run engine comparisons and recommend practical serving configurations for specific workloads
- Debug model quality or performance regressions during production rollouts
- Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, and cost per token
- Deploy, configure, benchmark, and extend inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo
- Build and productionize model-compression workflows including quantization, distillation, and accuracy recovery
- Implement speculative decoding, KV-cache optimization, prefix caching, and continuous batching
- Build reproducible benchmark harnesses for TTFT, TPOT, and other performance metrics
- Partner with GPU kernel and platform engineers to diagnose bottlenecks
- Write design docs, performance reports, and technical explanations
Requirements
- Strong Python and PyTorch engineering skills
- Hands-on experience deploying or optimizing LLM, VLM, or high-throughput transformer inference systems
- Practical knowledge of at least one modern inference stack (e.g., vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe)
- Strong understanding of transformer inference bottlenecks (KV cache, attention, memory bandwidth, batching, parallelism, long-context serving)
- Ability to reason quantitatively about latency, throughput, quality, utilization, and cost tradeoffs
- Strong communication skills and ability to collaborate with cross-functional teams
Skills & Technologies
