Nebius B.V. logo

Senior Machine Learning Engineer - Nebius B.V. - Palo Alto, United States

Machine Learning Engineer

Posted: July 24, 2026
Posted today
Last seen in crawl: July 23, 2026 (today)
Estimated Expiry: August 28, 2026
Role & Management
Role Level:Senior
Management Tier:No People Management
Job Type
Experience
5 years

Job Description

Nebius Token Factory is building an AI training and model post-training capability for frontier model improvement. This role owns the infrastructure that makes large-scale training and RL experiments possible, reliable, reproducible, and efficient. The work sits at the intersection of distributed systems, GPU performance, model training frameworks, RL pipelines, and production engineering. A Senior MLE owns substantial model and endpoint optimization projects end to end. They are deeply hands-on, can debug difficult serving problems independently, and can deliver measurable improvements without needing heavy supervision. Your responsibilities: - Own optimization work for specific model families, customer endpoints, or serving backends. - Run engine comparisons and recommend practical serving configurations for specific workloads. - Debug model quality or performance regressions during production rollouts. - Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, and cost per token. - Deploy, configure, benchmark, and extend inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, or similar systems. - Build and productionize model-compression workflows, including quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery. - Implement or integrate speculative decoding, draft-model approaches, KV-cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving. - Build reproducible benchmark harnesses for TTFT, TPOT, tokens per second per GPU, p95/p99 latency, GPU memory, reliability, and cost per token. - Partner with GPU kernel engineers and platform engineers to diagnose bottlenecks across model code, kernels, runtime, scheduler, gateway, and cluster layers. - Write clear design docs, performance reports, rollout plans, and customer-facing technical explanations. Must-haves: - Strong Python and PyTorch engineering skills. - Hands-on experience deploying or optimizing LLM, VLM, or high-throughput transformer inference systems. - Practical knowledge of at least one modern inference stack such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, or equivalent internal systems. - Strong understanding of transformer inference bottlenecks, including KV cache, attention, memory bandwidth, batching, parallelism, and long-context serving. - Ability to reason quantitatively about latency, throughput, quality, utilization, and cost tradeoffs. - Strong communication skills and ability to collaborate with research, kernel, infrastructure, product, and customer teams. Nice-to-haves: - Experience with quantization-aware training, post-training quantization, FP8, INT8, INT4, NVFP4, MXFP4, AWQ, GPTQ, SmoothQuant, or related techniques. - Experience with distillation, speculative decoding, EAGLE, Medusa, multi-token prediction, or other inference acceleration methods. - Experience with agentic workloads, including tool calling, structured outputs, streaming APIs, high concurrency, and multi-step orchestration. - CUDA or Triton familiarity, even if the role is not primarily a kernel-engineering role. - Open-source contributions to vLLM, SGLang, TensorRT-LLM, FlashInfer, LMCache, PyTorch, Triton, Ray, KServe, or related projects.

Company Information

Nebius B.V. logo
Technology
Headcount: 1,500
Current open roles at Nebius B.V. on JobCrawls
LocationActive listings
Remote - Global559
Remote - Europe57
Remote - Finland25
Remote - United States20
Amsterdam, Netherlands19
Berlin, Germany13
Helsinki, Finland11
Mäntsälä, Finland11
London, United Kingdom7
Amsterdam5
Israel4
Canada4
Singapore3
Remote3
London2
Dubai2
Abu Dhabi2
France, Paris2
Tel Aviv, Israel1
New Jersey, United States1
Remote - DACH1
London, UK1
San Francisco Bay Area, United States1
Kansas City, United States1
Remote - France1
Remote - Benelux1
Prague, Czech Republic1
Béthune, France1
Finland1
East London, United Kingdom1
Minnesota, United States1
Oklahoma, United States1
Abu Dhabi, Dubai1
Alabama, US1
Prague1
Remote - Asia1
Netherlands1
California, United States1
Béthune, Pas-de-Calais, France1
United Kingdom1
Remote - Singapore1
Berlin1
Canada, Remote - United States1
Paris, France1
Austin, United States1
Remote - Middle East1
Paris1
Dallas, United States1
New Jersey, US1
New York City, United States1
Czechia1
UK1
Philadelphia, United States1
Singapore, Singapore1
Remote - North America1
Abu Dhabi, United Arab Emirates1
Austin, Texas1
Current role mix at Nebius B.V. on JobCrawls
Role typeActive listings
Backend Engineer484
Software Engineer77
Account Executive76
Sales Representative4
Product Manager3
Data Center Operations Technician2
Backend engineers, Frontend engineers, Site reliability engineers2
Open Positions at Nebius2
Data Center Technician2
Backend Engineers1
Data Scientist1
VP of Developer Relations & Community1
Generalist1
Human Resources Specialist1
Accountant1
Head of Channel Marketing1
Operations Specialist1
Data Engineer1
Data Center Logistics Specialist1
Data Center IT Technician1
Data Center IT Manager1
System Engineer1
Current role-level mix at Nebius B.V. on JobCrawls
Role levelActive listings
Mid-Level561

Nebius B.V. appears in 788 indexed job postings in JobCrawls' Finland dataset since October 2023. In that historical index, the strongest location signals for this employer are Remote - Global, Remote - Europe, and Remote - Finland.

Data shown is based on historical job postings from our database.

Job Details

Responsibilities

  • Own optimization work for specific model families, customer endpoints, or serving backends
  • Run engine comparisons and recommend practical serving configurations for specific workloads
  • Debug model quality or performance regressions during production rollouts
  • Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, and cost per token
  • Deploy, configure, benchmark, and extend inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo
  • Build and productionize model-compression workflows including quantization, distillation, and accuracy recovery
  • Implement speculative decoding, KV-cache optimization, prefix caching, and continuous batching
  • Build reproducible benchmark harnesses for TTFT, TPOT, and other performance metrics
  • Partner with GPU kernel and platform engineers to diagnose bottlenecks
  • Write design docs, performance reports, and technical explanations

Requirements

  • Strong Python and PyTorch engineering skills
  • Hands-on experience deploying or optimizing LLM, VLM, or high-throughput transformer inference systems
  • Practical knowledge of at least one modern inference stack (e.g., vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe)
  • Strong understanding of transformer inference bottlenecks (KV cache, attention, memory bandwidth, batching, parallelism, long-context serving)
  • Ability to reason quantitatively about latency, throughput, quality, utilization, and cost tradeoffs
  • Strong communication skills and ability to collaborate with cross-functional teams

Skills & Technologies

PythonPyTorchvLLMSGLangTensorRT-LLMTriton Inference ServerNVIDIA DynamoRay ServeKServeCUDATritonQuantizationDistillationSpeculative Decoding
19 hours agoContent Complete

Help us improve JobCrawls — sign in to sync saved jobs across devices, or send feedback anytime.