Nebius B.V.

Est. expiry Nov 4

Last seen Sep 30

Posted Sep 30 · 0 days ago

Open 0 days. This role usually stays open about 22 days

Senior Site Reliability Engineer

Site Reliability Engineer
Remote - Europe
HybridEuropeFull-time, Senior
Estimated €6,849 - €8,646
At a glance

Senior SRE for Nebius Compute Node team in Amsterdam; hybrid work.

Original job description

The Role We are looking for a Senior Site Reliability Engineer (SRE) to join the Compute Node team at Nebius AI Cloud. The Compute Node team is responsible for building and operating the cluster scheduler and node-level services that run and manage virtual machines across all cloud regions. This role focuses on Linux systems engineering, virtualization and operational reliability. You will work close to the operating system and hypervisor, shaping how reliability and observability are embedded into the Compute platform. Your responsibilities will include: • Ensure reliability, availability and performance of compute nodes running VMs • Analyze and debug Linux systems across user space and kernel space, understanding capabilities, limitations and trade-offs at each layer • Troubleshoot complex production issues involving CPU, memory, NUMA, cgroups and scheduling • Work hands-on with virtualization and containerization, primarily using QEMU/KVM and Linux-native technologies • Design and evolve observability as a core capability of the node layer: metrics, logs, traces, alerts, SLIs and SLOs • Lead incident response, root-cause analysis, and postmortems, driving long-term reliability improvements • Collaborate closely with platform, kernel/hypervisor, GPU and infrastructure teams to improve system design and operability We expect you to have: Strong Linux expertise: deep understanding of Linux user space and kernel space knowledge of kernel subsystems (scheduler, memory management, filesystems, cgroups, namespaces) clear understanding of system boundaries and constraints at different layers Virtualization experience: hands-on experience with QEMU/KVM udnerstanding of VM lifecycle, performance characteristics and failure modes Containerization knowledge: practical experience with containers, namespaces and cgroups strong understanding of resource isolation and control Strong debugging skills: ability to reason about complex system failures structured, hypothesis-driven approach to incident analysis SRE mindset: clear understanding of the SRE role in system design and operations experience building and operating observability stacks, not just consuming them ability to turn system behavior into actionable reliability signals Nice to Have / Optional: Experience with Kubernetes internals or node-level components Hands-on experience with low-level Linux debugging tools (e.g. perf, eBPF, ftrace, strace, kernel crash dumps) Familiarity with large-scale compute or bare-metal platforms Contributions to open-source infrastructure or system software Experience debugging hardware and driver-level issues, including GPUs, NVLink, InfiniBand Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What it’s like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.

The text above is the employer's original job description, extracted as written. Other details on this page, like salary, responsibilities, and requirements, are interpreted from that text by our system, not values the employer explicitly confirmed, so treat them as our best interpretation rather than verified facts.

Skills & Technologies
LinuxQEMU/KVMContainersCgroupsNamespacesObservabilitySRE PracticesIncident Response
How this salary compares
Estimated €6,849 - €8,646Est. take-home €4,335 - €5,214

Estimated from 9 comparable listings

16% below Senior level roles
€4,446€7,494 median€14,104
Responsibilities
  • Ensure reliability, availability and performance of compute nodes running VMs
  • Analyze and debug Linux systems across user space and kernel space
  • Troubleshoot complex production issues involving CPU, memory, NUMA, cgroups and scheduling
  • Work hands-on with virtualization and containerization using QEMU/KVM
  • Design observability as a core capability of the node layer: metrics, logs, traces, alerts, SLIs and SLOs
  • Lead incident response, root-cause analysis, and postmortems
  • Collaborate with platform, kernel/hypervisor, GPU and infrastructure teams to improve system design
Requirements
  • Strong Linux expertise with deep understanding of Linux user space and kernel space
  • Knowledge of kernel subsystems (scheduler, memory management, filesystems, cgroups, namespaces)
  • Experience with QEMU/KVM virtualization
  • Practical experience with containers, namespaces and cgroups
  • Strong debugging skills and incident analysis
  • Experience building and operating observability stacks
Nebius B.V. logo
Nebius B.V. · 121 open roles
Top locations (all time): Remote - Global · 398 · Remote - Europe · 218 · Amsterdam, Netherlands · 88 · +69 other locations
View company
Current open roles at Nebius B.V. on JobCrawls
LocationActive listings
Remote - Global398
Remote - Europe218
Amsterdam, Netherlands88
London, United Kingdom31
Remote - United States30
Berlin, Germany23
Prague, Czech Republic20
Lappeenranta, Finland12
Remote - Finland12
Mäntsälä, Finland12
United Kingdom10
Helsinki, Finland9
Amsterdam8
Israel6
London5
Canada4
Tel Aviv, Israel4
Prague, Czechia3
Canada · Remote - United States3
Germany · Remote - Europe3
Béthune, Pas-de-Calais3
Singapore3
Remote · Singapore3
Philadelphia, Pennsylvania2
Abu Dhabi2
New York City, United States2
France · Paris2
Austin, United States2
Oklahoma2
Dubai2
France, Paris2
New Jersey2
Israel, Israel2
Abu Dhabi · Dubai2
Canada, Remote - United States1
East London, United Kingdom1
Remote - France1
Singapore, Singapore1
Minnesota, United States1
UK1
Remote - Sweden1
Berlin1
New Jersey, United States1
Remote - Germany1
Abu Dhabi, Dubai1
Prague1
New Jersey, US1
Remote - United Kingdom1
Finland1
Oklahoma, United States1
California, United States1
Abu Dhabi, United Arab Emirates1
Paris, France1
Czechia1
Alabama, US1
Béthune, France1
Mäntsälä1
Netherlands1
Remote - Netherlands1
Poland1
Austin, Texas1
San Francisco Bay Area, United States1
Kansas City, United States1
Béthune, Pas-de-Calais, France1
Minnesota1
Philadelphia, United States1
Dallas, United States1
Remote - Spain1
London, UK1
Belgrade, Serbia1
Remote - EU1
Paris1
Current role mix at Nebius B.V. on JobCrawls
Role typeActive listings
Backend Engineer295
Software Engineer87
Account Executive57
Sales Representative29
Product Manager22
Technical Program Manager8
Technical Project Manager7
System Engineer7
ML Engineer6
Site Reliability Engineer4
Senior Site Reliability Engineer4
Data Center Technician3
Senior HPC Cluster Engineer3
Senior Hypervisor Engineer3
Technical Product Manager3
Data Center Operations Technician3
IT Technician2
Customer Engineer2
Senior ML Engineer2
Data Scientist2
Project Manager2
Machine Learning Engineer2
ML Solutions Architect2
Delivery Manager2
Open Positions at Nebius2
Applied AI Researcher2
Product Designer2
Technical Due Diligence Manager2
Backend engineers, Frontend engineers, Site reliability engineers2
Security Engineer1
IT Support Manager1
Application Security Engineer1
MEP Engineer1
Datacenter IT Technician1
Partner Solutions Architect1
Solutions Architect Lead1
Solutions Architecture Leader1
Network Planning Project Manager1
Data Center IT Manager1
Technical Account Manager1
Support Engineer1
Automation Engineer1
ML Infrastructure Engineer1
Operations Specialist1
Backend Developer1
Data Center Operations Manager1
Data Center IT Technician1
Storage Product Manager1
Accountant1
Group Product Manager1
Pricing Lead1
GTM Recruiting Manager1
Hypervisor Engineer1
Technical Support Engineer1
Data Engineer1
Mechanical Data Center Technician1
Detection Engineering & Response Lead1
Instructional Designer1
Vulnerability Operations Center Lead1
IT Risk and Control Manager1
VP of Strategic Sales1
Human Resources Specialist1
Educational Content Author1
Offensive Security Lead1
Field Technical Lead1
Data Center Facilities Manager1
Backend Engineers1
Forward Deployment Engineer1
Head of Channel Marketing1
Site Selection & Colocation Manager1
Network Engineer1
Microsoft 365 Engineer1
Applied AI Solutions Engineer1
Senior Technical Product Manager1
Electrical Engineer1
Employee Relations Leader1
Cloud Solution Architect1
IT Infrastructure Engineer1
Structured Cabling Design Engineer1
VP of Developer Relations & Community1
Transportation Security Manager1
Regulatory Counsel1
Generalist1
Data Center Hardware Engineer1
Data Center Electrical Lead1
Product Growth Analytics Lead1
Infrastructure Security Engineer1
Partner GTM Planning and Analytics1
Senior Network Engineer1
Application Integration Developer1
Principal, EMEA GTM1
Specialist Solutions Architect1
Vendor Security & Standards Manager1
Financial Reporting Lead1
Communications Manager1
Physical Security Systems Technician1
Mechanical Engineer1
Hardware Engineer1
Software Developer1
Data Center Logistics Specialist1
Current role-level mix at Nebius B.V. on JobCrawls
Role levelActive listings
Mid-Level516
Senior106
Manager33
Executive6
Director1

Never miss a new Site Reliability Engineer job in Remote - Europe

Weekly or daily digest. Unsubscribe anytime.

Similar jobs

From JobCrawls search: same role title and primary location as this listing (this job excluded). Up to 8 results.