
Est. expiry Nov 4
Last seen Sep 30
Posted Sep 30 · 0 days ago
Open 0 days. This role usually stays open about 22 days
Senior Site Reliability Engineer
Senior SRE for Nebius Compute Node team in Amsterdam; hybrid work.
The Role We are looking for a Senior Site Reliability Engineer (SRE) to join the Compute Node team at Nebius AI Cloud. The Compute Node team is responsible for building and operating the cluster scheduler and node-level services that run and manage virtual machines across all cloud regions. This role focuses on Linux systems engineering, virtualization and operational reliability. You will work close to the operating system and hypervisor, shaping how reliability and observability are embedded into the Compute platform. Your responsibilities will include: • Ensure reliability, availability and performance of compute nodes running VMs • Analyze and debug Linux systems across user space and kernel space, understanding capabilities, limitations and trade-offs at each layer • Troubleshoot complex production issues involving CPU, memory, NUMA, cgroups and scheduling • Work hands-on with virtualization and containerization, primarily using QEMU/KVM and Linux-native technologies • Design and evolve observability as a core capability of the node layer: metrics, logs, traces, alerts, SLIs and SLOs • Lead incident response, root-cause analysis, and postmortems, driving long-term reliability improvements • Collaborate closely with platform, kernel/hypervisor, GPU and infrastructure teams to improve system design and operability We expect you to have: Strong Linux expertise: deep understanding of Linux user space and kernel space knowledge of kernel subsystems (scheduler, memory management, filesystems, cgroups, namespaces) clear understanding of system boundaries and constraints at different layers Virtualization experience: hands-on experience with QEMU/KVM udnerstanding of VM lifecycle, performance characteristics and failure modes Containerization knowledge: practical experience with containers, namespaces and cgroups strong understanding of resource isolation and control Strong debugging skills: ability to reason about complex system failures structured, hypothesis-driven approach to incident analysis SRE mindset: clear understanding of the SRE role in system design and operations experience building and operating observability stacks, not just consuming them ability to turn system behavior into actionable reliability signals Nice to Have / Optional: Experience with Kubernetes internals or node-level components Hands-on experience with low-level Linux debugging tools (e.g. perf, eBPF, ftrace, strace, kernel crash dumps) Familiarity with large-scale compute or bare-metal platforms Contributions to open-source infrastructure or system software Experience debugging hardware and driver-level issues, including GPUs, NVLink, InfiniBand Benefits & Perks: Competitive compensation Career growth and learning opportunities Flexibility and ownership Collaborative and innovative culture Opportunity to work on impactful AI projects International environment and talented teams What it’s like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.
The text above is the employer's original job description, extracted as written. Other details on this page, like salary, responsibilities, and requirements, are interpreted from that text by our system, not values the employer explicitly confirmed, so treat them as our best interpretation rather than verified facts.
Estimated from 9 comparable listings
- Ensure reliability, availability and performance of compute nodes running VMs
- Analyze and debug Linux systems across user space and kernel space
- Troubleshoot complex production issues involving CPU, memory, NUMA, cgroups and scheduling
- Work hands-on with virtualization and containerization using QEMU/KVM
- Design observability as a core capability of the node layer: metrics, logs, traces, alerts, SLIs and SLOs
- Lead incident response, root-cause analysis, and postmortems
- Collaborate with platform, kernel/hypervisor, GPU and infrastructure teams to improve system design
- Strong Linux expertise with deep understanding of Linux user space and kernel space
- Knowledge of kernel subsystems (scheduler, memory management, filesystems, cgroups, namespaces)
- Experience with QEMU/KVM virtualization
- Practical experience with containers, namespaces and cgroups
- Strong debugging skills and incident analysis
- Experience building and operating observability stacks
| Location | Active listings |
|---|---|
| Remote - Global | 398 |
| Remote - Europe | 218 |
| Amsterdam, Netherlands | 88 |
| London, United Kingdom | 31 |
| Remote - United States | 30 |
| Berlin, Germany | 23 |
| Prague, Czech Republic | 20 |
| Lappeenranta, Finland | 12 |
| Remote - Finland | 12 |
| Mäntsälä, Finland | 12 |
| United Kingdom | 10 |
| Helsinki, Finland | 9 |
| Amsterdam | 8 |
| Israel | 6 |
| London | 5 |
| Canada | 4 |
| Tel Aviv, Israel | 4 |
| Prague, Czechia | 3 |
| Canada · Remote - United States | 3 |
| Germany · Remote - Europe | 3 |
| Béthune, Pas-de-Calais | 3 |
| Singapore | 3 |
| Remote · Singapore | 3 |
| Philadelphia, Pennsylvania | 2 |
| Abu Dhabi | 2 |
| New York City, United States | 2 |
| France · Paris | 2 |
| Austin, United States | 2 |
| Oklahoma | 2 |
| Dubai | 2 |
| France, Paris | 2 |
| New Jersey | 2 |
| Israel, Israel | 2 |
| Abu Dhabi · Dubai | 2 |
| Canada, Remote - United States | 1 |
| East London, United Kingdom | 1 |
| Remote - France | 1 |
| Singapore, Singapore | 1 |
| Minnesota, United States | 1 |
| UK | 1 |
| Remote - Sweden | 1 |
| Berlin | 1 |
| New Jersey, United States | 1 |
| Remote - Germany | 1 |
| Abu Dhabi, Dubai | 1 |
| Prague | 1 |
| New Jersey, US | 1 |
| Remote - United Kingdom | 1 |
| Finland | 1 |
| Oklahoma, United States | 1 |
| California, United States | 1 |
| Abu Dhabi, United Arab Emirates | 1 |
| Paris, France | 1 |
| Czechia | 1 |
| Alabama, US | 1 |
| Béthune, France | 1 |
| Mäntsälä | 1 |
| Netherlands | 1 |
| Remote - Netherlands | 1 |
| Poland | 1 |
| Austin, Texas | 1 |
| San Francisco Bay Area, United States | 1 |
| Kansas City, United States | 1 |
| Béthune, Pas-de-Calais, France | 1 |
| Minnesota | 1 |
| Philadelphia, United States | 1 |
| Dallas, United States | 1 |
| Remote - Spain | 1 |
| London, UK | 1 |
| Belgrade, Serbia | 1 |
| Remote - EU | 1 |
| Paris | 1 |
| Role type | Active listings |
|---|---|
| Backend Engineer | 295 |
| Software Engineer | 87 |
| Account Executive | 57 |
| Sales Representative | 29 |
| Product Manager | 22 |
| Technical Program Manager | 8 |
| Technical Project Manager | 7 |
| System Engineer | 7 |
| ML Engineer | 6 |
| Site Reliability Engineer | 4 |
| Senior Site Reliability Engineer | 4 |
| Data Center Technician | 3 |
| Senior HPC Cluster Engineer | 3 |
| Senior Hypervisor Engineer | 3 |
| Technical Product Manager | 3 |
| Data Center Operations Technician | 3 |
| IT Technician | 2 |
| Customer Engineer | 2 |
| Senior ML Engineer | 2 |
| Data Scientist | 2 |
| Project Manager | 2 |
| Machine Learning Engineer | 2 |
| ML Solutions Architect | 2 |
| Delivery Manager | 2 |
| Open Positions at Nebius | 2 |
| Applied AI Researcher | 2 |
| Product Designer | 2 |
| Technical Due Diligence Manager | 2 |
| Backend engineers, Frontend engineers, Site reliability engineers | 2 |
| Security Engineer | 1 |
| IT Support Manager | 1 |
| Application Security Engineer | 1 |
| MEP Engineer | 1 |
| Datacenter IT Technician | 1 |
| Partner Solutions Architect | 1 |
| Solutions Architect Lead | 1 |
| Solutions Architecture Leader | 1 |
| Network Planning Project Manager | 1 |
| Data Center IT Manager | 1 |
| Technical Account Manager | 1 |
| Support Engineer | 1 |
| Automation Engineer | 1 |
| ML Infrastructure Engineer | 1 |
| Operations Specialist | 1 |
| Backend Developer | 1 |
| Data Center Operations Manager | 1 |
| Data Center IT Technician | 1 |
| Storage Product Manager | 1 |
| Accountant | 1 |
| Group Product Manager | 1 |
| Pricing Lead | 1 |
| GTM Recruiting Manager | 1 |
| Hypervisor Engineer | 1 |
| Technical Support Engineer | 1 |
| Data Engineer | 1 |
| Mechanical Data Center Technician | 1 |
| Detection Engineering & Response Lead | 1 |
| Instructional Designer | 1 |
| Vulnerability Operations Center Lead | 1 |
| IT Risk and Control Manager | 1 |
| VP of Strategic Sales | 1 |
| Human Resources Specialist | 1 |
| Educational Content Author | 1 |
| Offensive Security Lead | 1 |
| Field Technical Lead | 1 |
| Data Center Facilities Manager | 1 |
| Backend Engineers | 1 |
| Forward Deployment Engineer | 1 |
| Head of Channel Marketing | 1 |
| Site Selection & Colocation Manager | 1 |
| Network Engineer | 1 |
| Microsoft 365 Engineer | 1 |
| Applied AI Solutions Engineer | 1 |
| Senior Technical Product Manager | 1 |
| Electrical Engineer | 1 |
| Employee Relations Leader | 1 |
| Cloud Solution Architect | 1 |
| IT Infrastructure Engineer | 1 |
| Structured Cabling Design Engineer | 1 |
| VP of Developer Relations & Community | 1 |
| Transportation Security Manager | 1 |
| Regulatory Counsel | 1 |
| Generalist | 1 |
| Data Center Hardware Engineer | 1 |
| Data Center Electrical Lead | 1 |
| Product Growth Analytics Lead | 1 |
| Infrastructure Security Engineer | 1 |
| Partner GTM Planning and Analytics | 1 |
| Senior Network Engineer | 1 |
| Application Integration Developer | 1 |
| Principal, EMEA GTM | 1 |
| Specialist Solutions Architect | 1 |
| Vendor Security & Standards Manager | 1 |
| Financial Reporting Lead | 1 |
| Communications Manager | 1 |
| Physical Security Systems Technician | 1 |
| Mechanical Engineer | 1 |
| Hardware Engineer | 1 |
| Software Developer | 1 |
| Data Center Logistics Specialist | 1 |
| Role level | Active listings |
|---|---|
| Mid-Level | 516 |
| Senior | 106 |
| Manager | 33 |
| Executive | 6 |
| Director | 1 |
Never miss a new Site Reliability Engineer job in Remote - Europe
Weekly or daily digest. Unsubscribe anytime.
Similar jobs
From JobCrawls search: same role title and primary location as this listing (this job excluded). Up to 8 results.


