Nebius B.V. logo
Est. Monthly
Estimated €5,297 - €7,576
Posted August 26, 2026 · 1 day agoLast seen August 26, 2026Est. expiry September 30, 2026

Site Reliability Engineer

Remote - Europe
Remote · Technology
Full-time · Senior
English
No People Management
5 years experience
How this salary compares
Salary Context: Site Reliability Engineer

Hover or tap a row for full statistics (EUR / month on this chart).

Salary analysis

Compared with the selected benchmark ("In Remote - Europe: Site Reliability Engineer"), this listing's salary midpoint is about 91% lower. The offer sits below the benchmark range (€3,856–€8,670). The listed pay band (€441–€631) is tighter than the benchmark, which suggests lower salary variability. This benchmark is based on 1 comparable listings.

Monthly salary comparison for Site Reliability Engineer
MarketLower bound (25th percentile)MedianUpper bound (75th percentile)
Market Average: Site Reliability Engineer€4,828/per month€7,440/per month€13,569/per month
In Remote - Europe: Site Reliability Engineer€3,856/per month€6,263/per month€8,670/per month
This job's pay range — not quoted in ad (Senior)€441/per month€536/per month€631/per month
About the role

The Role We are looking for a Senior Site Reliability Engineer (SRE) to join the Compute Node team at Nebius AI Cloud. The Compute Node team is responsible for building and operating the cluster scheduler and node-level services that run and manage virtual machines across all cloud regions. This role focuses on Linux systems engineering, virtualization and operational reliability. You will work close to the operating system and hypervisor, shaping how reliability and observability are embedded into the Compute platform. Your responsibilities will include: Ensure reliability, availability and performance of compute nodes running VMs; Analyze and debug Linux systems across user space and kernel space, understanding capabilities, limitations and trade-offs at each layer; Troubleshoot complex production issues involving CPU, memory, NUMA, cgroups and scheduling; Work hands-on with virtualization and containerization, primarily using QEMU/KVM and Linux-native technologies; Design and evolve observability as a core capability of the node layer: metrics, logs, traces, alerts, SLIs and SLOs; Lead incident response, root-cause analysis, and postmortems, driving long-term reliability improvements; Collaborate closely with platform, kernel/hypervisor, GPU and infrastructure teams to improve system design and operability. We expect you to have: Strong Linux expertise: deep understanding of Linux user space and kernel space; knowledge of kernel subsystems (scheduler, memory management, filesystems, cgroups, namespaces); clear understanding of system boundaries and constraints at different layers; Virtualization experience: hands-on experience with QEMU/KVM; understanding of VM lifecycle, performance characteristics and failure modes; Containerization knowledge: practical experience with containers, namespaces and cgroups; strong understanding of resource isolation and control; Strong debugging skills: ability to reason about complex system failures; structured, hypothesis-driven approach to incident analysis; SRE mindset: clear understanding of the SRE role in system design and operations; experience building and operating observability stacks, not just consuming them; ability to turn system behavior into actionable reliability signals; Nice to Have / Optional: Experience with Kubernetes internals or node-level components; Hands-on experience with low-level Linux debugging tools (e.g. perf, eBPF, ftrace, strace, kernel crash dumps); Familiarity with large-scale compute or bare-metal platforms; Contributions to open-source infrastructure or system software; Experience debugging hardware and driver-level issues, including GPUs, NVLink, InfiniBand; Benefits & Perks: Competitive compensation; Career growth and learning opportunities; Flexibility and ownership; Collaborative and innovative culture; Opportunity to work on impactful AI projects; International environment and talented teams; What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI; Equal Opportunity Statement: Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law. Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.

Job Details

Responsibilities

  • Ensure reliability, availability and performance of compute nodes running VMs
  • Analyze and debug Linux systems across user space and kernel space
  • Troubleshoot complex production issues involving CPU, memory, NUMA, cgroups and scheduling
  • Work hands-on with virtualization and containerization, primarily using QEMU/KVM and Linux-native technologies
  • Design and evolve observability as a core capability of the node layer: metrics, logs, traces, alerts, SLIs and SLOs
  • Lead incident response, root-cause analysis, and postmortems, driving long-term reliability improvements
  • Collaborate closely with platform, kernel/hypervisor, GPU and infrastructure teams to improve system design and operability

Skills & Technologies

LinuxQEMU/KVMcontainerscgroupsnamespacesobservabilitySREKubernetes (optional)eBPF (nice to have)perf (nice to have)
Seen 1 day agoOpen ApplicationPartial Schema
Nebius B.V. logo
Nebius B.V. · 790 open roles
Top locations: Remote - Global · 423 · Remote - Europe · 119 · Amsterdam, Netherlands · 56+57 other locations
View company
Most-hired roles
Backend Engineer
331
Software Engineer
82
Account Executive
58
Technical Project Manager
6
Technical Product Manager
4
Role-level mix
Mid-Level (410)Senior (63)Manager (12)Executive (2)Director (1)

Help us improve JobCrawls — sign in to sync saved jobs across devices, or send feedback anytime.