
Site Reliability Engineer
Hover or tap a row for full statistics (EUR / month on this chart).
Salary analysis
Compared with the selected benchmark ("All roles in Remote - Europe"), this listing's salary midpoint is about 88% lower. The offer sits below the benchmark range (€1,692–€11,759). The listed pay band (€639–€1,026) is tighter than the benchmark, which suggests lower salary variability. This benchmark is based on 22 comparable listings.
| Market | Lower bound (25th percentile) | Median | Upper bound (75th percentile) |
|---|---|---|---|
| Market Average: Site Reliability Engineer | €6,006/per month | €7,294/per month | €11,397/per month |
| All roles in Remote - Europe | €1,692/per month | €4,842/per month | €11,759/per month |
| Pay in our data — not quoted in ad (Mid-Level) | €639/per month | €832/per month | €1,026/per month |
Site Reliability Engineer in Network Infrastructure Amsterdam, Netherlands; Remote - Europe About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The Role We’re looking for a Site Reliability Engineer to help build and run the fundamental part of Nebius - the Network - the infrastructure everything else depends on. This is an engineering-first SRE role: you’ll set clear reliability targets, build the tooling and automation to meet them, and make the network safer to operate as we scale quickly. Your responsibilities will include: - Define and own reliability goals for network services and critical paths (SLIs/SLOs, availability targets, error budgets where it makes sense) - Drive reliability improvements across the whole network: not only services, but also site readiness, inter-site connectivity (DCI), and operational standards - Own incident response for your areas, lead investigations/postmortems, and turn failures into durable fixes (not repeated firefighting) - Build and evolve observability: actionable metrics/logs/traces, alerting, and faster debug loops during and after incidents - Design safer change workflows: automation, CI/CD, test/staging environments, canarying, rollbacks, and auditability for network changes - Work closely with network engineers and platform teams to embed operability into designs and keep operations practical and fast We expect you to have: - Strong production Linux fundamentals and a structured approach to debugging complex systems - Solid understanding of networking basics and how real networks fail (control plane vs data plane, latency/loss, failure domains, etc.) - Hands-on experience operating high-availability systems and improving them over time (not just “keeping lights on”) - Ability to write and maintain software/automation (Go is common for us; Python is also welcome) - Experience with modern infrastructure tooling (e.g., IaC, CI/CD, container platforms) and comfort automating operational workflows It will be an added bonus if you have: - Experience with high-throughput traffic processing: load balancers, tunneling/decap, NAT64, or similar datapath-heavy systems - Low-level networking performance/debug background (eBPF/XDP, DPDK, perf/ftrace, kernel networking internals) - Experience building network-safe delivery pipelines (testing labs, staged rollouts, automated verification, drift detection) - Background with large-scale network observability/telemetry (e.g., routing/flow telemetry, regression detection at scale)
Job Details
Responsibilities
- Define and own reliability goals for network services (SLIs/SLOs, availability targets)
- Drive reliability improvements for site readiness and inter-site connectivity (DCI)
- Lead incident response, investigations, and postmortems
- Build and evolve observability using metrics, logs, and traces
- Design safer change workflows including automation, CI/CD, and canarying
- Collaborate with network engineers and platform teams to embed operability into designs
Requirements
- Strong production Linux fundamentals
- Structured approach to debugging complex systems
- Solid understanding of networking basics (control plane vs data plane, latency/loss, failure domains)
- Hands-on experience operating high-availability systems
- Ability to write and maintain software/automation in Go or Python
- Experience with modern infrastructure tooling (IaC, CI/CD, container platforms)
Skills & Technologies

Related Opportunities
Discover more opportunities that match your interests and skills