
Site Reliability Engineer
Hover or tap a row for full statistics (EUR / month on this chart).
Salary analysis
Compared with the selected benchmark ("In Remote - Europe: Site Reliability Engineer"), this listing's salary midpoint is about 91% lower. The offer sits below the benchmark range (€3,856–€8,670). The listed pay band (€480–€658) is tighter than the benchmark, which suggests lower salary variability. This benchmark is based on 1 comparable listings.
| Market | Lower bound (25th percentile) | Median | Upper bound (75th percentile) |
|---|---|---|---|
| Market Average: Site Reliability Engineer | €5,022/per month | €7,751/per month | €13,799/per month |
| In Remote - Europe: Site Reliability Engineer | €3,856/per month | €6,263/per month | €8,670/per month |
In this role you will own the reliability, performance, and observability of the entire inference stack. Your day starts with designing and refining telemetry pipelines — metrics, logs, and traces that turn hundreds of terabytes of signal into clear, actionable insight. From there you might tune Kubernetes autoscalers to squeeze more efficiency out of GPUs, craft Terraform modules that bake resilience into every new cluster, or harden our request-routing and retry logic so even transient failures go unnoticed by users. When incidents do arise, you’ll rely on the automation and runbooks you helped create to detect, isolate, and remediate problems in minutes, then drive the post-mortem culture that prevents recurrence. All of this effort points toward a single goal: scaling the platform smoothly while hitting aggressive cost and reliability targets. Success in the role calls for deep fluency with Kubernetes, Prometheus, Grafana, Terraform, and the craft of infrastructure-as-code. You script comfortably in Python or Bash, understand the nuances of alert design and SLOs for high-throughput APIs, and have spent enough time in production to know how distributed back-ends fail in the real world. Experience shepherding GPU-heavy workloads — whether with vLLM, Triton, Ray, or another accelerator stack — will serve you well, as will a background in MLOps or model-hosting platforms. Above all, you care about building self-healing systems, thrive on debugging performance from kernel to application layer, and enjoy collaborating with software engineers to turn reliability into a feature users never have to think about. If the idea of safeguarding the infrastructure that powers tomorrow’s multimodal AI energizes you, we’d love to hear your story.
Job Details
Responsibilities
- Own reliability, performance, and observability of the inference stack.
Requirements
- Strong Kubernetes, Prometheus, Grafana, Terraform, and IaC expertise.
Skills & Technologies

| Location | Active listings |
|---|---|
| Remote - Global | 423 |
| Remote - Europe | 119 |
| Amsterdam, Netherlands | 56 |
| London, United Kingdom | 22 |
| Remote - United States | 20 |
| Remote - Finland | 18 |
| Berlin, Germany | 15 |
| Mäntsälä, Finland | 12 |
| Helsinki, Finland | 10 |
| Lappeenranta, Finland | 9 |
| Prague, Czech Republic | 6 |
| Amsterdam | 5 |
| Israel | 5 |
| United Kingdom | 4 |
| Canada | 4 |
| Remote | 3 |
| Tel Aviv, Israel | 3 |
| Singapore | 3 |
| Remote - United Kingdom | 2 |
| Abu Dhabi | 2 |
| Dubai | 2 |
| Paris, France | 2 |
| Remote - France | 2 |
| New York City, United States | 2 |
| Austin, United States | 2 |
| France, Paris | 2 |
| Remote - Germany | 2 |
| London | 2 |
| Remote - Netherlands | 2 |
| Philadelphia, United States | 1 |
| Béthune, France | 1 |
| Abu Dhabi, Dubai | 1 |
| Singapore, Singapore | 1 |
| California, United States | 1 |
| Abu Dhabi, United Arab Emirates | 1 |
| Remote - EU | 1 |
| Munich, Germany | 1 |
| Minnesota, United States | 1 |
| Alabama, US | 1 |
| Prague, Czechia | 1 |
| East London, United Kingdom | 1 |
| Canada, Remote - United States | 1 |
| Berlin | 1 |
| Dallas, United States | 1 |
| London, UK | 1 |
| Oklahoma, United States | 1 |
| New Jersey, US | 1 |
| Austin, Texas | 1 |
| Kansas City, United States | 1 |
| San Francisco Bay Area, United States | 1 |
| Czechia | 1 |
| Finland | 1 |
| Remote - Sweden | 1 |
| UK | 1 |
| Prague | 1 |
| New Jersey, United States | 1 |
| Remote - Czech Republic | 1 |
| Netherlands | 1 |
| Paris | 1 |
| Béthune, Pas-de-Calais, France | 1 |
| Role type | Active listings |
|---|---|
| Backend Engineer | 331 |
| Software Engineer | 82 |
| Account Executive | 58 |
| Technical Project Manager | 6 |
| Site Reliability Engineer | 4 |
| Technical Product Manager | 4 |
| System Engineer | 4 |
| Sales Representative | 4 |
| Product Manager | 3 |
| Data Center Operations Technician | 3 |
| Data Center Technician | 3 |
| ML Engineer | 3 |
| Technical Program Manager | 3 |
| Delivery Manager | 2 |
| Applied AI Researcher | 2 |
| Hypervisor Engineer | 2 |
| Product Designer | 2 |
| IT Technician | 2 |
| Backend engineers, Frontend engineers, Site reliability engineers | 2 |
| Open Positions at Nebius | 2 |
| VP of Strategic Sales | 1 |
| Generalist | 1 |
| Offensive Security Lead | 1 |
| Head of Channel Marketing | 1 |
| Principal | 1 |
| Applied AI Solutions Engineer | 1 |
| Field Technical Lead | 1 |
| Compensation Analyst | 1 |
| Solutions Architecture Leader | 1 |
| Application Security Engineer | 1 |
| Human Resources Specialist | 1 |
| Solutions Architect | 1 |
| Mechanical Data Center Technician | 1 |
| Backend Developer | 1 |
| Senior Research Scientist | 1 |
| Cloud Solution Architect | 1 |
| Data Center IT Technician | 1 |
| Manager, ML Solutions Architecture | 1 |
| ML Solutions Architect | 1 |
| Network Planning Project Manager | 1 |
| Data Center Logistics Specialist | 1 |
| Technical Due Diligence Manager | 1 |
| Data Center Operations Manager | 1 |
| Senior Site Reliability Engineer | 1 |
| IT Support Manager | 1 |
| Technical Support Engineer | 1 |
| Data Engineer | 1 |
| Mechanical Engineer | 1 |
| Data Center IT Manager | 1 |
| Security Product Manager | 1 |
| Data Center Electrical Lead | 1 |
| GTM Recruiting Manager | 1 |
| Security Solutions Engineer | 1 |
| Physical Security Systems Technician | 1 |
| Senior HPC Engineer | 1 |
| Educational Content Author | 1 |
| Customer Engineer | 1 |
| Pricing Director | 1 |
| MEP Engineer | 1 |
| Instructional Designer | 1 |
| Partner Solutions Architect | 1 |
| Senior Software Developer | 1 |
| Forward Deployment Engineer | 1 |
| Financial Controller | 1 |
| Internal Control Business Partner | 1 |
| Data Center Facilities Manager | 1 |
| Group Product Manager | 1 |
| Structured Cabling Design Engineer | 1 |
| Site Selection & Colocation Manager | 1 |
| Product Growth Analytics Lead | 1 |
| IT Risk and Control Manager | 1 |
| Vulnerability Operations Center Lead | 1 |
| Electrical Engineer | 1 |
| Senior System Engineer | 1 |
| Machine Learning Engineer | 1 |
| ML Infrastructure Engineer | 1 |
| Data Scientist | 1 |
| Mechanical Design Engineer | 1 |
| Applied ML Engineer | 1 |
| Deal Initiation and Activation Manager | 1 |
| Operations Specialist | 1 |
| AI/ML Specialist Solutions Architect | 1 |
| HPC Engineer | 1 |
| Accountant | 1 |
| VP of Developer Relations & Community | 1 |
| Solutions Partner | 1 |
| Senior Support Engineer | 1 |
| Network Engineer | 1 |
| Project Development Manager | 1 |
| Backend Engineers | 1 |
| Role level | Active listings |
|---|---|
| Mid-Level | 410 |
| Senior | 63 |
| Manager | 12 |
| Executive | 2 |
| Director | 1 |
Related Opportunities
Discover more opportunities that match your interests and skills