
Site Reliability Engineer (SRE)
Join Social Discovery Group as a Site Reliability Engineer (SRE). You will own production reliability, automate deployments, and improve Kubernetes environments. Requires 3+ years in SRE/DevOps and strong Linux/Kubernetes skills. Remote work worldwide with generous vacation and benefits.
Social Discovery Group (SDG) is a group of social discovery companies. SDG solves the problems of loneliness, isolation, and disconnection - transforming virtual intimacy into the new normal. SDG’s products redefine the way people interact and connect with one another. Our portfolio includes social entertainment platforms designed to connect people online across different cultures and regions of the world. We bring together a team of like-minded people and IT professionals who specialize in creating and developing globally impactful social discovery products. Our international team of digital nomads works remotely from all over the world. We’re proud to be a two-time “Great Place to Work” winner (USA & Japan, 2024–2025) and a Top-5 Company for Work-From-Anywhere Jobs (FlexJobs, 2025). We are looking for a Site Reliability Engineer (SRE) passionate about infrastructure reliability, automation, and the development of scalable production systems. Your main tasks will be: Own and improve production infrastructure reliability and stability Prepare, execute, and support deployments and infrastructure changes Build and maintain Infrastructure-as-Code solutions using Ansible and Terraform Support and optimize Kubernetes-based and containerized environments Develop automation scripts and internal operational tooling Monitor system health, investigate incidents, and proactively improve observability Participate in CI/CD improvements together with Development, QA, DevOps, and SRE teams Work with monitoring and alerting systems to reduce downtime and improve system performance Maintain technical documentation, runbooks, and operational procedures Support DNS, WAF, CDN, and caching infrastructure where required We expect from you: 3+ years of experience in SRE, DevOps, System Administration, or Build/Release Engineering Strong Linux administration and troubleshooting skills Hands-on experience with Kubernetes and containerization technologies (Docker/Podman) Experience with CI/CD pipelines, preferably GitLab CI Practical experience with Infrastructure-as-Code and configuration management tools (Ansible and/or Terraform) Experience with observability and monitoring tools such as Prometheus, Grafana, Zabbix, or VictoriaMetrics Good understanding of networking fundamentals, DNS, HTTP/HTTPS, load balancing, and troubleshooting Experience with Git and modern software delivery workflows Ability to work independently, take ownership, and proactively improve infrastructure Fluent Russian level for technical documentation and team communication Nice to have: AWS or GCP experience RabbitMQ / AMQP experience Cloudflare, Akamai, WAF, CDN experience Experience with tracing and advanced observability tooling What do we offer: REMOTE OPPORTUNITY to work full-time; Vacation 28 calendar days per year; 7 wellness days per year (time off) that can be used to deal with household issues, to lie down and recover without taking sick leave; Bonuses up to $5000 for recommending successful applicants for positions in the company; 50% payment for professional training , international conferences, and meetings; Corporate discount for English lessons ; Health benefits. According to the paychecks, if you are not eligible for corporate medical insurance, the company will compensate you with up to $1,000 gross per year per employee. This can be spent on self-purchase of health insurance or on doctor’s fees for yourself and close relatives (spouse, children); Workplace organization. The company provides all employees with an equipped workplace and all the necessary equipment (table, armchair, wifi, etc.) in our offices or co-working locations. In the other locations, the company provides reimbursement of workplace costs up to $1000 gross once every 3 years, according to the paychecks. This money can be spent on the rent of the co-working room, on equipping the working place at home (desk, chair, Internet, etc.) during those 3 years; Internal gamified gratitude system: receive bonuses from colleagues and exchange them for our merchandise, team building activities, massage certificates, etc. Sounds good? Join us now! The initial pay level or pay range for this role will be shared with candidates during the recruitment process and before the commencement of employment.
The text above is the employer's original job description, extracted as written. Other details on this page, like salary, responsibilities, and requirements, are interpreted from that text by our system, not values the employer explicitly confirmed, so treat them as our best interpretation rather than verified facts.
Estimated from 7 comparable listings
- Own and improve production infrastructure reliability and stability
- Prepare, execute, and support deployments and infrastructure changes
- Build and maintain Infrastructure-as-Code solutions using Ansible and Terraform
- Support and optimize Kubernetes-based and containerized environments
- Develop automation scripts and internal operational tooling
- Monitor system health, investigate incidents, and proactively improve observability
- Participate in CI/CD improvements with Development, QA, DevOps, and SRE teams
- Work with monitoring and alerting systems to reduce downtime and improve system performance
- Maintain technical documentation, runbooks, and operational procedures
- Support DNS, WAF, CDN, and caching infrastructure where required
- 3+ years of experience in SRE, DevOps, System Administration, or Build/Release Engineering
- Strong Linux administration and troubleshooting skills
- Hands-on experience with Kubernetes and containerization technologies (Docker/Podman)
- Experience with CI/CD pipelines (GitLab CI preferred)
- Infrastructure-as-Code and configuration management tools (Ansible and/or Terraform)
- Experience with observability and monitoring tools such as Prometheus, Grafana, Zabbix, or VictoriaMetrics
- Good understanding of networking fundamentals, DNS, HTTP/HTTPS, load balancing, and troubleshooting
- Experience with Git and modern software delivery workflows
- Ability to work independently, take ownership, and proactively improve infrastructure
- Fluent Russian for technical documentation and team communication
| Location | Active listings |
|---|---|
| Remote - Global | 20 |
| WORLDWIDE | 3 |
| Role type | Active listings |
|---|---|
| Marketing Analyst | 2 |
| External Communications Manager | 1 |
| Site Reliability Engineer | 1 |
| Personal Assistant | 1 |
| Executive Search Partner | 1 |
| Executive Search Team Lead | 1 |
| Marketing Project Manager | 1 |
| Legal Director | 1 |
| Database Administrator | 1 |
| Atlassian Platform Engineer | 1 |
| CV Engineer | 1 |
| System Administrator | 1 |
| Deputy CDO | 1 |
| Cost Optimisation Specialist | 1 |
| QA Engineer | 1 |
| Net Developer | 1 |
| Customer Success Agent | 1 |
| HR Business Partner | 1 |
| Head of Design | 1 |
| Accounting Manager | 1 |
| Software Engineer | 1 |
| Executive Personal Assistant | 1 |
| Role level | Active listings |
|---|---|
| Senior | 13 |
| Mid-Level | 4 |
| Director | 2 |
| Executive | 2 |
| Manager | 2 |
Social Discovery Group (SDG) is a group of social discovery companies focused on connecting people worldwide and reducing loneliness. We operate globally with a fully remote team and have been recognized as a Great Place to Work and a top workplace for remote roles.
Never miss a new Site Reliability Engineer job in Remote - Global
Weekly or daily digest. Unsubscribe anytime.
Similar jobs
From JobCrawls search: same role title and primary location as this listing (this job excluded). Up to 8 results.




