Bentley Systems logo
Posted August 12, 2026 · 2 days agoLast seen August 13, 2026Est. expiry September 16, 2026

Site Reliability Engineering Manager

Manager, Site Reliability Engineering
Pune, India
Hybrid · Software Engineering
Full-time · Manager
English
No People Management
Bachelor
8 years experience
About the role

Position Summary: SRE Managers operate with significant autonomy and are accountable for the operational success of an entire service area, product platform, or collection of services. They balance people leadership, technical leadership, and operational ownership, regularly influencing priorities beyond their immediate team and partnering with engineering and product leaders to improve reliability, operational efficiency, and customer outcomes. They are trusted to make decisions that affect multiple teams and are expected to build systems, processes, and organizations that scale. Responsibilities: The Manager, SRE leads one or more teams of Site Reliability Engineers and is accountable for reliability, availability, operational excellence, automation strategy, and engineering effectiveness within their assigned service area. Responsibilities span four primary domains: Reliability Leadership Own availability, performance, scalability, and operational health of assigned services. Drive adoption and enforcement of SRE practices including SLOs, SLIs, error budgets, incident management, and operational reviews. Own reliability roadmaps and ensure investment is balanced between feature delivery, technical debt reduction, and operational sustainability. Establish measurable operational targets and continuously improve system resilience. Strategize and communicate with management of Support teams on releases, automation changes, SOPs, and ticket management. Engineering & People Leadership Build a high-performing team through hiring, coaching, mentoring, and career development. Set clear expectations and performance standards aligned to organizational objectives. Develop future technical and people leaders within the team. Ensure engineering excellence through design reviews, operational reviews, and technical governance. Collaborate with management of Development, QA, Product, and Release teams to ensure SRE principles, defensive coding practices, and release readiness posture are shifted left into planning, design, development, and delivery processes. Operational Excellence Lead response and recovery efforts for critical incidents. Drive systemic problem resolution through blameless postmortems and automation. Eliminate toil through software engineering and platform capabilities. Ensure operational readiness for new products, services, and major releases. Strategic Execution & Organizational Leadership Translate organizational strategy into actionable roadmaps and quarterly objectives. Work across development, product management, cloud operations, security, and support teams to align priorities. Guide architectural decisions that improve reliability, scalability, security, and cost efficiency. Lead multiple concurrent initiatives spanning products, platforms, or service areas. Establish operating mechanisms, standards, and best practices across teams. Represent SRE during executive reviews, customer escalations, and strategic planning discussions. Participate in budget planning, workforce planning, and organizational design discussions.

Job Details

Responsibilities

  • Lead one or more teams of Site Reliability Engineers
  • Accountable for reliability, availability, operational excellence, automation strategy, and engineering effectiveness within their service area
  • Own availability, performance, scalability, and operational health of assigned services
  • Drive adoption and enforcement of SRE practices including SLOs, SLIs, error budgets, incident management, and operational reviews
  • Own reliability roadmaps and ensure balanced investment between feature delivery, technical debt reduction, and operational sustainability
  • Establish measurable operational targets and continuously improve system resilience
  • Strategize and communicate with management of Support teams on releases, automation changes, SOPs, and ticket management
  • Build a high-performing team through hiring, coaching, mentoring, and career development
  • Set expectations and performance standards aligned to objectives
  • Develop future leaders within the team
  • Ensure engineering excellence through design reviews, operational reviews, and governance
  • Collaborate with Development, QA, Product, and Release teams to ensure SRE principles
  • Shift left readiness into planning, design, development, and delivery processes
  • Lead incident response and recovery efforts for critical incidents
  • Drive blameless postmortems and automation to resolve systemic problems
  • Eliminate toil through software engineering and platform capabilities
  • Ensure operational readiness for new products, services, and major releases
  • Translate organizational strategy into actionable roadmaps and quarterly objectives
  • Work with cross-functional teams to align priorities
  • Guide architectural decisions that improve reliability, scalability, security, and cost efficiency
  • Lead multiple initiatives across products and platforms
  • Establish operating mechanisms, standards, and best practices
  • Represent SRE in executive reviews and customer escalations
  • Participate in budget and workforce planning and organizational design

Requirements

  • Bachelor’s degree in Computer Science, Engineering, Information Systems, or equivalent practical experience
  • Typically 8+ years of experience across software engineering, cloud infrastructure, platform engineering, systems administration, or reliability engineering
  • Minimum 4 years of experience with cloud platforms (Azure, AWS, or GCP)
  • Infrastructure as Code (Terraform, Bicep, or Ansible)
  • Kubernetes and containerized workloads
  • Observability tooling and telemetry systems
  • CI/CD and software delivery practices
  • Minimum 3 years of experience leading engineers through formal people-leadership, technical leadership, or team management responsibilities
  • Experience operating and supporting business-critical production services
  • Experience leading major incidents, postmortems, and reliability improvement programs

Skills & Technologies

KubernetesCloud platforms (AzureAWSGCP)Terraform/Ansible/BicepObservability tooling

Education Level

Bachelor
Seen 18 hours agoPartial Schema
Bentley Systems logo
Bentley Systems · 72 open roles
Top locations: Remote - Global · 70 · Remote - Finland · 1 · Sweden, Finland, Norway · 1
View company
Most-hired roles
Software Engineer
7
Product Manager
6
Technical Support Engineer
5
Full-Stack Software Engineer
2
Application Engineer
2
Role-level mix
Mid-Level (69)Executive (1)

Help us improve JobCrawls — sign in to sync saved jobs across devices, or send feedback anytime.