
Beräknat utgångsdatum 29 okt.
Senast sedd 30 sep.
Publicerad 24 sep. · för 8 dagar sedan
Öppen i 8 dagar. Den här rollen är vanligtvis öppen i cirka 22 dagar
Senior Site Reliability Engineer
Senior SRE role at Develocity (remote-first, US-East Coast/NA).
We’re building a new SRE team and looking for founding members to help shape how we operate. You’ll be responsible for the reliability, performance, and availability of Develocity instances serving paying customers, open-source projects, and public-facing services, plus supporting infrastructure like artifact registries. You’ll work on our internally-built Cloud Application Platform, Kubernetes on AWS, and develop deep expertise in it. When incidents happen, you’ll troubleshoot issues across the stack, from application to infrastructure. You’ll collaborate with the Cloud Platform team to improve the tooling you depend on, and with engineering teams to build reliability into how we ship software. If you like automating things and hate doing the same task twice, you’ll fit in well. You’ll be part of a distributed, remote-first team that values asynchronous communication and written documentation. Strong self-direction and clear communication across time zones are essential. Responsibilities Operate and maintain all Develocity instances and supporting services. Participate in an on-call rotation, owning incident response and troubleshooting issues across the stack. Drive automation across application deployment, upgrades, monitoring, self-healing, and recovery. Build and maintain observability for all managed services (logging, metrics, tracing, and alerting). Work with engineering teams to build reliability into features from the start. Run incident response and retrospectives, and make sure we learn from them. Own disaster recovery, backups, and business continuity. Communicate with customers during incidents and maintenance windows. Optimize performance, resource usage, and costs. Help evolve our SaaS operations as we grow. Minimum qualifications 5+ years in SRE, DevOps, or equivalent role operating production services at scale. Strong Kubernetes experience in production environments. Cloud infrastructure expertise, preferably AWS (EKS, RDS, S3, EC2). Proficiency with observability tools (Prometheus, Grafana) and Infrastructure as Code (Terraform). Track record of incident management and response. Knowledge of SRE best practices (SLAs, SLOs). Scripting proficiency (Python, Bash) for automation. Experience with 24/7 on-call rotations. Strong written and verbal English communication. Preferred qualifications Experience operating SaaS platforms at scale. Familiarity with Develocity. JVM language experience (Java, Kotlin). Disaster recovery planning and execution experience. Customer-facing incident communication skills. Experience establishing SRE practices in new or growing teams. What We Offer A ground-floor role in a new SRE team—you'll shape how we do things, not inherit someone else’s decisions. Real ownership of production systems used by engineers at companies you've heard of. Direct interaction with customers when things go wrong (and when they go right). A culture that values automation over heroics. In-person meetings, such as our annual company offsite and team meetings. Work from home in a remote-first environment. Competitive salaries and equity grants. Compensation The US salary range for this position is $180,000-205,000 which reflects the target ranges for all US locations. Within this range, individual pay is determined by geographic location and additional factors including but not limited to experience, relevant skills, qualifications, seniority, performance, and travel requirements. Our recruiting team can share more information about the specific salary range for your location during the hiring process. Location Remote from anywhere in PST timezone. While our team works remotely and is spread across the globe, we deeply value daily interactions and collaboration. Perks & Benefits Competitive salary We offer all team members competitive salaries based on geographical location. Stock options We offer all team members company stock option plans. International benefits We provide benefits tailored to your country of work. Develocity World Meeting Once a year, we gather at a cool global destination for a week of in-person meetings and team-building events. Annual in-person team meeting In addition to the Develocity World Meeting, individual teams meet once a year to bond and collaborate in-person. Learning & development We offer an annual learning and development stipend and a monthly meeting-free day dedicated to learning. Home office stipend We offer a stipend to ensure you’re comfortably equipped to work remotely. Volunteer day We offer up to 8 hours of paid time each year for team members to give back to their local communities. Apply for the job
Texten ovan är arbetsgivarens ursprungliga arbetsbeskrivning, hämtad som den skrivits. Övriga uppgifter på sidan, som lön, ansvarsområden och krav, har vårt system tolkat ur samma text och är inte värden som arbetsgivaren uttryckligen har bekräftat. Betrakta dem därför som vår bästa tolkning, inte som verifierade fakta.
| Plats | Aktiva annonser |
|---|---|
| Distans - Europa | 8 |
| Rolltyp | Aktiva annonser |
|---|---|
| Site Reliability Engineer | 3 |
| FöretagskundsAccount Executive | 1 |
| Site reliability engineer | 1 |
| Försäljningsförhandsingenjör | 1 |
| Backend-ingenjör | 1 |
| Account Executive | 1 |
| Rollnivå | Aktiva annonser |
|---|---|
| Senior | 7 |
| Medelnivå | 1 |
Missa aldrig ett nytt Site Reliability Engineer-jobb i Distans - North America
Vecko- eller daglig sammanfattning. Avsluta när som helst.
Liknande jobb
Från JobCrawls sökning: samma rolltitel och primära plats som denna annons (denna tjänst exkluderad). Upp till 8 resultat.


