Staff Software Engineer - Databases SRE | UK | Remote
Grafana Labs
As Staff Software Engineer - SRE, you will own the reliability of Grafana Cloud databases (Mimir, Loki, Tempo, Pyroscope) delivered as a SaaS across AWS, GCP, and Azure. You’ll partner with embedded product engineering squads to meet high-SLA requirements and drive automation to scale reliability. You’ll define and evolve tenant-specific SLOs, lead incident response, and influence design for production scalability. This role offers a chance to shape resilient, scalable observability platforms for a global customer base.
Pay / Benefits- equity
- bonus (if applicable)
- remote work
- 30 days annual leave
- Grafana Shutdown Days for disconnect
- Own production reliability for high-SLA and complex customer environments
- Design and implement automation to scale reliability practices
- Ensure customers meet SLO targets and evolve per-tenant SLOs
- Proactively reduce SLO burn to prevent repeat incidents
- Serve as primary escalation point and on-call for incidents
- Lead customer-impacting incident response and post-incident reviews
- Contribute to design docs and code reviews; influence feature design for scalability
- Build automation to eliminate toil and improve observability and alerting
- Collaborate with engineering leaders to define roadmaps and technical designs
- Mentor engineers, advocate for SRE best practices in development
- 8+ years of engineering experience, 4+ in SRE/production engineering
- Formal customer reliability engineering experience preferred
- Strong Kubernetes experience in AWS, GCP, or Azure
- Familiarity with infrastructure-as-code tools (Helm, Terraform, Jsonnet)
- Experience leading teams and mentoring engineers
- Experience operating multi-tenant systems in production
- Strong experience designing and implementing SLOs
- Proficiency in programming languages (Go, Python, Java, etc)
- Knowledge of Linux internals, networking, cloud storage, and scaling
- Excellent problem-solving and troubleshooting abilities
- Experience with blame-free incident response, PIRs, and post-mortems
- Ability to reason about performance, scaling, and failure modes
- Autonomy and self-direction within a collaborative engineering team
- Comfortable partnering with product engineering teams
- Curiosity, transparency, action bias, and kindness
- problem-solving mindset
- strong collaboration and teamwork
- transparency and openness
- Kubernetes (AWS/GCP/Azure)
- Cloud platforms (AWS, GCP, Azure)
- Infrastructure as code (Helm, Terraform, Jsonnet)
Reference: WJ-747_30178307