IT & Software

Staff Software Engineer - Databases SRE | UK | Remote

Grafana Labs

Remote · Nationwide · United Kingdom

Overview

As Staff Software Engineer - SRE, you will own the reliability of Grafana Cloud databases (Mimir, Loki, Tempo, Pyroscope) delivered as a SaaS across AWS, GCP, and Azure. You’ll partner with embedded product engineering squads to meet high-SLA requirements and drive automation to scale reliability. You’ll define and evolve tenant-specific SLOs, lead incident response, and influence design for production scalability. This role offers a chance to shape resilient, scalable observability platforms for a global customer base.

Pay / Benefits
  • equity
  • bonus (if applicable)
  • remote work
  • 30 days annual leave
  • Grafana Shutdown Days for disconnect
Responsibilities
  • Own production reliability for high-SLA and complex customer environments
  • Design and implement automation to scale reliability practices
  • Ensure customers meet SLO targets and evolve per-tenant SLOs
  • Proactively reduce SLO burn to prevent repeat incidents
  • Serve as primary escalation point and on-call for incidents
  • Lead customer-impacting incident response and post-incident reviews
  • Contribute to design docs and code reviews; influence feature design for scalability
  • Build automation to eliminate toil and improve observability and alerting
  • Collaborate with engineering leaders to define roadmaps and technical designs
  • Mentor engineers, advocate for SRE best practices in development
Key requirements
  • 8+ years of engineering experience, 4+ in SRE/production engineering
  • Formal customer reliability engineering experience preferred
  • Strong Kubernetes experience in AWS, GCP, or Azure
  • Familiarity with infrastructure-as-code tools (Helm, Terraform, Jsonnet)
  • Experience leading teams and mentoring engineers
  • Experience operating multi-tenant systems in production
  • Strong experience designing and implementing SLOs
  • Proficiency in programming languages (Go, Python, Java, etc)
  • Knowledge of Linux internals, networking, cloud storage, and scaling
  • Excellent problem-solving and troubleshooting abilities
  • Experience with blame-free incident response, PIRs, and post-mortems
  • Ability to reason about performance, scaling, and failure modes
  • Autonomy and self-direction within a collaborative engineering team
  • Comfortable partnering with product engineering teams
  • Curiosity, transparency, action bias, and kindness
  • problem-solving mindset
  • strong collaboration and teamwork
  • transparency and openness
  • Kubernetes (AWS/GCP/Azure)
  • Cloud platforms (AWS, GCP, Azure)
  • Infrastructure as code (Helm, Terraform, Jsonnet)

Reference: WJ-747_30178307

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.