IT & Software

Lead Site Reliability Engineer

Spectrum IT Recruitment

Southampton · England · United Kingdom

Overview

Senior SRE role focused on keeping production cloud platforms observable, secure, scalable and reliable. You will lead automation and incident investigations within a growing Cloud Platform Engineering function, collaborating with Cloud Operations, Support, DevOps and Engineering teams. The role emphasizes Azure expertise, observability, and cost-conscious platform improvements to support mission-critical SaaS solutions. You will help shape architecture, governance and security practices while exploring AI-assisted tooling for automation and troubleshooting.

Pay / Benefits

  • Bonus
  • Medical Care

Responsibilities
  • Protect and improve production environments as part of the SRE team
  • Manage and prioritize a backlog of reliability, scalability and operational improvements
  • Lead investigations into outages, performance issues and cloud expenditure
  • Perform root-cause analysis and implement corrective actions
  • Automate repetitive operational activities
  • Provide technical leadership to Cloud Operations, Support, DevOps and Engineering teams
  • Establish and maintain SLOs, SLAs, SLIs and error budgets
  • Design and implement monitoring, alerting and dashboards across cloud and microservices
  • Deploy observability tools (Grafana, Prometheus, Azure Monitor, OpenTelemetry)
  • Develop custom metrics and queries for distributed microservices
  • Create reusable Bicep/Terraform modules for monitoring and cloud infra
  • Support and improve AKS-based Kubernetes environments
  • Review and optimize platform performance, security and cost
  • Contribute to cloud architecture and scalable platform solutions
  • Support continuous improvement of deployment and provisioning processes
  • Ensure security, governance and compliance requirements are met
  • Explore AI-assisted tooling to enhance automation and productivity

Key requirements
  • Six+ years in SRE/related cloud role
  • Strong hands-on Azure expertise
  • Production experience with Kubernetes/AKS
  • Extensive experience in observability and monitoring
  • Experience with Grafana, Prometheus, OpenTelemetry, Elasticsearch
  • Custom metrics, queries, dashboards for microservices
  • Scripting or development skills (PowerShell, Python, C#)
  • IaC experience (Bicep, ARM, Terraform)
  • Git or other VCS
  • Knowledge of SQL Server, Elasticsearch, YAML/JSON/XML
  • Understanding of microservices, cloud platforms and containerisation
  • Experience defining/woking with SLOs, SLAs, SLIs and error budgets
  • Troubleshooting and root-cause analysis skills
  • Security, governance and compliance awareness
  • Experience in transformation projects and live-service environments
  • Technical leadership
  • Collaborative cross-functional communication
  • Problem-solving and analytical mindset
  • Microsoft Azure
  • Kubernetes / AKS
  • Grafana

Reference: WJ-799_26559040

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.