IT & Software

Senior Site Reliability Engineer

Spectrum IT Recruitment

Southampton · Hampshire · United Kingdom

Overview

Senior Site Reliability Engineer to oversee production availability and health across cloud and on-prem environments. You will build tools to streamline platform infrastructure and improve reliability, performance, and delivery speed. You’ll partner with developers on testing, releases, and capacity planning, while guiding operational oversight for large distributed applications. This role offers impact through designing resilient systems and driving innovation in a fast-moving SaaS context.

Pay / Benefits
  • Life Insurance - 4 x Annual Salary
  • Private Medical Insurance
  • Employee Assistance Programme
  • Hybrid Working - 3 Days from Home
  • GP Online Assistance Portal
Responsibilities
  • Monitor system and application metrics to optimize performance and troubleshoot issues
  • Collaborate with developers to improve service quality through testing and structured releases
  • Participate in architectural discussions, manage platform operations, and forecast capacity
  • Design and implement automated solutions for scalable, dependable systems
  • Maintain focus on delivering new features while upholding SLAs and stability
Key requirements
  • 3–6 years in a similar role focused on systems engineering, automation, and reliability
  • Proficiency in at least one programming language (Python, Go, Java, or C#) and scripting (Bash/PowerShell)
  • Strong cloud experience with AWS and core services (EC2, ECS, Lambda, DynamoDB)
  • Hands-on with infrastructure-as-code tools (CloudFormation or Terraform)
  • Understanding of CI/CD practices and tools (Jenkins, GitLab CI/CD, CircleCI)
  • Experience with containers and microservices (Docker, Kubernetes)
  • Familiarity with observability/monitoring tools (Prometheus, Grafana, ELK, CloudWatch)
  • Incident management experience and blameless postmortems
  • Strong analytical and troubleshooting abilities
  • collaboration across cross-functional teams
  • problem solving under pressure
  • clear communication during incidents
  • Kubernetes
  • Grafana Observability Suite (Loki, Mimir, Tempo)
  • Splunk

Reference: WJ-747_30146960

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.