Senior Site Reliability Engineer
Spectrum IT Recruitment
As a Senior Site Reliability Engineer, you will own the production environment, ensuring availability and healthy system operation across cloud and on-prem deployments. You will build tooling to automate and strengthen platform infrastructure, targeting improved dependability, performance, and delivery speed. You’ll lead operational support for large distributed apps and work closely with developers to elevate service quality and reliability. This is a hands-on role focused on resilience, observability, and scalable architecture, with a meaningful impact on customer experiences and regulatory compliance.
Pay / Benefits- Life Insurance - 4 x Annual Salary
- Private Medical Insurance
- Employee Assistance Programme
- Hybrid Working - 3 Days from Home
- GP Online Assistance Portal
- Monitor system and application metrics to tune performance and troubleshoot issues
- Collaborate with developers to improve service quality through testing and structured releases
- Engage in architectural discussions and manage platform operations including capacity forecasting
- Design and implement automated, resilient, scalable solutions
- Maintain feature delivery while meeting SLAs and stability goals
- 3–6 years of hands-on SRE or systems engineering experience
- Proficiency in at least one programming language (Python, Go, Java, or C#) and scripting (Bash/PowerShell)
- Strong knowledge of AWS and core services (EC2, ECS, Lambda, DynamoDB)
- Experience with infrastructure-as-code tools (CloudFormation or Terraform)
- Solid understanding of CI/CD processes and tools (Jenkins, GitLab CI/CD, CircleCI)
- Experience with containerization and microservices (Docker, Kubernetes)
- Familiarity with observability/monitoring tools (Prometheus, Grafana, ELK, CloudWatch)
- Incident management experience with blameless postmortems
- Configuration management knowledge (Ansible, Puppet, Chef)
- Kubernetes administration experience; Kubernetes certifications are a bonus
- Analytical and troubleshooting mindset
- Cross-functional collaboration
- Communication during incidents and outages
- Kubernetes cluster management
- Grafana observability suite (Loki, Mimir, Tempo)
- Splunk, Datadog, PagerDuty, Rundeck
Reference: WJ-747_30378781