Senior AWS Site Reliability Engineer
Spectrum IT Recruitment
In this role you will safeguard production reliability and lead platform operations for large-scale distributed applications. You will build and automate tools to improve performance, availability, and delivery speed, aligning with the company’s focus on dependable software and fast feature delivery. You’ll collaborate with developers, participate in architectural discussions, and drive capacity planning and incident response. This opportunity suits a hands-on engineer who thrives in complex, cloud-native environments and seeks impact through scalable, observable systems.
Pay / Benefits- Life Insurance - 4 x Annual Salary
- Private Medical Insurance
- Bonus Scheme
- Employee Assistance Programme
- Hybrid Working - 3 Days from Home
- GP Online Assistance Portal
- Monitor system and application metrics to tune performance and troubleshoot issues
- Collaborate with developers to improve service quality via testing and structured releases
- Engage in architectural discussions, manage platform operations, and contribute to capacity forecasting
- Design and implement automated solutions for resilient, scalable systems
- Focus on delivering new features while maintaining stability and meeting service level goals
- 3-6 years of hands-on experience in a similar role with emphasis on systems engineering, automation, and reliability
- Proficient in at least one programming language (Python, Go, Java, or C#) with scripting in Bash or PowerShell
- Solid understanding of AWS and core services (EC2, ECS, Lambda, DynamoDB)
- Experience with infrastructure-as-code tools (CloudFormation or Terraform)
- Knowledge of CI/CD principles and tools (Jenkins, GitLab CI/CD, CircleCI)
- Strong understanding of containers and microservices (Docker, Kubernetes)
- Experience with observability/monitoring tools (Prometheus, Grafana, ELK, CloudWatch)
- Incident management experience and ability to lead cross-functional resolution
- Configuration management familiarity (Ansible, Puppet, or Chef)
- Professional cloud DevOps certifications (AWS/GCP) or equivalent
- Proficient in scripting and automation to support operations
- Collaborative mindset
- Analytical and problem-solving abilities
- Strong troubleshooting under pressure
- Kubernetes (large-scale clusters)
- Grafana Observability Suite (Loki, Mimir, Tempo)
- Splunk, Datadog, PagerDuty, Rundeck (monitoring/automation)
Reference: WJ-747_30201131