Senior Specialist Engineer (Specialist Site Reliability Engineer SRE)
National Health Service
Overview
In this role you will apply software and systems engineering to build and operate reliable, scalable production systems. You will focus on automation and CI/CD to maintain service reliability and performance in production. You’ll monitor across cloud services, drive observability improvements, and align with SLOs to minimize downtime. You’ll collaborate with cross-functional teams to promote reliable, efficient operations and share guidance on SRE practices.
Pay / Benefits- hybrid working model
- flexible working opportunities
- uk government salary framework with market pay supplement (MPS) up to 5000
- core HQ locations with modern facilities
- inclusive culture promoting equality
- Ensure services are stable, scalable, and performant through engineering best practices and system design
- Identify and address bottlenecks with advanced problem-solving and performance tuning
- Plan capacity to support current and future workloads
- Respond to production incidents and restore services quickly
- Perform root cause analysis and postmortems to prevent recurrence
- Design and implement monitoring/alerting systems using dashboards and tools
- Improve observability and reduce alert fatigue
- Develop automation to eliminate manual tasks and improve efficiency
- Write clear, maintainable code and drive IaC initiatives
- Contribute to SLO/SLI definition and continuous improvement of operational practices
- Advocate SRE principles and integrate reliability into development lifecycle
- Create and maintain technical documentation and provide training when appropriate
- Collaborate with software engineering, DevOps, and infrastructure teams to streamline deployment and operations
- Promote a culture of shared responsibility for service reliability
- Experience as Site Reliability Engineer, DevOps Engineer, Operations Engineer or similar
- Coding skills in Python, PowerShell or Bash
- Understanding of Linux/Unix & Windows systems, networking, and distributed systems
- Experience with observability tools (Prometheus, Grafana, Datadog) and alerting
- Understanding of infrastructure automation (Terraform, Ansible, PowerShell, Helm)
- Excellent communication and collaboration skills
- Problem-solving ability to respond to sudden demands
- communication
- collaboration
- problem solving
- Python
- PowerShell
- Bash
Reference: WJ-747_30454962