IT & Software

Senior Specialist Engineer (Specialist Site Reliability Engineer SRE)

National Health Service

Leeds · West Yorkshire · United Kingdom

Overview

In this role you will apply software and systems engineering to build and operate reliable, scalable production systems. You will focus on automation and CI/CD to maintain service reliability and performance in production. You’ll monitor across cloud services, drive observability improvements, and align with SLOs to minimize downtime. You’ll collaborate with cross-functional teams to promote reliable, efficient operations and share guidance on SRE practices.

Pay / Benefits
  • hybrid working model
  • flexible working opportunities
  • uk government salary framework with market pay supplement (MPS) up to 5000
  • core HQ locations with modern facilities
  • inclusive culture promoting equality
Responsibilities
  • Ensure services are stable, scalable, and performant through engineering best practices and system design
  • Identify and address bottlenecks with advanced problem-solving and performance tuning
  • Plan capacity to support current and future workloads
  • Respond to production incidents and restore services quickly
  • Perform root cause analysis and postmortems to prevent recurrence
  • Design and implement monitoring/alerting systems using dashboards and tools
  • Improve observability and reduce alert fatigue
  • Develop automation to eliminate manual tasks and improve efficiency
  • Write clear, maintainable code and drive IaC initiatives
  • Contribute to SLO/SLI definition and continuous improvement of operational practices
  • Advocate SRE principles and integrate reliability into development lifecycle
  • Create and maintain technical documentation and provide training when appropriate
  • Collaborate with software engineering, DevOps, and infrastructure teams to streamline deployment and operations
  • Promote a culture of shared responsibility for service reliability
Key requirements
  • Experience as Site Reliability Engineer, DevOps Engineer, Operations Engineer or similar
  • Coding skills in Python, PowerShell or Bash
  • Understanding of Linux/Unix & Windows systems, networking, and distributed systems
  • Experience with observability tools (Prometheus, Grafana, Datadog) and alerting
  • Understanding of infrastructure automation (Terraform, Ansible, PowerShell, Helm)
  • Excellent communication and collaboration skills
  • Problem-solving ability to respond to sudden demands
  • communication
  • collaboration
  • problem solving
  • Python
  • PowerShell
  • Bash

Reference: WJ-747_30454962

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.