IT & Software

Lead Cloud Site Reliability Engineer

Lloyds Banking Group

Halifax · Halifax County · Canada

Lead Site Reliability Engineer - Public Cloud Platform

Location: Manchester or Bristol

Working Pattern: Hybrid (2 days in office per week)

Salary Range £92,701 - £109,060

Salary: £92,701- £109,043

We support flexible working -

Flexible Working Options: Hybrid Working, Job Share

About this opportunity

At Lloyds Banking Group, our purpose is to Help Britain Prosper. As we continue our technology transformation, we're investing in cloud platforms, automation and engineering excellence to deliver secure, resilient and scalable services for millions of customers.

We're looking for a Site Reliability Engineer Lead to help strengthen reliability, observability and operational excellence across our Azure and Google Cloud Platform (GCP) environments.

You’ll lead a team of Site Reliability Engineers, helping to establish engineering standards, improve platform reliability and reduce operational complexity. Working closely with Product Owners, Engineering Leads and platform teams, you’ll influence how cloud services are designed, operated and continuously improved.

This role also includes leadership support for out-of-hours operational and incident management activities when required.

What You’ll Do

As a Site Reliability Engineer Lead, you will:

  • Lead and develop a team of Site Reliability Engineers, creating an inclusive environment that supports learning, collaboration and continuous improvement.
  • Partner with Product Owners, Engineering Leads and platform teams to balance reliability, operational resilience and feature delivery.
  • Use observability data, platform metrics and service insights to identify improvement opportunities and reduce operational risk.
  • Lead incident and problem management activities, promoting effective root cause analysis and continuous service improvement.
  • Champion Site Reliability Engineering practices including Service Level Objectives (SLOs), Service Level Indicators (SLIs) and error budgets.
  • Drive automation initiatives to reduce manual effort and improve platform reliability.
  • Contribute to engineering standards, operational best practices and platform strategy across cloud environments.
  • Support the ongoing evolution of resilient, scalable and secure cloud platforms.

What You’ll Bring

Essential Skills and Experience:

  • Designing, building or operating large-scale cloud platforms within Azure, GCP or comparable cloud environments.
  • Site Reliability Engineering (SRE), Platform Engineering, Infrastructure Engineering, Cloud Engineering or Production Operations.
  • Observability and monitoring practices, including metrics, logging and distributed tracing.
  • Incident management, problem management and service reliability improvement.
  • Service Level Objectives (SLOs), Service Level Indicators (SLIs) and error budgets.
  • Automation and reducing operational toil through engineering solutions.
  • Production readiness reviews, post-incident reviews and continuous improvement activities.
  • Leading technical teams and supporting the development of engineers.
  • Communicating technical concepts clearly to both technical and non-technical stakeholders.
  • Working collaboratively across multiple teams and disciplines.

Desirable Experience

  • Azure and GCP platform technologies.
  • Cloud-native architectures and distributed systems.
  • Infrastructure as Code (IaC) and platform automation.
  • Large-scale enterprise or regulated technology environments.
  • Operational resilience and availability engineering practices.

What We’re Looking For

We’re looking for someone who:

  • Enjoys solving complex reliability and operational challenges.
  • Takes a pragmatic, technology-agnostic approach to engineering decisions.
  • Values learning, knowledge sharing and continuous improvement.
  • Builds inclusive, collaborative and high-performing teams.
  • Is passionate about platform reliability, operational excellence and engineering quality.

Why Lloyds Banking Group?

You’ll join a technology organisation that is:

  • Modernising at scale through cloud adoption, AI-enabled operations and advanced observability.
  • Investing in engineering capability, career development and learning opportunities.
  • Encouraging innovation, experimentation and continuous improvement.
  • Committed to diversity, equity and inclusion.
  • Delivering technology that supports millions of customers across the UK.

This is an opportunity to help shape the future of cloud reliability and operations within one of the UK's largest financial services organisations.

Inclusion and Accessibility

We welcome applications from people with diverse backgrounds, experiences and perspectives.

#J-18808-Ljbffr

Reference: WJ-3875_12606323

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.