IT & Software

Senior Site Reliability Engineer (LON)

McNally Recruitment Ltd

London · England · United Kingdom

Senior Site Reliability Engineer (London)

We’re working in collaboration to source a Senior Site Reliability Engineer for a large UK client. The role is mostly working remotely, with only 1 day per week being required to work in the London office.

  • In this key role, you’ll improve, drive, and embed non-functional and operational characteristics such as availability, performance, efficiency, change management, monitoring, security, incident response, and capacity planning of our products and services
  • You’ll enjoy significant stakeholder interaction , working in collaboration with engineers to ensure a principled approach to deliver change in a safe and secure way
  • This is a chance to join an inclusive team with a collaborative ethos and a commitment to innovation and professional development
  • You’ll work from home some of the time, but you’ll also spend a significant amount of time working from an office or hub

What you’ll do

  • Work closely with our feature team and other colleagues to meet defined service level objectives and continually improve systems and environments.
  • Define error budgets that support finding the right balance between risk and reliability.
  • Provide structure and help to our release process , suggesting and making improvements where possible.
  • Help scale systems sustainably through mechanisms like automation , evolving them by pushing for changes that improve reliability and velocity.
  • Coach and provide guidance to colleagues and the wider team, leading where required.

In addition to this, you’ll:

  • Proactively contribute new ideas and innovations to meet short-term and longer-term goals
  • Continually balance and manage any potential risks
  • Be accountable for the day-to-day health of both production and non-production environments and respond to any incidents as required
  • Provide technical expertise and input to establish the risk tolerance of products and services
  • Communicate incident status updates clearly and frequently to other teams, customers and stakeholders

The skills you’ll need

  • At least 10 years of hands-on experience, including as a Senior SRE with a proactive approach to spotting problems, areas for improvement, and performance bottlenecks.
  • Experience working with cloud-native microservices, including containerisation, management of Kubernetes workloads and API management.
  • Hands-on experience with Azure, Infrastructure as Code (IaC), and technologies such as PowerShell, JSON, Azure Bicep, ARM and Azure DevOps.
  • The client is moving to Terraform, which is essential , moving from Bicep (desirable).
  • Experience with Full Stack Observability using tools such as Grafana Stack, Log Analytics, AppInsights
  • Excellent knowledge of DevOps processes and principles
  • Knowledge of IT Service Management and automation of IT fulfilment processes through Orchestration and ServiceNow
  • Strong communication skills with the ability to proactively engage with a wide range of stakeholders

SALARY INCLUDES 10% Benefits-As-Cash

#J-18808-Ljbffr

Reference: WJ-766_22550231

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.