IT & Software

Senior Site Reliability Engineer

VIQU IT Recruitment

Milton Keynes · England · United Kingdom

Senior Site Reliability Engineer

Up to £75,000 plus bonus and on call allowance

Milton Keynes (2 days on site a week)

VIQU have partnered with a well-established B2B SaaS company who are going through a significant platform transformation. and so are hiring for a Senior Site Reliability Engineer to build stability, respond to live incidents, and assist with system upkeep. The role will also play a key part in on implementing and adopting new tooling and processes surrounding the wider transformation.

This is a genuine opportunity to own and operate how the cloud function works, and progress into a team lead position as the team grows.

Experience required for the Senior Site Reliability Engineer

  • Previous experience as a Site Reliability Engineer or similar (cloud, infrastructure, DevOps or platform engineering) within a customer facing environment - e.g SaaS or MSP.
  • Strong hands-on experience with both Azure, and on-premise virtual machines.
  • Experience withInfrastructure as Code / Terraform, Container orchestration (Kubernetes or AKS), and Monitoring and observability tooling (Prometheus, Grafana, Datadog, or Azure Monitor).
  • Ability to implement new processes, and tools, ensuring the wider development and support teams adopts new ways of working.
  • Ability to communicate across internal teams and external customers.
  • Skilled in networking across both cloud (Azure) and on premise environments.
  • Either Windows or Linux systems administration skills (Linux preferred).
  • Previous use of AI tools to enhance efficiency.

Job Duties of the Senior Site Reliability Engineer

  • Utilise various technologies (Terraform, Kubernetes ect) to manage provision, and configure servers and networks, and automate application lifecycles.
  • Regularly use Datadog and other observability tools for application performance monitoring.
  • Implement new ways of working, helping to shape how the organisation responds and recovers to incidents.
  • Take ownership of incident resolutions.
  • Actively drive down key reliability metrics (MTTR, incident frequency, on-call toil) by evaluating key incidents.
  • Work on an a on call rota, ensuring you are available to respond to incidents during this time.
  • Identify areas for automation and help implement changes that raise the bar for reliability.

#J-18808-Ljbffr

Reference: WJ-766_21889243

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.