IT & Software

Senior Site Reliability Engineer

Malvern Panalytical

Remote · Nationwide · United Kingdom

Overview

As a Senior Site Reliability Engineer, you will ensure the reliability, availability, performance, and scalability of our Azure-based cloud platforms. You will lead proactive reliability initiatives, establish observability standards, define SLOs/SLIs/SLAs, and strengthen incident response while driving automation. You will collaborate with engineering, product, and operations to deliver new features quickly without sacrificing stability or security. You will champion SRE best practices and contribute to continuous improvement of cloud services and operational maturity. This role offers impact at a global scale within a growing, innovative tech organization.

Pay / Benefits
  • competitive salary
  • benefits package
  • professional development
  • certification support
  • wellbeing
  • career progression
Responsibilities
  • Monitor Azure cloud applications and services using telemetry, metrics, logging, and observability tools
  • Proactively identify issues and operational trends to prevent downtime and customer impact
  • Define and improve SLOs, SLIs, and SLAs with cross-functional teams
  • Lead incident response activities, post-incident reviews, and root cause analysis
  • Drive automation across monitoring, deployments, recovery processes, and runbooks
  • Collaborate to embed reliability and DevOps practices throughout the software development lifecycle
  • Identify opportunities to improve platform resilience, scalability, and operational maturity
  • Act as a technical leader for reliability initiatives and promote operational excellence
Key requirements
  • Significant experience in SRE, DevOps, Cloud Operations, or similar roles
  • Strong hands-on experience with production workloads on Microsoft Azure
  • Expertise in monitoring, alerting, observability platforms, and cloud-native practices
  • Deep understanding of SRE principles including SLOs, SLIs, SLAs, and reliability methodologies
  • Experience with automation and development using C#, Python, PowerShell, and IaC
  • Strong knowledge of incident management, root cause analysis, CI/CD pipelines, and best practices
  • Excellent communication and stakeholder management skills
  • Azure certifications such as Azure Administrator or Azure DevOps Engineer are desirable
  • Experience in regulated, scientific, or enterprise environments advantageous
  • communication
  • stakeholder management
  • collaboration
  • Microsoft Azure
  • SRE fundamentals (SLOs/SLOs/SLAs, error budgets)
  • Monitoring and observability

Reference: WJ-747_30139621

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.