Senior Site Reliability Engineer
Malvern Panalytical
As a Senior Site Reliability Engineer, you will ensure the reliability, availability, performance, and scalability of our Azure-based cloud platforms. You will lead proactive reliability initiatives, establish observability standards, define SLOs/SLIs/SLAs, and strengthen incident response while driving automation. You will collaborate with engineering, product, and operations to deliver new features quickly without sacrificing stability or security. You will champion SRE best practices and contribute to continuous improvement of cloud services and operational maturity. This role offers impact at a global scale within a growing, innovative tech organization.
Pay / Benefits- competitive salary
- benefits package
- professional development
- certification support
- wellbeing
- career progression
- Monitor Azure cloud applications and services using telemetry, metrics, logging, and observability tools
- Proactively identify issues and operational trends to prevent downtime and customer impact
- Define and improve SLOs, SLIs, and SLAs with cross-functional teams
- Lead incident response activities, post-incident reviews, and root cause analysis
- Drive automation across monitoring, deployments, recovery processes, and runbooks
- Collaborate to embed reliability and DevOps practices throughout the software development lifecycle
- Identify opportunities to improve platform resilience, scalability, and operational maturity
- Act as a technical leader for reliability initiatives and promote operational excellence
- Significant experience in SRE, DevOps, Cloud Operations, or similar roles
- Strong hands-on experience with production workloads on Microsoft Azure
- Expertise in monitoring, alerting, observability platforms, and cloud-native practices
- Deep understanding of SRE principles including SLOs, SLIs, SLAs, and reliability methodologies
- Experience with automation and development using C#, Python, PowerShell, and IaC
- Strong knowledge of incident management, root cause analysis, CI/CD pipelines, and best practices
- Excellent communication and stakeholder management skills
- Azure certifications such as Azure Administrator or Azure DevOps Engineer are desirable
- Experience in regulated, scientific, or enterprise environments advantageous
- communication
- stakeholder management
- collaboration
- Microsoft Azure
- SRE fundamentals (SLOs/SLOs/SLAs, error budgets)
- Monitoring and observability
Reference: WJ-747_30139621