Senior Site Reliability Engineer
Understanding Recruitment
Overview
In this Senior Site Reliability Engineer role, you will strengthen the reliability and operability of a latency-sensitive production platform. You will work with a small, elite team to raise standards, automate workflows, and improve tooling across production infrastructure. You will own incident diagnosis, monitoring and CI/CD improvements, and developer experience from local to production. This is a hands-on, engineering-led role focused on performance, security, and scalable operations.
Pay / Benefits- 150,000 - 200,000+ base salary
- Significant performance-based bonus + Equity
- Private healthcare
- UK visa sponsorship available
- Engineering-led organisation
- Direct influence over reliability, tooling and engineering practices
- Improve the reliability and operability of production systems
- Build and improve monitoring, logging, tracing, dashboards and alerting
- Improve incident diagnosis, root cause analysis and operational workflows
- Build safer and more repeatable deployment and rollback processes
- Automate repetitive operational and infrastructure work
- Improve CI/CD pipelines and release processes
- Develop internal tooling that helps engineers operate production systems more effectively
- Improve the developer experience from local development through production
- Work with Linux systems, networking, host configuration and resource contention
- Contribute to infrastructure security, access controls, secrets management and system hardening
- The systems are latency-sensitive, so the role can extend into host-level tuning, kernel settings, CPU isolation and networking behaviour
- Strong experience in Site Reliability Engineering, Platform Engineering, DevOps or Infrastructure Engineering
- Experience operating production infrastructure in cloud environments
- Strong Linux systems knowledge and understanding of networking fundamentals
- Experience with monitoring, observability and alerting
- Strong troubleshooting and root cause analysis skills
- Experience with CI/CD and infrastructure automation
- AWS, Terraform or Ansible experience would be advantageous
- Experience with high-performance, high-throughput or latency-sensitive systems would be particularly valuable
- Comfortable taking ownership of problems and driving improvements independently
- ownership mindset
- problem-solving and analytical thinking
- ability to work independently and drive improvements
- Linux systems
- Networking fundamentals
- Monitoring, observability and alerting
Reference: WJ-747_30137817