IT & Software

Senior Site Reliability Engineer

Understanding Recruitment

London · Greater London · United Kingdom

Overview

In this Senior Site Reliability Engineer role, you will strengthen the reliability and operability of a latency-sensitive production platform. You will work with a small, elite team to raise standards, automate workflows, and improve tooling across production infrastructure. You will own incident diagnosis, monitoring and CI/CD improvements, and developer experience from local to production. This is a hands-on, engineering-led role focused on performance, security, and scalable operations.

Pay / Benefits
  • 150,000 - 200,000+ base salary
  • Significant performance-based bonus + Equity
  • Private healthcare
  • UK visa sponsorship available
  • Engineering-led organisation
  • Direct influence over reliability, tooling and engineering practices
Responsibilities
  • Improve the reliability and operability of production systems
  • Build and improve monitoring, logging, tracing, dashboards and alerting
  • Improve incident diagnosis, root cause analysis and operational workflows
  • Build safer and more repeatable deployment and rollback processes
  • Automate repetitive operational and infrastructure work
  • Improve CI/CD pipelines and release processes
  • Develop internal tooling that helps engineers operate production systems more effectively
  • Improve the developer experience from local development through production
  • Work with Linux systems, networking, host configuration and resource contention
  • Contribute to infrastructure security, access controls, secrets management and system hardening
  • The systems are latency-sensitive, so the role can extend into host-level tuning, kernel settings, CPU isolation and networking behaviour
Key requirements
  • Strong experience in Site Reliability Engineering, Platform Engineering, DevOps or Infrastructure Engineering
  • Experience operating production infrastructure in cloud environments
  • Strong Linux systems knowledge and understanding of networking fundamentals
  • Experience with monitoring, observability and alerting
  • Strong troubleshooting and root cause analysis skills
  • Experience with CI/CD and infrastructure automation
  • AWS, Terraform or Ansible experience would be advantageous
  • Experience with high-performance, high-throughput or latency-sensitive systems would be particularly valuable
  • Comfortable taking ownership of problems and driving improvements independently
  • ownership mindset
  • problem-solving and analytical thinking
  • ability to work independently and drive improvements
  • Linux systems
  • Networking fundamentals
  • Monitoring, observability and alerting

Reference: WJ-747_30137817

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.