Lead Site Reliability Engineer
London Stock Exchange Group
In this Technical Lead SRE role you will shape the foundations of reliability across platforms within the Markets and Risk Intelligence division. You collaborate with architecture, engineering, security, and platform teams to embed reliability from day one. You will lead observability standards and drive incident reduction, performance tuning, and secure, cost-aware operations. This hands-on leadership position emphasizes ownership, proactive problem solving, and mentoring engineers to raise the bar on platform reliability.
Pay / Benefits- healthcare
- retirement planning
- paid volunteering days
- wellbeing initiatives
- Establish SRE foundations for new projects, ensuring readiness, monitoring, and alerting from day one
- Embed reliability, scalability, security, and observability into system design in collaboration with architecture and engineering
- Define and champion observability standards across metrics, logs, traces, and SLIs/SLOs
- Design and evolve monitoring and alerting to improve visibility and reduce toil
- Drive reliability improvements through incident reduction, tuning, and resilient patterns
- Partner with Security to meet compliance, security, and risk-management expectations
- Facilitate smooth handovers from delivery to BAU SRE operations with robust documentation and practices
- Influence cost optimization and efficiency through data-driven architectural decisions
- Mentor engineers, shape engineering standards, and foster continuous learning
- 10+ years hands-on experience in SRE, Platform Engineering, or related roles
- Strong AWS experience (EKS, ECS, EC2, networking, IAM, managed services)
- Deep hands-on Kubernetes and containerized platforms expertise
- Solid Linux systems administration background
- Proven observability platform design and operation experience (monitoring, logging, alerting)
- Hands-on Datadog experience for metrics, logs, APM, and alerting
- Strong understanding of SRE principles (SLOs, error budgets, incident management)
- Experience collaborating with architecture and engineering on system design and delivery
- Knowledge of cloud security principles and collaboration with security teams
- Experience with cloud cost optimization strategies and tooling
- Experience integrating AI with observability stacks (Prometheus, Grafana, ELK, OpenTelemetry)
- strong technical leadership
- collaborative mindset
- calm under pressure
- AWS (EKS, ECS, EC2, IAM)
- Kubernetes
- Linux administration
Reference: WJ-747_30141741