Tech Lead, Site Reliability Engineering
London Stock Exchange Group
In this role you will lead reliability for critical Risk Screening applications within the Risk Intelligence group, ensuring availability, performance, and security across cloud-native platforms. You will shape SRE practices, mentor engineers, and partner with product and engineering to align technology outcomes with business goals. You’ll drive incident management, automation, and resilient delivery across Azure and AWS environments in a UK-based team. This is an opportunity to influence engineering culture and deliver measurable improvements in service excellence.
Pay / Benefits- healthcare
- retirement planning
- paid volunteering days
- wellbeing initiatives
- global collaboration
- growth opportunities
- Own 24/7 reliability operations for Screening services (WC1 and World Check verify)
- Define and monitor SLOs/SLIs and maintain reliability scorecards
- Act as Major Incident Commander and drive blameless post-incident reviews
- Lead automation initiatives across incident response, deployment, compliance, and observability
- Embed non-functional requirements and secure-by-design principles into pipelines
- Co-own cloud reliability roadmap and standardize observability and incident communication tooling
- Ensure disaster recovery readiness and resilient patterns across services
- Lead and mentor SRE engineers and support career development
- Collaborate with HR to build local hiring pipelines and manage workforce planning
- Represent Nottingham site in global SRE forums and contribute to offshore strategy
- Support BCP/DR planning and site-level operational readiness for critical events
- 8+ years in production operations, SRE, or DevOps
- 3+ years in people management
- Cloud-native services with Azure and AWS (Azure SQL, Cosmos DB, App Gateway, Key Vault, Storage, DNS, Load Balancer, VMs, Azure ML, Sentinel; AWS Lambda, ECS, RDS, CloudWatch)
- Strong SRE understanding (SLOs, SLIs, error budgets, incident response)
- Container orchestration (Kubernetes, Docker)
- CI/CD and IaC tools (Terraform, GitHub Actions, Jenkins)
- Observability platforms (Datadog, BigPanda, OpenTelemetry)
- Experience with identity platforms and/or fraud detection systems
- Excellent communication and stakeholder management
- Analytical mindset and continuous improvement
- Empathetic, inclusive leadership and accountability
- Empathetic leadership
- Strategic thinking
- Calm under pressure
- Azure (Azure SQL, Cosmos DB, Key Vault, App Gateway, Storage, DNS, Load Balancer, VMs, Azure ML)
- AWS (Lambda, ECS, RDS)
- Terraform
Reference: WJ-747_30176307