IT & Software

Lead Site Reliability Engineer

JP Morgan Chase

London · Greater London · United Kingdom

Overview

In this Lead SRE role, you drive reliability and resilience across applications and platforms within the Infrastructure Platforms team. You lead incident management, mentor engineers, and steer data-driven efforts to meet service levels. You will shape AI-assisted reliability workflows and guide teams through complex problems, contributing to scalable, secure systems that support the firm’s objectives. This is a high-impact, leadership-focused opportunity in a globally recognized organization.

Responsibilities
  • Champions SRE culture and exerts technical influence across the team
  • Leads initiatives to improve reliability using data-driven analytics to enhance service levels
  • Collaborates to define service level indicators and establish SLOs and error budgets with stakeholders
  • Demonstrates deep technical expertise to resolve bottlenecks in assigned domains
  • Acts as point of contact during major incidents to accelerate resolution and reduce financial impact
  • Documents and shares knowledge internally via forums and communities of practice
  • Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices with traceability, resiliency, and security controls
Key requirements
  • Formal training or certification on SRE concepts and advanced applied experience
  • Design and code complex problems in public cloud like AWS
  • Deep proficiency in reliability, scalability, performance, security, enterprise architecture, toil reduction
  • Fluency in Python and knowledge of software processes with depth in technical disciplines
  • Proficiency and experience in observability using Grafana, Dynatrace, Prometheus, Datadog, Splunk
  • Proficiency in CI/CD tools (Jenkins, GitLab, Terraform) and container orchestration (ECS, Kubernetes, Docker)
  • Experience troubleshooting networking technologies
  • Ability to solve problems related to complex data structures and algorithms
  • Commitment to self-education, teaching new languages, and collaborating across stakeholder groups
  • Experience using enterprise-authorized AI capabilities to improve SRE workflows with proper validation and guardrails
  • leadership
  • mentoring
  • collaboration
  • AWS and cloud design
  • Python development
  • observability and monitoring tools (Grafana, Dynatrace, Prometheus, Datadog, Splunk)

Reference: WJ-747_30819864

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.