Lead SRE - Chase UK
JP Morgan Chase
As a Site Reliability Engineer within the International Consumer Bank, you will drive reliability and resilience for mission-critical digital banking services. You’ll collaborate with cross-functional teams to design for scale, automate toil, and implement meaningful service metrics and alerts. You will champion self-healing patterns, capacity planning, and performance tuning to ensure a seamless customer experience. This role offers the chance to work in a diverse, distributed team at a leading global financial institution, shaping the reliability foundation of Chase UK and Europe. You will apply AI-enabled approaches to accelerate incident response and improve operational effectiveness.
Responsibilities- Drive continuous improvement of reliability, monitoring, and alerting for mission-critical microservices
- Reduce operational toil through automation by building reliable infrastructure and tooling that expedites feature development
- Develop meaningful service metrics, dashboards, SLIs, SLOs, error budgets, and actionable alerts
- Engage with development teams throughout the software lifecycle to design for reliability and scale
- Design and implement self-healing and resiliency patterns (graceful degradation, rate limiting, circuit breakers, failover)
- Partner across engineering, product, and platform teams to promote reliability standards and adoption
- Execute performance testing and capacity planning to proactively identify bottlenecks
- Participate in feature planning to ensure metrics, alerting, logging, automation, resiliency, capacity, and performance needs are built in from the start
- Use approved AI tools to accelerate root-cause analysis, log and trace investigation, runbook drafting, post-incident analysis, test scaffolding, and documentation
- Continuously develop AI skills relevant to the role, including prompting, output validation, automation workflows, and safe usage patterns
- Formal training or certification on software engineering concepts and advanced applied experience
- Proven experience as a software engineer, with proficiency in Python, Go, or Java
- Experience designing, coding, testing, and delivering software in at least one technology stack
- Strong debugging and troubleshooting skills across distributed systems
- Experience as a Site Reliability Engineer or supporting production services
- Working knowledge of microservice infrastructure components (service discovery, ingress, networking, load balancing)
- Experience with Kubernetes
- Experience with cloud computing services
- Familiarity with observability and reliability toolchains (Grafana, Prometheus, Elasticsearch, Kibana, Jaeger)
- Ability to use AI-assisted engineering tools responsibly
- collaboration
- curiosity
- solution-oriented mindset
- Python
- Go
- Java
Reference: WJ-747_30181252