IT & Software

Lead Site Reliability Engineer

JP Morgan Chase

London · England · United Kingdom

Overview

In this Lead SRE role, you will shape reliability and observability across globally distributed trading platforms within the Asset Management Trading Technology group. You will collaborate closely with traders to align reliability with business goals, lead live-incident response and post-mortems, and drive automation and self-healing patterns. You will contribute production-grade code and help scale latency-sensitive systems, leveraging AI-assisted workflows to accelerate triage and analysis. This is an opportunity to transform SRE practices in a high-stakes, global trading environment.

Responsibilities

  • Engage with traders to understand workflows and reliability priorities across asset classes
  • Serve as trusted engineering partner to the desk to ensure stable, performant systems
  • Support live trading environments with incident response, RCA, and post-mortem leadership
  • Contribute to codebase (Java, Kotlin, Python) for reliability improvements and automation
  • Lead design and rollout of modern SRE patterns (automation, self-healing, resilience)
  • Utilize enterprise AI for major-incident triage and post-incident analysis with proper validation
  • Promote AI-assisted reliability workflows across SDLC (CI/CD checks, testing, readiness)
  • Improve latency, throughput, and stability of high-volume trading apps
  • Build and maintain monitoring, alerting, and distributed tracing tooling
  • Collaborate with infrastructure, networking, cloud, and cybersecurity teams for end-to-end reliability
  • Operate in a globally distributed org (EMEA, US, APAC) and interact with traders and senior stakeholders

Key requirements
  • Front office trading environment or similarly high-pressure, low-latency domain experience
  • Proficiency with SRE tooling: FIX messaging, Kafka, Grafana, Splunk, ITRS Geneos, Dynatrace, InfluxDB, MQ, Oracle DB
  • Experience using enterprise AI capabilities for SRE workflows with validation and data-sensitivity awareness
  • Ability to evaluate AI-driven recommendations for correctness and risk with guardrails
  • Deep knowledge of SRE principles: SLIs/SLOs, telemetry, DR, capacity, performance tuning
  • Experience designing observability frameworks for mission-critical systems
  • Proven incident leadership and long-term remediation能力
  • Strong programming skills in Python, Java, or Kotlin; production-grade coding
  • Experience with microservices, distributed systems, and event-driven architectures
  • Understanding of CI/CD pipelines, automated testing, and deployment strategies
  • Comfortable communicating with traders and senior stakeholders; strong written and verbal communication
  • Calm, decisive under high pressure; collaborative leadership
  • Excellent communication with technical and business stakeholders
  • Calm under pressure and decisive decision-making
  • Collaborative leadership with a partner mindset
  • Java, Kotlin, Python production-grade coding
  • SRE tooling: FIX, Kafka, Grafana, Splunk, ITRS Geneos, Dynatrace, InfluxDB, MQ, Oracle DB
  • Observability and tracing

Reference: WJ-799_20869151

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.