Lead SRE - AWS Platform
JP Morgan Chase
As Lead Site Reliability Engineer at JPMorganChase in Infrastructure Platforms, you will elevate reliability across critical applications and platforms. You’ll drive resiliency initiatives, design observability and incident response frameworks, and act as a technical authority for sizable products. The role emphasizes data-driven improvements, hands-on cloud and software architecture, and knowledge sharing to raise the team’s capabilities. You will lead major incidents, collaborate with cross-functional partners, and advance AI-assisted reliability workflows to streamline operations and ensure secure, auditable practices. This is a high-impact opportunity to shape reliability at scale within
Responsibilities- Champion site reliability culture and share knowledge via internal forums and communities of practice
- Lead reliability initiatives using data-driven analytics to improve service levels and resolve bottlenecks
- Collaborate to define comprehensive SLIs and partner with stakeholders to establish SLOs and error budgets
- Design and implement observability frameworks and alerting strategies, including monitoring and telemetry collection
- Serve as the primary incident contact for applications, diagnosing issues to minimize business impact
- Share deep technical expertise across domains and mentor peers
- Drive reuse-first adoption of AI-assisted reliability workflows across the SDLC and toolchain practices
- Utilize enterprise AI capabilities to accelerate major-incident triage, troubleshooting, and post-incident analysis with proper data sensitivity handling
- Formal training or certification on site reliability engineering concepts with advanced applied experience
- Hands-on AWS (deploying, operating, and maintaining resilient workloads)
- Proficiency in reliability, scalability, performance, and enterprise architecture; experience with resiliency design reviews
- Fluency in at least one programming language (Python, Java/Spring Boot, or .NET)
- Strong observability knowledge (white/black box monitoring, SLO alerting, telemetry)
- CI/CD tooling proficiency
- Container technologies and orchestration experience
- Networking troubleshooting ability
- Advanced knowledge across technical disciplines with ability to evaluate new technologies
- Experience using enterprise AI capabilities for SRE workflows and ability to validate outputs and guardrails
- Strong collaboration and teamwork
- Effective written and verbal communication
- Problem-solving mindset with data-driven approach
- AWS (hands-on, resilient workloads)
- Observability and telemetry (SLO-based alerting)
- CI/CD tooling and automation
Reference: WJ-747_30173395