Lead Site Reliability Engineer - Chief Technology Office
JP Morgan Chase
In this Lead SRE role, you will drive the reliability and resilience of critical applications within the CTO organization. You’ll guide incident response, design reviews, and cross-team collaboration to meet service level objectives and reduce toil. You will break down complex problems, mentor engineers, and own the end-to-end reliability of medium to large-scale products. This role offers impact across business outcomes, risk mitigation, and the opportunity to shape how high-performing systems are built and operated. You will work with senior stakeholders to elevate platform reliability and enable data-driven decision making.
Responsibilities- Champion site reliability culture and exert technical influence across the team
- Lead initiatives to improve reliability and stability using data-driven analytics
- Collaborate to define service level indicators and establish service level objectives with stakeholders
- Demonstrate deep technical expertise to identify and solve bottlenecks
- Act as the main incident contact to resolve issues quickly and avoid losses
- Document and share knowledge within the organization via forums and communities of practice
- Deep proficiency in reliability, scalability, performance, security, and enterprise system architecture
- Fluency in at least one programming language (e.g., Python, Java Spring Boot, .NET)
- Hands-on observability expertise with monitoring and telemetry tools
- Proficiency in CI/CD tools (e.g., Jenkins, GitLab, Terraform)
- Experience with container technologies and orchestration (e.g., ECS, Kubernetes, Docker)
- Experience troubleshooting common networking technologies
- Ability to analyze complex data structures and algorithms
- Strong collaboration and communication across stakeholder levels
- Mentoring and coaching engineers
- Effective cross-team communication
- Problem-solving under pressure
- Observability: white/black-box monitoring, SLO alerting, telemetry collection
- Incident management and post-incident reviews
- Cloud and infrastructure-as-code familiarity
Reference: WJ-747_30185857