Lead Site Reliability Engineer – Operations Excellence
JP Morgan Chase
As Lead Software Engineer on the AI and Machine Learning Platform, you will design, build, and operate scalable AI infrastructure for production LLM serving. You will own reliability, performance, and cost-efficiency across the LLM inference stack, from hosting and routing to observability and incident response. You’ll work with cloud, Kubernetes, and open-source serving stacks to modernize infrastructure and drive measurable platform improvements. This role offers meaningful impact by shaping secure, scalable AI platforms used across the organization.
Responsibilities- Design, develop and deliver secure, production-grade AI infrastructure software and services
- Build backend APIs to enable reliable AI infrastructure operation
- Operate and scale LLM serving stacks (e.g., vLLM, llm-d) in production
- Deploy and lifecycle-manage LLMs on Amazon EKS/SageMaker and on-prem GPU clusters
- Implement observability with logs, metrics and traces (Prometheus, Grafana, Alertmanager)
- Tune GPU/accelerator capacity and autoscaling for cost-efficient inference
- Lead reliability engineering for LLM endpoints, including capacity planning and incident response
- Participate in on-call rotations and post-incident analysis
- Automate recurring operational issues to improve stability and developer experience
- Build multi-agent orchestration systems where appropriate
- Foster an inclusive team culture and drive adoption of new technologies
- Formal training, certification, or equivalent practical experience in software engineering
- Hands-on system design, development, testing, and production stability
- Advanced proficiency in Python for production-grade services
- Experience with automation and continuous delivery
- Hands-on experience with AWS and Terraform for infrastructure delivery
- Strong understanding of SRE practices including incident management and runbooks
- Experience with observability (metrics, logs, traces)
- Production experience operating LLM inference servers like vLLM/llm-d
- Experience hosting LLMs on Amazon EKS and/or SageMaker and on GPU clusters
- Knowledge of latency/throughput, model versioning, and safe rollout patterns
- strong collaboration across teams
- focus on incident management and root-cause analysis
- clear communication in post-incident reports
- Python production-grade development
- AWS (SageMaker, EKS)
- Terraform
Reference: WJ-747_30175308