IT & Software

Vice President - Site Reliability Engineering (SRE) - The Core Engineering - Birmingham

Goldman Sachs

Birmingham · West Midlands · United Kingdom

Overview

In this VP role, you will shape and scale reliable production platforms by championing SRE practices across a large engineering organization. You will partner with engineering and product teams to define SLOs, design highly available systems, and reduce toil through automation. You’ll lead complex incident responses and blameless post-mortems, while promoting healthy on-call practices and resilient cloud-native architectures. This role offers impact at scale within a leading global financial services firm that values collaboration, rigor, and continuous improvement.

Pay / Benefits
  • benefits
  • wellness programs
  • personal finance offerings
  • mindfulness programs
  • training and development opportunities
Responsibilities
  • Establish SLOs/SLIs and error budgets with engineering leadership
  • Architect highly available, fault-tolerant systems with product teams
  • Conduct architectural reviews and apply patterns like circuit breakers and rate limiting
  • Reduce toil via automation, tooling, and self-service capabilities
  • Improve production readiness through load testing, capacity forecasting, and chaos engineering
  • Lead multi-system incident responses and drive blameless post-mortems
  • Design healthy on-call models and clear escalation paths
Key requirements
  • Strong proficiency in Java, Python, or Node.js
  • Infrastructure as Code (Terraform, Ansible, or CloudFormation)
  • Docker and Kubernetes (service meshes and ingress controllers)
  • Cloud providers: AWS, GCP, or Azure
  • Observability stacks (Prometheus, Grafana, Splunk, Datadog, OpenTelemetry, ELK, CloudWatch)
  • Automated testing and SDLC concepts; Linux environment; algorithms and data structures
  • Collaborates effectively with product developers
  • Influences architectural decisions and reduces toil
  • Translates complex technical issues into clear insights
  • Observability and telemetry
  • Distributed systems design and performance tuning
  • Network concepts (VPC, load balancing)

Reference: WJ-747_30154160

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.