IT & Software

Vice President, Site Reliability Engineering

The Bank of New York Mellon

London · Greater London · United Kingdom

Overview

In this role you will design and scale centralized engineering solutions to improve operational efficiency and resiliency for Production Services. You’ll build full‑stack platforms and internal tools, combining development, UI, and infrastructure automation. You will take solutions from concept through deployment and ongoing improvement, partnering with cross‑functional teams to meet enterprise standards. This is a hands‑on, senior role focused on reducing toil, accelerating incident recovery, and enabling productive engineering workflows. You’ll work at scale in a financial services environment and influence the evolution of monitoring, automation, and SRE practices.

Pay / Benefits
  • competitive compensation
  • broad benefits and wellbeing programs
  • paid leaves
  • paid volunteer time
  • global resources
  • strong culture of excellence
Responsibilities
  • Design, develop, and deploy centralized engineering solutions to improve efficiency and resiliency across Production Services
  • Build full‑stack applications and internal tools (backend services, APIs, automation, UI) using Python, Java, React or Angular
  • Create scalable solutions for self‑service tooling, dashboards, alert enrichment, incident reduction, and workflow automation
  • Develop reusable frameworks and components for broad adoption across Production Services
  • Automate infrastructure, deployment, configuration, and runtime support with Ansible and Kubernetes
  • Define and improve SLIs/SLOs and service health measures aligned to priorities
  • Build and optimize monitoring and observability with Prometheus, Grafana, AppDynamics, and Splunk
  • Apply AIOps to improve event correlation, anomaly detection, and proactive issue prevention
  • Collaborate with engineering, infrastructure, production support, security, and risk teams to ensure secure, scalable solutions
  • Identify and automate manual or fragmented processes across Production Services
Key requirements
  • Bachelor degree in Computer Science, Engineering, or related discipline or equivalent practical experience
  • Strong full‑stack development with Python and Java
  • Front‑end expertise with React or Angular for operational interfaces
  • Proven end‑to‑end solution design from development to production support
  • Experience in Site Reliability Engineering, Production/DevOps/Platform Engineering
  • Strong Linux/Unix administration, scripting, and infrastructure knowledge
  • Hands‑on experience with Ansible and Kubernetes in production
  • Ability to define and operationalize SLIs, SLOs, dashboards, alerts, and health indicators
  • Hands‑on experience with Prometheus, Grafana, AppDynamics, Splunk
  • Strong troubleshooting, analytical, and collaboration skills
  • Strong verbal and written communication
  • collaboration across technical and non‑technical stakeholders
  • problem‑solving under complex distributed environments
  • focus on continuous improvement
  • Python
  • Java
  • React

Reference: WJ-747_30135161

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.