IT & Software

Senior Lead Site Reliability / DevOps Engineer

JP Morgan Chase

Glasgow · Glasgow City · United Kingdom

Overview

As a Senior Lead Site Reliability Engineer, you drive reliability, observability, and performance across critical platforms in a fast-paced, agile environment. You guide teams on best practices, implement secure telemetry pipelines, and lead major incidents with blameless postmortems to prevent financial losses. You shape the design of observability architectures and ecosystem tooling to support scalable, secure production systems. This role blends technical leadership with hands-on engineering to deliver trusted, market-leading technology products.

Responsibilities
  • Provide technical guidance on site reliability practices to cross-functional teams and vendors
  • Develop secure production code for reliability tooling and telemetry pipelines; review others' code
  • Influence reliability design, observability architecture, and operational processes
  • Serve as a subject matter expert in site reliability, observability, or telemetry engineering
  • Lead resiliency design reviews and decompose complex reliability problems for teams
  • Act as main incident contact, drive rapid issue resolution and champion blameless postmortems
  • Collaborate to define service level indicators, objectives, and error budgets
  • Design and maintain OpenTelemetry pipelines across hybrid on-prem/cloud environments; support ingestion, processing, and backends (InfluxDB, Prometheus, Elasticsearch, OpenSearch)
  • Drive migration of legacy telemetry to standardized OpenTelemetry instrumentation to reduce technical debt
  • Contribute to engineering community and champion modern observability practices
  • Foster a culture of diversity, inclusion, and respect
Key requirements
  • Formal training or certification in software engineering concepts and applied experience in design, development, testing, and operations
  • Advanced knowledge of reliability, scalability, performance, security, and enterprise architecture with expertise in one or more technical disciplines (cloud, observability, distributed systems)
  • Advanced proficiency in programming languages (e.g., Java, Python, Go)
  • Advanced proficiency and experience in observability using tools like Grafana, Dynatrace, Prometheus, Datadog, Splunk, Elasticsearch, OpenSearch
  • Proficiency with CI/CD tools (e.g., Jenkins, GitLab, Terraform)
  • Experience with containers and orchestration (e.g., ECS, Kubernetes, Docker)
  • Hands-on experience with OpenTelemetry collectors in production
  • Ability to work independently with minimal oversight
  • Practical cloud-native experience
  • Ability to collaborate across levels and stakeholders
  • Leadership and technical mentorship
  • Strong collaboration and cross-functional communication
  • Problem-solving mindset and analytical thinking
  • OpenTelemetry instrumentation and collectors
  • Telemetry pipelines and backends (InfluxDB, Prometheus, Elasticsearch, OpenSearch)
  • White/black box monitoring and SLO/alerting

Reference: WJ-747_30161778

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.