IT & Software

Site Reliability Engineering (SRE) / Observability Technical Lead

NTT DATA

London · Greater London · United Kingdom

Overview

As an experienced SRE and Observability Lead, you will shape the strategy and execution of observability and reliability initiatives across clients. You’ll drive APM, IaC, automation, and distributed tracing using OpenTelemetry to achieve scalable, reliable systems. You’ll design and implement monitoring standards, mentor engineers, and collaborate with cross-functional teams to align with business goals. This role offers the opportunity to influence how we deliver secure, resilient, and high-performance solutions at scale.

Pay / Benefits
  • tailored benefits
  • learning and development opportunities
  • flexible work options
  • equal opportunities employer
  • Disability Confident Committed Employer
  • inclusive, diverse culture
Responsibilities
  • Lead observability and reliability frameworks across the organization
  • Design and implement monitoring and observability standards with engineering teams
  • Manage IaC initiatives using Terraform and coordinate with cloud/infrastructure teams
  • Drive automation of monitoring, alerting, and logging pipelines
  • Develop and maintain observability roadmaps for tracing, logging, and metrics
  • Collaborate with product management, sales, and pre-sales for technical solution design
  • Lead cross-functional teams to improve CI/CD pipelines and deployment reliability
  • Engage with vendors/partners to evaluate and integrate observability solutions
  • Mentor and develop junior engineers and analysts
Key requirements
  • 5+ years in SRE, Observability, or DevOps with leadership responsibilities
  • Proven experience with APM tools (New Relic, Datadog, AppDynamics, Dynatrace)
  • Hands-on OpenTelemetry (OTel) for distributed tracing
  • Strong IaC experience with Terraform
  • Cloud platform experience (AWS, GCP, or Azure)
  • Automation/configuration management experience (Ansible, Chef, or Puppet)
  • Deep knowledge of CI/CD pipelines/tools (GitHub Actions, Jenkins, or Azure DevOps)
  • Kubernetes and containerized environments (Docker, Helm)
  • Experience with log aggregation/analysis platforms (ELK Stack or Splunk)
  • Excellent leadership, communication, and collaboration skills
  • leadership
  • communication
  • collaboration
  • New Relic
  • Datadog
  • AppDynamics

Reference: WJ-747_30180617

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.