IT & Software

Senior Site Reliability Engineer - OpenTelemetry

EPAM Systems

London · Greater London · United Kingdom

Overview

In this role you will drive the design and expansion of a modern observability platform to boost monitoring, resilience, and operational stability for critical trading systems. You will work within the Production Engineering – Observability team to define standards, deliver enterprise-scale observability solutions, and promote best practices for operational excellence. The role involves hands-on instrumentation, data analytics, and dashboards to enable faster incident detection and reduced outages. This is a chance to shape observability at scale in a hybrid, finance-focused environment with a strong emphasis on automation and cross-team collaboration.

Responsibilities
  • Gather requirements and analyze current monitoring and observability setups
  • Define observability standards, telemetry strategies, and alerting frameworks
  • Implement OpenTelemetry instrumentation and OpenSearch-based solutions across apps and infrastructure
  • Design dashboards, analytics, and reporting to improve transparency and efficiency
  • Develop automation tools and processes to reduce manual overhead
  • Integrate observability frameworks with enterprise monitoring platforms like Geneos
  • Provide documentation and handover for long-term sustainability
  • Apply SRE principles to enhance stability, scalability and reduce incidents
Key requirements
  • Senior SRE or Observability Engineer with enterprise-scale implementation experience
  • Strong OpenTelemetry expertise including instrumentation and telemetry pipelines
  • Deep knowledge of OpenSearch for architecture, indexing, optimization, and analytics
  • Experience building dashboards and alerts with Grafana
  • Familiarity with enterprise monitoring tools such as Geneos
  • Practical understanding of SRE practices and automation
  • Background in high-availability or mission-critical environments; financial services experience is desirable
  • Knowledge of anomaly detection, alert correlation, and incident response automation
  • Exposure to observability in cloud-native or hybrid setups
  • collaboration
  • problem-solving
  • attention to detail
  • OpenTelemetry instrumentation
  • OpenTelemetry telemetry pipelines
  • OpenSearch data indexing and analytics

Reference: WJ-747_30859538

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.