IT & Software

Site Reliability Engineer

HCLTech

London · England · United Kingdom

HCLTech is a global technology company, home to 219,000+ people across 54 countries, delivering industry-leading capabilities centered on digital, engineering and cloud, powered by a broad portfolio of technology services and products. We work with clients across all major verticals, providing industry solutions for Financial Services, Manufacturing, Life Sciences and Healthcare, Technology and Services, Telecom and Media, Retail and CPG, and Public Services. Consolidated revenues as of $13+ billion.

For more information on how we process your personal data, please refer to HCLTech’ s Candidate Data Privacy Notice.

Resource will be part of Production Engineering - Observability team and will be required to deliver the strategic initiative of implementing and expanding the FIC Observability platform based on Open Telemetry and OpenSearch. The programme aims to improve the monitoring, reliability, and operational stability of critical FIC trading applications by providing a modern observability framework, enabling faster incident detection, reduced outage duration, and enhanced operational insights across the estate.

Scope of work would include

  • Gather requirements and conduct gap analysis across existing monitoring and observability platforms.
  • Design and define observability standards, telemetry collection strategies, dashboards, alerting frameworks, and monitoring best practices.
  • Build and implement OpenTelemetry-based instrumentation and OpenSearch solutions across critical FIC applications and infrastructure.
  • Develop automation, dashboards, analytics, and reporting capabilities to improve operational visibility and reduce manual overhead.
  • Provide knowledge transfer, documentation, and operational handover to ensure long-term supportability by the existing team.

Resource requirement

Senior SRE/Observability Developer with proven experience designing, implementing, and operating enterprise-scale observability solutions within mission-critical environments.

The candidate must have strong hands-on experience with:

  • Open Telemetry instrumentation, collectors, and telemetry pipelines.
  • OpenSearch architecture, indexing, data management, and analytics.
  • Grafana dashboard development, alerting, and operational reporting.
  • Enterprise monitoring platforms such as Geneos and related observability technologies.
  • Site Reliability Engineering (SRE) practices, operational stability, incident reduction, and platform automation.
  • Proven experience delivering OpenTelemetry implementations in complex enterprise environments.
  • Strong OpenSearch expertise including design, deployment, optimisation, and operational support.
  • Experience working within front-office, trading, or other high-availability financial services environments is highly desirable.

The engagement is expected to continue through 2027 to support the design, implementation, and rollout phases of the observability programme.

#J-18808-Ljbffr

Reference: WJ-766_22315438

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.