Site Reliability Engineer
HCLTech
HCLTech is a global technology company, home to 219,000+ people across 54 countries, delivering industry-leading capabilities centered on digital, engineering and cloud, powered by a broad portfolio of technology services and products. We work with clients across all major verticals, providing industry solutions for Financial Services, Manufacturing, Life Sciences and Healthcare, Technology and Services, Telecom and Media, Retail and CPG, and Public Services. Consolidated revenues as of $13+ billion.
For more information on how we process your personal data, please refer to HCLTech’ s Candidate Data Privacy Notice.
Resource will be part of Production Engineering - Observability team and will be required to deliver the strategic initiative of implementing and expanding the FIC Observability platform based on Open Telemetry and OpenSearch. The programme aims to improve the monitoring, reliability, and operational stability of critical FIC trading applications by providing a modern observability framework, enabling faster incident detection, reduced outage duration, and enhanced operational insights across the estate.
Scope of work would include
- Gather requirements and conduct gap analysis across existing monitoring and observability platforms.
- Design and define observability standards, telemetry collection strategies, dashboards, alerting frameworks, and monitoring best practices.
- Build and implement OpenTelemetry-based instrumentation and OpenSearch solutions across critical FIC applications and infrastructure.
- Develop automation, dashboards, analytics, and reporting capabilities to improve operational visibility and reduce manual overhead.
- Provide knowledge transfer, documentation, and operational handover to ensure long-term supportability by the existing team.
Resource requirement
Senior SRE/Observability Developer with proven experience designing, implementing, and operating enterprise-scale observability solutions within mission-critical environments.
The candidate must have strong hands-on experience with:
- Open Telemetry instrumentation, collectors, and telemetry pipelines.
- OpenSearch architecture, indexing, data management, and analytics.
- Grafana dashboard development, alerting, and operational reporting.
- Enterprise monitoring platforms such as Geneos and related observability technologies.
- Site Reliability Engineering (SRE) practices, operational stability, incident reduction, and platform automation.
- Proven experience delivering OpenTelemetry implementations in complex enterprise environments.
- Strong OpenSearch expertise including design, deployment, optimisation, and operational support.
- Experience working within front-office, trading, or other high-availability financial services environments is highly desirable.
The engagement is expected to continue through 2027 to support the design, implementation, and rollout phases of the observability programme.
#J-18808-LjbffrReference: WJ-766_22315438