IT & Software

Data Reliability Engineer

Ashdown Group

London · Greater London · United Kingdom

Overview

In this role you will own data reliability within a large data platform, focusing on data quality and incident reduction. You will work with the data and governance teams to design scalable observability and monitoring frameworks that ensure trusted data for critical business decisions. You’ll define and manage data SLAs/SLOs, implement automated validation and observability tooling, and lead root cause analysis when issues occur. This position offers impact across enterprise data pipelines and a chance to shift the organisation toward a proactive, reliability-led approach. The role suits a proactive problem-solver who thrives in collaboration and complex, cloud-based environments.

Responsibilities
  • Improve data quality, reduce incidents, and build scalable observability across a modern enterprise data platform
  • Own data reliability end-to-end from monitoring data health to enforcing standards across pipelines
  • Design and implement data health monitoring and anomaly detection frameworks
  • Define and manage data SLAs and SLOs; lead automated validation and observability tooling
  • Lead root cause analysis when data issues occur and improve systems at their root cause
  • Collaborate with Data Engineering and Data Governance teams to drive a proactive reliability culture
  • Experience with CI/CD and Infrastructure-as-Code is beneficial
  • Work in Azure data lake/ETL/ELT environments and modern cloud-based data environments
Key requirements
  • Experience in Data Engineering, Data Platform, or SRE-style roles
  • Strong SQL and Python skills
  • Hands-on with data observability tools (Grafana, Monte Carlo, Acceldata)
  • Experience with data governance/quality platforms (Informatica, Collibra, Microsoft Purview)
  • Azure ecosystem experience (data lakes, ETL/ELT) is a strong advantage
  • Familiarity with CI/CD and Infrastructure-as-Code is beneficial
  • Proactive, collaborative problem-solver with a focus on end-to-end reliability
  • Proactive problem-solving
  • Collaborative mindset
  • Root-cause analysis
  • SQL
  • Python
  • Data observability tools (Grafana, Monte Carlo, Acceldata)

Reference: WJ-747_30155885

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.