IT & Software

Site Reliability & Observability Engineer – Datadog / Azure

MYO Talent

Birmingham (Aston) · West Midlands · United Kingdom

Site Reliability & Observability Engineer / Datadog – Synthetic Monitoring, APM, RUM, Log Management, SLO's, Alerting / Azure / Azure DevOps / Cloudflare / 6-month contract / Hybrid – West Midlands / Remote / £450 – 600 per day Inside IR35.


One of our leading clients is seeking a Lead Site Reliability & Observability Engineer to build and operate a world-class monitoring, synthetic testing, and reliability platform.


Location – West Midlands / Remote – 5 days per week with 1-2 days per week onsite

Duration – 6 months +

Day rate – £450 – 600 per day Inside IR35


This role will lead the implementation of Datadog across Azure and Cloudflare, creating a comprehensive early warning system that continuously validates APIs, integrations, and customer user journeys in production.


Key Responsibilities:

·? Own and evolve the Datadog observability platform.

· Design and maintain synthetic monitoring for critical API and UI workflows.

· Build continuous production validation covering business-critical customer journeys.

· Integrate monitoring, testing, dashboards, and alerting into Azure DevOps and GitHub pipelines.

· Develop monitoring-as-code and testing-as-code practices using Terraform.

· Create actionable dashboards, SLOs, SLIs, alerts, and anomaly detection.

· Integrate Datadog with Azure, Cloudflare, and modern SaaS architectures.

· Drive reliability, performance, and root-cause analysis across production systems.


Required Experience:

· Strong hands-on Datadog expertise, including:

o Synthetic Monitoring

o APM

o RUM

o Log Management

o SLOs and Alerting

· Experience operating large-scale global SaaS platforms.

· Deep Azure experience.

· Experience integrating Cloudflare services.

·? Strong CI/CD experience with Azure DevOps and GitHub.

·? Expertise in API, integration, and browser-based testing.

· Infrastructure as Code experience using Terraform.

· Experience with distributed systems, microservices, and cloud-native architectures.

Desirable:

· Datadog certifications.

· Azure certifications.

·? Cloudflare administration experience.

· Background in Site Reliability Engineering (SRE) or Platform Engineering leadership roles.


JBRP1_UKTJ

Reference: WJ-747_30945636

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.