IT & Software

Site Reliability Engineer, Studios

iMG world

London · Greater London · United Kingdom

Overview

As Site Reliability Engineer at IMG, you design, build, and operate resilient platforms underpinning our digital, cloud, and live-broadcast services. You will improve reliability and observability while driving automation and disaster recovery readiness across on-prem and cloud environments. You collaborate with engineering and operations to raise release quality and incident response standards. You play a key role in scaling systems for live, business-critical workloads and supporting high-availability workflows.

Responsibilities
  • Design, build, and maintain reliable, scalable infrastructure across on-prem and cloud environments
  • Improve availability, latency, and efficiency through reliability engineering practices
  • Enhance observability with monitoring, logging, alerting, dashboards, and service health indicators
  • Define SLIs/SLOs, alerting standards, and runbooks for critical services
  • Automate provisioning, configuration, deployment, and recovery using IaC and scripting
  • Collaborate with software, platform, and broadcast engineering teams to improve resilience
  • Serve as escalation point for production incidents and drive post-incident follow-up
  • Lead root cause analysis and preventive actions for incidents
  • Support high availability, backup, failover, and disaster recovery design and testing
  • Enforce security, access control, patching, and best practices across infrastructure
  • Optimize capacity, cost, and performance; produce technical documentation
  • Support live events and critical operational workflows requiring rapid response and clear communication
  • Contribute to planning for new services, migrations, and platform enhancements with resilience in mind
  • Improve platform reliability, stability, and recovery across IMG services
  • Drive reduced mean time to detect/resolve incidents via observability and automation
  • Promote operational ownership and service standards across environments
  • Strengthen resilience for live client-facing workflows through tested failover approaches
Key requirements
  • Proven experience as Site Reliability Engineer, DevOps Engineer, Platform Engineer, or similar
  • Strong knowledge of Linux
  • Hands-on experience with AWS, Azure, or Google Cloud
  • Experience with Docker and Kubernetes
  • Experience with CI/CD tooling and modern software delivery
  • Hands-on with Infrastructure as Code tools like Terraform or CloudFormation
  • Experience with monitoring, logging, alerting, and observability design
  • Solid understanding of networking, security, architecture, and distributed systems
  • Scripting or programming in Python or Bash
  • Experience in high-availability, live production, or business-critical environments
  • Strong troubleshooting, calm decision-making under pressure, and continuous improvement mindset
  • Excellent communication and collaboration with technical and non-technical stakeholders
  • calm under pressure
  • collaboration
  • proactive ownership
  • Linux fundamentals
  • AWS/Azure/Google Cloud
  • Docker

Reference: WJ-747_30173459

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.