Site Reliability Engineer, Studios
iMG world
As Site Reliability Engineer at IMG, you design, build, and operate resilient platforms underpinning our digital, cloud, and live-broadcast services. You will improve reliability and observability while driving automation and disaster recovery readiness across on-prem and cloud environments. You collaborate with engineering and operations to raise release quality and incident response standards. You play a key role in scaling systems for live, business-critical workloads and supporting high-availability workflows.
Responsibilities- Design, build, and maintain reliable, scalable infrastructure across on-prem and cloud environments
- Improve availability, latency, and efficiency through reliability engineering practices
- Enhance observability with monitoring, logging, alerting, dashboards, and service health indicators
- Define SLIs/SLOs, alerting standards, and runbooks for critical services
- Automate provisioning, configuration, deployment, and recovery using IaC and scripting
- Collaborate with software, platform, and broadcast engineering teams to improve resilience
- Serve as escalation point for production incidents and drive post-incident follow-up
- Lead root cause analysis and preventive actions for incidents
- Support high availability, backup, failover, and disaster recovery design and testing
- Enforce security, access control, patching, and best practices across infrastructure
- Optimize capacity, cost, and performance; produce technical documentation
- Support live events and critical operational workflows requiring rapid response and clear communication
- Contribute to planning for new services, migrations, and platform enhancements with resilience in mind
- Improve platform reliability, stability, and recovery across IMG services
- Drive reduced mean time to detect/resolve incidents via observability and automation
- Promote operational ownership and service standards across environments
- Strengthen resilience for live client-facing workflows through tested failover approaches
- Proven experience as Site Reliability Engineer, DevOps Engineer, Platform Engineer, or similar
- Strong knowledge of Linux
- Hands-on experience with AWS, Azure, or Google Cloud
- Experience with Docker and Kubernetes
- Experience with CI/CD tooling and modern software delivery
- Hands-on with Infrastructure as Code tools like Terraform or CloudFormation
- Experience with monitoring, logging, alerting, and observability design
- Solid understanding of networking, security, architecture, and distributed systems
- Scripting or programming in Python or Bash
- Experience in high-availability, live production, or business-critical environments
- Strong troubleshooting, calm decision-making under pressure, and continuous improvement mindset
- Excellent communication and collaboration with technical and non-technical stakeholders
- calm under pressure
- collaboration
- proactive ownership
- Linux fundamentals
- AWS/Azure/Google Cloud
- Docker
Reference: WJ-747_30173459