Manager, System and Platform Operations
Publicis Media
As the System and Platform Operations Director, you will lead reliability and stability for Epsilon Retail Media production environments. You’ll drive the operational vision, combining development and run initiatives to ensure premium support and full-service delivery. You’ll align with Engineering, Product, and Security to safeguard production systems while enabling rapid, safe releases. This role offers impact at scale, shaping how live platforms perform and how teams collaborate to exceed customer expectations. You’ll join a values-driven, cross-functional environment that prizes collaboration, growth, and innovation.
Pay / Benefits- competitive compensation
- great benefits package
- hybrid working opportunities
- career advancement opportunities
- inclusive and diverse culture
- recognition of employee impact
- Establish and manage operational practices for a future-fit support model
- Implement proactive incident detection, response, remediation, and continuous improvement
- Own operational integrity of all production environments
- Monitor services, report on health metrics, and exceed individual and customer SLAs
- Own incident management and on-call responses, leading major resolutions
- Oversee change management and assess customer impact
- Enable rapid yet reliable delivery of new products and fixes
- Collaborate with Engineering, Product, Delivery, and Security on system reliability
- Document processes and maintain ITSM lifecycle (Change, Config, Service Levels, Performance, Incident, Problem)
- 5+ years in Site Reliability or equivalent operational roles
- Docker and Kubernetes expertise
- Terraform experience
- Solid networking, security, and system architecture knowledge
- Scripting in Java, Golang, Python, Bash (or similar)
- Monitoring/observability tools: DataDog, Prometheus, Grafana
- Database knowledge: PostgreSQL, Bigtable
- API and microservices understanding
- People leadership with experience guiding technical teams
- Experience with ITSM, high-availability environments, and SaaS/cloud backends
- excellent communication
- collaborative mindset
- problem-solving
- Docker
- Kubernetes
- Terraform
Reference: WJ-747_30299770