Head of Site Reliability Engineering (SRE)
Computershare
In this role you will lead and mature the global Site Reliability Engineering function within Technology Services, guiding reliability strategy and practices. You will partner with Engineering, Infrastructure Operations, and Security to improve service resilience and performance across critical platforms. As a senior leader, you’ll drive observability, automation, and continuous improvement at scale, embedding reliability into the product lifecycle. This is a chance to shape an SRE operating model in a fast-changing, global technology environment.
Pay / Benefits- hybrid work model
- flexible work
- health and wellbeing rewards
- employee share purchase plan
- company contributions
- recognition awards
- Drive adoption of SRE principles (SLOs, error budgets, toil reduction)
- Establish observability and monitoring standards
- Lead automation-first operations
- Improve incident and problem management maturity
- Partner with software and infrastructure engineering to embed reliability into the product lifecycle
- Establish SRE governance, standards, and operating model
- Experienced SRE leader with a track record of building and developing SRE/Production Engineering teams
- Strong understanding of SLOs, SLIs, error budgets, and toil reduction
- Extensive automation experience with scripting and tools (Python, PowerShell, Bash, Terraform, Ansible Automation Platform)
- Hands-on knowledge of observability/monitoring platforms (Dynatrace, Prometheus, Grafana, Splunk)
- Good understanding of CI/CD tooling (Jenkins, GitLab CI, Azure DevOps)
- Background spanning software engineering and technology operations
- Professional cloud/SRE-related certifications
- Passion for reliability, resilience, automation and continuous improvement
- leadership
- strategic thinking
- collaboration
- SRE principles (SLOs/SLIs, error budgets, toil reduction)
- Observability and monitoring
- Automation and scripting
Reference: WJ-747_30178281