Site Reliability Engineering Manager
Akkodis
Manager, Site Reliability Engineering (SRE) Akkodis is seeking an experienced Manager, Site Reliability Engineering (SRE) to lead a team responsible for the reliability, availability, performance, and scalability of enterprise applications and platforms.
The successful candidate will drive Site Reliability Engineering best practices, observability initiatives, automation, incident management, and continuous improvement across cloud and on-premises environments. This opportunity is ideal for a technical leader with strong experience in cloud operations, production support, automation, and people leadership.
Work Model:
Hybrid
About the Opportunity Akkodis is seeking an experienced Manager, Site Reliability Engineering (SRE) to lead a team responsible for the reliability, availability, performance, and scalability of enterprise applications and platforms.
What Will the Successful Candidate Do? The Manager, Site Reliability Engineering (SRE)'s primary responsibilities include, but are not limited to:
Lead, mentor, and develop a team of Site Reliability Engineers.
Drive reliability, availability, performance, and scalability across critical applications and platforms.
Implement and operationalize SRE practices, including SLIs, SLOs, error budgets, and post-incident reviews.
Oversee production support operations, on-call processes, incident management, and escalations.
Partner with Development, Infrastructure, Security, and DevOps teams to improve service reliability.
Drive observability and monitoring initiatives across the organization.
Reduce operational effort through automation and self-healing solutions.
Lead major incident response activities and support problem management processes.
Support capacity planning, resiliency testing, and disaster recovery initiatives.
Develop operational standards, runbooks, and knowledge management practices.
Recruit, hire, onboard, and develop SRE talent.
Promote a culture of continuous improvement, collaboration, and operational excellence.
What the Successful Candidate Needs to Succeed Must-Have Skills
Site Reliability Engineering (SRE)
Infrastructure as Code (IaC)
Automation & Scripting
Observability & Monitoring Tools
Production Support Operations
Required Experience
8+ years supporting enterprise applications and distributed systems.
3+ years of experience leading technical teams.
Experience implementing Site Reliability Engineering practices.
Experience supporting mission-critical production environments.
Experience with incident response and problem management.
Experience driving automation and operational improvements.
Strong stakeholder management and collaboration skills.
Required Qualifications
Bachelor’s Degree in Computer Science, Software Engineering, or equivalent experience.
Strong understanding of SRE principles, SLIs, SLOs, and error budgets.
Hands-on experience with Azure and Kubernetes.
Experience with Infrastructure as Code and automation frameworks.
Experience with observability platforms such as Dynatrace, Datadog, New Relic, or AppDynamics.
Strong scripting or programming skills.
Excellent communication and leadership abilities.
Preferred Qualifications
Experience in fintech, payment processing, or regulated environments.
Experience with change management and compliance practices.
Knowledge of SDLC and DevOps best practices.
Experience with disaster recovery and resiliency testing.
Technical Skills
Site Reliability Engineering (SRE) - Must Have
Infrastructure as Code (Terraform, ARM, Bicep, etc.)
Automation & Scripting
DevOps Practices
Soft Skills
Strong leadership and coaching abilities
Strong problem-solving capabilities
Ability to work effectively with cross-functional teams
Additional Requirements
Ability to work in a hybrid environment in Toronto.
Experience managing production support and operational teams.
Ability to lead through major incidents and operational challenges.
Accessibility At Akkodis, part of The Adecco Group, our purpose is simple: to make the future work for everyone. We foster a workplace where diversity is celebrated and every voice matters.
We encourage applications from individuals of all backgrounds and identities.
#SRE #SiteReliabilityEngineering #Azure #Kubernetes #CloudEngineering #EngineeringManager #DevOps #PlatformEngineering #Observability #Automation #InfrastructureAsCode #TorontoJobs #HybridJobs #TechnologyJobs #HiringNow #Akkodis #CloudOperations #ProductionSupport #LeadershipJobs
#J-18808-Ljbffr
The successful candidate will drive Site Reliability Engineering best practices, observability initiatives, automation, incident management, and continuous improvement across cloud and on-premises environments. This opportunity is ideal for a technical leader with strong experience in cloud operations, production support, automation, and people leadership.
Work Model:
Hybrid
About the Opportunity Akkodis is seeking an experienced Manager, Site Reliability Engineering (SRE) to lead a team responsible for the reliability, availability, performance, and scalability of enterprise applications and platforms.
What Will the Successful Candidate Do? The Manager, Site Reliability Engineering (SRE)'s primary responsibilities include, but are not limited to:
Lead, mentor, and develop a team of Site Reliability Engineers.
Drive reliability, availability, performance, and scalability across critical applications and platforms.
Implement and operationalize SRE practices, including SLIs, SLOs, error budgets, and post-incident reviews.
Oversee production support operations, on-call processes, incident management, and escalations.
Partner with Development, Infrastructure, Security, and DevOps teams to improve service reliability.
Drive observability and monitoring initiatives across the organization.
Reduce operational effort through automation and self-healing solutions.
Lead major incident response activities and support problem management processes.
Support capacity planning, resiliency testing, and disaster recovery initiatives.
Develop operational standards, runbooks, and knowledge management practices.
Recruit, hire, onboard, and develop SRE talent.
Promote a culture of continuous improvement, collaboration, and operational excellence.
What the Successful Candidate Needs to Succeed Must-Have Skills
Site Reliability Engineering (SRE)
Infrastructure as Code (IaC)
Automation & Scripting
Observability & Monitoring Tools
Production Support Operations
Required Experience
8+ years supporting enterprise applications and distributed systems.
3+ years of experience leading technical teams.
Experience implementing Site Reliability Engineering practices.
Experience supporting mission-critical production environments.
Experience with incident response and problem management.
Experience driving automation and operational improvements.
Strong stakeholder management and collaboration skills.
Required Qualifications
Bachelor’s Degree in Computer Science, Software Engineering, or equivalent experience.
Strong understanding of SRE principles, SLIs, SLOs, and error budgets.
Hands-on experience with Azure and Kubernetes.
Experience with Infrastructure as Code and automation frameworks.
Experience with observability platforms such as Dynatrace, Datadog, New Relic, or AppDynamics.
Strong scripting or programming skills.
Excellent communication and leadership abilities.
Preferred Qualifications
Experience in fintech, payment processing, or regulated environments.
Experience with change management and compliance practices.
Knowledge of SDLC and DevOps best practices.
Experience with disaster recovery and resiliency testing.
Technical Skills
Site Reliability Engineering (SRE) - Must Have
Infrastructure as Code (Terraform, ARM, Bicep, etc.)
Automation & Scripting
DevOps Practices
Soft Skills
Strong leadership and coaching abilities
Strong problem-solving capabilities
Ability to work effectively with cross-functional teams
Additional Requirements
Ability to work in a hybrid environment in Toronto.
Experience managing production support and operational teams.
Ability to lead through major incidents and operational challenges.
Accessibility At Akkodis, part of The Adecco Group, our purpose is simple: to make the future work for everyone. We foster a workplace where diversity is celebrated and every voice matters.
We encourage applications from individuals of all backgrounds and identities.
#SRE #SiteReliabilityEngineering #Azure #Kubernetes #CloudEngineering #EngineeringManager #DevOps #PlatformEngineering #Observability #Automation #InfrastructureAsCode #TorontoJobs #HybridJobs #TechnologyJobs #HiringNow #Akkodis #CloudOperations #ProductionSupport #LeadershipJobs
#J-18808-Ljbffr
Reference: WJ-3875_13061892