Incident Lead
Infoya
Infoya is a global IT solutions provider specializing in transforming complex challenges into streamlined, AI-powered outcomes. Through proprietary technology accelerators and full-scale enterprise services, Infoya automates workflows, enhances operational efficiency, and drives digital transformation across industries. With a presence in Canada, the US, India, and Costa Rica, we blend technical depth with creative problem-solving to deliver measurable impact.
Job Description
About the Job: We are seeking an experienced Incident Lead withScrum Master capabilities to lead the response, coordination, and governance ofproduction incidents across cross-functional technology teams. The successfulcandidate will own critical incident execution, SWAT queue health, stakeholdercommunication, and service restoration while applying Agile practices to improve team flow, accountability, and continuous improvement.
Office Location: Toronto
Employment Type: Permanent
Work Arrangement: Hybrid (2 days in office per week)
PositionResponsibilities:
- Lead the end-to-end management of criticalproduction incidents from initial triage through service restoration,stakeholder communication, root-cause review, and closure.
- Establish incident command, confirm severity andbusiness impact, assign clear ownership, and coordinate application,engineering, infrastructure, security, product, and vendor teams.
- Drive timely resolution of critical ticketswithin agreed SLAs and elevate risks, blockers, and resource constraintsappropriately.
- Maintain accurate incident timelines, decisions,actions, dependencies, and recovery updates throughout the incident lifecycle.
- Remove production support bottlenecks and enablerapid decision-making during high-priority incidents.
- Own daily ticket triage and the SWAT queue,ensuring incidents and support tickets are correctly categorized, prioritized,assigned, and progressed.
- Monitor ticket ageing, stalled work, recurringissues, capacity constraints, and ownership gaps to maintain a manageablebacklog.
- Balance urgent restoration work with defects,service requests, technical debt, and preventive improvement initiatives.
- Improve ticket throughput and backlog hygienewhile maintaining quality, compliance, and operational controls.
Scrum Master &Agile Delivery Responsibilities
- Facilitate daily SWAT stand-ups, sprintplanning, backlog refinement, retrospectives, service reviews, and operationalgovernance meetings.
- Coach support and engineering teams on Scrum andAgile practices suited to production support and interrupt-driven work.
- Partner with product owners and service owners to maintain a prioritized, transparent backlog with clear acceptance criteriaand ownership.
- Identify and remove team impediments, managedependencies, support capacity planning, and improve delivery flow acrossteams.
- Use retrospectives and operational data to implement measurable improvements in incident response and support delivery.
Operational Metrics,Reporting & Governance
- Track and report SLA compliance, mean time toacknowledge, mean time to resolution (MTTR), ticket ageing, throughput, backloghealth, critical incident volume, and recurrence trends.
- Prepare dashboards and scorecards that provideleadership with clear visibility into service performance, operational risks,bottlenecks, and improvement actions.
- Facilitate incident and operational governance reviews, ensuring decisions, escalations, risks, and action items are documented and closed on time.
- Promote cross-team accountability through clear owners, target dates, escalation paths, and transparent follow-through.
Problem Management& Operational Excellence
- Lead post-incident reviews and root-causeanalysis for major and recurring incidents without creating a blame-focusedenvironment.
- Ensure corrective and preventive actions are prioritized, tracked, and implemented to reduce recurring incidents.
- Identify trends and systemic weaknesses, then partner with technology teams to improve resilience, monitoring, automation,and support readiness.
- Drive continuous improvement in incident processes, escalation models, runbooks, communications, and service management practices.
Requirements
RequiredQualifications:
- 8+ years of experience leading production support and incident management teams, including coordinating the triage, prioritization, and resolution of software incidents in an enterprise technology environment.
- Demonstrated Scrum Master experience, including facilitation of Agile ceremonies, backlog governance, impediment removal, coaching, and continuous improvement.
- Proven ability to coordinate high-severity incidents across application, engineering, infrastructure, security, product, business, and vendor teams.
- Hands-on experience with ticket triage, incident queues, escalation management, root-cause analysis, and corrective-action tracking.
- Working knowledge of SLA, MTTR, ticket ageing, throughput, backlog health, and other production support metrics.
- Strong stakeholder communication, facilitation, decision-making, conflict-resolution, and executive reporting skills.
- Ability to remain composed, establish accountability, and drive outcomes in high-pressure and time-sensitive situations.
- Experience managing cross-functional and geographically distributed teams.
PreferredQualifications:
- Experience supporting enterprise applications, microservices, integrations, and cloud environments such as AWS, Microsoft Azure, or Google Cloud Platform.
- Familiarity with ITIL practices, DevOps, CI/CD pipelines, observability, monitoring, and modern production support workflows.
- Experience building operational dashboards and scorecards using data from service management and delivery platforms.
- Proficiency with tools such as ServiceNow, Confluence, or similar incident and collaboration platforms.
- Preferred certifications include ITIL, Certified Scrum Master (CSM), Professional Scrum Master (PSM), SAFe Scrum Master, PMP, or PRINCE2.
Salary Range: $90,000 to $93,000CAD/ year
The final compensation offeredwill depend on local market conditions and geographic location, as well asjob-related factors such as the candidate’s knowledge, skills, qualifications,relevant experience, and education/training. Compensation may also includeadditional components such as benefits, and/or other incentives, whereapplicable. In accordance with new employment standards requirements, we retaincopies of this job posting and applicant information for three (3) years afterthe posting is removed. We do not use AI technology; all applications are also reviewed by our recruitment team.
Infoya is an equal opportunityemployer committed to diversity and inclusion. We welcome applications from allqualified individuals, regardless of race, color, religion, sex, sexualorientation, gender identity, national origin, age, disability, protected veteranstatus, aboriginal status, or any other legally protected factors.
Reference: WJ-291_10898350