IT & Software

Site Reliability Engineer

MaintainX

Toronto · On · Canada

MaintainX is a leading mobile-first work execution platform

for industrial and frontline teams.

More than 13,000 customers , including Duracell, McDonald's, Shell, DHL and Volvo, use MaintainX to cut unplanned downtime and run better operations, across 13.9 million managed assets and 79.5 million completed work orders.

In August 2026 MaintainX became part of Autodesk, joining

Autodesk Operations Solutions

, the organization unifying Autodesk's operations platform alongside Tandem, FlexSim and Fusion Operations.

Autodesk's strategy is to converge design, make and operate into one continuous lifecycle: design an asset, build it, run it, then feed what you learn running it back into the next design. Autodesk had design and make. Operate is the phase that tells you what actually happened, and it is ours.

We're looking for a

Site Reliability Engineer

to help advance MaintainX's reliability, observability, and developer autonomy as we scale our platform.

In this role, you'll partner closely with product and platform development teams to improve the stability, resilience, and operational readiness of our services. You'll work alongside teams to design for reliability from the start, establish clear ownership and standards, and build shared tooling that enables teams to operate their services with confidence.

You'll also contribute to company-wide initiatives that define how MaintainX approaches reliability software development, including observability standards, incident response practices, and service health metrics, helping the organization adopt proven industry practices at scale.

This role is well-suited for an developer who enjoys working across teams, influencing technical direction through strong development practices, and turning reliability principles into practical, scalable systems.

What You'll Do:

Assess service maturity and provide insights to development teams

Partner with development teams to implement observability best practices

Enable development teams to become autonomous with their service deployment, support, and infrastructure

Mentor developers on reliability practices, focusing on making them self-sufficient

Act as the bridge, ear and eyes of the Platform Division teams to drive tooling and practice adoption across development teams

About You:

Deep understanding of observability practices in a distributed system environment and how it influences system design and team behaviour

Practical experience with SRE concepts (SLOs, error budgets, incident management)

3-5+ years in software development, SRE, DevOps, or production development roles with experience operating production systems

Proficient in cloud-native platforms and infrastructure-as-code concepts and tools

Working knowledge of at least one programming language (TypeScript/Node.js is a plus)

Excellent communication and collaboration abilities across technical and non-technical teams

Ability to translate complex reliability concepts into actionable guidance

#J-18808-Ljbffr

Reference: WJ-3875_13222947

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.