IT & Software

Senior Site Reliability Engineer, SRE

Jobtailor

Toronto · On · Canada

  • Design, build, and maintain scalable, highly available, and fault-tolerant distributed systems
  • Partner with development teams as a reliability consultant and influence architectural decisions
  • Write code to automate operational tasks and CI/CD pipelines
  • Build internal tools, libraries, and frameworks for self-service observability
  • Participate in a 24/7 on-call rotation and act as incident commander during critical disruptions
  • Conduct blameless root cause analyses and implement corrective actions
  • Monitor, measure, and optimize system performance, latency, and capacity
  • Forecast capacity needs using usage patterns and historical data
  • Build and integrate AIOps solutions, including automated responses and self-healing systems
  • Use AI-assisted coding tools such as Claude Code and Cursor
  • Develop and document runbooks and procedural guides for the observability knowledge base
  • Analyze telemetry data, build predictive capacity models, and identify bottlenecks and failure modes

Requirements

  • Bachelor's degree in Computer Science, a related technical field, or equivalent practical experience
  • 5+ years of professional experience in a Site Reliability Engineering, DevOps, or Software Engineering role focused on infrastructure and operations
  • Strong programming proficiency in one or more high-level languages such as Rust, Go, Python, or Typescript
  • Comfortable writing, testing, and deploying production-grade code
  • Deep knowledge of AWS services, especially networking, IAM, EKS, ALBs/NLBs, Route 53, and CloudWatch
  • Proven experience with Kubernetes in production, including service exposure, networking, and availability engineering
  • Solid understanding of Linux/Unix operating systems, TCP/IP, DNS, HTTP, and modern distributed systems architecture

Core Competencies

Demonstrates expertise in designing and maintaining scalable distributed systems, with a strong focus on automation, observability, and incident management. Proficient in programming and cloud services, particularly in AWS and Kubernetes, to optimize system performance and reliability.

Highest-signal resume keywords

  • Site Reliability Engineering
  • AWS Services
  • Kubernetes
  • Programming Proficiency
  • Automation

Hard Skills

  • Rust
  • Go
  • Python
  • Typescript
  • Linux/Unix
  • TCP/IP
  • DNS
  • HTTP
  • Distributed Systems Architecture
  • CI/CD

Soft Skills

  • Incident Management
  • Root Cause Analysis
  • Collaboration

Certifications & Qualifications

  • Bachelor's Degree in Computer Science

Industry Keywords

  • Infrastructure
  • Operations
  • Observability
  • Capacity Forecasting
  • Self-Healing Systems

Tools & Technologies

  • AIOps Solutions
  • Claude Code
  • Cursor
  • CloudWatch
  • EKS
  • ALBs/NLBs
  • Route 53

#J-18808-Ljbffr

Reference: WJ-3875_12649730

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.