IT & Software

Site Reliability Engineer

PagerDuty

Toronto · On · Canada

  • As a Site Reliability Engineer I on the Core Infrastructure team you’ll help build and operate the foundational infrastructure that powers PagerDuty’s real-time digital operations platform
  • Our systems support millions of events and alerts daily, enabling customers to detect, respond to, and resolve incidents quickly and reliably
  • You’ll work at the intersection of platform evolution and operational excellence, building and evolving foundational network, compute, and ingress infrastructure while scaling and hardening existing systems
  • Your work will directly impact the reliability, scalability, and security of the services our customers rely on to keep their businesses running as PagerDuty continues to grow across products, regions, and customer use cases
  • Support and improve foundational infrastructure, including networking, compute
    platforms, Kubernetes clusters, and ingress/traffic management systems
  • Contribute to the reliability and scalability of PagerDuty’s core platform by
    hardening existing systems and supporting the rollout of new infrastructure
    capabilities
  • Participate in agile rituals (standups, planning, retros) and communicate
    progress/risks early
  • You stay current on technical trends to suggest innovative tools and approaches
    to interesting problems
  • Monitor system health using metrics, logs, and alerts, and participate in 24/7
    on-call rotations to help detect, respond to, and resolve incidents

Benefits

  • 20 hours per year of paid volunteer time
  • Health insurance
  • Wellness Days and mid-year Wellness Week: extra time off for whole company to unplug and recharge at the same time
  • Generous paid parental leave and return to work policy to help with transition back
  • Generous paid time off
  • Flexible workplace/WFH
  • Hands-on career and leadership development programs

Qualifications

  • Working knowledge of networking fundamentals, such as load balancing, DNS,TLS, and ingress traffic flow
  • Hands-on experience operating Linux-based systems in productionenvironments
  • Experience with Infrastructure as Code (e.g., Terraform, CloudFormation)0 to 1+ years of experience in Site Reliability Engineering, DevOps, or PlatformEngineering roles
  • Experience with container orchestration (e.g., EKS, Kubernetes)
  • Experience working on cloud-native infrastructure (e.g., AWS, GCP, Azure),including networking and compute concepts
  • Proficiency in at least one programming language (e.g., Python, Ruby, Go, etc.)
  • Experience with AWS cloud networking concepts such as VPCs, subnets,routing, security groups, and load balancers
  • Experience operating or contributing to production Kubernetes platforms (e.g.,EKS), including cluster upgrades, networking, or ingress configuration
  • Experience with monitoring, observability, and logging platforms (e.g., DataDog,New Relic, SumoLogic, Splunk, Prometheus, Grafana)
  • Familiarity with service meshes, ingress controllers, or API gateways (e.g.,Envoy, Istio, NGINX)

#J-18808-Ljbffr

Reference: WJ-3875_12947171

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.