Site Reliability Engineer
Infinity Quest
Site Reliability Engineer Responsibilities :
- Proactively identify reliability risks and independently drive initiatives to address them before they become incidents
- Participate in and continuously improve our on-call rotation, including incident response, triage, and leading blameless post-incident reviews
- Define and implement monitoring, logging, and distributed tracing strategies; build and maintain dashboards; set meaningful alerts; and drive SLO/SLI/SLA and error budget adoption across services
- Scope technical projects and break them down into user stories and tasks, driving them to completion with minimal oversight
- Make sound technical decisions, leveraging input from teammates and contributing to technical conversations across engineering teams
- Automate the provisioning and management of infrastructure using Infrastructure as Code (IaC) tools such as Terraform
You may be a good fit if :
- You have at least 4 years of experience working in a professional environment as a Software Engineer (with some SRE or operations responsibilities)
- You have contributed to the design and build of cloud-native applications written in Python, Java or Go
- You have strong hands-on experience with observability
- You understand the difference between monitoring and observability, and can articulate how metrics, logs, and traces work together
- You have worked with Infrastructure As Code tooling, for example Terraform
- You have participated in on-call rotations and are comfortable leading incident response under pressure, communicating clearly with stakeholders throughout
- You are comfortable taking ownership of initiatives or projects independently, from scoping through to delivery
- You build effective working relationships, give and receive constructive feedback openly, and are trusted by colleagues at all levels
Technologies we use include :
- Python, Java, and Go are our primary server languages
- Our browser applications are based on Angular and React –
- Code lives in GitHub and flows to production through a CI/CD pipeline built on GitHub Actions, with some workloads on Jenkins
- Infrastructure runs on AWS (EC2) with workloads on Kubernetes-managed Docker containers
- Datadog is our primary observability platform — experience with Datadog APM, dashboards, monitors, and RUM is a plus - Infrastructure is managed as code using Terraform
Reference: WJ-747_30291444