Site Reliability Engineer
Jobtailor
Take on ambiguous reliability, scalability, and efficiency challenges and drive solutions across SRE and development teams
Build and run large-scale, massively distributed, fault-tolerant systems supporting the Genesis platform
Optimize existing systems and build infrastructure
Eliminate toil through automation to improve uptime and rate of change
Cultivate a culture of reliability throughout the organization
Guide technical decisions balancing system health with product priorities
Ensure long-term health, maintainability, and reliability of services
Perform capacity planning and performance analysis
Proactively prevent incidents
Work across teams to build robust, reusable solutions
Requirements
- Strong software engineering skills in Python, Go, or similar
- Extensive experience designing, analyzing, and troubleshooting distributed systems
- Deep expertise with cloud computing platforms, including Kubernetes and Cloud Functions
- Expertise in Non-Abstract Large Systems Design (NALSD)
- Experience leading complex, large-scale technical projects
- Experience providing technical leadership across teams
- Ability to apply coding, algorithms, and complexity analysis to solve ambiguous problems at scale with minimal disruption
- Collaborative and intellectually curious mindset
- Comfortable working across a wide variety of backgrounds and bringing cross-team perspective
- Personal website or GitHub, LinkedIn, and resume fields are available in the application form; the website, LinkedIn, and resume are not required
Core Competencies
Demonstrates strong software engineering skills in Python and Go, with extensive experience in designing and troubleshooting distributed systems. Proven ability to lead complex technical projects and optimize large-scale, fault-tolerant systems while fostering a culture of reliability and collaboration.
Highest-signal resume keywords
- Python Programming
- Go Programming
- Distributed Systems Design
- Kubernetes Expertise
- Technical Leadership
Hard Skills
- Software Engineering
- System Optimization
- Capacity Planning
- Performance Analysis
- Automation
- Algorithms
- Complexity Analysis
- Fault-Tolerant Systems
- Non-Abstract Large Systems Design
- Incident Prevention
Soft Skills
- Collaborative Mindset
- Intellectual Curiosity
Industry Keywords
- Reliability
- Scalability
- Efficiency
- Cross-Team Collaboration
Tools & Technologies
- Cloud Computing Platforms
- Cloud Functions
Reference: WJ-766_21880700