Senior Site Reliability Engineer- Remote
ClickHouse
In this role you will lead the reliability strategy for ClickHouse Cloud, shaping scalable, secure, and highly available systems. You’ll collaborate with Control Plane, Data Plane, Core, Security, Support and Operations to drive fault tolerance and performance improvements. You will own incident management, post-mortems, and continuous improvement, using software engineering to build tools that boost operational efficiency. This is a high-impact opportunity to enable elastic, high-performance cloud services at scale.
Pay / Benefits- remote-friendly
- healthcare contributions
- equity/stock options
- flexible time off
- $500 home office setup
- global gatherings
- Design and implement scalable, secure, and highly available systems for ClickHouse Cloud
- Establish and manage SLOs/SLAs
- Ensure monitoring and alerting across infrastructure components (Data Plane, Control Plane, Core)
- Lead incident response, post-mortems, and blameless analysis for outages
- Drive reliability and performance improvements across ClickHouse services
- Plan and drive chaos initiatives across engineering teams
- Manage on-call processes and coordinate escalation to minimize downtime
- 8+ years of experience in Site Reliability Engineering or related field
- Bachelor’s or Master’s degree in Computer Science or related field
- Hands-on experience with Go and/or Python
- Strong knowledge of AWS, Azure, or Google Cloud Platform
- Experience with Kubernetes or Docker Swarm
- Experience with automation/configuration tools such as Ansible, Terraform, or Puppet
- Strong debugging skills and ownership mindset
- Excellent communication and interpersonal abilities
- Interest in efficiency, availability, scalability, and data governance
- strong problem-solver
- ownership and accountability
- collaboration across cross-functional teams
- Go and/or Python
- Cloud platforms (AWS/Azure/GCP)
- Kubernetes or Docker Swarm
Reference: WJ-747_30140234