Senior Site Reliability Engineer - Platform Reliability (Resilience)
elastic
Overview
In this Site Reliability Engineer role, you will join the Platform Engineering SRE team to design, build, and scale the global Elastic infrastructure. You will automate engineering efforts to ensure reliability across multi-cloud hosting of Elastic Cloud services. The role emphasizes collaboration, operational excellence, and proactive incident management to minimize customer impact. This is an opportunity to shape scalable, reliable platforms that support Elastic’s AI-driven search solutions and services.
Pay / Benefits- Competitive pay
- Health coverage for you and your family
- Flexible locations and schedules
- Generous vacation days
- Donation matching up to $2000
- Paid volunteer time
- Lead technical initiatives to automate system engineering for global reliability
- Expand and maintain platform infrastructure to meet scaling demands with software, tooling, and automation
- Foster an inclusive, collaborative environment focused on operational excellence
- Respond to and prevent customer-impacting incidents; participate in follow-the-sun on-call rotation
- Background in software engineering
- Experience with public cloud and managed Kubernetes is advantageous
- Customer-first approach to solving operational problems
- Experience working in distributed teams or remote environments
- Embrace progress-oriented mindset toward platform reliability
- collaboration
- inclusive communication
- coaching and mentoring
- Public cloud platforms
- Managed Kubernetes
- Infrastructure-as-Code tooling (Crossplane, Terraform)
Reference: WJ-747_30488823