Senior Site Reliability Engineer (Observability & Analytics, Platform Infra)
Elastic
- Platform Observability & Analytics runs the infrastructure that tells Elastic the truth about its own platform
- The observability clusters show Cloud engineers how production is behaving right now, and the analytics pipelines show the business how the platform and the products get used over time
- This role sits on the observability side. We run 200+ hosted deployments across every supported cloud region, ingesting logs, metrics and traces for all of Elastic Cloud, plus the SLA and SLO monitoring for ESS and Serverless
- When Cloud engineering needs to know what production is doing, they’re looking at something we run
- Owning end-to-end delivery of moderate-to-high complexity projects on the team’s roadmap, with minimal day-to-day direction
- Operating and hardening shared Elastic Cloud infrastructure (ECH, ECE, and ECK) as Infrastructure as Code - writing and reviewing the Terraform, Python, and Go that other engineers depend on
- Carrying a 24/7 on-call rotation: responding to incidents, driving them to resolution, and writing clear RCAs/postmortems that lead to lasting fixes rather than repeat pages
- Reviewing others’ code and designs, and being a trusted second set of eyes on production changes to critical infrastructure
- Mentoring less experienced engineers, and proactively raising risks, ideas, and improvements in team discussions
- Improving runbooks, documentation, and operational processes so the on-call load gets lighter over time
Benefits
- Toast to your health: Fully paid health coverage for you and your family, in many locations.
- Craft your calendar: Flexible location and schedule for most roles.
- Create space for you: Distributed by design workforce, plus generous number of vacation days each year.
- Embrace parenthood: Minimum of 16 weeks of parental leave, plus generous family formation benefits.
- Give back your time: 40 hours each year to use toward volunteering with organizations and causes you’re passionate about.
- Amplify your impact: Double your charitable giving — we match donations up to $1500 USD (or local currency equivalent).
- A track record of consistently delivering end-to-end projects of moderate-to-high complexity with minimal oversight, and being a valuable code/design reviewer for your team5+ years of SRE, platform engineering, or infrastructure engineering experience
- Strong software engineering fundamentals in Python; comfort with Go is a plus
- Comfort working across time zones, in both real-time and asynchronous contexts
- Comfort thinking about the security implications of the infrastructure you build, not just its reliability - you don't need to be a security specialist, but you default to a security-conscious mindset
- Experience carrying a 24/7 on-call rotation, resolving incidents under pressure, and writing RCAs that hold up under review
- Clear written and verbal communication - you document what you build and can explain it to both engineers and non-engineers
- A pattern of mentoring less experienced engineers and speaking up with ideas and concerns in team discussions
- Proficiency with Terraform; comfortable owning large, multi-workspace configurations in a team setting
- Deep Linux systems knowledge and experience operating containerized workloads in production
- Experience with the Elastic Stack (Elasticsearch, Logstash, Beats, Kibana) in production
- Experience with GitOps-style deployment tooling (ArgoCD, Helm) or policy-as-code frameworks (e.g., Kyverno) on Kubernetes
- Experience with secrets management (Vault) or access-control/bastion tooling (Teleport)
- Experience with configuration management tools (e.g., Puppet, Ansible) at fleet scale
- Exposure to FedRAMP, GovCloud, or other regulated/compliance-driven infrastructure
Reference: WJ-3875_13128102