Senior Site Reliability Engineer d/f/m
RWE
As a Senior SRE for the Kubernetes Platform, you will own and operate a central, automated EKS-based Kubernetes service across RWEST, enabling development teams to focus on applications. You will shape platform standards, lifecycle, security, and observability, driving reliability and ease of consumption. Join a cross-functional team tasked with evolving the platform to meet enterprise-scale needs. This role offers impact through building secure, self-service tooling and guardrails that scale across multiple squads.
Responsibilities- Operate and improve the central Amazon EKS platform, including day-2 operations, upgrades and support
- Provide a secure, self-service Kubernetes platform with automated guardrails for developers
- Contribute to platform standards across cluster config, monitoring, networking, security and lifecycle
- Maintain 24x7 operational model for Kubernetes and related services (GitHub Enterprise, Azure DevOps, JFrog, Elastic)
- Build automation, IaC and reusable platform patterns to reduce manual work
- Enhance observability through monitoring, logging, alerting and operational visibility
- Strengthen security and compliance via guardrails, governance and automated controls
- Collaborate with development squads for onboarding and platform adoption
- Contribute to incident response, RCAs and continuous platform improvement
- Document standards and establish the central SRE Platform Team as owner of shared services
- Extensive experience designing, operating and troubleshooting Kubernetes in production
- Strong understanding of cloud-native architectures and distributed systems
- Experience operating AWS and Kubernetes-based platforms at scale
- Experience designing secure, resilient, highly available production environments
- Strong knowledge of Kubernetes architecture, networking, storage, security and scheduling
- Experience managing platform lifecycle activities (upgrades, maintenance, readiness)
- Strong automation mindset with IaC and configuration as code
- Experience applying software engineering principles to infrastructure and operations
- Experience designing reusable, cross-team solutions
- Strong observability knowledge (metrics, logs, traces)
- Experience integrating security, governance and compliance
- Experience supporting production environments, incident response and continuous improvement
- Knowledge of modern software delivery and engineering automation
- Excellent analytical, troubleshooting and communication skills
- ownership mindset
- problem solving
- clear communication
- Kubernetes production operations
- AWS
- IaC (infrastructure as code)
Reference: WJ-747_31001161