Senior Site Reliability Engineer (SRE)
fortice
Overview
In this hybrid Senior Site Reliability Engineer role, you will lead a team on a Defence-focused data migration to a new cloud platform. You’ll split your time between operations and automation, building systems and dashboards in a cross-functional setting. The position emphasizes a metrics-driven approach within an agile programme, delivering reliable IT operations for a high-impact Defence initiative. This is a global, development‑friendly role with strong progression and training opportunities.
Pay / Benefits- training and certifications
- significant funding for development opportunities
- global opportunity
- career progression opportunities
- hybrid working model
- potential sponsorship for UK Government Security Clearance
- Maintain availability of critical services using monitoring and alerting
- Implement Infrastructure as Code and CI/CD pipelines with automation frameworks
- Manage AWS Console/CLI usage and integrate Prometheus & Grafana for observability
- Apply privileged access management controls
- Work with container orchestration (Docker & Kubernetes) and observability services
- Troubleshoot and debug issues across cloud environments
- Set up dashboards and performance metrics for the programme
- Experience maintaining service availability with monitoring/alerting
- Knowledge of Infrastructure as Code and CI/CD pipelines
- Proficiency with AWS Console/CLI, Prometheus & Grafana
- Experience with privileged access management processes & technologies
- Understanding of Docker and Kubernetes and related observability services
- Troubleshooting cloud environment issues
- Metrics-driven mindset
- Collaborative with agile teams
- Problem-solving and troubleshooting
- Infrastructure as Code
- CI/CD pipelines
- AWS Console/CLI
Reference: WJ-747_30161564