Site Reliability Engineer – 11863CF
Proactive Appointments
As a Site Reliability Engineer, you will ensure the stability and performance of critical systems on a contract basis. You will work on-site in Gloucester four days a week, supporting DV-cleared environments and cloud-based services. You will leverage modern configuration management, container orchestration, and CI/CD pipelines to keep systems reliable and scalable. You will collaborate with cross-functional teams to integrate monitoring, security, and deployment practices that meet stringent UK government standards. This role offers hands-on responsibility and the chance to shape reliability at scale in a security-conscious setting.
Responsibilities- Maintain and improve system reliability and performance across environments
- Manage configuration with tools like Ansible or Chef
- Orchestrate containers using Kubernetes/OpenShift/Docker Swarm
- Develop and sustain CI/CD pipelines (e.g., Jenkins)
- Monitor systems with InfluxDB, Prometheus, Grafana and respond to incidents
- Integrate MQ messaging (RabbitMQ or equivalent) and ensure data flows
- Work with relational databases and SQL to support services
- Administer Linux systems and shell scripting
- Implement security best practices across cloud hosting (AWS EC2/RDS/S3/Lambda)
- DV cleared (live)
- Experience with Terraform
- Container orchestration experience
- CI/CD tooling proficiency
- Monitoring and observability tooling experience
- MQ messaging integration
- Relational databases and SQL knowledge
- Linux administration and scripting
- Cloud hosting experience (AWS)
- Network security protocol understanding
- collaborative
- proactive
- clear communication
- Ansible
- Chef
- Terraform
Reference: WJ-747_30472642