NOC Engineer
Spectrum IT Recruitment
Overview
In this role you will help build and operate resilient cloud platforms on AWS, focusing on prevention through automation. You’ll join an engineering-led team to improve reliability and observability while tackling complex production challenges. You’ll contribute to post-incident reviews and drive continuous service improvements in a large-scale environment. This is a hands-on opportunity to shape highly available cloud services with cross-functional collaboration.
Pay / Benefits- Bonus
- Pension
- Healthcare
- Fully Remote (UK)
- Monitor and maintain highly available production platforms on AWS
- Respond to and manage 24/7 production incidents
- Investigate complex issues and restore services quickly
- Develop automation to reduce manual tasks and boost resilience
- Improve monitoring, alerting and observability across cloud environments
- Collaborate with Software, Platform, Cloud and Security Engineers to enhance reliability
- Contribute to post-incident reviews and continuous service improvements
- Support containerised workloads using Kubernetes and Docker
- Linux systems administration
- AWS cloud infrastructure
- Kubernetes and Docker
- Production support and incident management
- Python, Bash or Go scripting
- Monitoring and observability tools such as Grafana, Prometheus, Datadog, Splunk or CloudWatch
- Networking fundamentals (DNS, TCP/IP, load balancing)
- Passion for automation and continuous improvement
- Experience with Infrastructure as Code (Terraform), SRE principles (SLIs, SLOs) is beneficial
- Problem-solving mindset
- Cross-functional collaboration
- Continuous improvement orientation
- AWS
- Kubernetes
- Docker
Reference: WJ-747_30204139