Cloud Infrastructure Engineer (Open LMS) UK, Remote
Learning Technologies Group
In this role you will design, build, and operate a multi-tenant SaaS hosting platform on AWS, enabling scalable Moodle LMS deployments. You’ll own infrastructure across AWS, configuration management, and observability, collaborating with cross-functional teams to improve reliability and deployment workflows. You’ll apply deep Linux expertise and distributed systems thinking to solve complex platform challenges. This is a hands-on, architecture-influencing role with a strong focus on scale, reliability, and operational excellence.
Responsibilities- Design, build, and maintain AWS infrastructure using Terraform (EC2, RDS, S3, SQS, Lambda, ALB, ElastiCache, Route 53, VPC)
- Develop and maintain Puppet modules to manage fleets of EC2 instances across auto-scaling groups
- Extend Python-based automation and tooling supporting platform operations
- Operate and improve distributed service discovery and configuration management (etcd)
- Manage and tune multi-tier caching (Varnish, Redis/Valkey, PHP OPcache)
- Run and scale observability stack (Prometheus, Grafana, Loki, Fluentd, PagerDuty) and participate in on-call rotations
- Evaluate and implement distributed storage solutions as the platform evolves
- Improve deployment workflows and release processes
- Collaborate with internal teams on API contracts, integration patterns, and operator tooling
- Participate in incident response, root cause analysis, and platform reliability improvements
- Strong production experience with AWS services (EC2, RDS, S3, SQS, Lambda, ALB, ElastiCache, Route 53, IAM, VPC)
- Proficiency in creating and maintaining Terraform modules for production infra
- Proficiency in creating and maintaining Puppet modules (or equivalent) for fleet management
- Solid Python skills for writing/maintaining production daemons
- Deep Linux systems knowledge (Ubuntu) including Apache/Nginx, PHP-FPM, Varnish, systemd, mounts, networking
- Understanding of distributed systems concepts (consensus, leader election, etcd, eventual consistency)
- Experience building/maintaining observability pipelines (Prometheus, Grafana, Loki, Fluentd) in production
- Comfort with GitLab-based CI/CD workflows
- Clear communicator for architectural decisions and tradeoffs
- Clear communicator
- documentation mindset
- cross-functional collaboration
- Terraform
- Puppet (or equivalent)
- Python
Reference: WJ-747_30179303