AI Infra SRE: Scale Kubernetes, Linux & 24/7 Ops
Radiant
Radiant is seeking a senior Site Reliability Engineer in the UK to design, deploy, and operate scalable AI-native infrastructure. You will own Kubernetes clusters, tune Linux and I/O, and drive automation across the platform.
You will champion ITSM practices, maintain Prometheus/Grafana monitoring, and participate in 24x7 on-call support. Mentoring and cross-training with Platform SRE and HPC teams are key parts of the role.
#J-18808-LjbffrReference: WJ-766_21987450