Senior SRE - AI Inference Platform (GPU/Kubernetes)
Jobgether
Jobgether is seeking a Senior Site Reliability Engineer for a tokens-based inference platform in the UK. You will own reliability, performance, and observability for large-scale GPU-heavy workloads, designing telemetry, monitoring, and automation to ensure high availability.
You'll collaborate with software and infra teams to build self-healing systems, improve SLOs, and drive incident response with robust post-mortems.
#J-18808-LjbffrReference: WJ-766_22268023