Senior SRE: Build Fault-Tolerant AI Cloud Infra
Nebius
Nebius is seeking an experienced Site Reliability Engineer to maintain and grow our AI cloud platform. You will work on a massive monorepo, modify CI/CD tooling, and help shape metrics to capture user problems and reduce friction.
Ideal candidates blend SRE and software engineering skills across Java/Kotlin, Go, Python, and Ruby, with a deep understanding of Unix and the JVM. You will thrive in a fast‑moving, ownership‑driven environment and impact real AI workloads.
#J-18808-LjbffrReference: WJ-5107_14301832