IT & Software

Staff Software Engineer, Kubernetes Platform

Humanloop

London · Greater London · United Kingdom

Overview

In this role you will own and extend Anthropic's Kubernetes scheduler and control plane to support vast accelerator fleets and multi-cloud deployments. You will shape core cluster services and build operators that enable reliable, scalable ML workloads at scale. You’ll partner with research, training, and inference teams to translate workload needs into platform capabilities and work with cloud providers to align features. This is a hands-on position that combines systems engineering with incident response, aimed at keeping frontier AI training reliable and performant.

Pay / Benefits
  • competitive compensation and benefits
  • optional equity donation matching
  • generous vacation and parental leave
  • flexible working hours
  • lovely office space
Responsibilities
  • Own, operate, and extend the Kubernetes scheduler for accelerator fleets, including custom scheduling plugins and policies for gang scheduling, topology awareness, and preemption
  • Scale the Kubernetes control plane (apiserver, etcd, controller-manager) to support clusters far beyond typical limits, and identify bottlenecks early
  • Design, build, and operate core cluster services such as service discovery that every workload in the fleet depends on
  • Build and maintain custom controllers, operators, and CRDs
  • Partner with research, training, and inference to understand workload shapes and turn their requirements into platform capabilities
  • Collaborate with cloud providers on required features and escalations
  • Participate in on-call, lead incident response, and design processes (postmortems, runbooks, SLOs) that help the team avoid repeating failures
Key requirements
  • Significant software engineering experience building and operating production distributed systems
  • Proficiency in at least one systems-appropriate language (e.g., Go, Python, Rust, or C++)
  • Deep, hands-on Kubernetes experience (well beyond "user of”) into scheduler, controllers, apiserver, or operating large multi-tenant clusters
  • Demonstrated ability to debug complex issues across the stack, from API behavior down to node and network-level root causes
  • A track record of designing for reliability, correctness, and clear failure semantics in systems other engineers depend on
  • Strong written and verbal communication; comfort building consensus with internal stakeholders
  • Strong written and verbal communication
  • ability to build consensus with internal stakeholders
  • problem solving under ambiguity
  • Kubernetes internals (scheduler, apiserver, etcd, controllers)
  • Custom controllers, operators, and CRDs
  • Cluster scaling and multi-tenant deployments

Reference: WJ-747_30145537

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.