IT & Software

Staff Infrastructure Engineer, Cluster Infrastructure

Humanloop

London · Greater London · United Kingdom

Overview

In this role you will own the technical direction for agent-driven cluster lifecycle management and scale compute to support Anthropic’s growing AI research and safety experiments. You will partner with teams across security, cloud providers, and research to ingest capacity on time and deliver secure-by-default, high-bandwidth inter-connectivity. You’ll shape long-term compute, data, and infrastructure strategy while promoting operational excellence and mentoring engineers. This is a chance to influence infrastructure at hyperscale, enabling reliable AI development and safety work.

Pay / Benefits
  • Competitive compensation and benefits
  • Optional equity donation matching
  • Generous vacation and parental leave
  • Flexible working hours
  • Office space for collaboration
  • Visa sponsorship available
Responsibilities
  • Own the strategy and roadmap for cluster lifecycle management (provisioning, updates, decommissioning)
  • Coordinate with teams to ingest new compute capacity on schedule
  • Align physical build-out with cloud-based connectivity for high bandwidth inter-cluster links
  • Collaborate with security to ensure secure-by-default provisioning
  • Define and drive scalability, homogeneity, and fault tolerance of clusters
  • Work with cloud providers and internal teams to shape long-term compute and infra strategy
  • Establish operational-excellence practices: incident response, postmortems, on-call health
  • Mentor and coach engineers in the team
Key requirements
  • Deep expertise in distributed systems, reliability, and cloud platforms (Kubernetes, IaC, AWS/GCP/Azure)
  • Proficiency in at least one systems language (Rust, Go, or Python) and Terraform IaC
  • Track record leading complex, multi-quarter technical initiatives across teams
  • Ability to align senior stakeholders and communicate effectively at all levels
  • Strategic thinking
  • Cross-functional collaboration
  • Clear written and verbal communication
  • Kubernetes internals
  • Cluster provisioning and management systems
  • Cluster orchestration (Mesos, Borg-like)

Reference: WJ-747_30169883

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.