IT & Software

Senior Software Engineer (Server Fleet Infrastructure)

CoreWeave

York And North Yorkshire · England · United Kingdom

  • At CoreWeave, infrastructure isn’t just a foundation, it’s a product
  • We build scalable, high-performance computing systems that power the largest AI workloads in the world
  • We’re looking for Engineers that thrive at the intersection of software and systems, deploying and managing large scale bare metal compute
  • Within this domain, you’ll design and build software that manages complex infrastructure across globally distributed datacenters
  • Working in Go, Python/Ansible, deep in Linux environments, observability/monitoring stacks, and leveraging technologies like gRPC and Kubernetes CRs/Controllers/Operators
  • Whether you’re automating bare metal, building fleet lifecycle management services, solving multi-layer integration challenges, or observing our globally distributed fleet, your work will be critical to the company’s delivery of reliable and efficient infrastructure
  • You’ll join a high-performing, communicative, and supportive team, solving complex problems at hyperscale, deploying purpose built AI across the world
  • Design and implement solutions to problems of scale for multi-site deployment and management of CoreWeave’s global server hardware fleet
  • Build and maintain backend services and APIs (gRPC/REST) in Go or Python to interact with Kubernetes and other infrastructure systems
  • Develop provisioning services, automation workflows, and fleet management tools that span from bare metal to container orchestration
  • Write and maintain Kubernetes custom controllers and operators to automate infrastructure behavior
  • Design and implement observability solutions for large-scale server monitoring to improve system stability and insight
  • Adapt and extend open source tooling to enhance visibility into system metrics, performance, and health
  • Create test plans, deployment automation, dashboards, alerts, and insights into our fleet operations
  • Resolve integration challenges across the entire infrastructure stack, from data center hardware to orchestration platforms
  • Participate in an on-call rotation
  • Familiarity with CI/CD tools like Argo, Flux, and GitHub Actions
  • Strong understanding of Linux internals
  • Proficiency in Go and/or Python software development
  • 5+ years of experience in software or infrastructure engineering
  • Experience designing, implementing, and monitoring Kubernetes operators for custom resource definitions
  • Experience with infrastructure automation and configuration management tools like Ansible, Puppet, Chef, Salt
  • Experience with distributed cloud computing principles, including testing strategies, observability, error budgets, and fault-tolerant design
  • Experience implementing metrics pipelines, custom alerts, and monitoring strategies
  • Ability to break down complex problems into achievable tasks and collaborate with teammates to execute them
  • Willingness and ability to thrive in a fast-paced startup environment

#J-18808-Ljbffr

Reference: WJ-766_22195708

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.