IT & Software

Staff Software Engineer (MetalDev)

CoreWeave

York And North Yorkshire · England · United Kingdom

  • As a Staff Software Engineer within our Compute Architecture organization, you will help build the software systems that operate the backbone of our large-scale GPU data centers
  • The METALDEV team builds Go-based distributed services that bring new infrastructure online, manage hardware lifecycle workflows, monitor production health, and automate safe operations across fleets of GPU servers and rack-scale systems
  • This is a software-first role at the intersection of distributed systems, production reliability, and hardware-aware automation, where your work directly improves the reliability, safety, and scalability of real-world infrastructure.\
  • Design, build, and operate Go-based services that manage the lifecycle of large-scale GPU data center infrastructure
  • Build automation for data center bring-up, hardware discovery, health monitoring, remediation, and production operations
  • Develop reliable APIs, services, and workflows for managing BMCs, firmware state, server health, and rack-level infrastructure
  • Improve observability, alerting, and operational tooling so production issues can be detected, understood, and resolved quickly
  • Translate incidents and hardware failure modes into software improvements that make the platform more resilient
  • Partner with hardware-adjacent, infrastructure, operations, and software teams to design systems that work safely at fleet scale
  • Provide technical leadership through design reviews, code reviews, architectural guidance, and mentorship
  • Make pragmatic architecture decisions that balance reliability, simplicity, scalability, and operational burden

Skilled in applying a data-driven approach to reliability, optimization, and continuous improvementB.S., M.S., or PhD in Computer Science or related field, or equivalent experience8+ years of software engineering experience with a strong focus on infrastructure, cloud engineering, and distributed databases—particularly within large-scale datacenter and cloud environmentsExpertise in Go and proven experience building REST/gRPC APIs for mission-critical platformsExcellent communicator able to work effectively with both technical and non-technical stakeholdersTrack record of leading incident response, postmortems, and driving robust service reliabilityProven success in mentoring engineers, leading technical projects, and influencing engineering strategy across teamsHands-on experience with observability stacks (Prometheus, Grafana, PromQL), CI/CD pipelines, and operating large fleets of GPU serversStrong background in architecting and scaling cloud-native Kubernetes infrastructure and distributed servicesExperience contributing to and collaborating with open source communitiesWorking knowledge of Kafka, ClickHouse and CRDBDMTF, RedFish APIs, and GPU servers

#J-18808-Ljbffr

Reference: WJ-766_22197820

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.