Staff Software Engineer (MetalDev)
CoreWeave
- As a Staff Software Engineer within our Compute Architecture organization, you will help build the software systems that operate the backbone of our large-scale GPU data centers
- The METALDEV team builds Go-based distributed services that bring new infrastructure online, manage hardware lifecycle workflows, monitor production health, and automate safe operations across fleets of GPU servers and rack-scale systems
- This is a software-first role at the intersection of distributed systems, production reliability, and hardware-aware automation, where your work directly improves the reliability, safety, and scalability of real-world infrastructure.\
- Design, build, and operate Go-based services that manage the lifecycle of large-scale GPU data center infrastructure
- Build automation for data center bring-up, hardware discovery, health monitoring, remediation, and production operations
- Develop reliable APIs, services, and workflows for managing BMCs, firmware state, server health, and rack-level infrastructure
- Improve observability, alerting, and operational tooling so production issues can be detected, understood, and resolved quickly
- Translate incidents and hardware failure modes into software improvements that make the platform more resilient
- Partner with hardware-adjacent, infrastructure, operations, and software teams to design systems that work safely at fleet scale
- Provide technical leadership through design reviews, code reviews, architectural guidance, and mentorship
- Make pragmatic architecture decisions that balance reliability, simplicity, scalability, and operational burden
Skilled in applying a data-driven approach to reliability, optimization, and continuous improvementB.S., M.S., or PhD in Computer Science or related field, or equivalent experience8+ years of software engineering experience with a strong focus on infrastructure, cloud engineering, and distributed databases—particularly within large-scale datacenter and cloud environmentsExpertise in Go and proven experience building REST/gRPC APIs for mission-critical platformsExcellent communicator able to work effectively with both technical and non-technical stakeholdersTrack record of leading incident response, postmortems, and driving robust service reliabilityProven success in mentoring engineers, leading technical projects, and influencing engineering strategy across teamsHands-on experience with observability stacks (Prometheus, Grafana, PromQL), CI/CD pipelines, and operating large fleets of GPU serversStrong background in architecting and scaling cloud-native Kubernetes infrastructure and distributed servicesExperience contributing to and collaborating with open source communitiesWorking knowledge of Kafka, ClickHouse and CRDBDMTF, RedFish APIs, and GPU servers
#J-18808-LjbffrReference: WJ-766_22197820