Senior Software Engineer (Server Fleet Infrastructure)
CoreWeave
- At CoreWeave, infrastructure isn’t just a foundation, it’s a product
- We build scalable, high-performance computing systems that power the largest AI workloads in the world
- We’re looking for Engineers that thrive at the intersection of software and systems, deploying and managing large scale bare metal compute
- Within this domain, you’ll design and build software that manages complex infrastructure across globally distributed datacenters
- Working in Go, Python/Ansible, deep in Linux environments, observability/monitoring stacks, and leveraging technologies like gRPC and Kubernetes CRs/Controllers/Operators
- Whether you’re automating bare metal, building fleet lifecycle management services, solving multi-layer integration challenges, or observing our globally distributed fleet, your work will be critical to the company’s delivery of reliable and efficient infrastructure
- You’ll join a high-performing, communicative, and supportive team, solving complex problems at hyperscale, deploying purpose built AI across the world
- Design and implement solutions to problems of scale for multi-site deployment and management of CoreWeave’s global server hardware fleet
- Build and maintain backend services and APIs (gRPC/REST) in Go or Python to interact with Kubernetes and other infrastructure systems
- Develop provisioning services, automation workflows, and fleet management tools that span from bare metal to container orchestration
- Write and maintain Kubernetes custom controllers and operators to automate infrastructure behavior
- Design and implement observability solutions for large-scale server monitoring to improve system stability and insight
- Adapt and extend open source tooling to enhance visibility into system metrics, performance, and health
- Create test plans, deployment automation, dashboards, alerts, and insights into our fleet operations
- Resolve integration challenges across the entire infrastructure stack, from data center hardware to orchestration platforms
- Participate in an on-call rotation
- Familiarity with CI/CD tools like Argo, Flux, and GitHub Actions
- Strong understanding of Linux internals
- Proficiency in Go and/or Python software development
- 5+ years of experience in software or infrastructure engineering
- Experience designing, implementing, and monitoring Kubernetes operators for custom resource definitions
- Experience with infrastructure automation and configuration management tools like Ansible, Puppet, Chef, Salt
- Experience with distributed cloud computing principles, including testing strategies, observability, error budgets, and fault-tolerant design
- Experience implementing metrics pipelines, custom alerts, and monitoring strategies
- Ability to break down complex problems into achievable tasks and collaborate with teammates to execute them
- Willingness and ability to thrive in a fast-paced startup environment
Reference: WJ-766_22195708