Staff Software Engineer, Compute Platform
GitHub
In this Staff Software Engineer role, you will help lead GitHub’s Kubernetes-based Compute Platform that supports production workloads across hundreds of teams. You’ll shape the platform’s direction, improve reliability and scalability, and reduce toil through automation and guardrails. The role emphasizes collaboration with a distributed team, strong engineering fundamentals, and hands-on production experience with Kubernetes and cloud-native systems, ideally including AKS on Azure. You will mentor engineers and own critical platform initiatives that enable safe migrations and scalable operations.
Pay / Benefits- remote-first
- competitive pay
- learning and growth opportunities
- excellent benefits
- Design, build, and operate Kubernetes-based platform systems that support production services
- Improve reliability, scalability, and operability across the Kubernetes fleet
- Lead cross-team platform initiatives including AKS migration, fleet management, control-plane reliability, and capacity management
- Define safe migration paths, operational standards, and long-term platform direction with partners
- Debug complex production issues across Kubernetes, cloud infra, Linux, containers, networking, and observability
- Build automation and guardrails for workload quality, capacity efficiency, cluster lifecycle, and incident response
- Mentor engineers through design/code reviews and technical direction
- Write, review, test, and maintain reliable platform software and automation
- Participate in on-call and incident response to drive prevention of repeat issues
- 9+ years in software engineering or related field with production software experience and multiple programming languages (C, C++, C#, Java, JavaScript, Go, Ruby, Rust, Python)
- OR_Associate’s Degree with 8+ years experience in production software
- OR_Bachelor’s Degree with 7+ years experience
- OR_Master’s Degree with 5+ years experience
- OR_Doctorate with 3+ years experience
- 2+ years experience with large-scale Kubernetes fleet operations, cluster lifecycle, capacity management, or workload migration
- Experience in production-grade cloud infrastructure and distributed systems
- strong written communication for a distributed team
- strong written communication
- mentorship and leadership
- problem solving and adaptability
- Kubernetes and cloud-native platforms
- Azure Kubernetes Service (AKS) experience
- Kubernetes controllers, operators, CRDs, admission control
Reference: WJ-747_30154319