Systems Engineering Manager, Site Reliability Engineering, ML Compute
In this role you lead a team of software and systems engineers to ensure uptime and study-wide availability of key services. You own end-to-end reliability, performance, and automation to prevent recurrence of issues. You mentor engineers, manage follow-the-sun on-call rotations, and design software that improves latency, scalability, and efficiency for Google's services. You will work across cross-functional teams on large-scale infrastructure and ML compute platforms to keep systems running smoothly and safely.
Responsibilities- Lead and mentor a team of software/systems engineers
- Own availability and performance of key services
- Automate response to non-exceptional service conditions
- Manage on-call rotations across continents (follow-the-sun)
- Design, implement and deliver software to improve reliability, scalability, latency and efficiency
- Collaborate with cross-functional teams and stakeholders to sustain uptime and capacity planning
- Develop and apply automation to prevent problem recurrence
- Support ML compute infrastructure ensuring TPU/GPUs are supported and ML jobs run reliably
- Bachelor's degree in Computer Science or related field or equivalent experience
- 5 years programming experience in one or more languages
- 3 years people management experience
- 3 years leading projects and working with administration (filesystems, inodes, system calls) or networking (TCP/IP, routing, SDN)
- Mentoring and leadership
- Problem-solving mindset
- Collaborative and cross-functional teamwork
- Programming in one or more languages
- Large-scale distributed systems design
- Systems administration (filesystems, inodes, system calls) or networking (TCP/IP, routing, SDN)
Reference: WJ-747_30156677