IT & Software

Systems Engineering Manager, Site Reliability Engineering, ML Compute

Google

London · Greater London · United Kingdom

Overview

In this role you lead a team of software and systems engineers to ensure uptime and study-wide availability of key services. You own end-to-end reliability, performance, and automation to prevent recurrence of issues. You mentor engineers, manage follow-the-sun on-call rotations, and design software that improves latency, scalability, and efficiency for Google's services. You will work across cross-functional teams on large-scale infrastructure and ML compute platforms to keep systems running smoothly and safely.

Responsibilities
  • Lead and mentor a team of software/systems engineers
  • Own availability and performance of key services
  • Automate response to non-exceptional service conditions
  • Manage on-call rotations across continents (follow-the-sun)
  • Design, implement and deliver software to improve reliability, scalability, latency and efficiency
  • Collaborate with cross-functional teams and stakeholders to sustain uptime and capacity planning
  • Develop and apply automation to prevent problem recurrence
  • Support ML compute infrastructure ensuring TPU/GPUs are supported and ML jobs run reliably
Key requirements
  • Bachelor's degree in Computer Science or related field or equivalent experience
  • 5 years programming experience in one or more languages
  • 3 years people management experience
  • 3 years leading projects and working with administration (filesystems, inodes, system calls) or networking (TCP/IP, routing, SDN)
  • Mentoring and leadership
  • Problem-solving mindset
  • Collaborative and cross-functional teamwork
  • Programming in one or more languages
  • Large-scale distributed systems design
  • Systems administration (filesystems, inodes, system calls) or networking (TCP/IP, routing, SDN)

Reference: WJ-747_30156677

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.