IT & Software

Compute Orchestration & Scheduling

Microsoft

London · Greater London · United Kingdom

Overview

In this role you will help build the compute infrastructure powering frontier-model development. You will work on a team that manages cluster orchestration, resource allocation, and fault recovery for a large accelerator fleet. You’ll partner with researchers, model engineers, and hardware teams to translate model requirements into scalable platform capabilities. You will drive improvements in scheduling efficiency, reliability, and observability at scale, shaping how AI research runs on next‑gen GPUs. This is an opportunity to impact AI training and inference for cutting‑edge research and products.

Responsibilities
  • Develop and tune the compute infrastructure to allocate NVIDIA Grace Blackwell and Vera Rubin resources to AI workloads
  • Scale GPU clusters across multiple hardware generations to thousands of accelerators
  • Use operational data to inform the compute roadmap for large-scale AI research
  • Collaborate with model-development teams to improve training and serving infrastructure
  • Identify and remove roadblocks, delivering rapid iterative improvements with users
  • Demonstrate Microsoft’s culture and values
Key requirements
  • Bachelor in Computer Science or related field with at least six years of engineering experience
  • Proficiency in languages such as C, C++, Python, Go, or JavaScript; or equivalent practical experience
  • navigate ambiguity
  • collaboration with cross-functional teams
  • fast-paced, design-driven execution
  • C, C++, Python, Go, JavaScript
  • distributed systems
  • GPU compute infrastructure

Reference: WJ-747_31001565

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.