Lead AI Infrastructure & Distributed Systems Engineer
LinuxRecruit
In this early-stage robotics startup, you will own the core model training infrastructure, building scalable systems to train foundation-model style AI for physical tasks. You’ll optimize GPU clusters, orchestration, and data pipelines to remove bottlenecks and accelerate research. This hands-on role sits at the crossroads of HPC, MLOps, and high-performance networking, with full architectural ownership from day one. You’ll work in person in central London as part of a small elite team, with substantial equity upside as the company scales. You drive impact by turning ambitious research into production-ready infrastructure.
Pay / Benefits- equity upside
- in-person role in central London
- founding-level team
- small elite group
- startup environment
- Design scalable distributed training infrastructure from scratch
- Eliminate bottlenecks in training pipelines and accelerate research cycles
- Optimize GPU compute, cluster scheduling, and high-performance networking
- Maintain and improve MLOps and AI infrastructure practices
- Work hands-on with AWS, Kubernetes, Slurm, and PyTorch in production environments
- Own architectural decisions and drive system reliability and performance
- Strong production background with AWS, Kubernetes, Slurm, PyTorch, and distributed training frameworks
- Deep hands-on experience with GPU compute optimisation, cluster scheduling, and high performance networking
- Experience in MLOps, AI infrastructure, or HPC
- Prior robotics experience not required; CV pipelines are advantageous
- In-person, founding-level collaborator in central London
- Equity upside potential
- high agency
- ownership and accountability
- fast-paced startup mindset
- AWS
- Kubernetes
- Slurm
Reference: WJ-747_30158183