Software Engineer (Compute Efficiency), London
Isomorphic Labs
Overview
In this role you will help scale IsoLabs’ planet-wide AI accelerator fleet by designing robust observability and recovery systems. You’ll work with ML platforms and researchers to maximize throughput and efficiency across training and inference. The role sits at the intersection of biology, machine learning, and infrastructure, enabling faster drug discovery at digital speed. You will be part of a collaborative culture solving high-impact problems at scale.
Pay / Benefits- hybrid working
- Design, deploy, and scale observability systems and telemetry pipelines across distributed clusters to monitor compute efficiency and hardware health
- Improve node reliability and integrate new hardware to leverage advancements across the accelerator fleet
- Identify compute waste and bottlenecks, partnering with ML and platform teams to optimize utilization and throughput
- Collaborate with AI org and ML Infrastructure to instrument and improve canonical efficiency metrics for training and inference
- Contribute to reliability improvements of ML runs across development, staging, and production
- Operate, maintain, and harden cloud infrastructure and cluster deployments (R&D and production)
- Partner with diverse teams including science, research, product, business development and operations
- Contribute to core technical decisions on tooling, infrastructure, and architectural design
- Real-world experience operating, monitoring, and debugging infrastructure for large-scale AI/ML workloads
- Experience in cloud compute infrastructure design, preferably GCP
- Strong programming skills
- Significant experience deploying in Kubernetes at scale
- Familiarity with Nvidia GPU generations
- Proven track record building production observability and telemetry stacks
- Cross-functional collaboration
- Problem-solving drive
- Proactive communication
- GCP
- Kubernetes
- Nvidia GPU generations
Reference: WJ-747_30994400