IT & Software

Software Engineer, Model Inference

Google

London · England · United Kingdom

Salary: £160,000 - 200,000 per year

Requirements:
  • We require a bachelors degree or equivalent practical experience.
  • We require 8 years of experience in software development.
  • We require 2 years of experience deploying and maintaining machine learning models in a live production environment.
  • We require experience profiling, configuring, or executing ML workloads directly on hardware accelerators such as GPUs or TPUs.
  • We require experience designing, building, or optimizing model serving infrastructure or inference backends.
Responsibilities:
  • We collaborate closely with research teams to understand next-generation modeling approaches and ensure they are designed and implemented with production considerations in mind.
  • We work with infrastructure teams to deliver serving infrastructure that is designed for maximum efficiency and performance, addressing bottlenecks in speed, scale, and quality.
  • We identify opportunities to automate tasks, eliminate redundancies, build performant tests, and improve the overall velocity of model releases.
  • We gain a deep understanding of serving frameworks, pre-processing pipelines, caching mechanisms, and other relevant technologies.
  • We leverage roofline analysis, hardware-level profiling, and systems analysis to identify and eliminate performance bottlenecks across ML frameworks, compilers, custom kernels, and serving infrastructure on hardware accelerators.
Technologies:
  • AI
  • Hardware
  • Machine Learning
  • Model Serving
  • 3D
  • AI Agents
  • CUDA
  • LLM
  • PyTorch

More:

At Google DeepMind, we are building the worlds first general-purpose learning agent and advancing AI to solve complex global challenges and accelerate high-quality product innovation for billions of users. We work across multiple teams and offer opportunities for both individual contributor and technical lead growth, with openings for Software Engineering and Research Engineering backgrounds. This role involves bringing AI research to life by working directly with researchers and engineers to optimize and deploy large language models such as Gemini onto Googles production infrastructure. We offer learning opportunities, varied career pathways, and benefits at Google, and the position is open to preferred working locations in London, UK or Mountain View, CA, USA.

last updated 36 week of 2026

#J-18808-Ljbffr

Reference: WJ-766_21886642

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.