IT & Software

Senior ML Engineer: Inference & Latency Optimization

Nebius Group

London · England · United Kingdom

Nebius is seeking a Senior Machine Learning Engineer to own model and endpoint optimization from artifacts to production deployment on our Applied AI team. You will improve latency, throughput, memory efficiency, GPU utilization, and cost per token while maintaining model quality.

This hands-on role involves diagnosing complex serving problems, evaluating configurations, and delivering measurable production improvements in collaboration with kernel and platform engineers.

#J-18808-Ljbffr

Reference: WJ-766_22381625

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.