IT & Software

ML Runtime Engineer: Scale Inference Engines

Fractile

West Of England · England · United Kingdom

Fractile is hiring a ML Runtime Engineer to integrate our AI accelerators with the latest inference frameworks and to build a scalable runtime stack. You will tackle challenges such as KV cache management and multi-user inference, focusing on transformer architectures in a collaborative environment.

The role requires deep experience in ML inference at scale, knowledge of paged attention and vLLM, and strong software engineering skills to maintain robust, high-performance systems.

#J-18808-Ljbffr

Reference: WJ-766_21508948

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.