ML Runtime Engineer: Scale Inference Engines
Fractile
Fractile is hiring a ML Runtime Engineer to integrate our AI accelerators with the latest inference frameworks and to build a scalable runtime stack. You will tackle challenges such as KV cache management and multi-user inference, focusing on transformer architectures in a collaborative environment.
The role requires deep experience in ML inference at scale, knowledge of paged attention and vLLM, and strong software engineering skills to maintain robust, high-performance systems.
#J-18808-LjbffrReference: WJ-766_21508948