Senior Inference Systems & Performance Engineer
Callosum
Callosum in London is hiring for a role owning end-to-end performance for our inference platforms. You will manage KV caches, batching, memory, and scheduling across heterogeneous hardware to scale model serving.
The candidate should have deep LLM inference knowledge, strong distributed system experience, and low-level debugging skills on GPUs, networks, and Linux. Visa sponsorship and relocation are available; on-site in London.
#J-18808-LjbffrReference: WJ-766_22013347