IT & Software

Backend Engineer — Model Inference

Best AI Tools Wiki

Toronto · On · Canada

Cohere is hiring a Backend Engineer to optimize model inference systems. You will build low-latency serving infrastructure, implement batching strategies, and develop the backend services that power Cohere's enterprise API. Requirements

5+ years of backend engineering experience Strong proficiency in Go, Rust, or C++ Experience with high-throughput, low-latency systems Knowledge of model serving and inference optimization Experience with gRPC, REST APIs, and microservices Nice to Have

Experience with vLLM, TensorRT-LLM, or similar Familiarity with GPU memory management Experience with continuous batching techniques Equity in a well-funded AI startup Premium health and dental Remote-first culture Home office budget Annual learning stipend Flexible PTO Skills

Go Rust Model Serving Backend Engineering gRPC Inference Optimization This listing was posted on and is almost certainly closed. It is kept for reference — check the employer’s own careers page for current openings. Vincony has all 400+ AI models in one place — compare responses, AI debate, Image/Video/Voice generator, and 20 more tools to help you learn and build with AI.

#J-18808-Ljbffr

Reference: WJ-3875_13222894

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.