Backend Engineer — Model Inference
Best AI Tools Wiki
Cohere is hiring a Backend Engineer to optimize model inference systems. You will build low-latency serving infrastructure, implement batching strategies, and develop the backend services that power Cohere's enterprise API.
Requirements
- 5+ years of backend engineering experience
- Strong proficiency in Go, Rust, or C++
- Experience with high-throughput, low-latency systems
- Knowledge of model serving and inference optimization
- Experience with gRPC, REST APIs, and microservices
Nice to Have
- Experience with vLLM, TensorRT-LLM, or similar
- Familiarity with GPU memory management
- Experience with continuous batching techniques
Premium health and dental
Remote-first culture
Home office budget
Annual learning stipend
Flexible PTO
Skills
Go Rust Model Serving Backend Engineering gRPC Inference Optimization
This listing was posted on and is almost certainly closed. It is kept for reference — check the employer’s own careers page for current openings.
Vincony has all 400+ AI models in one place — compare responses, AI debate, Image/Video/Voice generator, and 20 more tools to help you learn and build with AI.
#J-18808-LjbffrReference: WJ-4483_1288108