Senior Software Engineer
Intellectual Capital Resources
In this role you deploy, extend, and optimize a leading open-source LLM inference engine across scaling and regulated-environment initiatives. You will work with cross-functional teams to improve performance, accuracy, and explainability, while contributing upstream to OSS communities. You help shape how advanced models are served and monitored in production, with a focus on scalable, safe AI infrastructure. This is a hands-on, impact-driven position with a clear emphasis on OSS collaboration and hardware-aware optimization.
Responsibilities- Deploy, instrument, and monitor open-weight models served via the inference engine
- Build new engine features to support novel hardware architectures
- Extend the engine for accuracy, explainability, and accountability in regulated settings
- Contribute upstream to the relevant open-source community
- Optimize inference performance (latency, throughput, hardware utilisation)
- Senior-level Python + low-level programming (C/C++, Rust, CUDA)
- Direct contributions to a major inference/serving engine or similar ML infra project (e.g. PyTorch, Ray)
- Strong grasp of LLM inference internals (KV caching, batching, quantisation)
- Experience upstreaming to active OSS communities
- Hardware accelerator optimisation experience (GPU/TPU/etc.)
- Genuine enthusiasm for AI-assisted development (LLMs/agents as core tools)
- Python
- C/C++
- Rust
- CUDA
- LLM inference internals (KV caching, batching, quantisation)
- Hardware accelerator optimisation (GPU/TPU)
Reference: WJ-747_30139399