CUDA Engineer
Fuse Energy Supply
As a CUDA Engineer at Fuse Energy, you will design and optimize low-level GPU kernels that power transformer inference workloads in a fast-growing renewable energy and AI compute environment. You will push throughput and latency, tune memory hierarchies, and implement quantisation-aware, mixed-precision kernels across cutting-edge GPUs. Your work directly supports energy-efficient AI inference at scale, helping Fuse meet rapid compute demands while aligning with the company’s mission to accelerate clean energy via advanced technology. You will collaborate with cross-functional teams to improve performance, reliability, and efficiency of our inference pipelines.
Pay / Benefits- Competitive salary and an equity sign-on bonus
- Biannual bonus scheme
- Fully expensed tech to match your needs
- Breakfast and dinner allowance for office based employees
- Write and optimize custom CUDA kernels for core transformer inference operations
- Profile kernels to identify bottlenecks in occupancy, memory throughput, and warp divergence
- Apply kernel fusion to reduce memory round-trips and launch overhead across inference pipelines
- Optimize memory access patterns and manage the memory hierarchy for maximum bandwidth utilization
- Implement quantisation-aware kernels and mixed-precision arithmetic to reduce latency and memory footprint
- Build and tune caching mechanisms for efficient autoregressive decoding
- Tune kernel launch configurations for target GPU architectures
- Benchmark kernels against baselines and drive throughput/latency improvements
- Write tests for CUDA code to catch performance and correctness regressions
- Maintain internal CUDA libraries and contribute to team coding standards and documentation
- 4+ years writing production CUDA code with a track record of shipping performance-critical kernels
- Deep understanding of GPU microarchitecture, warps, occupancy, register pressure, and memory hierarchy
- Strong CUDA C++ skills, including streams and asynchronous execution
- Hands-on experience profiling to diagnose compute-bound vs memory-bound bottlenecks
- Experience with kernel fusion, memory coalescing, and avoiding warp divergence
- Experience writing quantised and mixed-precision kernels
- Solid grasp of parallel algorithm design and numerical precision tradeoffs
- CUDA C++
- GPU profiling
- kernel fusion
- memory hierarchy optimization
- quantisation-aware kernels
- mixed-precision arithmetic
Reference: WJ-747_30135841