Senior Solution Engineer, GPU & AI Infrastructure
Hamilton Barnes ?
A high-performance neocloud provider, purpose-built for AI, HPC and cloud-native infrastructure, is looking for a Senior Solution Engineer to act as the primary technical architect on its largest AI and HPC customer engagements. This is a hands-on, customer-facing role designing state-of-the-art NVIDIA GPU clusters for training and inferencing foundation models at scale, spanning bare-metal and Kubernetes-based orchestration on the latest NVIDIA Blackwell architecture. The successful candidate will sit at the intersection of customer business need and ultra-high-performance hardware execution, owning designs from whiteboard to Bill of Materials.
You'll be architecting infrastructure for some of the largest GPU cluster deployments serving major global AI labs, working directly on the bleeding edge of NVIDIA Blackwell (B300, GB300NVL) rather than legacy hyperscaler stacks.
The business is scaling fast as a genuine alternative to the traditional hyperscalers, stripping out legacy cloud overhead to deliver bare-metal GPU performance and predictable pricing, and this role sits right at the centre of that growth story.
Fully remote with no fixed office, a four-day working week as standard, uncapped holiday and no set hours: this is flexible, autonomous working built around output rather than presenteeism.
Responsibilities:
Author comprehensive High-Level and Low-Level Design documentation for enterprise-scale GPU supercomputing clusters
Produce detailed Bills of Materials covering compute nodes, NVLink switches, network fabrics, cabling, cooling, power distribution and high-performance storage
Architect scale-up and scale-out topologies (NVLink/NVSwitch, Fat-Tree, Rail-Optimised) for NVIDIA Blackwell B300 and GB300NVL platforms
Design high-throughput, low-latency fabrics using InfiniBand and RoCE/RoCEv2, including lossless Ethernet mechanisms such as PFC, ECN and Adaptive Routing
Act as technical lead alongside sales and commercial teams on high-value AI infrastructure opportunities, engaging directly with customer CTOs, Chief AI Officers and ML engineering leads
Architect and oversee proof-of-concept deployments, benchmarking with tools such as NCCL tests, GPUDirect RDMA and MLPerf to validate real-world workload performance
Experience Required (NOT ALL ESSENTIAL):
Experience: 5+ years in Solution Architecture, Systems Engineering or Technical Pre-Sales, focused on high-performance cloud, HPC or AI infrastructure
Core Tech/Domain: Deep hands-on knowledge of NVIDIA HGX/DGX platforms, NVLink/NVSwitch fabrics and Blackwell architectures (B300, GB300NVL, GB200 NVL72/NVL36)
Methodology/Protocols: Expert-level InfiniBand (Quantum-2/Quantum-X800, Adaptive Routing) and RoCE/RoCEv2 (Spectrum-X/Spectrum-4), plus Kubernetes orchestration (NVIDIA GPU Operator, MPI Operator) and bare-metal tooling (Slurm, Ansible, Terraform)
Soft Skills: Strong technical leadership and presentation skills, able to translate complex hardware and network trade-offs for executive stakeholders; strong spoken and written English is essential
Four-day working week as standard (five-day only when attending events)
Uncapped holiday
Fully remote, flexible working with no fixed hours
Direct exposure to some of the largest GPU cluster deployments for major global AI labs
Collaborative, inclusive culture with genuine autonomy
#J-18808-LjbffrReference: WJ-766_22031865