IT & Software

Senior Solution Engineer – GPU and AI Infrastructure (CAF673F)

Referment

London · England · United Kingdom

Referment is working with a cloud infrastructure provider that supports enterprise AI, high-performance computing and cloud-native workloads. The business is looking for a senior technical specialist to shape large-scale GPU platforms for customers across the UK.

The Role

  • Design enterprise GPU clusters and produce clear high-level and low-level designs, rack and network diagrams, and detailed bills of materials.
  • Shape scale-up and scale-out architectures across NVIDIA GPU platforms, NVLink and NVSwitch, high-speed InfiniBand and RoCE fabrics, storage, power and cooling.
  • Translate customer workload, performance and commercial requirements into practical bare-metal, Slurm or Kubernetes-based solutions.
  • Lead technical discovery sessions, architecture workshops and executive presentations, and contribute to complex proposals and RFP responses.
  • Plan and oversee proofs of concept, using appropriate benchmarking to validate throughput, latency and distributed training performance.
  • Work closely with enterprise customers, hardware vendors and internal engineering teams, feeding technical insight back into the platform roadmap.

What We’re Looking For

  • At least five years in solution architecture, systems engineering or technical pre-sales focused on HPC, AI infrastructure or high-performance cloud platforms.
  • Deep knowledge of NVIDIA HGX or DGX systems, GPU interconnects and modern rack-scale GPU architectures.
  • Strong experience designing InfiniBand and/or RoCE networks, including congestion management and GPU-direct technologies.
  • Practical understanding of GPU workload orchestration through Kubernetes and associated NVIDIA operators, or bare-metal environments using technologies such as Slurm, Ansible and Terraform.
  • A track record of producing robust HLDs, LLDs, network diagrams and itemised infrastructure designs.
  • Confident communication with CTOs, infrastructure leaders and ML engineers, with the judgement to explain hardware, networking and cost trade-offs.
  • Awareness of high-density data-centre power, cooling and storage considerations.

Relevant Desirable Experience

NVIDIA AI infrastructure, networking or InfiniBand certifications would be useful, as would experience with performance testing for distributed AI workloads. A relevant degree is welcome, although equivalent practical experience is equally valuable.

This could suit a Senior Solutions Architect, HPC Systems Engineer or technical pre-sales specialist who has designed GPU clusters and wants to remain close to both customers and engineering. You must be based in the UK.

#Referment

#J-18808-Ljbffr

Reference: WJ-766_22344016

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.