IT & Software

GPU Systems Engineer

Hudson River Trading

London · Greater London · United Kingdom

Overview

In this role you will design, build, and optimize large-scale distributed GPU compute clusters to support HPC/AI research at HRT. You will work with cross-functional teams across compute, storage, OS, and automation to keep global research and trading operations running 24/7. You’ll profile, benchmark, and tune GPU workloads, identify bottlenecks, and automate deployment and monitoring across thousands of nodes. The role offers high impact through infrastructure projects that scale with cutting-edge technology and global data centers, shaping the performance of our AI-driven research ecosystem.

Pay / Benefits
  • base salary 150,000–300,000 USD/year
  • discretionary bonuses
  • medical/dental/vision
  • basic life insurance
  • retirement savings plans
  • sick and parental leave/Different PTO including 20 vacation days and 10 holidays
Responsibilities
  • Design, build, and optimize large-scale distributed GPU compute clusters
  • Identify and resolve GPU workload bottlenecks across compute, storage, and networking layers
  • Collaborate with research/engineering teams to profile, benchmark, and fine-tune GPU workloads
  • Automate system deployment, monitoring, and troubleshooting across thousands of nodes
  • Collaborate with teams to support evolving workloads
  • Own critical infrastructure projects from concept to implementation and support
  • Test and deploy new hardware and software; partner with vendors to resolve complex issues
Key requirements
  • 5+ years of experience in large-scale Linux systems engineering in HPC, AI or distributed infrastructure roles
  • Extensive experience in Linux system installation, performance tuning, and troubleshooting
  • Expertise in troubleshooting distributed GPU workloads
  • Deep knowledge around GPU optimization and performance
  • Proficiency in Python scripting and automation frameworks
  • CUDA or C/C++ experience is a plus
  • Experience with NVIDIA technologies beyond CUDA (NCCL, GPUDirect RDMA, NVLink)
  • Familiarity with configuration management tools (Salt, Ansible, Puppet, Chef)
  • Comfort diagnosing complex system issues at hardware, OS, and network levels
  • Strong communication and organizational skills; collaboration across diverse teams
  • Thrive in fast-paced, high-impact environments
  • 强良好的沟通能力
  • 组织能力
  • 跨团队协作
  • Linux HPC/AI系统架构
  • GPU性能优化与调优
  • GPU工作负载排错

Reference: WJ-747_30169655

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.