GPU Systems Engineer
Hudson River Trading
In this role you will design, build, and optimize large-scale distributed GPU compute clusters to support HPC/AI research at HRT. You will work with cross-functional teams across compute, storage, OS, and automation to keep global research and trading operations running 24/7. You’ll profile, benchmark, and tune GPU workloads, identify bottlenecks, and automate deployment and monitoring across thousands of nodes. The role offers high impact through infrastructure projects that scale with cutting-edge technology and global data centers, shaping the performance of our AI-driven research ecosystem.
Pay / Benefits- base salary 150,000–300,000 USD/year
- discretionary bonuses
- medical/dental/vision
- basic life insurance
- retirement savings plans
- sick and parental leave/Different PTO including 20 vacation days and 10 holidays
- Design, build, and optimize large-scale distributed GPU compute clusters
- Identify and resolve GPU workload bottlenecks across compute, storage, and networking layers
- Collaborate with research/engineering teams to profile, benchmark, and fine-tune GPU workloads
- Automate system deployment, monitoring, and troubleshooting across thousands of nodes
- Collaborate with teams to support evolving workloads
- Own critical infrastructure projects from concept to implementation and support
- Test and deploy new hardware and software; partner with vendors to resolve complex issues
- 5+ years of experience in large-scale Linux systems engineering in HPC, AI or distributed infrastructure roles
- Extensive experience in Linux system installation, performance tuning, and troubleshooting
- Expertise in troubleshooting distributed GPU workloads
- Deep knowledge around GPU optimization and performance
- Proficiency in Python scripting and automation frameworks
- CUDA or C/C++ experience is a plus
- Experience with NVIDIA technologies beyond CUDA (NCCL, GPUDirect RDMA, NVLink)
- Familiarity with configuration management tools (Salt, Ansible, Puppet, Chef)
- Comfort diagnosing complex system issues at hardware, OS, and network levels
- Strong communication and organizational skills; collaboration across diverse teams
- Thrive in fast-paced, high-impact environments
- 强良好的沟通能力
- 组织能力
- 跨团队协作
- Linux HPC/AI系统架构
- GPU性能优化与调优
- GPU工作负载排错
Reference: WJ-747_30169655