Senior HPC Engineer
Hays
As Senior HPC Engineer you will own and evolve a Linux-based HPC platform that underpins data-intensive workloads. You will collaborate with technical users to deliver scalable, secure compute services and influence the future direction of the infrastructure. The role blends systems engineering, automation and performance tuning to improve reliability and user experience. You will work on on-premise compute environments with CPU and GPU capabilities, supported by a culture of collaboration and technical leadership.
Responsibilities- Administer and optimise Linux-based HPC infrastructure
- Manage compute, storage, networking and platform services
- Support and maintain HPC scheduling platforms, including SLURM
- Build and support containerised environments using Docker and Singularity/Apptainer
- Troubleshoot performance issues across infrastructure and user workloads
- Develop automation and monitoring solutions to improve platform reliability
- Work directly with users to understand requirements and deliver technical solutions
- Produce technical documentation, training materials and operational procedures
- Linux systems administration
- HPC cluster administration and support
- Job scheduling technologies such as SLURM
- Bash scripting and automation
- Docker, Singularity or Apptainer
- Git or similar version control tools
- Troubleshooting across compute, storage, networking and operating systems
- Collaboration
- Stakeholder engagement and communication
- Customer-focused problem solving
- GPU infrastructure and CUDA
- AI/ML platforms
- Scientific software deployment
Reference: WJ-747_30454080