HPC Support Engineer
Hays
In this role you will support and enhance the institution's high-performance computing environment, including HPC and GPU clusters, research storage, high-speed networking and hybrid cloud infrastructure. You will administer Linux-based clusters, manage scientific software stacks and help researchers maximise HPC resources. You will automate deployment and monitoring using Python, Bash, Ansible, Docker, Kubernetes and CI/CD tools, while collaborating with researchers and technical teams. This is your chance to contribute to cutting-edge computational research and AI-driven projects at a prestigious, research-led university in London.
Responsibilities- Administer and support HPC clusters, GPU systems and specialist research hardware
- Manage workload scheduling with Slurm or OpenPBS
- Maintain scientific software, compilers and libraries using Spack or EasyBuild
- Automate the lifecycle of compute and GPU nodes
- Troubleshoot complex Linux, storage, networking and performance issues
- Support containerised and cloud-based research workloads
- Help researchers optimise applications for parallel computing and GPU acceleration
- Deliver technical guidance, documentation and training
- Advanced Linux cluster administration experience
- Experience with HPC schedulers such as Slurm or OpenPBS
- Knowledge of CUDA or OpenCL
- Strong scripting and automation skills using Python, Bash or Ansible
- Experience with Docker, Kubernetes, Git and CI/CD practices
- Knowledge of monitoring and observability tools such as Grafana, ELK Stack or Splunk
- Experience using HPC software build frameworks such as Spack or EasyBuild
- Knowledge of scientific software stacks, compilers and libraries
- Experience managing large-scale storage and file systems
- The ability to work directly with researchers and translate technical requirements into practical solutions
- Experience with cloud-based HPC, infrastructure-as-code tools, research computing environments or higher education would be advantageous
- collaboration
- clear communication with researchers
- problem-solving
- Advanced Linux cluster administration
- HPC schedulers: Slurm or OpenPBS
- CUDA or OpenCL
Reference: WJ-747_31001623