IT & Software

HPC Operations Lead

LinuxRecruit

London · Greater London · United Kingdom

Overview

As HPC Operations Lead, you will steer the operational performance of a large-scale HPC and storage environment within a world-leading research institute. You’ll lead a specialist team, coordinate service delivery, and bridge technical and scientific users to enable discovery. The role shapes long-term technology direction while ensuring systems are robust, accessible, and continuously improving. You’ll work at the intersection of technology and discovery, contributing to research outcomes at scale.

Responsibilities
  • Own and improve the operational performance of the HPC and storage environment
  • Lead and manage a specialist team and coordinate across technical and scientific stakeholders
  • Oversee incident management, service performance, and continuous improvement of services
  • Influence long-term technology direction and strategy for research computing platforms
  • Engage with researchers to translate needs into usable platforms
  • Design and operate high-performance storage services for internal and external workloads
  • Work with large-scale Linux-based systems, Slurm workload scheduler, Infiniband networking, and GPFS storage
  • Ensure accessibility and usability of complex infrastructure across teams
Key requirements
  • Proven leadership experience in operations of complex, large-scale computing or storage services
  • Strong operational awareness with ability to manage competing priorities and limited resources
  • Ability to collaborate across teams and communicate technical concepts to researchers
  • Experience with HPC environments and storage technologies at petabyte scale (preferred)
  • Familiarity with Linux-based systems, workload schedulers (e.g., Slurm), high-performance networking and parallel file systems
  • Understanding of automation, data centre environments or networking
  • Collaborative mindset
  • Clear communication with non-technical stakeholders
  • Problem-solving under pressure
  • HPC clusters and environments
  • Linux system administration
  • Slurm workload manager

Reference: WJ-747_30147835

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.