IT & Software

Senior IT Systems Administrator (Linux & HPC)

KBR

Surrey · Surrey · United Kingdom

Overview

In this role you will own the Linux and HPC platform stack, delivering secure, reliable enterprise and research-computing environments. You will manage Linux systems, HPC clusters, and workload scheduling while ensuring performance and resilience at scale. You’ll work with cross-functional teams to diagnose cross-domain issues and drive continuous improvement. The role offers hands-on technical ownership and operates across on-premise and cloud-enabled HPC platforms.

Pay / Benefits
  • competitive benefits
  • professional development
Responsibilities
  • Administer, patch, harden and upgrade Linux server platforms (Red Hat Enterprise Linux or equivalent)
  • Manage core Linux services (identity, SSH, DNS, time, repos, filesystems, logging, scheduled tasks)
  • Automate tasks with shell scripts and configuration-management/orchestration tools
  • Monitor performance, availability and security; resolve cross-system issues
  • Maintain build standards, documentation and runbooks
  • Operate and support HPC clusters (SLURM: queues, partitions, policies, jobs, accounting)
  • Support NVIDIA Base Command Manager and Azure CycleCloud for cluster provisioning and lifecycle management
  • Collaborate with engineers and users to diagnose job, compiler and MPI issues; plan maintenance activities
  • Administer hardware lifecycle, Cisco compute platforms, and NetApp storage integration
  • Troubleshoot networking, storage connectivity and end-to-end dependencies
  • Apply secure configuration, backup/DR recovery, incident response and change management processes
  • Provide technical guidance and knowledge transfer to colleagues and service-desk teams
Key requirements
  • Hands-on Linux administration experience in complex enterprise or research environments
  • Experience supporting HPC clusters and diagnosing cross-layer issues
  • Strong SLURM administration and workload troubleshooting
  • Bash or similar scripting for task automation
  • Experience with Linux performance, patching, security hardening and vulnerability remediation
  • Experience with shared file services (NFS, permissions, throughput)
  • Networking fundamentals and distributed-system dependencies
  • Experience delivering controlled technical change and maintaining documentation in ITSM
  • Bachelor's degree in computing, engineering or related field or equivalent
  • Desirable: experience with NVIDIA Base Command Manager, Bright Cluster Manager, NetApp, InfiniBand, MPI workloads, Ansible, Git, virtualization/containers
  • Desirable: certifications in Linux, HPC, Cisco, NVIDIA (advantageous)
  • Independent working style
  • Strong communication to specialists and non-specialists
  • Attention to detail and operational discipline
  • Linux administration (RHEL or equivalent)
  • HPC cluster management
  • SLURM workload scheduler

Reference: WJ-747_30300026

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.