Senior IT Systems Administrator (Linux & HPC)
KBR
Overview
In this role you will own the Linux and HPC platform stack, delivering secure, reliable enterprise and research-computing environments. You will manage Linux systems, HPC clusters, and workload scheduling while ensuring performance and resilience at scale. You’ll work with cross-functional teams to diagnose cross-domain issues and drive continuous improvement. The role offers hands-on technical ownership and operates across on-premise and cloud-enabled HPC platforms.
Pay / Benefits- competitive benefits
- professional development
- Administer, patch, harden and upgrade Linux server platforms (Red Hat Enterprise Linux or equivalent)
- Manage core Linux services (identity, SSH, DNS, time, repos, filesystems, logging, scheduled tasks)
- Automate tasks with shell scripts and configuration-management/orchestration tools
- Monitor performance, availability and security; resolve cross-system issues
- Maintain build standards, documentation and runbooks
- Operate and support HPC clusters (SLURM: queues, partitions, policies, jobs, accounting)
- Support NVIDIA Base Command Manager and Azure CycleCloud for cluster provisioning and lifecycle management
- Collaborate with engineers and users to diagnose job, compiler and MPI issues; plan maintenance activities
- Administer hardware lifecycle, Cisco compute platforms, and NetApp storage integration
- Troubleshoot networking, storage connectivity and end-to-end dependencies
- Apply secure configuration, backup/DR recovery, incident response and change management processes
- Provide technical guidance and knowledge transfer to colleagues and service-desk teams
- Hands-on Linux administration experience in complex enterprise or research environments
- Experience supporting HPC clusters and diagnosing cross-layer issues
- Strong SLURM administration and workload troubleshooting
- Bash or similar scripting for task automation
- Experience with Linux performance, patching, security hardening and vulnerability remediation
- Experience with shared file services (NFS, permissions, throughput)
- Networking fundamentals and distributed-system dependencies
- Experience delivering controlled technical change and maintaining documentation in ITSM
- Bachelor's degree in computing, engineering or related field or equivalent
- Desirable: experience with NVIDIA Base Command Manager, Bright Cluster Manager, NetApp, InfiniBand, MPI workloads, Ansible, Git, virtualization/containers
- Desirable: certifications in Linux, HPC, Cisco, NVIDIA (advantageous)
- Independent working style
- Strong communication to specialists and non-specialists
- Attention to detail and operational discipline
- Linux administration (RHEL or equivalent)
- HPC cluster management
- SLURM workload scheduler
Reference: WJ-747_30300026