Senior HPC Systems Administrator
University of Oxford
As Senior HPC Systems Administrator you will lead the design, deployment, and ongoing evolution of the labs HPC infrastructure to support world-leading AI research. You will collaborate with researchers, software engineers, industry partners, and IT teams to deliver scalable, secure compute, storage, and networking. The role combines hands-on administration with strategic planning and mentoring of researchers and students. Join the BOLD Lab initiative to shape cutting-edge AI research through robust, high-performance computing at Oxford.
Responsibilities- Design, build, and maintain HPC clusters and related infrastructure
- Manage Linux compute and storage environments
- Support GPU-enabled research computing systems
- Maintain high-performance storage and backup solutions
- Monitor performance, security, and availability
- Support and mentor researchers, software engineers, and postgraduate students
- Develop documentation, training materials, and best practices
- Collaborate with departmental and university-wide IT teams
- Evaluate and deploy new technologies to meet evolving research needs
- Extensive experience managing HPC infrastructure in research or technical environments
- Strong Linux systems administration expertise
- Experience with GPU servers and high-performance networking technologies
- Scripting and automation skills (e.g., Bash, Python)
- Knowledge of storage systems and backup/archive procedures
- Strong troubleshooting and systems integration abilities
- Excellent written and verbal communication skills
- Ability to work independently and collaboratively in a research environment
- clear communication
- collaboration
- problem-solving
- GPU computing environments
- Linux HPC administration
- Scripting and automation (Bash, Python)
Reference: WJ-747_30933506