HPC Infrastructure Engineer – Linux / Cluster
Intellectual Capital Resources
Overview
In this role, you will take ownership of an on-premise HPC/compute environment, assess current infrastructure, and implement improvements to reliability, manageability, monitoring, and automation. You will work hands-on to identify weaknesses and leave the cluster robust and easy to support. You’ll collaborate with stakeholders to plan upgrades and ensure smooth handovers. This is a practical, impact-focused opportunity to enhance HPC performance and resilience.
Responsibilities- Take technical ownership of an existing on-premise HPC/compute environment
- Assess current infrastructure and identify weaknesses
- Implement improvements across reliability, manageability, monitoring and automation
- Perform cluster administration, troubleshooting and remediation
- Enhance resilience and reliability of the HPC environment
- Develop technical documentation and perform knowledge transfer
- Support infrastructure upgrades, migrations or remediation programmes
- Collaborate with cross-functional teams to improve monitoring and automation
- Linux HPC / compute infrastructure
- HPC cluster administration
- Linux server and node management
- Cluster troubleshooting and remediation
- Infrastructure monitoring and management
- Networking and storage within compute environments
- Infrastructure automation
- Improving resilience and reliability
- Technical documentation
- Knowledge transfer
- Experience with workload schedulers, infrastructure-as-code, configuration management, high-performance networking or distributed storage (beneficial)
- Hands-on/problem-solving orientation
- Clear communication
- Ability to transfer knowledge
- Linux
- HPC cluster administration
- Infrastructure automation
Reference: WJ-747_31000472