Senior Linux Administrator
Riverlane
In this role you will own and grow Riverlane’s core HPC and Linux infrastructure to support cutting-edge quantum error correction research. You will lead day-to-day HPC cluster administration while shaping long-term plans for a scalable, fault-tolerant environment. You’ll work closely with the Infrastructure Team to ensure reliability, performance, and security for high-demand compute workloads. This is a chance to impact foundational software and hardware integration in a fast-paced, mission-driven company.
Pay / Benefits- annual bonus plan
- private medical insurance
- life insurance
- contributory pension scheme
- equity
- 28 days annual leave
- Administer and maintain the HPC cluster (compute nodes, storage, networking)
- Evolve cluster topology, node configuration and resource utilisation as the estate grows
- Deploy and troubleshoot HPC tooling, especially Slurm (queues and scheduling policies)
- Monitor performance and ensure high availability using observability tools (Prometheus)
- Manage network file storage for throughput and latency
- Patch, support and troubleshoot core Linux infrastructure (RHEL)
- Escalation point for complex issues; support HPC users
- Maintain Synopsys EDA tooling including licence management and performance tuning
- Improve self-service capabilities for users (e.g., password resets, VNC sessions)
- Automate routine tasks with Bash, Python, Ansible; plan improvements with Infrastructure Team
- Own work packages, meet sprint goals; build and maintain system documentation
- Ensure security, hardening, lifecycle management and disaster recovery readiness
- Maintain off-site backup strategy and data retention compliance
- Provide stakeholder updates on progress and improvements
- Extensive experience administering infrastructure in fast-paced engineering environments
- Enterprise Linux expertise (RHEL) including hardening and lifecycle management
- Hands-on with Slurm or similar job schedulers in distributed compute environments
- Strong understanding of network file storage and workload optimization
- Proficiency with containers and virtualization infrastructure
- Automation using Ansible, Python, Bash; infrastructure-as-code practices
- Experience with backup/recovery strategies and DR planning
- Experience managing licensed EDA tools (e.g., Synopsys) is a strong plus
- Excellent communication skills for technical and non-technical users
- strong communication
- supportive with users
- team collaboration
- Slurm scheduling
- RHEL administration and hardening
- HPC cluster architecture
Reference: WJ-747_30990275