IT & Software

Senior Linux Administrator

Riverlane

Cambridge · Cambridgeshire · United Kingdom

Overview

In this role you will own and grow Riverlane’s core HPC and Linux infrastructure to support cutting-edge quantum error correction research. You will lead day-to-day HPC cluster administration while shaping long-term plans for a scalable, fault-tolerant environment. You’ll work closely with the Infrastructure Team to ensure reliability, performance, and security for high-demand compute workloads. This is a chance to impact foundational software and hardware integration in a fast-paced, mission-driven company.

Pay / Benefits
  • annual bonus plan
  • private medical insurance
  • life insurance
  • contributory pension scheme
  • equity
  • 28 days annual leave
Responsibilities
  • Administer and maintain the HPC cluster (compute nodes, storage, networking)
  • Evolve cluster topology, node configuration and resource utilisation as the estate grows
  • Deploy and troubleshoot HPC tooling, especially Slurm (queues and scheduling policies)
  • Monitor performance and ensure high availability using observability tools (Prometheus)
  • Manage network file storage for throughput and latency
  • Patch, support and troubleshoot core Linux infrastructure (RHEL)
  • Escalation point for complex issues; support HPC users
  • Maintain Synopsys EDA tooling including licence management and performance tuning
  • Improve self-service capabilities for users (e.g., password resets, VNC sessions)
  • Automate routine tasks with Bash, Python, Ansible; plan improvements with Infrastructure Team
  • Own work packages, meet sprint goals; build and maintain system documentation
  • Ensure security, hardening, lifecycle management and disaster recovery readiness
  • Maintain off-site backup strategy and data retention compliance
  • Provide stakeholder updates on progress and improvements
Key requirements
  • Extensive experience administering infrastructure in fast-paced engineering environments
  • Enterprise Linux expertise (RHEL) including hardening and lifecycle management
  • Hands-on with Slurm or similar job schedulers in distributed compute environments
  • Strong understanding of network file storage and workload optimization
  • Proficiency with containers and virtualization infrastructure
  • Automation using Ansible, Python, Bash; infrastructure-as-code practices
  • Experience with backup/recovery strategies and DR planning
  • Experience managing licensed EDA tools (e.g., Synopsys) is a strong plus
  • Excellent communication skills for technical and non-technical users
  • strong communication
  • supportive with users
  • team collaboration
  • Slurm scheduling
  • RHEL administration and hardening
  • HPC cluster architecture

Reference: WJ-747_30990275

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.