IT & Software

HPC Network Engineer (Private Cloud)

Whitehall Resources

Cambridge · England · United Kingdom

Whitehall Resources require an HPC Network Engineer (Private Cloud) to work with a key client on a 6 month initial contract.

*Inside IR35. *2 days per week on site in Cambridge.

Job Overview:

This is a fantastic opportunity to join a newly formed Private Cloud team. We are responsible for the design, deployment and operations of a new greenfield on-prem private cloud based on OpenStack.

The Private Cloud HPC Network Engineer will help design, build and evolve the high-performance network architecture underpinning our Private Cloud platform that will be used for HPC workloads.

The role combines deep data-centre and HPC networking expertise with modern cloud networking, software-defined infrastructure and automation. As part of the Private Cloud team, you will collaborate closely with compute, storage, platform engineering teams to deliver a scalable, resilient and highly automated cloud infrastructure.

Responsibilities:

  • * Design and evolve network architecture, standards and patterns for large-scale Private Cloud Openstack and HPC environments. Develop scalable Leaf-Spine architectures using technologies including BGP, EVPN, VXLAN and ECMP.
  • * Design networking for high-performance compute, storage and latency-sensitive engineering workloads. Evaluate technologies including SR-IOV, SmartNICs, DPUs and high-performance Ethernet.
  • * Integrate physical network infrastructure with OpenStack networking services including Neutron, OVN and ML2. Define network architecture for virtual machines, bare-metal workloads and Kubernetes platforms.
  • * Perform performance analysis, network optimisation and lead technical evaluation and benchmarking of networking technologies and vendors.
  • * Lead network security architecture and threat modelling while also leading complex cross-domain issues spanning compute, storage, networking and HPC applications.

Required Skills and Experience:

  • * Strong experience designing and supporting high-performance or large-scale compute environments using RoCE, RDMA or Infiniband in a Cisco NX-OS and/or Arista EOS .
  • * Deep understanding of BGP routing and network architectures, Leaf-Spine network, Layer 2 and Layer 3, ECMP, EVPN and VXLAN and network resilience. With exposure to newer technologies such as SmartNICs, DPUs & NVMe-over-Fabrics..
  • * Practical experience of OpenStack network services and developing integrated network automation and API driven interfaces.
  • * Strong HPC network observability, troubleshooting and performance tuning skills.

“Nice To Have” Skills and Experience:

  • * Strong Linux knowledge and Kubernetes networking.
  • * Exposure to SR-IOV and DPDK.
  • * Demonstrable hands-on Python, Ansible and Terraform experience.
  • * Using SRE principles in the creation of SLI/SLO and error budgets

#J-18808-Ljbffr

Reference: WJ-766_22171675

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.