Senior Engineer, Platform Access
ARM
In this role you design, deploy, and operate a secure, scalable Linux VDI environment to enable global engineering teams to access high-performance computing. You will work within the Engineering Platform Access team to improve availability, security, and user experience while automating operations and enhancing observability. The role offers the chance to influence the evolution of engineering access services through self-service and policy-driven automation. You will collaborate with cross-functional teams to tackle complex infrastructure challenges and drive dependable platform performance.
Pay / Benefits- hybrid working
- accommodations during recruitment
- inclusive environment
- professional growth
- equal opportunities
- Support Linux VDI or remote visualization platforms (ETX, Citrix, OpenText) and related protocols (ThinX, HDX, PCoIP).
- Manage distributed Linux environments across on-prem, HPC, and multi-/hybrid-cloud setups.
- Develop, deploy, and improve the VDI infrastructure with focus on availability, security, performance, and user experience.
- Lead root-cause analysis, post-incident reviews, and implement automation to reduce operational toil.
- Implement and enhance monitoring, logging, and observability to improve platform dependability.
- Contribute to self-service workflows, policy-aware automation, AI-assisted operations, and secure cloud access patterns.
- Strong Linux systems experience in distributed, production environments (HPC/enterprise).
- Experience supporting Linux VDI/remote visualization platforms (ETX, Citrix, OpenText) with related protocols.
- Scripting/automation proficiency (Python, Bash, Ansible).
- Automation and IaC experience using Terraform, Ansible, Python, Bash, GitLab CI/CD, REST APIs.
- Experience with monitoring/logging/observability tools (Dynatrace, Grafana, ELK/OpenSearch, Splunk, CloudWatch).
- Familiarity with HPC workload managers (IBM Spectrum LSF) and enterprise engineering access platforms.
- Production support practices: incident response, root-cause analysis, documentation, continuous improvement.
- collaboration
- problem-solving
- clear communication
- Terraform
- Ansible
- Python
Reference: WJ-747_30949045