Staff ML Engineer - Developer Tools
ARM
As Staff ML Engineer in Arm’s Developer Platforms, you will help build production-grade ML tooling that analyzes, optimizes, and prepares models for efficient execution on Arm-based hardware. You’ll shape the architecture and roadmap for ML developer tooling and ensure a seamless experience across ML frameworks, runtimes, and hardware. You work with distributed teams to deliver intuitive tools for edge developers and drive high-quality software through automated testing and strong engineering processes. This role offers the chance to influence how developers ship performant ML models on Arm from day one, in a fast-paced, collaborative environment.
Responsibilities- Define, design and deliver edge-focused model analysis and optimization tooling via web services and desktop apps
- Provide technical leadership for ML tooling, shaping architecture and roadmap
- Collaborate across ML frameworks, runtimes and Arm hardware to create intuitive developer experiences
- Build relationships with distributed teams, product managers and UX specialists to align tooling with developer needs
- Ensure engineering quality through maintainable code, automated tests, CI, processes and customer feedback
- Solid understanding of neural-network architecture and model execution including computation graphs, operator dispatch, memory planning, and heterogeneous hardware execution
- Proven experience in hardware-aware model optimization techniques (quantisation, graph optimisation, operator fusion, precision reduction)
- Experience with multiple ML frameworks and runtimes (PyTorch, ONNX/ONNX Runtime, ExecuTorch, TensorFlow/LiteRT, OpenVINO)
- Experience analyzing, profiling and debugging ML workloads (latency, memory, model size, accuracy)
- Passion for building high-quality, developer-oriented tooling and workflows
- Strong collaboration across distributed teams
- Customer-focused mindset and ability to translate developer needs into tooling
- Proactive communication and stakeholder management
- Neural-network architecture and model execution concepts (graphs, dispatch, memory planning)
- Hardware-aware optimization (quantisation, graph optimisation, operator fusion, precision reduction)
- Experience across ML frameworks/runtimes: PyTorch, ONNX/ONNX Runtime, ExecuTorch, TensorFlow/LiteRT, OpenVINO
Reference: WJ-747_30174775