Staff Software Engineer, Inference
Humanloop
In this role you will help operate Claude’s inference stack at scale, enabling reliable and high-performance model serving for millions of users. You’ll own end-to-end infrastructure work—from intelligent request routing to fleet-wide orchestration across accelerators—balancing compute efficiency with research acceleration. You’ll work with a cross-functional team to remove blockers, optimize performance, and deploy multi-accelerator deployments. This is a chance to influence scalable AI infrastructure at a company focused on safe, beneficial AI.
Pay / Benefits- competitive compensation and benefits
- optional equity donation matching
- generous vacation and parental leave
- flexible working hours
- office space for collaboration
- Identify and tackle infrastructure blockers to enable large-scale model serving
- Drive performance optimization across distributed systems and inference pipelines
- Implement and optimize load balancing, request routing, and traffic management
- Support multi-accelerator deployments and cloud infrastructure in AWS/GCP
- Build and maintain production-grade deployment pipelines for new models and features
- Collaborate with research teams to enable high-performance inference for evolving architectures
- Analyze observability data to tune performance in production
- Manage multi-region deployments and geographic routing across global users
- Significant software engineering experience with distributed systems
- Familiarity with performance optimization and large-scale service orchestration
- Experience with LLM inference optimization, batching, and caching (encouraged)
- Experience with load balancing, request routing, or traffic management systems
- Knowledge of Kubernetes and cloud infrastructure (AWS, GCP)
- Proficiency in Python or Rust
- Ability to work across academia and industry with impact-focused mindset
- Interest in machine learning systems and infrastructure
- Care about societal impacts of AI
- results-oriented
- flexibility and impact
- strong collaboration and communication
- LLM inference optimization
- batching strategies
- multi-accelerator deployments
Reference: WJ-747_30135434