Senior Software Engineer, ML Infrastructure
Roku
As a Senior Software Engineer, you will own and optimize production ML infrastructure across ranking, model delivery, and LLM-agent behavior to deliver reliable conversational experiences at scale. You’ll collaborate with ML, product, data, and platform partners to improve training, evaluation, deployment, and observability, while managing latency and cost. You’ll help shape the agent-based Roku Entertainment Assistant and build systems that scale with user demand. This role offers hands-on architecture, ML lifecycle decisions, and a tangible impact on millions of TV streamers worldwide.
Pay / Benefits- global mental health and financial wellness resources
- healthcare (medical, dental, vision)
- life, disability and accident insurance
- retirement options (401(k)/pension)
- commuter benefits
- time off and local leave policies
- Own fulfilment-ranking pipelines from training through deployment and online validation
- Build and operate an offline-evaluation platform combining LLM-as-judge harnesses with deterministic checks
- Develop the Roku Entertainment Assistant's LLM agent, including tool routing, retrieval, guardrails and answer caching
- Design caching as a latency and cost lever for high-volume services
- Build observability for ML and LLM systems, including latency attribution, quality metrics, tracing and per-request cost
- Improve reliability and operability of distributed systems and diagnose cross-service failures
- Drive accountable AI-assisted engineering practices across the team and release workflows
- Collaborate across ML, product, data and platform teams to turn ambiguous problems into measurable outcomes
- Strong production software-engineering experience with distributed services
- Experience with ML infrastructure or production model delivery and demonstrated impact
- Practical understanding of the ML lifecycle: training, evaluation, versioning, deployment, online validation, monitoring, safe rollback
- Experience designing for latency, resilience, observability and cost in production
- Experience with LLMs, embeddings, semantic search, retrieval, or agent systems and how to evaluate quality and manage failure modes
- Experience with caching systems and low-latency trade-offs
- Commitment to automation, CI/CD, code quality, and evidence-led engineering
- Ability to communicate trade-offs across disciplines and work with cross-functional partners
- Degree in Computer Science, Electrical Engineering or a related field, or equivalent practical experience
- Experience with AWS or GCP, Kubernetes, SQL warehouses and batch-processing technologies such as Trino, Presto or Spark, Kafka or other streaming systems; Go or Python; Node.js is a plus
- Collaborative approach
- Ability to explain technical trade-offs to non-technical partners
- Curiosity and adaptability
- ML infrastructure
- LLMs, embeddings, semantic search, retrieval
- Agent systems and tool routing
Reference: WJ-747_30204226