IT & Software

Software Engineer, ML Infra, Distributed Systems – Staff, Principal

Jobtailor

Toronto · On · Canada

Design and build scalable, high throughput, and low latency distributed systems using Scala

Build reusable components and services that serve various ML applications like Personalization, Search, Ads and Exploration

Partner closely with ML engineers to understand their challenges and limitations and develop scalable solutions to address them. Proactively recommend solutions to keep our ML Inference stack state of the art.

Take a data driven approach to identifying & optimizing latency, cost, and efficiency of our infra. Lead large scale cross functional refactorings if necessary

Mentor other engineers on the team on system design, effective incident management, interviewing, leveraging LLMs for work, etc.

Collaborate with ML, Product, and cross functional engineering teams to define the long term vision and architecture for ML Infrastructure at Tubi.

Requirements

Experience designing and building scalable, distributed systems in any modern backend language (e.g., Scala, Java, Python, Go, C++); experience with Scala or JVM based language is a plus.

Strong experience with AWS or an equivalent cloud platform

Experience building online microservices at scale with low latency serving

Experience with both SQL (e.g. Postgres) and NoSQL databases (e.g. Cassandra), message brokers (e.g. Kafka), and caches (e.g. Redis)

Experience with containerization technologies, such as Docker or Kubernetes

Led the response and resolution efforts for multiple major, large-scale incidents.

#J-18808-Ljbffr

Reference: WJ-3875_12811585

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.