Lead Platform Engineer
Lorien Resourcing
As a Lead Platform Engineer, you will shape and run an MLOps platform that enables AI and data science teams to deploy and run models reliably in production. You will lead through deep technical expertise, building tooling, workflows, and ops foundations on top of Kubernetes to make the platform usable, secure, and scalable. You’ll work closely with data scientists and ML engineers to ensure production readiness and operational reliability. This role combines hands-on delivery with architectural input, offering tangible impact in a high-stakes environment.
Responsibilities- Provide technical leadership across platform, DevOps, and MLOps
- Design, build, and operate a Kubernetes-based MLOps platform supporting the full model lifecycle
- Implement and run MLOps tooling for model experimentation, packaging, versioning, and deployment
- Develop model serving and inference platforms within Kubernetes
- Collaborate with data scientists to ensure usability and alignment with real workflows
- Own platform operability, reliability, security, and supportability in production
- Troubleshoot complex issues across Kubernetes, platform services, and MLOps layers
- Contribute to architectural decisions while remaining hands-on with delivery
- Apply pragmatic engineering judgment to AI workloads and infrastructure constraints
- Senior or Lead Platform/DevOps Engineer with strong platform background
- Deep hands-on experience building and operating Kubernetes-based platforms
- Practical experience with Helm and Terraform (Infrastructure as Code)
- Experience extending Kubernetes with higher-level platforms and services
- Strong understanding of monitoring, logging, incident response, reliability, and maintenance
- Experience supporting production workloads with engineers and data scientists
- MLOps experience with tools like Kubeflow or similar
- Experience running model serving and inference platforms (e.g., KServe, vLLM)
- Notebook-based environments in secure platforms (e.g., JupyterHub)
- Exposure to emerging MLOps and AI governance tooling
- Strong collaboration with engineers and data scientists
- Technical leadership through depth and judgment
- Pragmatic problem-solving under production constraints
- Kubernetes
- Helm
- Terraform (IaC)
Reference: WJ-747_30346813