Senior IT Application Specialist, DevOps & AI Operations - SaaS Client
S.I. Systems Ltd.
Senior IT Application Specialist, DevOps & AI Operations - SaaS Client
Job Type: Permanent
Positions to fill: 1
Start Date: Oct 26, 2026
Job End Date: NA
Job ID:
Senior IT Application Specialist, DevOps & AI Operations - SaaS Client
Type: FTE/Permanent
Location: Oakville, ON - Hybrid if local
Overview
Our client is seeking a Senior IT Application Specialist with a strong DevOps background to manage enterprise applications and cloud platforms at scale. This hands-on role combines Kubernetes platform operations, infrastructure automation and observability with daily AI use to improve monitoring, log investigation, troubleshooting and operational efficiency.
Responsibilities
- Own the lifecycle of enterprise applications and supporting platforms, including design, deployment, ongoing support, performance tuning and upgrades.
- Manage containerized workloads using Docker, Kubernetes and Helm, with Terraform for infrastructure automation and GitOps/ArgoCD for application delivery.
- Maintain monitoring and logging using Grafana, Elasticsearch and VictoriaMetrics; automate health checks and improve incident detection and investigation.
- Use AI tools in daily operations to assist with coding, log analysis, troubleshooting, documentation and workflow automation.
- Evaluate and implement AI tools and integrations with appropriate security, testing and governance.
- Lead complex incident resolution, collaborate with business and technical teams, mentor colleagues, and maintain operational documentation and runbooks.
Must Haves
- 5-8 years of relevant experience managing production cloud platforms and/or enterprise applications, with substantial hands-on DevOps responsibility.
- Experience with Docker, Kubernetes and Helm , including deploying, configuring and troubleshooting production workloads.
- Hands-on experience with Terraform, GitOps and ArgoCD for infrastructure provisioning and application deployment.
- Experience with Grafana, Elasticsearch and VictoriaMetrics for monitoring, logging and platform health.
- Production cloud experience with Google Cloud or AWS . Strong AWS experience will be considered transferable.
- Full willingness to adopt AI in day-to-day operations , including learning and applying AI tools to improve investigation, automation and support.
- Expert-level Linux administration and strong scripting/programming skills in Python, Bash and PowerShell .
- Strong knowledge of SSO methodologies, including SAML and LDAPS, and Zero Trust security principles .
- Ability to troubleshoot complex production issues, communicate clearly with technical and business stakeholders, and take ownership through resolution.
- Availability to participate in a 24×7 rotating on-call schedule .
- Relevant post-secondary education or an equivalent combination of education and demonstrable technical experience.
Nice to Haves
- Direct experience with Google Cloud Platform and Google Kubernetes Engine (GKE) .
- Existing daily use of Claude Code, Codex, OpenCode, Gemini CLI or comparable AI tools for coding and operational work.
- Practical experience using AI for monitoring, log investigation, anomaly detection, root-cause analysis or incident response .
- Experience building workflow automations and AI integrations using LLM APIs, n8n, GCP Cloud Functions or AWS Lambda , including governed agent workflows.
Disclaimer:
AI may be used in evaluating candidates.
This posting is for an existing vacancy.
#J-18808-LjbffrReference: WJ-2742_347242