IT & Software

Software Engineering Tech Lead (SRE + AI)

CISCO Systems

London · Greater London · United Kingdom

Overview

As Technical Leader, you’ll shape the architectural vision for an AI-powered Production Intelligence platform that elevates reliability and automation for Cisco’s collaboration SaaS across global datacenters. You will blend Site Reliability Engineering with agentic AI to improve monitoring, diagnostics, and auto-remediation. You’ll lead architecture, build AI agents and MCP integrations, and mentor engineers while partnering with cross-functional teams to deliver scalable, secure production capabilities.

Responsibilities
  • Define technical roadmap and architecture for AI-assisted observability, automated incident response, and self-healing cloud infrastructure
  • Design and build production-grade AI agents, MCP integrations, and deterministic evaluation pipelines
  • Architect ingestion and correlation pipelines for logs, metrics, traces, events, and runbooks to speed MTTD/ MTTR
  • Develop proactive anomaly detection and HITL remediation workflows with safety, security, and quality guardrails
  • Partner with application and infrastructure teams to define SLIs/SLOs, manage error budgets, and lead PIRs
  • Mentor senior and mid-level engineers, establish engineering best practices, and align global teams
Key requirements
  • Bachelor’s degree + 8 years of related experience, Master’s + 6 years, or PhD + 3 years in CS, Software Engineering, or related field
  • Proven record as Technical Lead or Lead SRE/Software Engineer delivering distributed, high-availability SaaS platforms at scale
  • Strong proficiency in Python, Go, Java, or C++ with experience in microservices, APIs, and production automation
  • Deep experience with Kubernetes, Docker, and container orchestration in large-scale multi-cluster environments
  • Proven background in SRE practices: SLI/SLO design, observability platforms (metrics/logs/traces), incident management, automated RCA
  • Mentorship
  • Cross-team collaboration
  • Effective communication
  • AI & Agentic Systems: LLM pipelines, AI Agents, MCP servers/clients, RAG architectures
  • Observability & Telemetry: OpenTelemetry, Prometheus, Grafana, Splunk, distributed tracing
  • Cloud & Infrastructure: AWS, GCP, Azure, Terraform/IaC, GitOps/CI/CD (Jenkins, GitHub Actions)

Reference: WJ-747_30143057

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.