Software Engineering Tech Lead (SRE + AI)
CISCO Systems
As Technical Leader, you’ll shape the architectural vision for an AI-powered Production Intelligence platform that elevates reliability and automation for Cisco’s collaboration SaaS across global datacenters. You will blend Site Reliability Engineering with agentic AI to improve monitoring, diagnostics, and auto-remediation. You’ll lead architecture, build AI agents and MCP integrations, and mentor engineers while partnering with cross-functional teams to deliver scalable, secure production capabilities.
Responsibilities- Define technical roadmap and architecture for AI-assisted observability, automated incident response, and self-healing cloud infrastructure
- Design and build production-grade AI agents, MCP integrations, and deterministic evaluation pipelines
- Architect ingestion and correlation pipelines for logs, metrics, traces, events, and runbooks to speed MTTD/ MTTR
- Develop proactive anomaly detection and HITL remediation workflows with safety, security, and quality guardrails
- Partner with application and infrastructure teams to define SLIs/SLOs, manage error budgets, and lead PIRs
- Mentor senior and mid-level engineers, establish engineering best practices, and align global teams
- Bachelor’s degree + 8 years of related experience, Master’s + 6 years, or PhD + 3 years in CS, Software Engineering, or related field
- Proven record as Technical Lead or Lead SRE/Software Engineer delivering distributed, high-availability SaaS platforms at scale
- Strong proficiency in Python, Go, Java, or C++ with experience in microservices, APIs, and production automation
- Deep experience with Kubernetes, Docker, and container orchestration in large-scale multi-cluster environments
- Proven background in SRE practices: SLI/SLO design, observability platforms (metrics/logs/traces), incident management, automated RCA
- Mentorship
- Cross-team collaboration
- Effective communication
- AI & Agentic Systems: LLM pipelines, AI Agents, MCP servers/clients, RAG architectures
- Observability & Telemetry: OpenTelemetry, Prometheus, Grafana, Splunk, distributed tracing
- Cloud & Infrastructure: AWS, GCP, Azure, Terraform/IaC, GitOps/CI/CD (Jenkins, GitHub Actions)
Reference: WJ-747_30143057