IT & Software

Site Reliability Engineer (DevSecOps), Production Engineering - ThousandEyes

CISCO Systems

London · Greater London · United Kingdom

Overview

In this role, you will design and operate large-scale, cloud-based systems to ensure reliability, security, and performance of the ThousandEyes platform. You’ll collaborate with software teams across a multi-region, SaaS environment to optimize architectures, automate operations, and strengthen security. The role emphasizes incident response, scalable tooling, and leveraging cloud-native technologies to sustain a highly available service. You will work within a cross-functional team to drive reliability in a fast-paced, enterprise setting.

Responsibilities
  • Collaborate with software engineers to optimize architecture for availability, latency, performance, and reliability using cloud-native tools
  • Design and deploy scalable operations tooling to support platform growth across regions
  • Maintain elastic and resilient AWS cloud-native services
  • Participate in 24x7 incident response and on-call rotation
  • Leverage CNCF solutions like Kubernetes, Service Mesh, Prometheus, OpenTelemetry, and ArgoCD to improve reliability
  • Automate production operations with guardrails for continuous platform operation
  • Develop automation for scalable service/platform operations including deployment, scale testing, graceful failure, and chaos testing
  • Stay updated on industry best practices to improve platform scalability
  • Identify and address obstacles hindering operational excellence across engineering teams
  • Generalize and standardize solutions for repeated success across a multi-region microservice platform
  • Contribute to platform scale testing and collaboration with application teams to improve system reliability
  • Manage growing infrastructure with emphasis on infrastructure as code and data handling
Key requirements
  • Hands-on experience deploying and troubleshooting containerized workloads in production Kubernetes environments
  • Experience diagnosing and administering Linux/Unix systems (process management, file systems, networking)
  • Experience developing automation tooling or backend services using Python or Go
  • Practical experience building hardened container images and integrating automated security scanning tools into CI/CD pipelines (SAST, DAST, container scanners)
  • Excellent communication and documentation skills
  • Strong sense of ownership, drive, and attention to detail
  • Kubernetes in production
  • Linux/Unix systems administration
  • Python or Go for automation
  • CI/CD with security scanning (SAST/DAST)

Reference: WJ-747_31000467

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.