Site Reliability Engineer (DevSecOps), Production Engineering - ThousandEyes
CISCO Systems
In this role, you will design and operate large-scale, cloud-based systems to ensure reliability, security, and performance of the ThousandEyes platform. You’ll collaborate with software teams across a multi-region, SaaS environment to optimize architectures, automate operations, and strengthen security. The role emphasizes incident response, scalable tooling, and leveraging cloud-native technologies to sustain a highly available service. You will work within a cross-functional team to drive reliability in a fast-paced, enterprise setting.
Responsibilities- Collaborate with software engineers to optimize architecture for availability, latency, performance, and reliability using cloud-native tools
- Design and deploy scalable operations tooling to support platform growth across regions
- Maintain elastic and resilient AWS cloud-native services
- Participate in 24x7 incident response and on-call rotation
- Leverage CNCF solutions like Kubernetes, Service Mesh, Prometheus, OpenTelemetry, and ArgoCD to improve reliability
- Automate production operations with guardrails for continuous platform operation
- Develop automation for scalable service/platform operations including deployment, scale testing, graceful failure, and chaos testing
- Stay updated on industry best practices to improve platform scalability
- Identify and address obstacles hindering operational excellence across engineering teams
- Generalize and standardize solutions for repeated success across a multi-region microservice platform
- Contribute to platform scale testing and collaboration with application teams to improve system reliability
- Manage growing infrastructure with emphasis on infrastructure as code and data handling
- Hands-on experience deploying and troubleshooting containerized workloads in production Kubernetes environments
- Experience diagnosing and administering Linux/Unix systems (process management, file systems, networking)
- Experience developing automation tooling or backend services using Python or Go
- Practical experience building hardened container images and integrating automated security scanning tools into CI/CD pipelines (SAST, DAST, container scanners)
- Excellent communication and documentation skills
- Strong sense of ownership, drive, and attention to detail
- Kubernetes in production
- Linux/Unix systems administration
- Python or Go for automation
- CI/CD with security scanning (SAST/DAST)
Reference: WJ-747_31000467