Sr. Site Reliability Engineer/ SWE
Visa
In this role you will support Visa’s cloud platform to empower engineers to focus on innovation by improving reliability and developer productivity. You will drive observability, automate recurring issues, and collaborate with software teams to ensure security, availability, and performance. The position combines software engineering with site reliability responsibilities, including on-call incident response. You’ll work with core DevTools to streamline workflows and scale CI/CD across a global platform, offering a meaningful impact at scale.
Responsibilities- Primary DevTools support for GitHub, Jenkins, Jira, and Artifactory; troubleshoot tool-related issues to minimize downtime
- Maintain and optimize CI/CD pipelines and integrations for reliability and scalability
- Collaborate with development teams to improve workflows and automation
- Design, implement, and maintain systems for high availability, scalability, and performance
- Monitor application reliability and lead incident response and root cause analysis
- Develop and maintain observability solutions (metrics, logging, tracing)
- Participate in on-call rotations and advocate automation and self-service to reduce operational overhead
- Document processes, troubleshooting guides, and playbooks
- Promote automation to reduce manual operational burdens
- Bachelor's degree in IT, CS or related field or 3+ years of relevant experience in IT operations and delivery
- 3-8 years in SRE and/or DevTools support roles
- Proficiency in at least one programming/scripting language (Python, Java, Go, PowerShell, JavaScript, Terraform, Ansible, Helm, etc.)
- Strong knowledge of Linux and/or Windows, distributed computing
- Experience with CI/CD tooling (Jenkins, GitHub, ArgoCD, Artifactory, Azure DevOps) in large-scale environments
- Experience with observability tools (Grafana, Prometheus, Splunk, Datadog, New Relic, DynaTrace, Sentry) in large-scale environments
- Experience supporting relational and non-relational databases (MySQL, MongoDB, PostgreSQL)
- Experience managing distributed container platforms and container infrastructure (Docker, Kubernetes)
- Hands-on cloud platform experience
- On-call support experience
- Strong problem-solving, systems thinking, and reliability mindset
- problem-solving
- systems thinking
- self-starter
- GitHub
- Jenkins
- ArgoCD
Reference: WJ-747_30139696