Site Reliability Engineer 2
Oracle Corporation
Overview
In this role you will help Oracle Analytics remain reliable and scalable by diagnosing issues across distributed systems, building automation, and supporting 24/7 operations. You will partner with analytics developers and other teams to reduce downtime and improve incident response. You’ll shape runbooks, dashboards, and AI-assisted tooling to enable proactive, data-driven cloud operations. This is a hands-on SRE role in a fast-paced, customer-focused environment.
Pay / Benefits- competitive benefits
- flexible medical
- life insurance
- retirement options
- volunteer programs
- accessibility accommodations
- Provide 24x7 operational support for Oracle Analytics services and release cycles
- Respond to incidents, diagnose complex production issues, and drive mitigation and post-incident actions
- Read and troubleshoot existing application/service code to identify root causes and safe remediation
- Develop deep product expertise to prevent regressions and reduce recurring incidents
- Build and maintain operational tooling, dashboards, monitoring, runbooks, and knowledge-base content
- Develop AI-assisted automation tools to improve incident investigation, monitoring, and reporting
- Create scripts/services for monitoring, telemetry collection, capacity analysis, patching, and remediation
- Analyze service health, workloads, and capacity trends to identify reliability risks
- Improve CI/CD processes, deployment practices, and operational readiness
- Collaborate with Development, Support, Product Management and other teams on issue resolution
- Share operational knowledge and continuously improve team processes
- Follow security, compliance, change-management, and operational procedures
- BS or MS in Computer Science, Engineering, or equivalent practical experience
- Experience supporting cloud infrastructure and operational processes
- Strong understanding of networking basics (DNS, HTTP/HTTPS, TLS, load balancing)
- Linux/Unix administration experience
- Experience with cloud services and large-scale distributed applications in production
- Ability to troubleshoot complex issues by reading existing code and systems
- Experience documenting runbooks and operational guides
- Experience working in agile environments
- Strong written and verbal communication, including remote team collaboration
- Ability to work independently and participate in on-call and after-hours support
- Strong communication skills
- Collaborative mindset with remote teams
- Analytical and methodical problem-solving approach
- Networking and TCP/IP fundamentals
- Linux/Unix system administration
- Python, Bash, JavaScript/Node.js
Reference: WJ-747_30867222