Site Reliability Engineer 2
Oracle Corporation
In this role you will support and own operational reliability of Oracle's cloud networking stack, focusing on Linux-based troubleshooting and real-time incident response. You will work with the Virtual Networking team in a 24/7 on-call context to diagnose, resolve, and document incidents, while building automation scripts to improve live troubleshooting and service resilience. You will collaborate with global teams to design scalable, secure network solutions across VCNs, VPNs, and load balancing, shaping the delivery of Oracle Cloud services. This is a hands-on, high-impact position that combines network engineering, scripting, and incident management to sustain critical customer services.
Pay / Benefits- flexible medical
- life insurance
- retirement options
- volunteer programs
- disability accommodation support
- inclusive workplace
- On-call operational support for production services and incident triage
- End-to-end performance and operability accountability of the service stack
- Collaborate with global development teams to improve service architecture
- Articulate service characteristics to guide engineering of Oracle Cloud capabilities
- Understand scale, security, capacity, and performance requirements of services
- Escalate complex issues and develop SOPs; document incidents, changes, and procedures in Jira/Confluence
- Maintain technical documentation and contribute to post-incident reviews
- Work within OCI Virtual Networking specialization to support non-routine, high-complexity tasks
- Linux troubleshooting and system understanding
- Python and Bash scripting for live troubleshooting and automation
- Experience with remote access technologies (Ethernet, VPN, Load Balancing, BGP)
- Strong TCP/IP knowledge and networking fundamentals
- Containerisation and orchestration experience
- Understanding of Virtual Cloud Networks (VCNs) in public cloud environments
- Familiarity with CI/CD and release automation tools
- Experience with infrastructure automation tools (Terraform, Chef) is a plus
- Telemetry data manipulation and visualization (Grafana, MQL)
- Experience with public cloud providers (OCI or equivalent)
- Familiarity with Jira and Confluence for incident tracking and documentation
- Collaborative mindset
- Under pressure problem-solving
- Professional curiosity
- Linux system processes, memory, disk and log management
- TCP/IP stack and routing concepts
- BGP, VPN, and load balancing
Reference: WJ-747_30172720