Principal Site Reliability Engineer
Oracle Corporation
Overview
In this role you will own and operate a mission-critical stack within Oracle Cloud Infrastructure, focusing on secure, scalable virtual networking and performance. You will work with cross-functional teams in a 24/7 on-call environment to triage incidents, drive improvements, and document standard procedures. The role emphasizes security, resilience, and efficient operations, with close collaboration with global development teams. A clear hook is the opportunity to shape distributed systems and telemetry-driven improvements in a leading AI-enabled cloud platform.
Pay / Benefits
- competitive benefits
- flexible medical
- life insurance
- retirement options
- volunteer programs
- accessibility support
Responsibilities
- Share full-stack ownership of services and provide operational support in an on-call rotation
- Understand end-to-end service configuration, dependencies, and behavior; deliver secure, scalable performance
- Define and implement improvements in service architecture with global development teams
- Articulate technical characteristics to guide product delivery within Oracle Cloud
- Assess scale, capacity, security, and performance requirements of the stack
- Apply automation and orchestration principles in daily work
- Escalate complex issues and develop SOPs for unknown problems
- Explain impact of architectural decisions on distributed systems and dig into telemetry and dependencies
- Maintain high-quality incident/problem/change documentation using Jira and Confluence
- Operate in non-routine, highly complex environments within OCI Virtual Networking
- Demonstrate curiosity and deep technical understanding of services and technologies
- Collaborate to ensure documentation and knowledge sharing across teams
Key requirements
- Strong understanding of virtual networks, security, and automation
- TCP/IP stack and routing concepts in Linux; IPSEC, VPNs, and BGP
- Containerisation and orchestration experience
- Solid understanding of VCNs in public cloud environments
- CI/CD and release automation experience
- Scripting in Python or Shell
- Infrastructure automation tools like Terraform and Chef
- Leadership experience for changes and upgrades based on technical analysis
- Ability to support network segmentation
- Telemetry data manipulation using Grafana and MQL
- Experience with OCI or other major public clouds
- Experience with Jira and Confluence for incident tracking and knowledge management
- Leadership and ownership
- Problem-solving under pressure
- Collaborative and cross-functional communication
- Virtual network architecture and security
- TCP/IP, Linux routing, VPN/IPSEC/BGP
- Containerisation and orchestration
Reference: WJ-799_20847547