IT & Software

Principal Core Infrastructure Engineer

Oracle Corporation

Remote · Nationwide · United Kingdom

Overview

In this role you will lead development and architect scalable distributed systems, setting elasticity and performance targets. You will define scalability requirements, optimize data paths, and leverage data plane platforms for large-scale data operations. You will design fault-tolerant, in-service-upgradable components with robust failover and rate-limiting strategies, and establish KPIs, telemetry, dashboards, and proactive alerts. You will diagnose production issues, mentor peers, and drive security, compliance, and IaC-driven automation within change-management plans. This role anchors Oracle's AI- and cloud-powered solutions that impact billions of lives.

Pay / Benefits
  • competitive benefits
  • flexible medical
  • life insurance
  • retirement options
  • volunteer programs
Responsibilities
  • Lead development and implementation of scalable distributed systems with horizontal and vertical scaling
  • Optimize code and systems for large-scale data processing and high-throughput requirements
  • Define scalability requirements and ensure design/implementation meets them
  • Design elastic systems that scale up/down
  • Leverage data plane platforms for large-scale retrieval, storage, and processing
  • Design fault-tolerant components with redundancy, replication, and automatic failover
  • Handle network partitions with trade-offs among consistency, availability, and partition tolerance
  • Implement load-shedding, throttling, and rate-limiting mechanisms
  • Define KPIs and telemetry; build dashboards and alerting for system health
  • Design validation scenarios (fault injection, brownouts) and data replication/synchronization
  • Troubleshoot, mentor peers, and maintain operational readiness
  • Implement security controls and remediation; ensure regulatory compliance
  • Develop automation and IaC; enable safe patching, updates, and rollbacks in change management
  • Plan and execute moderately complex tasks and coordinate multiple projects
  • Collaborate across organization to align on expectations
  • Participate in incident management and root-cause analysis
  • Contribute to talent development through interviews and mentoring
Key requirements
  • Experience designing scalable distributed systems with elastic scaling
  • Strong knowledge of fault tolerance, replication and failover
  • Experience with load-shedding, throttling, rate-limiting
  • Proficiency in defining KPIs, telemetry, dashboards, and alerting
  • Familiarity with data replication and synchronization techniques
  • Operational troubleshooting across production environments
  • Security and compliance in multi-tenant cloud environments
  • Infrastructure as Code (IaC) and automation
  • Change management for patching, updates, and rollbacks
  • Mentoring and leadership in engineering contexts
  • Collaborative mindset
  • Problem solving and debugging under pressure
  • Mentoring and coaching
  • Distributed systems design
  • System scalability and elasticity
  • System reliability design

Reference: WJ-747_30158284

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.