Site Reliability Engineer III
CME- Group
In this role you will engineer reliability for CME Group's Google Cloud-based platforms and middleware, enabling high-concurrency, ultra-low-latency applications. You will partner with senior engineers, mentor juniors, and drive cloud transformation and automated resilience across production systems. You’ll shape observability, incident response, and disaster recovery as part of a cross-functional team. This is a code-first, reliability-driven opportunity to impact a leading derivatives marketplace at scale.
Pay / Benefits- Bonus Programme
- Equity Programme
- Employee Stock Purchase Plan (ESPP)
- Private Medical and Dental coverage
- Mental Health Benefit Programme
- Hybrid Working
- Architect, operate, and migrate platform workloads to Google Cloud Platform, including messaging, service discovery, and data distribution components
- Design, scale, and maintain observability with OpenTelemetry, Splunk, Prometheus, Grafana; define SLIs/SLOs
- Lead incident response, own minor incidents, conduct post-mortems, and ensure rapid recovery
- Identify and reduce toil via automation and platform improvements
- Contribute to disaster recovery strategies and resilient system testing
- Lead technical discussions, present options, and mentor junior SREs across teams
- Python, Go, Java, or Bash for production-grade tooling
- Linux-based systems, distributed systems, Kubernetes/GKE, and GCP/GCE
- Terraform, Ansible, or Kubernetes Config Connector (KCC) for IaC
- Networking basics: TCP/IP, UDP, HTTP, DNS, load balancing, messaging protocols
- Awareness of Generative AI and agent-based automation (e.g., Gemini)
- Analytical problem-solving in fast-paced, high-consequence trading environment
- Strong communication and adaptability for cross-functional collaboration
- strong communication
- collaboration
- mentoring juniors
- Python
- Go
- Java
Reference: WJ-747_30173142