Site Reliability Engineering Professional
BT Group
As a Site Reliability Engineering Professional you help manage BT International’s global core platforms to ensure reliability, security, and scalability. You will operate within the 2nd Line Operations team, supporting critical services across international networks and digital platforms. You’ll drive incident resolution, continuous improvement, and automation to enhance service availability. This role offers the opportunity to shape platform reliability through observability, SRE practices, and collaboration with engineering and product teams. You will work in a global, fast-paced environment focused on delivering exceptional customer experience.
Pay / Benefits- 10% on target annual bonus
- online private GP 24/7 for you and your immediate family
- paid carers leave – up to 2 weeks
- maternity, paternity, and adoption leave – 18 weeks full pay and 8 weeks half pay
- discounted EE and BT products
- pension scheme – 5% from you and 10% from us
- Support 24x7 operation of core network and platform services with high availability and performance
- Monitor proactively to identify risks and prevent customer-impacting incidents
- Lead technical restoration during major incidents and act as escalation point for complex issues
- Meet operational KPIs, SLAs, OLAs, and customer targets
- Drive continuous improvement via root cause analysis, problem management, automation, and defect reduction
- Collaborate with Engineering and Product to improve reliability and readiness of platforms
- Create and maintain runbooks, service maps, playbooks, and handover processes
- Enhance monitoring, observability, and operational tooling
- Support automation and CI/CD adoption to boost efficiency
- Coach colleagues and customer-facing teams to maintain service excellence
- Collaborate with global operational teams, suppliers, and stakeholders to deliver outcomes
- Experience in 24x7 Operations, NOC, Service Operations, or SRE environment
- Strong knowledge of IP/Optical networking and WAN tech (MPLS, BGP, OSPF, Ethernet, SD-WAN, Internet, Cloud connectivity)
- Experience troubleshooting end-to-end across network, platform, cloud, and application domains
- Understanding of SRE principles (reliability, availability, automation, observability, resilience)
- Hands-on with monitoring/observability tools (Dynatrace, Splunk, Grafana, ELK, Prometheus, or similar)
- Familiarity with Incident, Problem, Change, Major Incident Management and RCA
- Analytical and problem-solving skills; ability to decide under pressure
- Strong communication and stakeholder management; customer-focused
- Ability to work across global, cross-functional teams in a fast-paced environment
- Passion for continuous improvement, innovation, and operational transformation
- strong communication
- stakeholder management
- collaboration
- Dynatrace
- Splunk
- Grafana
Reference: WJ-747_30181780