Service Reliability Engineer - London
Fitch Ratings
In this role you will embed with Fitch Ratings development squads to ensure reliable, scalable services and modern cloud operations. You will lead SRE practices across AWS and Azure, shaping CI/CD and observability while championing AI-enabled operations. You’ll collaborate with cross-functional teams to implement guardrails, security, and resilient deployment patterns. This position offers impact across global squads and a path to influence platform strategies in a leading financial information services group.
Pay / Benefits- Hybrid work environment
- Dedicated trainings and leadership development
- Retirement planning and tuition reimbursement
- Comprehensive healthcare offerings
- Generous parental leave
- Volunteer days and charitable giving programs
- Lead delivery of reliable, scalable services for Fitch Ratings
- Guide squads on Kubernetes and modern deployment patterns
- Mentor associate engineers and establish best practices
- Collaborate with Development Squads and Operations to design service builds and DevOps tooling
- Architect and govern GitHub Actions CI/CD with quality gates, canary/blue-green strategies, and AI-assisted redeploy checks
- Own observability in Datadog with SLIs/SLOs, dashboards, alerting, and telemetry-driven automation
- Champion AI-enabled operations using AWS Bedrock/SageMaker and MCP for log analysis and incident triage
- Define cloud guardrails and security controls with IAM boundaries, OPA policies, and centralized logging
- Influence cross-functional roadmaps and lead complex release planning across CI&PE; serve as escalation point and participate in L3 on-call rotation
- Deep hands-on experience in SRE, DevOps, or Platform Engineering across AWS and Azure
- Production experience with Docker and Kubernetes
- Linux and Windows administration; IIS/.NET and Java Spring Boot applications experience
- Built and maintained CI/CD pipelines (GitHub Actions; Bamboo a plus) with DevSecOps practices
- Scripting in Python, PowerShell, or Bash
- Cloud security best practices (IAM, secrets management, container/image scanning) and core infrastructure knowledge
- APM/telemetry tooling familiarity
- Agile delivery experience and active participation in stand-ups and sprint ceremonies
- Cross-functional collaboration and stakeholder influence
- Strong problem-solving and incident triage mindset
- AWS and Azure cloud platforms
- Docker and Kubernetes in production
- Linux and Windows administration
Reference: WJ-747_30165198