Site Reliability Engineer
Trainline
Overview
Join Trainline as a mid-level Site Reliability Engineer within ReliabilityOps. You’ll help keep a cloud-first platform observable, scalable, and resilient, partnering with product teams to enable safe delivery. You’ll handle incident response, post-incident reviews, and tooling to improve MTTR and SLO adherence. This role offers hands-on architecture influence and shared ownership in a fast-growing fintech-like travel platform.
Pay / Benefits- private healthcare & dental insurance
- work from abroad policy
- 2-for-1 share purchase plans
- EV Scheme to reduce carbon emissions
- extra festive time off
- family-friendly benefits
- Develop understanding of system architecture, dependencies, and failure modes across the platform
- Participate in production incident response, investigations, mitigation, and service restoration
- Contribute to post-incident reviews and follow-up actions for reliability, scalability, and resilience
- Take part in on-call rotation
- Design, build, and maintain observability using metrics, logs, events, and traces
- Improve monitoring and alerting aligned with business impact to reduce noise and MTTD
- Surface operational data quickly during live incidents
- Make informed tooling and technology choices using SRE principles
- Support AWS-hosted infrastructure and shared platform services using IaC and CI/CD tooling
- Collaborate with product engineering to ensure services are deployment-ready and operationally safe
- Advise on reliability and resilience practices
- Write and maintain reliable, well-structured code and scripts
- Prioritise work effectively using agile processes
- Contribute to broader reliability discussions and ongoing improvements
- SRE concepts such as SLI, SLO and error budgets
- Hands-on observability tooling experience (New Relic, ELK, Influx, Grafana)
- Experience with AWS or similar cloud providers
- Troubleshooting Linux operating systems
- Scripting in at least one language (preferably Python)
- Understanding of load balancing, reverse proxy concepts, upstream config/health checks
- Knowledge of application architecture concepts (threading, queues, readiness/health checks, circuit breakers, backoff, throttling)
- Experience with time series data management (retention, cardinality, moving averages)
- Experience with GitHub Actions and Terraform
- Strong collaboration and agile workflow
- collaboration
- growth mindset
- ability to challenge and be challenged
- SRE fundamentals (SLI/SLO, error budgets)
- Observability tooling: New Relic, Elastic/ELK, Grafana, Influx
- AWS and cloud-native infrastructure
Reference: WJ-747_30182297