IT & Software

Site Reliability Engineer

Trainline

London · Greater London · United Kingdom

Overview

Join Trainline as a mid-level Site Reliability Engineer within ReliabilityOps. You’ll help keep a cloud-first platform observable, scalable, and resilient, partnering with product teams to enable safe delivery. You’ll handle incident response, post-incident reviews, and tooling to improve MTTR and SLO adherence. This role offers hands-on architecture influence and shared ownership in a fast-growing fintech-like travel platform.

Pay / Benefits
  • private healthcare & dental insurance
  • work from abroad policy
  • 2-for-1 share purchase plans
  • EV Scheme to reduce carbon emissions
  • extra festive time off
  • family-friendly benefits
Responsibilities
  • Develop understanding of system architecture, dependencies, and failure modes across the platform
  • Participate in production incident response, investigations, mitigation, and service restoration
  • Contribute to post-incident reviews and follow-up actions for reliability, scalability, and resilience
  • Take part in on-call rotation
  • Design, build, and maintain observability using metrics, logs, events, and traces
  • Improve monitoring and alerting aligned with business impact to reduce noise and MTTD
  • Surface operational data quickly during live incidents
  • Make informed tooling and technology choices using SRE principles
  • Support AWS-hosted infrastructure and shared platform services using IaC and CI/CD tooling
  • Collaborate with product engineering to ensure services are deployment-ready and operationally safe
  • Advise on reliability and resilience practices
  • Write and maintain reliable, well-structured code and scripts
  • Prioritise work effectively using agile processes
  • Contribute to broader reliability discussions and ongoing improvements
Key requirements
  • SRE concepts such as SLI, SLO and error budgets
  • Hands-on observability tooling experience (New Relic, ELK, Influx, Grafana)
  • Experience with AWS or similar cloud providers
  • Troubleshooting Linux operating systems
  • Scripting in at least one language (preferably Python)
  • Understanding of load balancing, reverse proxy concepts, upstream config/health checks
  • Knowledge of application architecture concepts (threading, queues, readiness/health checks, circuit breakers, backoff, throttling)
  • Experience with time series data management (retention, cardinality, moving averages)
  • Experience with GitHub Actions and Terraform
  • Strong collaboration and agile workflow
  • collaboration
  • growth mindset
  • ability to challenge and be challenged
  • SRE fundamentals (SLI/SLO, error budgets)
  • Observability tooling: New Relic, Elastic/ELK, Grafana, Influx
  • AWS and cloud-native infrastructure

Reference: WJ-747_30182297

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.