Senior Site Reliability Engineer
VIQU IT Recruitment
Senior Site Reliability Engineer
Up to £75,000 plus bonus and on call allowance
Milton Keynes (2 days on site a week)
VIQU have partnered with a well-established B2B SaaS company who are going through a significant platform transformation. and so are hiring for a Senior Site Reliability Engineer to build stability, respond to live incidents, and assist with system upkeep. The role will also play a key part in on implementing and adopting new tooling and processes surrounding the wider transformation.
This is a genuine opportunity to own and operate how the cloud function works, and progress into a team lead position as the team grows.
Experience required for the Senior Site Reliability Engineer
- Previous experience as a Site Reliability Engineer or similar (cloud, infrastructure, DevOps or platform engineering) within a customer facing environment - e.g SaaS or MSP.
- Strong hands-on experience with both Azure, and on-premise virtual machines.
- Experience withInfrastructure as Code / Terraform, Container orchestration (Kubernetes or AKS), and Monitoring and observability tooling (Prometheus, Grafana, Datadog, or Azure Monitor).
- Ability to implement new processes, and tools, ensuring the wider development and support teams adopts new ways of working.
- Ability to communicate across internal teams and external customers.
- Skilled in networking across both cloud (Azure) and on premise environments.
- Either Windows or Linux systems administration skills (Linux preferred).
- Previous use of AI tools to enhance efficiency.
Job Duties of the Senior Site Reliability Engineer
- Utilise various technologies (Terraform, Kubernetes ect) to manage provision, and configure servers and networks, and automate application lifecycles.
- Regularly use Datadog and other observability tools for application performance monitoring.
- Implement new ways of working, helping to shape how the organisation responds and recovers to incidents.
- Take ownership of incident resolutions.
- Actively drive down key reliability metrics (MTTR, incident frequency, on-call toil) by evaluating key incidents.
- Work on an a on call rota, ensuring you are available to respond to incidents during this time.
- Identify areas for automation and help implement changes that raise the bar for reliability.
Reference: WJ-766_21889243