Site Reliability Engineer - Frontend
Capital On Tap
Overview
In this SRE role, you help keep Capital On Tap’s platforms fast, reliable, and scalable. You’ll partner with XP to provide foundational tooling and infrastructure that product teams rely on. You design, build, and monitor systems, and tackle toil with automation. This is a hands-on role at a growing fintech, focused on improving observability and reliability across cloud, Kubernetes, and serverless environments. Join a collaborative, values‑driven team shaping the next phase of our platform.
Pay / Benefits- Private Healthcare including dental and opticians through Vitality
- Worldwide travel insurance through Vitality
- Anniversary Rewards (sabbatical options)
- Salary Sacrifice Pension Scheme up to 7% match
- 28 days holiday plus bank holidays
- Annual Learning and Wellbeing Budget
- Manage and automate resources in Azure, Datadog, NGINX and Cloudflare
- Develop, deploy and monitor Kubernetes and Serverless resources
- Build, manage and evolve infrastructure as code using Terraform, Helm and Go CRDs
- Improve systems, processes and technologies by consulting stakeholders to boost platform performance
- Design and maintain monitoring and alerting strategies using Datadog
- Contribute to new application architecture and design processes
- Reduce toil by automating repetitive tasks and streamlining workflows
- Create SLIs and SLOs and align with product team on core service objectives
- Collaborate with foundational teams to build reusable, automated solutions
- Optimize CI/CD pipelines and developer workflows (Azure DevOps, Github, Octopus Deploy, MirrorD, Flux)
- Lead incident response, including communications, investigations, remediation and post-mortems to protect the customer experience
- Experience managing public cloud environments
- Proficiency with IaC (Terraform)
- Experience with CI/CD, pipelines, templates, troubleshooting
- Proficiency with containers (Kubernetes, Docker)
- Experience building/deploying frontend apps with CodeMagic
- Knowledge or experience with Google Firebase
- Experience with a cloud monitoring solution
- Proficiency in at least one scripting language (Python, PowerShell, Go)
- Strong communication and collaboration skills
- Software development background
- Strong communication
- Collaborative work style
- Problem-solving mindset
- Azure
- Datadog
- NGINX
Reference: WJ-747_30186060