IT & Software

Lead DevOps Engineer

Speak to Kit

London · England · United Kingdom

Overview

A leading blockchain analytics business is hiring a Lead DevOps Engineer based in London on a hybrid basis. The role is permanent and full-time, with a salary of £120,000 to £150,000 base. Benefits include private health insurance, 25 days annual leave plus bank holidays and a birthday day off, enhanced parental leave of 16 weeks fully paid, a £500 remote working budget, a $1,000 learning and development budget, mental health support, life assurance at four times salary, and a cycle to work scheme. You can also work from almost anywhere for up to 90 days per year.

This role exists because the business is ready to move its DevOps practice from AI-assisted to genuinely autonomous, and it needs someone who has already done that in production to lead the way. The team is technically strong and already uses AI across Terraform, Helm, PR reviews, runbooks, and alert triage. The next step is a fully agent-driven routine infra path: an engineer describes a change, an agent writes the code, tests it, opens the PR, and promotes through canary without a human touching it unless a policy gate or error budget trips. Scaling, cert rotation, drift correction, and known alert patterns follow the same path. The human line stays deliberate, drawn around blast radius decisions, and this hire's job is to build the guardrails that let the team move it with confidence. You will arrive as the centre of gravity on agentic systems, contribute significantly to the design of that capability, and bring the team up to production-level fluency alongside you. This is a high-expectation environment where decisions are defended on evidence in front of peers who will push back, and that is precisely what makes it a genuinely interesting place to do this work.

Key Responsibilities

  • Own the DevOps and platform roadmap, including Kubernetes platform evolution, application packaging, migration to EKS, and enabling engineering teams to ship reliably to production
  • Lead by doing: engineer, review, and enhance Kubernetes and CNCF-aligned infrastructure, setting technical standards for the team
  • Architect multi-cluster, multi-region environments using Istio or Linkerd, Cluster API, and Kyverno
  • Build progressive delivery frameworks with Flux and Flagger for GitOps-driven, canary, and automated releases
  • Implement modern provisioning with controllers such as Crossplane and ACK for Kubernetes-native cloud integration
  • Define and enforce Zero Trust architecture with Vault, Boundary, service identity, and mTLS-secured meshes
  • Engineer policy-driven automation and compliance using OPA, Kyverno, and secure supply chain configurations
  • Establish IaC and GitOps standards with automated testing on every infrastructure change
  • Prototype agentic infrastructure components, including deployment and observability platforms in service meshes
  • Contribute to the Kong AI Gateway, including Dataplane deployments, ACM/SSL integration, and observability via DataDog
  • Champion DevSecOps maturity by embedding SAST/DAST, chaos engineering, and error budget monitoring
  • Collaborate with Security, Data, and AI teams to shape DevOps and AI platform architectures with regulatory compliance in mind
  • Stay ahead of CNCF and AI ecosystem developments, from eBPF observability to agent-aware orchestration

Requirements

Must-haves

  • Experience leading or mentoring engineering teams, setting direction hands-on
  • Strong Kubernetes knowledge: cluster lifecycle, API extensions, Operators, Helm, CNCF ecosystem (Cilium, ExternalDNS, Kyverno, Gatekeeper)
  • Multi-cluster, multi-region Kubernetes platform design with Istio, Consul, or Linkerd
  • Infrastructure-as-Code with Terraform on AWS or GCP, modular design, GitOps integration, automated testing
  • GitOps pipelines with ArgoCD or FluxCD for progressive delivery and drift correction
  • Containerised, serverless, or event-driven systems with strong observability (DataDog, Splunk, or OpenTelemetry)
  • Vault-based secret management, least privilege access, compliance automation
  • CI/CD workflows including SAST, DAST, policy enforcement, and performance telemetry
  • Reliability and resilience through SLOs, error budgets, and chaos engineering
  • Working knowledge of LLM-based services and AI infrastructure: deploying, securing, and operating AI Gateways and services
  • Already built repeatable autonomous systems in production infrastructure settings, with a defined rule for what happens when they go wrong
  • Taken a team through the shift to agentic or automated workflows and got it to stick

Nice-to-haves

  • Platform modernisation or reliability initiatives in scale-up or regulated environments
  • Operator development, CRD automation, eBPF, or Cilium for observability
  • Policy-as-code using OPA or Kyverno within secure supply chain or CSPM frameworks
  • Familiarity with MCP and A2A orchestration patterns in Kubernetes service mesh environments
  • Agent Gateways and Registries connecting microservices and AI agents
  • Secure containers, sandboxing, or confidential computing for regulated workloads
  • Data-intensive systems such as Spark, Databricks, or Data Mesh
  • Programming experience in Go, Python, or TypeScript
  • Open-source or CNCF community contributions

What Success Looks Like

  • A fully autonomous routine infra path is in production: agent writes the code, tests it, opens the PR, promotes through canary, and closes the loop without a human unless a policy gate or error budget trips
  • Scaling, cert rotation, drift correction, and known alert patterns run without manual intervention
  • The team's capability density around production agentic systems has measurably increased, with engineers able to build and operate autonomous workflows, not just use AI as a writing aid
  • Guardrails are in place that define clearly where the autonomous line sits and what triggers human review, making it safe to move that line further over time
  • Deployment velocity has increased without introducing regressions, drift, or policy violations

Team and Culture

  • A technically strong team that already works at a high level, with high expectations of one another and no appetite for trial and error
  • Opinions are expected to be backed by evidence and defended in front of peers who will push back, that is the normal mode of discourse, not an exception
  • AI is already embedded in daily work across the team, and the direction is toward more autonomy, not less
  • The culture rewards people who arrive with conviction and raise the bar, not those who wait to see which way the wind blows

Challenges

  • The technical bar is already high, and this hire needs to arrive as the centre of gravity on agentic systems rather than growing into that position over time
  • The environment does not tolerate a long runway of experimentation: decisions about where to move the autonomous line carry real consequences, and the reasoning behind them needs to hold up
  • Building the guardrails that make production autonomy safe is as demanding as building the autonomy itself, and both need to happen in parallel
  • The team needs to be brought up to production-level fluency on agentic workflows, which means carrying people through a genuine shift in how they work, not just introducing new tooling

#J-18808-Ljbffr

Reference: WJ-766_22186436

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.