Senior Platform & Observability Engineer
Koda Tech
London | Hybrid (3 days per week in office)
We're hiring a Senior Platform & Observability Engineer to help shape and mature the observability strategy for a growing technology organisation operating large-scale business-critical systems.
This is an opportunity for someone who enjoys solving complex operational challenges, influencing technical direction, and building practical solutions that improve reliability, visibility and engineering effectiveness.
Rather than simply maintaining monitoring tools, you'll help define how observability should work across the organisation, working closely with engineering, infrastructure and support teams to establish standards, improve incident response, and create a clearer picture of system health.
What you'll be doing
- Assessing the current monitoring and observability landscape across applications, infrastructure and services
- Identifying gaps in existing monitoring, alerting and operational workflows
- Defining observability standards and best practices across the business
- Designing a target-state observability architecture and roadmap
- Working with software engineers to improve telemetry, metrics, logging and tracing
- Driving adoption of structured logging practices
- Helping establish standards for alerting, incident detection and service health monitoring
- Collaborating with platform, infrastructure and operations teams to automate observability capabilities through Infrastructure as Code
- Acting as a trusted advisor on reliability, monitoring and operational excellence
- Supporting the wider Platform and DevOps function beyond the observability programme
What we're looking for
We're more interested in experience solving real operational problems than specific tools.
You'll likely have experience in several of the following:
- Platform Engineering
- DevOps Engineering
- Site Reliability Engineering (SRE)
- Infrastructure Engineering
- Observability Engineering
We'd love to speak with people who have:
- Built or significantly improved observability capabilities within a business
- Defined monitoring, logging or reliability standards rather than simply operating existing tools
- Experience working across engineering, infrastructure and support teams
- Strong understanding of metrics, logs, alerting and service health concepts
- Experience introducing or scaling structured logging practices
- Strong Infrastructure as Code experience
- Experience working with hybrid environments, including both cloud and physical infrastructure, is highly desirable
- Confidence challenging existing approaches and driving technical change
- Excellent stakeholder communication skills
Nice to have
Experience with technologies such as:
- OpenTelemetry
- Datadog
- Grafana
- Prometheus
- Elastic
- Splunk
- Terraform
- Kubernetes
- Public cloud platforms
No single technology is required. We're interested in how you've used observability to solve problems rather than which tools you've worked with.
The type of person who succeeds here
This role would suit someone who has worked in a startup or scale-up environment and has helped build processes, standards or platforms from the ground up.
You'll be comfortable navigating ambiguity, influencing technical decisions, and helping teams move towards a more mature operational model.
If you've previously inherited a fragmented monitoring landscape and successfully implemented a clearer, more effective approach to observability, we'd like to hear from you.
Reference: WJ-747_30256058