IT & Software

Staff Software Engineer, Observability & Profiling

Humanloop

London · Greater London · United Kingdom

Overview

In this role you will design and evolve scalable observability infrastructure to give engineers deep visibility into multi-cluster systems. You’ll work with cross-functional teams to improve reliability, reduce incident response time, and scale telemetry from kernel to application level. The role focuses on building high-throughput data pipelines, instrumentation, and AI-assisted diagnostic tooling to support rapid, data-driven decisions. You’ll contribute to a mission-driven company that aims to make AI safe, steerable, and beneficial at scale.

Pay / Benefits
  • competitive compensation
  • visa sponsorship available
  • office space in San Francisco
  • generous vacation and parental leave
  • flexible working hours
  • equity donation matching (optional)
Responsibilities
  • Design and build scalable telemetry ingest and storage pipelines for metrics, logs, traces, and error data across Anthropic's multi-cluster infrastructure
  • Develop observability solutions that provide low-overhead visibility into system behavior across the fleet
  • Own and evolve core observability platforms, driving migrations and architectural improvements for reliability and cost efficiency
  • Create instrumentation libraries, SDKs, and eBPF-based auto-instrumentation to emit high-quality telemetry
  • Reduce MTTR through cross-signal correlation, unified query interfaces, and AI-assisted diagnostic tooling
  • Drive fleet-wide efficiency via continuous profiling and utilization telemetry to guide optimizations across CPUs, memory, and accelerators
  • Collaborate with Research, Inference, Product, and Infrastructure teams to tailor observability solutions to diverse needs
Key requirements
  • Hands-on experience building and operating large-scale observability or monitoring infrastructure
  • Deep hands-on experience with observability signals from instrumentation through ingest to query and analysis
  • Understanding of high-throughput telemetry pipelines and scale tradeoffs for data collection, storage, and querying
  • Comfort digging below the application layer into kernel, network stack, or hardware
  • Excellent communication skills and ability to partner with internal teams on visibility and incident response
  • Interest in building foundational infrastructure and navigating ambiguous, high-impact technical challenges
  • Strong collaboration skills
  • Problem-solving orientation
  • Proactive communication
  • eBPF-based observability in production
  • OpenTelemetry instrumentation, collector pipelines, and tail-based sampling
  • Continuous profiling at fleet scale and symbolization

Reference: WJ-747_30182553

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.