Staff Software Engineer, Observability & Profiling
Humanloop
In this role you will design and evolve scalable observability infrastructure to give engineers deep visibility into multi-cluster systems. You’ll work with cross-functional teams to improve reliability, reduce incident response time, and scale telemetry from kernel to application level. The role focuses on building high-throughput data pipelines, instrumentation, and AI-assisted diagnostic tooling to support rapid, data-driven decisions. You’ll contribute to a mission-driven company that aims to make AI safe, steerable, and beneficial at scale.
Pay / Benefits- competitive compensation
- visa sponsorship available
- office space in San Francisco
- generous vacation and parental leave
- flexible working hours
- equity donation matching (optional)
- Design and build scalable telemetry ingest and storage pipelines for metrics, logs, traces, and error data across Anthropic's multi-cluster infrastructure
- Develop observability solutions that provide low-overhead visibility into system behavior across the fleet
- Own and evolve core observability platforms, driving migrations and architectural improvements for reliability and cost efficiency
- Create instrumentation libraries, SDKs, and eBPF-based auto-instrumentation to emit high-quality telemetry
- Reduce MTTR through cross-signal correlation, unified query interfaces, and AI-assisted diagnostic tooling
- Drive fleet-wide efficiency via continuous profiling and utilization telemetry to guide optimizations across CPUs, memory, and accelerators
- Collaborate with Research, Inference, Product, and Infrastructure teams to tailor observability solutions to diverse needs
- Hands-on experience building and operating large-scale observability or monitoring infrastructure
- Deep hands-on experience with observability signals from instrumentation through ingest to query and analysis
- Understanding of high-throughput telemetry pipelines and scale tradeoffs for data collection, storage, and querying
- Comfort digging below the application layer into kernel, network stack, or hardware
- Excellent communication skills and ability to partner with internal teams on visibility and incident response
- Interest in building foundational infrastructure and navigating ambiguous, high-impact technical challenges
- Strong collaboration skills
- Problem-solving orientation
- Proactive communication
- eBPF-based observability in production
- OpenTelemetry instrumentation, collector pipelines, and tail-based sampling
- Continuous profiling at fleet scale and symbolization
Reference: WJ-747_30182553