Director, Site Reliability Engineering
Omnicell
In this role you will build and lead Omnicell’s global SRE organization to ensure cloud platforms are highly available, scalable, and observable. You will establish enterprise reliability standards, governance, and automation, partnering with Product Engineering, Cloud Platform Engineering, and Security to embed reliability across the software lifecycle. The position blends executive strategy with hands-on engineering leadership, driving production readiness and resilient architectures at enterprise scale. This is a transformational leadership opportunity to shape reliability engineering for Omnicell’s cloud-first future.
Responsibilities- Build and scale a global SRE organization and mentor senior technical leaders
- Define engineering standards, career frameworks, and leadership expectations
- Develop enterprise reliability strategy including governance for SRE, Production Engineering, and related domains
- Establish reliability framework with SLIs, SLOs, error budgets, and production readiness practices
- Partner with Product Engineering, Cloud Platform Engineering, Security, and EA to embed reliability in the SDLC
- Define and execute multi-year roadmaps to improve platform reliability and deployment confidence
- Lead observability strategy covering metrics, tracing, logging, and dashboards
- Champion automation to reduce toil, enable self-healing capabilities, and improve deployment safety
- Govern production engineering activities focusing on performance, scalability, capacity planning, and resilience
- Provide executive-level reporting on reliability metrics and strategic investments
- Experience leading a global SRE organization and defining enterprise reliability standards
- Ability to converse with senior leadership on strategy and with engineers on architecture (e.g., Kubernetes)
- Expertise in production design reviews, observability, and automation
- Strong partnership skills across Product Engineering, Cloud Platform Engineering, Security, and EA
- Track record of driving reliability improvements and reducing operational toil
- Strategic leadership
- Cross-functional collaboration
- Mentoring and talent development
- Kubernetes architectures
- Observability engineering (OpenTelemetry, metrics, distributed tracing, centralized logging, APM, synthetic monitoring)
- Production engineering practices (availability, scalability, capacity planning, resilience testing)
Reference: WJ-747_30157885