Principal Site Reliability Engineer, Platform Engineering: Dedicated
GitLab
In this Principal Engineer role, you will shape the architecture and direction of GitLab Dedicated, our fully managed single-tenant SaaS offering. You’ll drive platform transformations to improve resilience, security, and compliance while scaling isolated environments. You’ll mentor senior engineers and elevate engineering maturity across teams, aligning Dedicated with GitLab’s modular, cell-based architecture. This position offers high-impact leadership at scale and opportunities to influence cross-team decisions and best practices.
Pay / Benefits- Flexible Paid Time Off
- Equity Compensation & Employee Stock Purchase Plan
- Growth and Development Fund
- Parental Leave
- Team Member Resource Groups
- Set technical direction for GitLab Dedicated and guide architecture as the environment footprint grows
- Lead platform transformations across resilience, failover, tenant orchestration, change management, self-service tooling, and platform integrations
- Drive modular architecture aligned with the Cells strategy while maintaining security, isolation, and compliance
- Strengthen service ownership and operational maturity for production systems owned by engineering teams
- Identify and mitigate systemic reliability and scalability risks using production signals and architectural insight
- Establish reusable platform patterns and automation to enable scalable growth
- Lead complex technical decisions across teams balancing reliability, security, cost, maintainability, and customer needs
- Advance engineering excellence through architectural leadership, mentorship, and influence with senior engineers and leaders
- Deep expertise in Site Reliability, Platform, Infrastructure, or Backend Engineering with experience operating large-scale production systems
- Hands-on experience with cloud infrastructure, automation, observability, infrastructure as code, and modern production practices
- Strong software engineering fundamentals with production systems or infrastructure tooling in Go, Ruby, Python, or similar
- Strong distributed systems and systems-design expertise for reliability, scalability, and operational simplicity
- Track record of technical leadership across multiple teams and driving complex initiatives
- Experience leading platform or infrastructure transformations, including modernization or scaling
- Experience improving ownership, reliability, and accountability at scale
- Exceptional technical communication and influence to align teams and guide architectural decisions
- strong communication and collaboration
- influencing without authority
- mentorship of senior engineers
- Go
- Ruby
- Python
Reference: WJ-747_30468536