IT & Software

Senior Site Reliability Engineer

GoDaddy

London · England · United Kingdom

  • The Commerce Site Reliability Engineering team is responsible for the reliability, scalability, and day-to-day operation of the platforms that power GoDaddy's Commerce ecosystem. We build and operate shared infrastructure, support critical production systems, and partner closely with engineering teams to ensure services remain secure, resilient, and highly available
  • As a Senior Site Reliability Engineer, you’ll join a team that values ownership, operational excellence, and continuous improvement. Engineers are empowered to identify problems, drive meaningful change, and influence how reliability is delivered across the broader Commerce organisation
  • From improving operational maturity and reducing toil to modernising delivery platforms and strengthening incident response practices, this team plays a key role in enabling engineering teams to move quickly and safely
  • You’ll work closely with engineers across infrastructure, cloud, security, networking, and application teams while helping shape the future of reliability engineering at GoDaddy. This role offers significant opportunity to broaden your impact, develop technical leadership skills, and grow toward Staff and Principal engineering positions over time
  • Lead reliability and operational improvement initiatives across GoDaddy’s Commerce platform, helping engineering teams build and operate services safely and at scale
  • Own critical production systems, drive incident response and post-incident improvements, and continuously raise the bar for operational excellence
  • Design, build, and enhance cloud infrastructure, automation, observability, and deployment platforms that improve reliability, scalability, and developer productivity
  • Partner with engineering, infrastructure, security, and product teams to solve complex technical challenges, manage operational risk, and support business-critical services
  • Use automation, AI-assisted engineering tools, and data-driven insights to reduce operational toil, improve diagnostics, and accelerate delivery
  • Mentor engineers, share knowledge, and influence engineering practices that improve reliability across the broader Commerce organisation
  • Contribute to the team’s technical direction by identifying opportunities to improve systems, processes, and operational maturity
  • Experience building and maintaining Infrastructure as Code, automation solutions, and CI/CD pipelines that improve reliability, scalability, and delivery confidence
  • Experience using observability data, monitoring, and operational metrics to identify issues, improve system performance, and support data-driven decision making
  • Strong expertise in AWS, Linux, container platforms such as Kubernetes, and modern infrastructure engineering practices
  • A proven track record of leading or owning production incidents, driving root cause analysis, and implementing long-term reliability improvements
  • Strong software engineering or scripting skills using languages such as Python, Go, TypeScript, or similar technologies
  • Significant experience 5 years + operating, troubleshooting, and improving large-scale production systems in cloud-based environments
  • Experience using AI-assisted engineering tools to improve productivity, accelerate troubleshooting, automate repetitive tasks, and enhance operational workflows
  • Experience mentoring engineers, influencing technical decisions, and helping raise operational and engineering standards within a team
  • A continuous improvement mindset with a focus on automation, reducing operational toil, and leaving systems and processes better than you found them
  • Demonstrated ability to independently lead complex technical initiatives, manage competing priorities, and deliver outcomes across multiple teams or stakeholders
  • Experience contributing to architectural strategy and roadmap planning

We encourage you to apply even if your experience or skillset doesn't align perfectly with every requirement. We value a wide range of backgrounds and transferable skills, and we are excited to support learning and growth

  • Experience supporting eCommerce, payments, fintech, or other high-availability customer-facing platforms
  • Experience with SaltStack, Ansible, Terraform, Pulumi, CloudFormation, or AWS CDK
  • Experience leveraging AI-assisted tools to improve engineering productivity and operational effectiveness
  • Experience operating shared or multi-tenant infrastructure services at scale
  • Experience designing observability, SLO, SLI, or operational readiness frameworks

#J-18808-Ljbffr

Reference: WJ-766_21889123

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.