IT & Software

Site Reliability Engineer

Trimble Navigation

Newcastle Upon Tyne · Tyne And Wear · United Kingdom

Overview

In this role you will serve as the backbone of the Project Delivery Cloud Platform, driving reliability and scalable cloud environments. You will build and maintain IaC, enhance observability, and manage CI/CD pipelines to streamline delivery. You’ll lead incident responses and develop runbooks, partnering with cross-functional teams to embed SRE practices. This is an opportunity to shape production systems at Trimble’s AECO unit and contribute to a culture of continuous improvement and innovation.

Responsibilities
  • Develop and maintain infrastructure-as-code with Terraform for reliable cloud environments
  • Enhance observability using New Relic, DataDog, Sumologic and Splunk for monitoring and logging
  • Manage CI/CD pipelines with Azure DevOps, GitHub, Terraform and related tooling
  • Automate routine tasks to boost operational efficiency
  • Evaluate designs for reliability, performance, security, and efficiency
  • Lead incident response and perform root cause analysis with long-term fixes
  • Develop runbooks and procedures for incident response and operations
  • Collaborate with cross-functional teams to review technical designs for SRE alignment
  • Participate in on-call rotations and handle critical incidents
  • Improve documentation and promote knowledge sharing
Key requirements
  • Bachelor’s degree in Computer Engineering or related field
  • 5+ years of production infrastructure ownership
  • Strong collaboration and cross-functional work experience
  • Proven success managing production infrastructure
  • Capacity planning and cost optimization expertise
  • Extensive experience with cloud providers (Azure or AWS)
  • Proficiency in Python and Terraform, plus containerization
  • Experience with Kubernetes or other containerization tech
  • Familiarity with CI/CD tools (Azure DevOps, Jenkins, Argo CD, Helm, GitHub)
  • Experience with monitoring and incident management (Prometheus, Grafana, New Relic, DataDog, Splunk, CloudWatch, Sumologic)
  • Solid understanding of networking and security concepts
  • Collaboration and cross-functional teamwork
  • Strong incident management calm under pressure
  • Proactive problem-solving and communication
  • Terraform (IaC)
  • Cloud platforms: Microsoft Azure and AWS
  • Python scripting

Reference: WJ-747_30977329

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.