IT & Software

Senior Site Reliability Engineer

NICE Systems

Remote · Nationwide · United Kingdom

Overview

As a production and reliability-focused engineer, you’ll monitor and manage the health of a complex, distributed platform. You’ll build and automate infrastructure and applications, driving reliability, quality, and faster delivery of NiCE’s software solutions. You’ll analyze metrics, partner with development teams, and design scalable systems while lifting automation and SLO-driven operations. This role emphasizes cross-functional collaboration, incident resilience, and continuous improvement in a fast-paced, hybrid environment.

Pay / Benefits
  • NICE-FLEX hybrid model (2 days in office, 3 days remote)
  • flexible hybrid work arrangement
  • collaborative and innovative culture
Responsibilities
  • Run production environments with a holistic view of system health
  • Build software and systems to manage platform infrastructure and applications
  • Improve reliability, quality, and time-to-market of software solutions
  • Measure and optimize system performance and anticipate customer needs
  • Provide primary operational support for multiple large distributed applications
  • Analyze metrics for performance tuning and fault finding
  • Partner with development teams to improve services through testing and release procedures
  • Participate in system design consulting, platform management, and capacity planning
  • Create sustainable systems through automation and uplift
  • Balance feature speed and reliability with defined service level objectives
Key requirements
  • 3-6 years in a similar role focused on systems engineering, automation, and reliability
  • Proficiency in at least one programming language (Python, Go, Java, C#) and scripting (Bash, PowerShell)
  • Deep understanding of cloud platforms (AWS) and services (EC2, ECS, Lambda, DynamoDB)
  • Experience with infrastructure as code tools (CloudFormation, Terraform)
  • CI/CD concepts and tools (Jenkins, GitLab CI/CD, CircleCI)
  • Containerization and microservices (Docker, Kubernetes)
  • Monitoring/observability tools (Prometheus, Grafana, ELK, CloudWatch)
  • Incident management and blameless postmortems with cross-functional communication
  • Strong communication
  • Team player
  • Fast learner
  • Python
  • Go
  • Java

Reference: WJ-747_30139192

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.