Senior Infrastructure SRE
Jobtailor
Requirements
- 5+ years of hands-on experience operating and designing cloud infrastructure
- Expert-level understanding of the core services of either Azure or AWS
- Working proficiency in at least one additional platform (Azure, AWS, or GCP)
- Experience designing and supporting production infrastructure spanning multiple cloud platforms
- 3+ years of production experience with Infrastructure as Code using Terraform, Pulumi, or CloudFormation
- Ability to design scalable, reusable IaC modules and enforce GitOps workflows
- Experience managing IaC across multiple cloud providers
- Strong proficiency in Python, Go, or Bash for production automation
- Demonstrated ability to write tested, maintainable automation and tooling
- Practical application of SRE principles in production environments
- Experience defining and managing SLIs/SLOs, error budgets, and toil metrics
- Track record of improving system reliability
- Expert-level experience running Kubernetes and containerized workloads in production on managed Kubernetes, plus VM-based compute
- Practical experience operating a service mesh in production
- Strong proficiency with SAML, OAuth/OIDC, LDAP, and cloud IAM
- Working knowledge of PingFederate, Entra ID, Okta, or ADFS
- Strong proficiency with object, block, and file storage and Kubernetes persistent volumes
- Working knowledge of infrastructure in regulated environments such as HIPAA, SOC 2, PCI, or FedRAMP
- Familiarity with audit evidence, access controls, encryption, and data residency constraints
- Proven track record of reducing operational toil through automation
- Strong communication and documentation skills
- Bachelor's degree in Computer Science, Computer Engineering, Information Technology, or related technical field, or equivalent practical experience
- Evidence of continuous learning and staying current with SRE and cloud-native trends
- Nice-to-have: 2+ years in healthcare technology or highly regulated SaaS environments
- Nice-to-have: relevant cloud certifications
- Nice-to-have: Kubernetes-at-scale experience or CKA/CKAD certification
- Nice-to-have: messaging and event-streaming platforms
- Nice-to-have: CI/CD pipelines and deployment automation
- Nice-to-have: AI-assisted engineering and operations tooling
- Nice-to-have: contribution to open-source SRE tools or infrastructure projects
Core Competencies
Demonstrates expertise in designing and implementing highly available cloud infrastructure solutions, with a strong focus on Infrastructure as Code using Terraform or Pulumi. Proven ability to manage SLIs, SLOs, and improve system reliability while mentoring teams and driving multi-team initiatives.
Highest-signal resume keywords
- Cloud Infrastructure Design
- Infrastructure as Code (Terraform, Pulumi)
- SLI/SLO Management
- Kubernetes Management
- Automation Proficiency (Python, Go, Bash)
ATS Optimization Keywords
Hard Skills
- Cloud Infrastructure Design
- Infrastructure as Code
- SLI/SLO Management
- Kubernetes Management
- Python
- Go
- Bash
- Service Mesh Operation
- Cloud IAM
- Storage Management
Soft Skills
- Strong Communication
- Documentation Skills
- Mentoring
Certifications & Qualifications
- Cloud Certifications
- CKA Certification
- CKAD Certification
Industry Keywords
- HIPAA
- SOC 2
- PCI
- FedRAMP
- Healthcare Technology
- Regulated Environments
Tools & Technologies
- Terraform
- Pulumi
- Kubernetes
- CloudFormation
- PingFederate
- Entra ID
- Okta
- ADFS
- CI/CD Pipelines
- AI-assisted Tooling
Reference: WJ-3875_12648583