Site Reliability Engineer
Dabster
- Manage, monitor, and optimize large-scale Kubernetes clusters hosted on public cloud platforms (Google GKE).
- Implement and maintain infrastructure as code using tools such as Terraform.
- Collaborate with development and operations teams to improve system reliability and deployment automation.
- Build and maintain CI/CD pipelines using Jenkins or similar tools.
- Troubleshoot production issues, conduct root cause analysis, and implement preventive measures.
- Automate operational tasks using Python or other scripting languages.
- Contribute to observability and monitoring improvements using modern tools and best practices.
- Participate in on-call rotations and incident response processes.
Required Skills and Experience:
- 4-8 years of experience in Site Reliability Engineering, DevOps, or Cloud Infrastructure roles.
- Strong hands-on experience managing Kubernetes clusters in production (GKE).
- Proficiency with Terraform and cloud infrastructure automation.
- Practical experience with Jenkins and CI/CD pipeline management.
- Sound understanding of SRE principles (incident management, blameless postmortems, capacity planning, error budgets, etc.).
- Good programming or scripting skills in Python (preferred) or similar languages.
- Strong analytical, troubleshooting, and problem-solving abilities.
- Excellent written and verbal communication skills.
- Experience with Prometheus, Grafana, or OpenTelemetry for observability.
- Exposure to GitOps practices and tools (e.g. Flux).
Reference: WJ-766_22286003