Senior Site Reliability Engineer
Carta
Overview
As a Senior Site Reliability Engineer at Carta, you will scale and maintain our internal platform offerings to ensure reliability and performance of applications. You’ll design monitoring, alerting, and incident response, while guiding software engineers to build scalable solutions. You’ll drive improvement across global infrastructure and help shape the platform as Carta expands. Join a collaborative, innovation-driven team shaping the future of private market software.
Responsibilities- Build and scale internal platform services (compute, storage, networking) for reliability and performance
- Design and implement monitoring, alerting, and incident response systems
- Collaborate with application engineers to ensure scalable designs
- Act as a change agent to incrementally improve systems as the company grows globally
- Extensive cloud platform experience (AWS, Google Cloud Platform, or Azure) with services like EC2, S3, RDS, Lambda
- Kubernetes or other container orchestration experience (preferred)
- Infrastructure as Code with Terraform, Ansible, or CloudFormation
- Networking knowledge including CNI, network policies; experience with proxies/service mesh a plus
- Monitoring and observability with Prometheus, Grafana, ELK/Datadog
- Proficiency in Python; experience designing and maintaining API services (REST/GraphQL)
- AI fluency and ability to build agents to reduce toil; CI/CD experience expected though not essential
- Strong communication
- Collaborative problem solving
- Comfort with ambiguity and change
- Python
- Terraform
- gRPC
Reference: WJ-747_30155285