Site Reliability Engineer, Big Data (Remote, International)
PulsePoint
In this role you will own the design, deployment and operation of streaming and storage infrastructure at scale within WebMD's Data Platform. You’ll work across Kafka and Ceph-based systems on hybrid on-prem and cloud environments, shaping architecture, capacity planning and incident response. You’ll partner with multiple engineering teams to deliver reliable, observable platforms that handle billions of events daily. This is a chance to influence platform standards and tackle large-scale distributed systems challenges in a remote-friendly setting.
Pay / Benefits- remote work possible
- large-scale infrastructure experience
- opportunity to define platform architecture
- Design and optimize Kafka architecture, topic governance, partition strategy and throughput
- Operate Ceph storage, pool design and capacity planning
- Develop and implement operational automation to reduce manual work and speed incident response
- Maintain SQL Server backup and recovery pipelines and support basic clustering
- Build data tooling with self-service capabilities and observability for the team
- 5+ years operating distributed systems at scale in production
- Deep expertise in Kafka, Ceph, or similar distributed infrastructure
- Proven ability to design for scale and reliability
- Experience mentoring engineers and making technical decisions
- Willingness to work 9am-6pm ET U.S. hours
- Mentoring and technical leadership
- Ownership across system layers
- Ability to simplify complex systems
- Apache Kafka
- Ceph
- SQL Server backup and recovery
Reference: WJ-747_30146303