Embedded Data Engineer - ML
Trainline
Overview
As Embedded Data Engineer in ML, you design scalable data pipelines and models to power analytics and ML workflows. You deploy cloud-native data apps on AWS and maintain production reliability with observability. You collaborate with ML Engineers and Data Scientists to deliver well-structured datasets and share best practices across the data community. This role sits at the intersection of data engineering and ML, enabling real-time analytics and ML-powered decisions at scale.
Pay / Benefits- private healthcare & dental insurance
- work from abroad policy
- 2-for-1 share purchase plans
- EV Scheme to reduce carbon emissions
- extra festive time off
- family-friendly benefits
- Design and build scalable data pipelines, data models, and feature stores for analytics and ML workloads in the ML domain
- Deploy and maintain cloud-native data applications on AWS with CI/CD automation
- Maintain production data pipeline quality, performance, and reliability through observability and best practices
- Collaborate with ML Engineers and Data Scientists to create reliable, well-structured datasets for ML use cases
- Engage with the wider Data Engineering, Data Platform, and analytics community to share knowledge and align on best practices
- Python and SQL proficiency
- Experience building data pipelines for downstream ML workloads (feature engineering and model training workflows)
- Cloud data modelling and data marts/warehouses in AWS
- Data pipelines with Spark and Airflow (or similar) in a cloud environment
- Experience with real-time and batch data workloads and modern transformation/orchestration patterns
- Knowledge of Ray (optional) and modern data formats like Parquet and Iceberg
- Infra as Code and containerisation (Terraform, Docker) is helpful
- CI/CD experience (Jenkins or GitHub Actions) for production data systems
- Strong collaborative problem-solving skills
- collaboration
- problem-solving
- communication
- Python
- SQL
- Spark
Reference: WJ-747_30165022