Software Engineer - Data, Lakehouse and AI Data Platform Engineer - Analyst/Associate - London
Goldman Sachs
Overview
In this role you design, build, test and support data pipelines and curated datasets on the firm’s modern data platform to enable reliable analytics and AI use cases. You contribute across ingestion, transformation, modelling and data quality to deliver scalable data products. You’ll work with cross-functional teams to shape data assets that power operational decisions and emerging AI. This opportunity combines production delivery with modern data technologies at a leading financial services firm.
Pay / Benefits- benefits, wellness and personal finance offerings
- mindfulness programs
- Build, enhance and support batch and streaming data pipelines on the Lakehouse and AI data platform
- Refactor or modernise existing data flows to boost reliability, performance and maintainability
- Ensure pipelines are production-ready, tested and supportable
- Develop raw, refined and curated datasets for analytics, reporting and AI
- Apply data modelling to represent entities and historical changes
- Collaborate with consumers to shape usable, well-documented data products
- Implement data quality controls to ensure completeness and accuracy
- Use reconciliation approaches to validate production outputs and investigate issues
- Deliver outcomes in partnership with engineers, platform teams and data consumers
- Communicate progress, risks, dependencies and design choices
- Bachelor’s or master’s degree in a relevant field or equivalent practical experience with strong quantitative/data engineering capabilities
- Hands-on programming in Python or Java
- Strong SQL skills including troubleshooting and optimisation
- Ability to learn new tools and delivery workflows quickly
- Familiarity with software engineering basics: version control, testing, release discipline, CI/CD
- Understanding of temporal data modelling and history handling
- Knowledge of schema design/evolution and data compatibility considerations
- Experience with partitioning, clustering and performance optimisation at scale
- Practical approach to data quality and root-cause analysis
- Experience building or supporting production data pipelines in a collaborative engineering environment
- Experience with distributed data processing frameworks like Apache Spark
- Working knowledge of JSON, Avro and Parquet
- Strong problem-solving and attention to detail
- Clear, structured communication with stakeholders
- Willingness to collaborate with partner teams
- Python
- Java
- SQL
Reference: WJ-747_30158304