Data Engineer (Mid‑Level), Global
Vantage Data Centers
Overview
As a Mid-Level Data Engineer in London, you will build, operate, and scale our enterprise data platform to support analytics, reporting, and AI-enabled use cases. You’ll work within the Data Engineering & Business Intelligence team, collaborating with analysts and stakeholders while taking ownership of data pipelines. The role emphasizes independence, reliability, and fast execution in a dynamic environment. You will contribute to a scalable, governed data platform on Azure, enabling data-driven decisions. This is a hands-on role with a clear path to impact across analytics and AI initiatives.
Pay / Benefits
- above market total compensation package
- comprehensive health and welfare
- retirement benefits
- paid leave
- training and development opportunities
- recognition and global collaboration
Responsibilities
- Design, build, and maintain scalable data pipelines using Python and PySpark on Azure
- Develop and operate batch and incremental ETL pipelines with Azure Data Factory and store data in Azure Data Lake Storage Gen2
- Implement SQL- and Spark-based transformations to create curated datasets for reporting and analytics
- Own assigned pipelines and datasets, including monitoring, troubleshooting, and performance tuning in production
- Work with Azure Synapse to support analytical workloads and data consumption patterns
- Collaborate with analysts and stakeholders to translate data requirements into practical solutions
- Prepare data for advanced analytics and AI use cases, ensuring quality, consistency, and documentation
- Apply data governance, security, and engineering standards for maintainable and scalable solutions
- Participate in code reviews and platform improvement initiatives
- Identify data quality issues and pipeline risks, communicating them in a fast-paced environment
- Develop and maintain PySpark notebooks and jobs for ingestion, transformation, and curation
- Create and modify Azure Data Factory pipelines for batch/incremental ingestion
- Implement Spark transformations that write to Azure Data Lake Gen2 with established structures
- Create SQL views and tables in Azure Synapse to support analytics
- Respond to pipeline failures, data validation issues, and operational alerts
- Perform basic Spark performance tuning within architectural patterns
- Validate data outputs with business partners and address defects
- Commit code with Git, follow branching standards, and participate in PR reviews
- Update pipeline and runbook documentation and manage backlog items in sprints
Key requirements
- Bachelor’s degree in Engineering, Computer Science, Data Analytics, or related field
- 3–5 years of data engineering or analytics engineering experience
- Proficiency in Python for data pipelines and PySpark
- Proficiency in SQL for data querying and transformation
- Strong understanding of ETL/ELT, data transformations, and data integration
- Experience analyzing enterprise data sources to identify relationships and business rules
- Experience building solutions on Microsoft Azure with Azure Data Factory, Azure Synapse, and Azure Data Lake Storage Gen2
- Experience with source control and CI/CD (GitHub or Azure DevOps)
- Knowledge of data modeling (fact and dimension tables)
- Strong communication and collaboration skills in a fast-paced environment
- Experience working in Agile environments
- Experience with Jira or similar project tracking tools
- Travel up to 10% (may increase over time)
- strong communication
- collaboration across teams
- self-starter mindset
- Python
- PySpark
- SQL
Reference: WJ-799_20873168