Principal Data Architect - Databricks & AI
Wolters Kluwer
In this role you will define and own the target data architecture and roadmap for a strategic, enterprise data capability. You will design a Databricks-based lakehouse and scalable data patterns across Bronze–Gold layers to enable secure, governed data use for analytics and AI. You’ll prototype, review designs, and guide complex engineering challenges while mentoring engineers, with scope to lead a small team as the capability grows. This is an opportunity to shape a high-impact data foundation that powers cross-product insights and advanced analytics.
Responsibilities- Define target data architecture, standards, reference patterns, and the technical roadmap for an enterprise data capability
- Design a modern Databricks lakehouse foundation for secure, scalable data use across a complex product landscape
- Establish scalable medallion patterns (Bronze, Silver, Gold) including ingestion, transformation, storage, serving and consumption
- Design canonical, dimensional and semantic data models for cross-domain interoperability
- Architect secure data integration (batch, streaming, APIs, CDC, event-driven)
- Design privacy-preserving data capabilities (masking, anonymisation, tokenisation) for AI/analytics
- Create AI-ready data capabilities including model-data pipelines, vector search, retrieval-augmented generation where relevant
- Develop prototypes/POCs, review designs and code, optimize performance, resolve complex issues
- Build production-ready solutions with governance, lineage, quality, observability, security, reliability and cost control
- Collaborate with product, engineering, AI, security, platform and business leaders to translate needs into architecture and delivery plans
- Mentor data engineers and architects to promote reusable platform capabilities and standards
- Significant experience as a Data Architect, Lead Data Architect, Principal Data Engineer or comparable senior role
- Advanced hands-on Databricks experience in production environments (lakehouse, Delta Lake, governance, workflow, SQL, optimization)
- Proven ownership of modern data platforms or data products from architecture through production
- Strong data privacy and security experience (masking, anonymisation, tokenisation)
- Strong data-modelling expertise across conceptual, logical, physical, dimensional and canonical models
- Experience with batch and real-time integration (APIs, CDC, streaming, event-driven pipelines)
- Strong software development foundation (Python, SQL, Scala, Java and/or Spark)
- Deep experience with Azure data services; AWS experience advantageous
- Experience with governance, cataloguing, lineage, quality, metadata, access control, secure multi-tenant/customer data
- Credibility with engineers and stakeholders, strong judgement on speed, scale, governance, cost, usability
- collaboration
- strong judgement
- leadership
- Databricks
- lakehouse
- Delta Lake
Reference: WJ-747_30167154