Senior Data Scientist I - LeapSpace
Elsevier
In this role you drive the development and evaluation of advanced search and generative AI systems within Elsevier’s Search & AI Evaluation team. You own complex problems end-to-end, shaping retrieval and RAG pipelines and contributing to the team’s technical direction. You’ll work with cross-functional partners to deliver scalable, production-ready solutions that improve relevance and user outcomes. This is a hands-on senior IC role with growing technical leadership, focused on impactful research-integrated products.
Pay / Benefits- flexible working hours
- health benefits and private medical benefits
- pension scheme
- share option scheme
- parliamentary leave and sabbaticals
- study assistance
- Lead design and optimization of lexical, vector, and hybrid retrieval at scale
- Architect and improve RAG pipelines including retrieval strategies and prompt design
- Experiment with embeddings, re-ranking models, and retrieval architectures to boost relevance
- Collaborate with engineering for robust, production-ready implementations
- Define and evolve evaluation strategies for search and GenAI across products
- Design frameworks for IR and GenAI evaluation, grounding, and hallucination detection
- Contribute to evaluation datasets, gold standards, and annotation strategies
- Guide experimental design, offline evaluation, and A/B testing with statistical rigor
- Promote responsible AI practices including bias, fairness, and risk evaluation
- Apply NLP, embeddings, and GenAI techniques to production use cases
- Contribute to knowledge graphs and semantic enrichment for retrieval systems
- Work with domain experts to integrate scientific taxonomies and ontologies into retrieval systems
- Incorporate structured data (datasets, chemicals, genes, drugs, trials, outcomes) into AI pipelines
- Advance Elsevier’s knowledge graph and metadata integration strategy
- Present findings clearly to technical and non-technical stakeholders
- Take ownership from problem definition through deployment
- Master’s or PhD in Computer Science, Data Science, Machine Learning, or related field (or equivalent practical experience)
- Experience in data science, machine learning, or applied NLP
- Strong hands-on experience with: Search and retrieval systems (lexical, vector, hybrid); RAG pipelines and LLM-based systems; Evaluation methodologies for ML/IR/GenAI
- Advanced programming skills in Python
- Experience with ML/NLP frameworks (PyTorch, Hugging Face, LangChain, LangGraph, Haystack)
- Experience with Databricks or similar distributed data/ML platforms
- Strong understanding of experimentation design and statistical analysis
- Cross-functional collaboration
- Clear communication with both technical and non-technical stakeholders
- Problem ownership and autonomy
- Python programming
- PyTorch
- Hugging Face
Reference: WJ-747_30140554