IT & Software

Staff Data Engineer (Emerald)

H1

York And North Yorkshire · England · United Kingdom

  • As a Staff Data Engineer on the Emerald team, you will play a critical role in shaping the architecture, scalability, and technical direction of H1’s healthcare entity resolution platform. EMERALD is responsible for linking large-scale external healthcare datasets, including PubMed, clinical trials, conferences, ct.gov, and web-collected data to H1’s canonical physician and organization profiles
  • This role sits at the intersection of distributed data engineering, entity matching, identity resolution, and large-scale healthcare data processing. You will lead a small team of engineers while remaining deeply hands‑on technically, owning the systems and pipelines powering automatching, grouping logic, identity mapping, deduplication, and enrichment workflows processing tens of millions of records
  • You will partner closely with Product, AI/ML, Analytics, and Engineering teams to improve platform accuracy, scalability, reliability, and operational efficiency across one of H1’s most critical data platforms
  • Lead the design, optimization, and scalability of distributed Spark/PySpark pipelines powering entity resolution and large-scale healthcare data processing
  • Own systems supporting automatching, identity mapping, grouping logic, deduplication, enrichment, and auto‑approval workflows across healthcare provider and organization datasets
  • Build and maintain scalable processing frameworks for PubMed, clinical trial, ct.gov, conference, and other healthcare data sources
  • Drive infrastructure optimization initiatives focused on improving throughput, runtime, observability, and cloud compute cost efficiency
  • Partner closely with AI/ML teams to integrate matching and resolution models into EMERALD and improve matching precision and recall
  • Lead complex technical initiatives from architecture and design through deployment, monitoring, and long‑term production support
  • Serve as a technical leader and mentor across the team through code reviews, technical guidance, and engineering best practices
  • Collaborate directly with Product and business stakeholders to align technical solutions with operational and customer needs
  • Support production operations, incident response, troubleshooting, and ongoing platform reliability

Benefits

  • Flexible work hours
  • Commuter benefits
  • Stock options
  • Computer setup
  • Work from home opportunities
  • Health & life insurance
  • Retirement options
  • Unlimited PTO
  • Flex Give & Flex Spend
  • Impactful BRGs

Experience building scalable ETL/ELT frameworks across both batch and streaming architecturesExtensive experience with Apache Spark and AWS-based big data technologies including EMR, S3, and distributed compute environmentsExperience optimizing large-scale Spark workloads for performance, scalability, and infrastructure cost efficiencyExperience with containerization and infrastructure technologies such as Docker, Kubernetes, and TerraformExperience working with relational or distributed databases such as PostgreSQL or RedshiftYou bring strong hands‑on engineering expertise across distributed computing, large‑scale data processing, and infrastructure optimization while also helping guide technical direction and mentor engineers across the organizationStrong grasp of software engineering fundamentals including distributed systems, data structures, concurrency, and system designExperience working with healthcare, life sciences, Real World Evidence (RWE), or large‑scale healthcare datasets is strongly preferredStrong coding experience in Python (PySpark), Scala, Java, or equivalent languages used for distributed processing systemsDeep expertise with distributed data processing frameworks such as Apache Spark and Hadoop, particularly within AWS environmentsStrong communication and collaboration skills across both technical and non‑technical stakeholdersExperience improving performance, scalability, observability, and infrastructure efficiency within distributed systemsExperience with streaming and event‑driven architectures using technologies such as Kafka or Spark StreamingExperience with entity resolution, identity mapping, automatching, deduplication, or large‑scale matching systems is strongly preferredProven ability to operate effectively within highly scalable, production‑grade distributed systemsFamiliarity with modern development and infrastructure tooling including Git, CI/CD pipelines, Docker, Kubernetes, Terraform, Argo, Hudi, and JIRA8+ years of experience building and maintaining large‑scale distributed data systems and pipelinesExperience with streaming technologies such as Kafka, Spark Streaming, or KSQLExperience performing root cause analysis across large‑scale distributed systems and complex data pipelinesStrong proficiency in Python (PySpark), Scala, Java, or other modern programming languages used for large‑scale distributed processingDemonstrated technical leadership experience mentoring engineers and driving complex technical initiativesExperience with orchestration and lakehouse technologies such as Argo and Hudi or comparable platformsYou are an experienced data engineer with deep expertise building and optimizing distributed data systems in cloud‑native environments. You thrive solving complex scalability and performance challenges across high‑volume data processing systems and enjoy operating in highly technical, fast‑paced engineering environmentsAbility to write clean, maintainable, modular, and production‑grade codeStrong understanding of distributed file formats including Apache Parquet and Apache AVRO

#J-18808-Ljbffr

Reference: WJ-766_22157625

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.