IT & Software

Data Engineer

CLR3

Toronto · Ontario · Canada

Build the pipelines behind datastore: decoding years of on-chain history into clean, versioned Parquet that researchers can trust.

datastore sells something unusual: files, not API access. Customers download decoded on-chain history as Parquet and run their own queries. That only works if the data is actually right, which makes correctness, lineage and reproducibility the product.

You will build and run the pipelines that decode Solana and Hyperliquid history at scale: backfills over billions of rows, schema design, checksums, manifests and the quality checks that let a quant trust a file they did not produce themselves.

What you will do

  • Design and run large-scale decoding and backfill pipelines
  • Model typed schemas for instructions, events and state across dozens of protocols
  • Build validation that catches bad data before customers do
  • Keep versioning, checksums, manifests and lineage docs accurate on every delivery
  • Tune storage layout and partitioning so files query fast in DuckDB and Polars
  • Add new protocols and chains to the catalogue
  • Support customer questions about schemas and coverage

What we are looking for

  • Experience building production data pipelines at meaningful scale
  • Strong SQL plus one of Python, Rust or Go
  • Real familiarity with columnar formats, ideally Parquet, and query engines like DuckDB, Polars or Spark
  • Care for data correctness: testing, validation and reconciliation
  • Comfort owning pipelines in production, including when they break at night
  • Able to work from our Toronto office part of the week

Nice to have

  • Experience with blockchain data or other messy, high-volume event streams
  • Familiarity with warehouse ecosystems your customers use, like Snowflake, BigQuery or ClickHouse
  • Background in quantitative research support or backtesting infrastructure
  • Experience with orchestration tools and with knowing when a cron job is enough

Who you are

  • You think an unverified number is worse than no number
  • You write pipelines you would be happy to debug at 2am, so they rarely need it
  • You like schemas that make the next person's query obvious
  • You get satisfaction from a backfill that reconciles to the last row

How we hire

  • 1 Intro call with an engineer, about 30 minutes
  • 2 Short take-home assignment working with a real decoded dataset
  • 3 Technical conversation about your assignment and pipelines you have run
  • 4 Conversation with the founders
  • 5 Offer
#J-18808-Ljbffr

Reference: WJ-2742_347862

Apply now

Continue on the employer's official application - the same link they use for every candidate.

More jobs

Find more on GigBlows

This role is listed on GigBlows for discovery and search. Hiring decisions and applications are handled by the employer or their chosen application system.