data-pipeline

the plumber of truth

a pipeline is a pure function per partition: the same inputs and code give the same output, and a rerun replaces the partition instead of appending to it. land the exact source bytes, derive the partition from event time, and gate every boundary. backfills, corrections, and crash recovery then all become one operation: rerun the partition.

idempotent loadsschemas are contractslate data is normal

Use when designing, building, reviewing, or debugging a data pipeline -- batch vs micro-batch vs streaming, ELT vs ETL, medallion/layered raw -> staged -> curated, lakehouse vs warehouse, idempotency, deterministic partitioning, exactly-once vs at-least-once + dedup, upserts/merge, late and out-of-order data, event time vs processing time, watermarks, backfills and reprocessing, schema drift and data contracts, CDC (Debezium), slowly changing dimensions, time zones and timestamp units, REST pagination/rate limits/retries, websocket ingest, dlt, Kafka/Redpanda basics, data-quality gates (freshness, row count, gaps, reconciliation), lineage, observability, and orchestrator choice (Dagster vs Airflow vs Prefect). Flagship case -- a crypto market-data pipeline (Binance klines + funding -> Parquet -> DuckDB -> ML features).

methodology

  1. write the contract: consumers, grain, freshness SLA, and the partition key (event-time day × entity). default to batch. stream only for sub-minute SLAs or several consumers.
  2. land raw bytes immutably with checksums. keep raw un-coerced. decide types, units, and time zones once, in staging.
  3. make every write an atomic partition replace with a declared natural key. prove it by running twice and comparing hashes.
  4. declare lookback windows and late-data handling. detect source changes (sha, CDC, updated_at) and rerun the changed partition plus its downstream window.
  5. add blocking gates: completeness and gaps, uniqueness, validity, reconciliation vs source, and freshness. fail one on purpose to prove downstream stops.
  6. orchestrate as assets and partitions (dagster). transform in SQL/dbt (dbt). store per olap.
  7. review with the design checklist. record lessons in trained.

contents

related gurus: dbt (transforms), olap (DuckDB, ClickHouse, lakehouse), postgres (OLTP sources and serving), feature-engineering (point-in-time features), ml (training on the output), crypto (market-data meaning).

search pages

go to any page