dagster

the asset gardener

declare the data that should exist (assets), how it splits (partitions), and what makes it valid (checks). dagster then derives the graph, runs, staleness, and backfills. keep storage explicit, keep resources injectable, and verify everything headless before opening the ui.

declares what existslineage firstmaterialises on purpose

Use when building, testing, running, or debugging Dagster -- software-defined assets, asset checks (blocking gates), partitions and partition mappings, backfills, schedules, sensors, declarative automation, FreshnessPolicy, resources and config, IO managers, jobs and ops, executors, Definitions and code locations, the dg CLI and components, dagster-dbt, dagster-duckdb, dagster-dlt, Pipes, unit tests with materialize(), dagster dev, and Dagster+ deployment.

methodology

  1. model assets and a time partition first. add entity loops inside the asset. declare lookback with partition mappings.
  2. write storage explicitly and return MaterializeResult when the layout is a contract. use an io manager only for in-process handoffs.
  3. put every external dependency in a resource. put gates in @asset_check(blocking=True).
  4. pick the trigger: partitioned schedule for "yesterday", a sensor with a cursor and run_key for source changes, AutomationCondition for declarative upkeep.
  5. pin in_process_executor for single-writer stores. for range backfills give every asset a single-run policy and read partition_time_window.
  6. prove it headless: dagster definitions validate or dg check defs, cli materialize, pytest with materialize() and build_sensor_context / build_schedule_context. verify api names against the installed version.
  7. record lessons in trained. for pipeline design questions see data-pipeline.

contents

related gurus: data-pipeline (design, correctness, airflow/prefect comparison), dbt, olap (duckdb), feature-engineering.

search pages

go to any page