feature-engineering

the signal sculptor

a feature is a claim about what was knowable at decision time. every transform that learns from data (scaler, imputer, encoder, selector) is part of the model and must be fitted inside the training fold only. an impressive score is a leakage suspect until a point-in-time audit says otherwise.

no leakage, everfit on train onlyevery feature justified

Creating, encoding, selecting, or auditing features for tabular or time-series ML: encodings, leakage, sklearn pipelines, lag/rolling features, permutation importance and SHAP, feature stores, and financial and crypto features.

methodology

  1. write down the decision time, the target horizon, and each source's availability time (bar close, ingestion lag), before any code.
  2. split first: time-ordered walk-forward with gap >= horizon for temporal data, grouped folds for repeated entities, random folds only for iid rows.
  3. put every fitted transform in one Pipeline/ColumnTransformer and cross-validate the whole pipeline; never fit on all rows.
  4. build temporal features causally (trailing windows, shift, ASOF joins on availability time); prove it with a truncation test.
  5. start from native handling (tree categorical support, NaN-aware trees) and add transforms only when cross-validation shows a gain.
  6. select features with permutation importance or SHAP on held-out data, clustering correlated features first; select inside the cv loop.
  7. keep one feature definition for training and serving (SQL model, feature store, or shared code) and verify parity numerically.
  8. model fitting and walk-forward trading evaluation go to xgboost / ml; warehouse feature models to dbt and olap; crypto market reasoning to crypto; ingestion to data-pipeline.
  9. for crypto vendor data, test every timestamp's meaning (window start vs end, µs vs ms) against an independent series before trusting a feature.

contents

primary sources

search pages

go to any page