feature-engineering
the signal sculptor
a feature is a claim about what was knowable at decision time. every transform that learns from data (scaler, imputer, encoder, selector) is part of the model and must be fitted inside the training fold only. an impressive score is a leakage suspect until a point-in-time audit says otherwise.
no leakage, everfit on train onlyevery feature justified
Creating, encoding, selecting, or auditing features for tabular or time-series ML: encodings, leakage, sklearn pipelines, lag/rolling features, permutation importance and SHAP, feature stores, and financial and crypto features.
methodology
- write down the decision time, the target horizon, and each source's availability time (bar close, ingestion lag), before any code.
- split first: time-ordered walk-forward with
gap >= horizonfor temporal data, grouped folds for repeated entities, random folds only for iid rows. - put every fitted transform in one
Pipeline/ColumnTransformerand cross-validate the whole pipeline; neverfiton all rows. - build temporal features causally (trailing windows,
shift, ASOF joins on availability time); prove it with a truncation test. - start from native handling (tree categorical support, NaN-aware trees) and add transforms only when cross-validation shows a gain.
- select features with permutation importance or SHAP on held-out data, clustering correlated features first; select inside the cv loop.
- keep one feature definition for training and serving (SQL model, feature store, or shared code) and verify parity numerically.
- model fitting and walk-forward trading evaluation go to xgboost / ml; warehouse feature models to dbt and olap; crypto market reasoning to crypto; ingestion to data-pipeline.
- for crypto vendor data, test every timestamp's meaning (window start vs end, µs vs ms) against an independent series before trusting a feature.
contents
- trainedlearned layer, measured cases, gotchas; read first
- common-pitfallsofficial: inconsistent preprocessing, data leakage, randomness
- preprocessingofficial: scalers, power transforms, encoders incl. TargetEncoder cross fitting, binning, splines
- composeofficial: Pipeline, ColumnTransformer, TransformedTargetRegressor
- imputeofficial: imputers and missing indicators
- feature-selectionofficial: univariate, RFE, SelectFromModel, sequential
- permutation-importanceofficial: permutation importance and the correlated-feature trap
- cross-validationofficial: TimeSeriesSplit gap, group folds
- point-in-time-joinsofficial: Feast point-in-time joins and ttl
- window-functionsofficial: over and group-wise features; rolling windows and as-of joins are sibling guides
- examplesrunnable, measured: encodings + target leakage, OHLCV look-ahead, fracdiff/triple barrier, DuckDB SQL parity
- crypto-featurescrypto: funding/OI/basis/long-short, RV estimators, dollar bars, OFI, BTC beta/residuals, survivorship, on-chain lag; measured IC per family
- binance-public-dataofficial Binance data archive layouts (klines, aggTrades, µs switch)
- kdense-scikit-learncommunity sklearn skill (preprocessing, pipelines)
- kdense-polarscommunity polars skill (pandas migration, performance)
- kdense-shapcommunity SHAP skill (explainers, maskers for correlated features)
- tradermonty-crypto-regimecommunity crypto regime score (descriptive, unvalidated)
- awesome-feature-engineeringcurated technique list
- sourcesthe upstream libraries and docs behind this guide