feature-engineering trained

Every number on this page is measured, on python 3.12.14, scikit-learn 1.9.1, pandas 3.0.6, polars 1.44.2, duckdb 1.5.5, and numpy 2.5.3.

mental model

raw rows
splitby time, group, or random
train fold
fit
transforms
model
test fold
transform only
score
the two definitions everything else follows from
feature(t) = f(data available at or before decision time t)
target(t) = g(t+1 .. t+h)
  • A learned transform (scaler mean, encoder category means, imputer median, selected column set) is model state. Fitting it on test rows leaks; the scikit-learn common pitfalls guide says the same.
  • Temporal leakage has three forms: a feature window that reaches forward (look-ahead), a validation split that shares overlapping labels (contamination), and a join keyed on event time instead of availability time (point-in-time violation; the Feast point-in-time joins guide covers it).
  • Trees need little numeric preprocessing but benefit from good categorical handling. Linear models and distance methods need scaling, encoding, and explicit interactions (Kuhn & Johnson ch.1-2).

examples

The core shapes, each runnable end to end:

# tabular: every fitted step inside the cv loop (examples/tabular-encodings/run.py)
pipe = make_pipeline(
ColumnTransformer([("te", TargetEncoder(target_type="binary", cv=KFold(5, shuffle=True, random_state=0)), cat_cols)],
remainder="passthrough"),
HistGradientBoostingClassifier(categorical_features="from_dtype"))
cross_val_score(pipe, X, y, cv=StratifiedKFold(5, shuffle=True, random_state=0), scoring="roc_auc")
# time series: trailing windows only, target from the future, walk-forward with gap (examples/ohlcv-leakage/features.py)
df.with_columns(ret_1=pl.col("close").log().diff()).with_columns(
vol_24=pl.col("ret_1").rolling_std(24), # bars t-23..t
fwd_ret=pl.col("close").log().shift(-24) - pl.col("close").log()) # target only
TimeSeriesSplit(n_splits=5, gap=24)
-- same features in duckdb, null warm-up mirrored with a count guard (examples/duckdb-parity/run.py)
case when count(ret_1) over w24 = 24 then stddev_samp(ret_1) over w24 end as vol_24
window w24 as (order by ts rows between 23 preceding and current row)
-- point-in-time: join on availability, not bar open
from events e asof join feat f on e.event_ts >= f.ts + interval 1 hour

best practices

  • Split before any preprocessing; cross-validate the whole Pipeline (the scikit-learn common pitfalls guide).
  • Train TargetEncoder only through fit_transform(X_train, y_train) (cross fitting). A plain fit then transform on the same rows leaks; the scikit-learn preprocessing guide states it in one line: "fit(X, y).transform(X) does not equal fit_transform".
  • For gradient boosting, start with native categoricals (categorical_features="from_dtype" on a pandas category dtype) or OrdinalEncoder(handle_unknown="use_encoded_value"). On adult all three tree encodings tied at 0.929 AUC, so choose by cardinality and serving simplicity, not by habit.
  • Linear models: OneHotEncoder(handle_unknown="infrequent_if_exist", min_frequency=...) plus a scaler.
  • Temporal cv: TimeSeriesSplit(gap=h) where h is the label horizon. For overlapping labels use purging/embargo (López de Prado, AFML ch.7). Repeated entities need GroupKFold (the scikit-learn cross-validation guide).
  • Stamp every feature row with available_at, and ASOF-join labels on it using backward direction and an inclusive >=.
  • Cyclical time: encode hour/day as sin/cos pairs, or use SplineTransformer(extrapolation="periodic") (the scikit-learn preprocessing guide).
  • Financial targets: model returns or triple-barrier outcomes, not price levels. Scale barriers by volatility known at t. Modelling belongs in xgboost.
  • One feature definition for train and serve (a dbt model or Feast feature view). Assert numeric parity when two engines compute the same feature.

strengths

  • sklearn Pipeline + ColumnTransformer makes leakage-safe preprocessing the default and serializes with the model.
  • TargetEncoder (sklearn >= 1.3) compresses high-cardinality categoricals and cross-fits by design.
  • polars expressions (rolling_*, over, join_asof) and DuckDB window/ASOF SQL compute the same point-in-time features. They matched to 1.3e-13 in practice, which makes SQL-first feature definitions viable.
  • Feast provides point-in-time retrieval with a ttl bound. Hand-rolled joins get this wrong.

weaknesses / pain points

  • Permutation importance under-reports correlated features. Cluster them and keep one per cluster first (the scikit-learn permutation importance guide).
  • Impurity (MDI) importance is biased toward high-cardinality features and is computed on training data. Prefer permutation importance or SHAP on held-out data.
  • Feature stores add infrastructure. For one model on one warehouse, a dbt model plus an ASOF join is cheaper.
  • Rolling features have a warm-up period. Dropping it silently removes the earliest regime, and filling it invents data.
  • Financial features are mostly weak. Honest point-in-time AUC on BTC 24h direction was 0.523 (base rate 0.520).

gotchas

  1. Global target-mean encoding leaks even for pure noise. It measured +0.030 AUC on adult, and the cross-fitted TargetEncoder removed it.
  2. Shuffled KFold on overlapping temporal labels raised BTC AUC from 0.523 to 0.759 with no model change.
  3. center=True rolling windows (polars/pandas) look forward half a window. The BTC AUC was 0.813, and the truncation test failed.
  4. Binance kline ts is the bar open time. The bar becomes known one interval later. A naive ASOF join on open time used unfinished bars for 5 of 5 events.
  5. Binance spot timestamps switched to microseconds on 2025-01-01 (binance-public-data README). Normalize with where(x > 1e14, x // 1000).
  6. DuckDB returns timestamptz in the session TimeZone (the machine's local zone) and as µs to polars. Joins then fail with a dtype mismatch. SET TimeZone = 'UTC' and cast.
  7. Rolling std nulls vs SQL windows: polars rolling_std(n) stays null until n values exist. A SQL ROWS n-1 PRECEDING window emits partial-window values. Add a count(...) over w = n guard. Both default to sample std (ddof=1).
  8. TargetEncoder(random_state=..., shuffle=...) is deprecated in 1.9 (removal in 1.11). Pass cv=KFold(5, shuffle=True, random_state=0).
  9. Full-sample scaling barely moved AUC (0.5188 vs 0.5189, logreg). This is not proof of safety: contamination bias depends on drift. Keep the pipeline rule.
  10. OpenMP oversubscription: on a machine at load 280, HistGradientBoosting on 20k x 10 took more than 60s. With OMP_NUM_THREADS=1 it took 1.8s. Cap threads when many jobs share a host.
  11. TargetEncoder treats NaN as its own category and encodes unseen categories as the target mean (the scikit-learn preprocessing guide).
  12. join_asof needs both sides sorted on the key. Within by groups, polars did not check sortedness until the fix in pola-rs/polars#21693.

known bugs

issueeffectworkaround
duckdb/duckdb#25536 (closed 2026-09)ASOF JOIN picks a nondeterministic row when build rows tie on the ordering keydedupe to one row per (entity, available_at) before joining; upgrade past the fix
duckdb/duckdb#21244 (closed 2026-03)ASOF JOIN with LIMIT hangs on >= 64 rowsupgrade; avoid LIMIT directly on the ASOF join
pola-rs/polars#21693 (closed 2025-03)join_asof did not check sortedness within by groups, giving silently wrong matchessort by [by, key] explicitly
scikit-learn#28881 (open)TargetEncoder ignores sample_weightdo not rely on weighted encodings; encode in SQL if weights matter
scikit-learn#30952 (open)TargetEncoder.transform is slow for single rows with many categoriesprecompute a lookup table from encodings_ for online serving
feast-dev/feast#3426 (open)BigQuery offline store fails on complex feature viewssplit feature views; verify the retrieval SQL

troubleshooting

symptomroot causefix
cv score far above holdout or liveencoder, scaler, or selector fitted on all rowsmove it into the Pipeline, and cv the pipeline
time-series AUC 0.7+ on a liquid marketshuffled folds, centered windows, or shift(-k) in featureswalk-forward with gap >= h; truncation test; grep for negative shifts
truncation test changes past rowsfeature depends on future rowsmake windows trailing; compute from data <= t
SQL and pandas features differ on early rowspartial-window emission in SQLcount() over w = n guard
ASOF join returns a bar that was still openkeyed on the bar open timekey on available_at = open + interval
polars join SchemaError datetime[ms, UTC] vs datetime[us, Asia/...]DuckDB session TimeZoneSET TimeZone = 'UTC', cast to a common unit
sklearn cv or HGB crawlsOpenMP threads times cores on a busy hostOMP_NUM_THREADS=1 (or small), n_jobs at the cv level
fracdiff series still trendsd too smallraise d until stationary (ADF), keeping the highest level correlation; d=0.4 kept corr 0.917

practiced cases

  • Tabular encodings and target leakage. OpenML adult (1590), 48,842 rows, 5-fold stratified, ROC AUC, scikit-learn 1.9.1.
    • onehot + scaler + logreg: 0.9066 ± 0.0019.
    • ordinal + HGB: 0.9287.
    • native categorical HGB: 0.9289.
    • TargetEncoder + HGB: 0.9290.
    • Leakage demo, adding a 10k-level pure-noise id: baseline 0.9289. Global target mean (bug) gave 0.9586 (+0.0297). The cross-fitted TargetEncoder in the pipeline gave 0.9289 (+0.0000).
  • OHLCV look-ahead. BTCUSDT 1h, 2024-01 through 2026-08, 23,376 bars, polars 1.44.2, 24h direction target, base rate 0.520, HGB.
    • Walk-forward with gap 24: 0.5230.
    • Shuffled KFold: 0.7588.
    • Centered 168-bar z-score: 0.8128.
    • Full-sample scaling, logreg: 0.5188 vs 0.5189 fit-on-train.
    • Truncation test: PIT features identical; the centered feature changed.
  • AFML features.
    • fracdiff weights above 1e-4 at d=0.4: 282.
    • Correlation with log price by d: 0.2 gave 0.984, 0.4 gave 0.917, 0.6 gave 0.716, 1.0 gave 0.004.
    • Triple barrier (h=24, k=1 x daily-scaled 168h vol): -1: 5084, 0: 13223, +1: 4877.
    • Stationarity (ADF) was not tested.
  • DuckDB SQL parity. DuckDB 1.5.5 window SQL vs polars on 7 features.
    • Max abs diff at most 1.25e-13, with identical null masks (1/24/24/168/167/167/24).
    • The open-time ASOF join used unfinished bars for 5/5 events. The availability-time join was correct, and polars join_asof(strategy="backward") matched DuckDB.
  • Crypto feature layer. 18-symbol perp panel, 6 families, purged walk-forward: see the section below.
  • Not exercised in these runs: Feast runtime, dbt feature models, VIF and mutual information, SHAP runs, text features, hashing encoder, cross-sectional ranks across symbols.

crypto feature layer

Setup. Measured with polars 1.44.2, scikit-learn 1.9.1, scipy, and python 3.12.

  • Universe: 18 USD-M perps fixed as of 2024-01, including FTMUSDT and MKRUSDT, which were later delisted.
  • Panel: hourly, 2025-01-01 to 2026-08-31, 239,611 symbol-hours.
  • Decision time is bar close. Targets are close-to-close log returns at 1h and 1d.
  • Survival rule: |t| >= 2 on non-overlapping samples, and the same sign in at least 4 of 5 time folds.
  • All figures are gross IC, before fees and spread.

Crypto reasoning (why BTC moved, regimes) belongs to crypto. Ingestion and scheduling belong to data-pipeline.

what survived (rank IC, t)

familyfeature1h1dverdict
derivativesfunding rate, 8h-normalized, xs-0.013 (t -5.6)-0.025 (t -2.1, folds unstable)weak contrarian; 1h only
derivativesfunding z-score (90 settlements) xs / ts-0.010 (t -4.2) / -0.018 (t -2.1)n.s.weak, 1h only
derivativescumulative 7d fundingn.s.n.s.dead
derivativesperp-spot basis z (168h) xs-0.032 (t -14.6)n.s.1h only; likely perp-premium mean reversion plus bounce
derivativesOI change 24h, long/short account ratio, taker long/short (fixed), liquidation proxyn.s.n.s.dead
derivativesOI / 30d quote volume xsn.s.+0.028 (t +2.1)marginal
derivativestop-trader long/short position xs-0.007 (t -3.0)n.s.marginal
microstructureVWAP(24h) deviation xs / ts-0.033 (t -12.5) / -0.032 (t -3.7)xs -0.037 (t -2.8)strongest single feature: short-term reversal
microstructurekline taker-flow OFI 1h / 24hxs -0.005 / ts -0.026n.s.weak reversal, not continuation
microstructuretrade-count zxs -0.008 (t -3.3)n.s.marginal
microstructure1-minute trade OFI -> next minute+0.001-dead (+0.658 contemporaneous only)
cross-sectional24h momentum xs-0.027 (t -10.1)-0.028 (t -2.1)reversal, not momentum
cross-sectional7d / 30d-skip-1d momentum-0.008 / n.s.n.s.no momentum premium in 2025-26
cross-assetBTC residual 24h xs-0.024 (t -9.1)n.s.idiosyncratic reversal
cross-asset30d BTC beta xs-0.012 (t -3.9)-0.043 (t -2.6)high-beta alts underperformed (low-beta anomaly), sample-specific
cross-assetdominance proxies (BTC-minus-alts 7d, BTC volume share)--0.074 (t -1.8) / n.s.not significant
quarterly basisBTC annualized basis -> 1d / 7d-+0.042 (t +1.1) / +0.049 (t +0.5)not significant; mean 5.97%/yr, corr 0.80 with funding
on-chain7d change in active addresses, lagged 1d-+0.032 (se 0.032)not significant
on-chain30d change in hash rate: same-day vs lagged 1d-+0.068 vs +0.022the "signal" existed only without the publication lag

Combined model. A purged walk-forward ridge on 18 cross-sectionally ranked features scored OOS xs rank IC +0.0195 (t 7.4) at 1h and +0.038 (t 2.7) at 1d. Everything that survived is a 1h reversal or crowding effect, and 1h close-to-close reversal is partly bid-ask bounce. Pass it through xgboost walk-forward with fees before calling it alpha. No trend, momentum, OI, liquidation, or on-chain feature survived at 1d.

Volatility is where features clearly work. Target: log next-day RV from 5-minute BTC spot returns over 966 days; the estimator is fit on the first half and scored on the second (OOS R2):

estimator1 day7-day mean
close-to-close0.1350.382
Parkinson0.3470.452
Garman-Klass0.3420.433
Yang-Zhang-0.429
RV 5-minute0.4030.440

Range estimators beat close-to-close by 2.6x on a single day. A 7-day Parkinson estimate was the best forecaster, even beating 5-minute RV.

Bars (AFML). 7 days of BTCUSDT perp aggTrades (5.31M trades), sampled to equal bar counts (2,016 bars each):

bar typeexcess kurtosisJarque-Bera
time (5m)6.934,305
tick0.1314
volume0.205
dollar-0.082

Activity-clocked returns are close to normal, as López de Prado claims. Lag-1 autocorrelation stayed below 0.02 for every bar type.

crypto gotchas

  1. metrics.create_time is the window START. Measured: taker long/short at create_time T correlates 0.999 with the kline covering [T, T+5m) and 0.121 with [T-5m, T). Joining on create_time produced a fake taker_ls IC of +0.178 at 1h (t 20). With availability set to create_time + 5m it fell to -0.003. (examples/crypto-features/check_metrics_time.py). Treat every vendor timestamp as unknown until tested.
  2. Delisted perps keep publishing klines with volume 0 and a frozen close. FTMUSDT has 14,462 flat hours after 2025-01-06 and MKRUSDT has 8,583 after 2025-09-08. Cut each symbol at its last bar with volume > 0, or it enters the cross-section as 0% returns. Its funding file ran until 2025-06-19, months after trading stopped.
  3. Spot timestamps are µs from 2025-01-01; futures stay ms (the binance-public-data repository). Normalize by magnitude.
  4. Funding calc_time carries +1..+5 ms jitter (...000001), and the REST fundingTime does too. Truncate to the hour before joining; an exact-equality join drops rows.
  5. Funding intervals differ by symbol and change over time. fapi/v1/fundingInfo lists 4h symbols (e.g. LPTUSDT). Normalize to an 8h-equivalent rate (rate * 8 / interval_h) before comparing or z-scoring, using the funding_interval_hours column in the file. All 18 names in this universe stayed at 8h in 2024-26.
  6. Survivorship. Delisted symbols disappear from exchangeInfo (or show SETTLING), so a "current top-N" universe is survivor-biased. Here the as-of vs survivors-only difference in 30d momentum was small: IC -0.0168 vs -0.0177, and top-minus-bottom -4.4 vs -4.6 bp/day. The window held only 2 delistings. Choosing today's large caps is itself the bigger bias.
  7. Yang-Zhang degenerates on 24/7 markets. Open equals the previous close, so the overnight term is about 0 and it collapses toward Rogers-Satchell.
  8. Quarterly basis must be rolled. Use the nearest contract with >= 14 days to expiry, and annualize by days to expiry (expiry 08:00 UTC per the symbol suffix). Near expiry the basis converges to 0 and fakes a regime change.
  9. On-chain publication lag. A CoinMetrics day-D value exists only after D closes. Hash rate is an estimate from block times and gets revised. The same-day value showed +0.068 IC, but only +0.022 once lagged. Revision history was not observable from the community API (unverified), so store fetch-time snapshots.
  10. Contemporaneous flow is not a feature. 1m OFI vs the same-minute return had rank corr +0.658; against the next minute it was +0.001.
  11. The data.binance.vision CDN resets bursts of parallel connections. Use about 6 workers with exponential backoff, since 16 workers failed with ECONNRESET.

crypto sources

ecosystem

  • sklearn owns ColumnTransformer, TargetEncoder, SplineTransformer, KBinsDiscretizer, permutation_importance, and the mutual_info_* helpers.
  • polars and pandas each have their own guide; kdense-polars covers the pandas migration. DuckDB and ClickHouse belong to olap; dbt feature models to dbt.
  • Feature stores: Feast.
  • Explainability: the kdense-shap guide; XGBoost pred_contribs in xgboost.
  • Finance: López de Prado AFML (triple barrier, fracdiff, purged cv). Model training and walk-forward evaluation are in xgboost and ml.
  • Learning: Kaggle course, FES book, awesome-feature-engineering.

search pages

go to any page