feature-engineering trained
Every number on this page is measured, on python 3.12.14, scikit-learn 1.9.1, pandas 3.0.6, polars 1.44.2, duckdb 1.5.5, and numpy 2.5.3.
mental model
feature(t) = f(data available at or before decision time t)target(t) = g(t+1 .. t+h)
- A learned transform (scaler mean, encoder category means, imputer median, selected column set) is model state. Fitting it on test rows leaks; the scikit-learn common pitfalls guide says the same.
- Temporal leakage has three forms: a feature window that reaches forward (look-ahead), a validation split that shares overlapping labels (contamination), and a join keyed on event time instead of availability time (point-in-time violation; the Feast point-in-time joins guide covers it).
- Trees need little numeric preprocessing but benefit from good categorical handling. Linear models and distance methods need scaling, encoding, and explicit interactions (Kuhn & Johnson ch.1-2).
examples
The core shapes, each runnable end to end:
# tabular: every fitted step inside the cv loop (examples/tabular-encodings/run.py)pipe = make_pipeline(ColumnTransformer([("te", TargetEncoder(target_type="binary", cv=KFold(5, shuffle=True, random_state=0)), cat_cols)],remainder="passthrough"),HistGradientBoostingClassifier(categorical_features="from_dtype"))cross_val_score(pipe, X, y, cv=StratifiedKFold(5, shuffle=True, random_state=0), scoring="roc_auc")
# time series: trailing windows only, target from the future, walk-forward with gap (examples/ohlcv-leakage/features.py)df.with_columns(ret_1=pl.col("close").log().diff()).with_columns(vol_24=pl.col("ret_1").rolling_std(24), # bars t-23..tfwd_ret=pl.col("close").log().shift(-24) - pl.col("close").log()) # target onlyTimeSeriesSplit(n_splits=5, gap=24)
-- same features in duckdb, null warm-up mirrored with a count guard (examples/duckdb-parity/run.py)case when count(ret_1) over w24 = 24 then stddev_samp(ret_1) over w24 end as vol_24window w24 as (order by ts rows between 23 preceding and current row)-- point-in-time: join on availability, not bar openfrom events e asof join feat f on e.event_ts >= f.ts + interval 1 hour
best practices
- Split before any preprocessing; cross-validate the whole
Pipeline(the scikit-learn common pitfalls guide). - Train
TargetEncoderonly throughfit_transform(X_train, y_train)(cross fitting). A plainfitthentransformon the same rows leaks; the scikit-learn preprocessing guide states it in one line: "fit(X, y).transform(X) does not equal fit_transform". - For gradient boosting, start with native categoricals (
categorical_features="from_dtype"on a pandascategorydtype) orOrdinalEncoder(handle_unknown="use_encoded_value"). On adult all three tree encodings tied at 0.929 AUC, so choose by cardinality and serving simplicity, not by habit. - Linear models:
OneHotEncoder(handle_unknown="infrequent_if_exist", min_frequency=...)plus a scaler. - Temporal cv:
TimeSeriesSplit(gap=h)wherehis the label horizon. For overlapping labels use purging/embargo (López de Prado, AFML ch.7). Repeated entities needGroupKFold(the scikit-learn cross-validation guide). - Stamp every feature row with
available_at, and ASOF-join labels on it using backward direction and an inclusive>=. - Cyclical time: encode hour/day as sin/cos pairs, or use
SplineTransformer(extrapolation="periodic")(the scikit-learn preprocessing guide). - Financial targets: model returns or triple-barrier outcomes, not price levels. Scale barriers by volatility known at
t. Modelling belongs in xgboost. - One feature definition for train and serve (a dbt model or Feast feature view). Assert numeric parity when two engines compute the same feature.
strengths
- sklearn
Pipeline+ColumnTransformermakes leakage-safe preprocessing the default and serializes with the model. TargetEncoder(sklearn >= 1.3) compresses high-cardinality categoricals and cross-fits by design.- polars expressions (
rolling_*,over,join_asof) and DuckDB window/ASOF SQL compute the same point-in-time features. They matched to 1.3e-13 in practice, which makes SQL-first feature definitions viable. - Feast provides point-in-time retrieval with a
ttlbound. Hand-rolled joins get this wrong.
weaknesses / pain points
- Permutation importance under-reports correlated features. Cluster them and keep one per cluster first (the scikit-learn permutation importance guide).
- Impurity (MDI) importance is biased toward high-cardinality features and is computed on training data. Prefer permutation importance or SHAP on held-out data.
- Feature stores add infrastructure. For one model on one warehouse, a dbt model plus an ASOF join is cheaper.
- Rolling features have a warm-up period. Dropping it silently removes the earliest regime, and filling it invents data.
- Financial features are mostly weak. Honest point-in-time AUC on BTC 24h direction was 0.523 (base rate 0.520).
gotchas
- Global target-mean encoding leaks even for pure noise. It measured +0.030 AUC on adult, and the cross-fitted
TargetEncoderremoved it. - Shuffled KFold on overlapping temporal labels raised BTC AUC from 0.523 to 0.759 with no model change.
center=Truerolling windows (polars/pandas) look forward half a window. The BTC AUC was 0.813, and the truncation test failed.- Binance kline
tsis the bar open time. The bar becomes known one interval later. A naive ASOF join on open time used unfinished bars for 5 of 5 events. - Binance spot timestamps switched to microseconds on 2025-01-01 (binance-public-data README). Normalize with
where(x > 1e14, x // 1000). - DuckDB returns
timestamptzin the session TimeZone (the machine's local zone) and as µs to polars. Joins then fail with a dtype mismatch.SET TimeZone = 'UTC'and cast. - Rolling std nulls vs SQL windows: polars
rolling_std(n)stays null until n values exist. A SQLROWS n-1 PRECEDINGwindow emits partial-window values. Add acount(...) over w = nguard. Both default to sample std (ddof=1). TargetEncoder(random_state=..., shuffle=...)is deprecated in 1.9 (removal in 1.11). Passcv=KFold(5, shuffle=True, random_state=0).- Full-sample scaling barely moved AUC (0.5188 vs 0.5189, logreg). This is not proof of safety: contamination bias depends on drift. Keep the pipeline rule.
- OpenMP oversubscription: on a machine at load 280,
HistGradientBoostingon 20k x 10 took more than 60s. WithOMP_NUM_THREADS=1it took 1.8s. Cap threads when many jobs share a host. TargetEncodertreats NaN as its own category and encodes unseen categories as the target mean (the scikit-learn preprocessing guide).join_asofneeds both sides sorted on the key. Withinbygroups, polars did not check sortedness until the fix in pola-rs/polars#21693.
known bugs
| issue | effect | workaround |
|---|---|---|
| duckdb/duckdb#25536 (closed 2026-09) | ASOF JOIN picks a nondeterministic row when build rows tie on the ordering key | dedupe to one row per (entity, available_at) before joining; upgrade past the fix |
| duckdb/duckdb#21244 (closed 2026-03) | ASOF JOIN with LIMIT hangs on >= 64 rows | upgrade; avoid LIMIT directly on the ASOF join |
| pola-rs/polars#21693 (closed 2025-03) | join_asof did not check sortedness within by groups, giving silently wrong matches | sort by [by, key] explicitly |
| scikit-learn#28881 (open) | TargetEncoder ignores sample_weight | do not rely on weighted encodings; encode in SQL if weights matter |
| scikit-learn#30952 (open) | TargetEncoder.transform is slow for single rows with many categories | precompute a lookup table from encodings_ for online serving |
| feast-dev/feast#3426 (open) | BigQuery offline store fails on complex feature views | split feature views; verify the retrieval SQL |
troubleshooting
| symptom | root cause | fix |
|---|---|---|
| cv score far above holdout or live | encoder, scaler, or selector fitted on all rows | move it into the Pipeline, and cv the pipeline |
| time-series AUC 0.7+ on a liquid market | shuffled folds, centered windows, or shift(-k) in features | walk-forward with gap >= h; truncation test; grep for negative shifts |
| truncation test changes past rows | feature depends on future rows | make windows trailing; compute from data <= t |
| SQL and pandas features differ on early rows | partial-window emission in SQL | count() over w = n guard |
| ASOF join returns a bar that was still open | keyed on the bar open time | key on available_at = open + interval |
polars join SchemaError datetime[ms, UTC] vs datetime[us, Asia/...] | DuckDB session TimeZone | SET TimeZone = 'UTC', cast to a common unit |
| sklearn cv or HGB crawls | OpenMP threads times cores on a busy host | OMP_NUM_THREADS=1 (or small), n_jobs at the cv level |
| fracdiff series still trends | d too small | raise d until stationary (ADF), keeping the highest level correlation; d=0.4 kept corr 0.917 |
practiced cases
- Tabular encodings and target leakage. OpenML adult (1590), 48,842 rows, 5-fold stratified, ROC AUC, scikit-learn 1.9.1.
- onehot + scaler + logreg: 0.9066 ± 0.0019.
- ordinal + HGB: 0.9287.
- native categorical HGB: 0.9289.
- TargetEncoder + HGB: 0.9290.
- Leakage demo, adding a 10k-level pure-noise id: baseline 0.9289. Global target mean (bug) gave 0.9586 (+0.0297). The cross-fitted TargetEncoder in the pipeline gave 0.9289 (+0.0000).
- OHLCV look-ahead. BTCUSDT 1h, 2024-01 through 2026-08, 23,376 bars, polars 1.44.2, 24h direction target, base rate 0.520, HGB.
- Walk-forward with gap 24: 0.5230.
- Shuffled KFold: 0.7588.
- Centered 168-bar z-score: 0.8128.
- Full-sample scaling, logreg: 0.5188 vs 0.5189 fit-on-train.
- Truncation test: PIT features identical; the centered feature changed.
- AFML features.
- fracdiff weights above 1e-4 at d=0.4: 282.
- Correlation with log price by d: 0.2 gave 0.984, 0.4 gave 0.917, 0.6 gave 0.716, 1.0 gave 0.004.
- Triple barrier (h=24, k=1 x daily-scaled 168h vol): -1: 5084, 0: 13223, +1: 4877.
- Stationarity (ADF) was not tested.
- DuckDB SQL parity. DuckDB 1.5.5 window SQL vs polars on 7 features.
- Max abs diff at most 1.25e-13, with identical null masks (1/24/24/168/167/167/24).
- The open-time ASOF join used unfinished bars for 5/5 events. The availability-time join was correct, and polars
join_asof(strategy="backward")matched DuckDB.
- Crypto feature layer. 18-symbol perp panel, 6 families, purged walk-forward: see the section below.
- Not exercised in these runs: Feast runtime, dbt feature models, VIF and mutual information, SHAP runs, text features, hashing encoder, cross-sectional ranks across symbols.
crypto feature layer
Setup. Measured with polars 1.44.2, scikit-learn 1.9.1, scipy, and python 3.12.
- Universe: 18 USD-M perps fixed as of 2024-01, including FTMUSDT and MKRUSDT, which were later delisted.
- Panel: hourly, 2025-01-01 to 2026-08-31, 239,611 symbol-hours.
- Decision time is bar close. Targets are close-to-close log returns at 1h and 1d.
- Survival rule: |t| >= 2 on non-overlapping samples, and the same sign in at least 4 of 5 time folds.
- All figures are gross IC, before fees and spread.
Crypto reasoning (why BTC moved, regimes) belongs to crypto. Ingestion and scheduling belong to data-pipeline.
what survived (rank IC, t)
| family | feature | 1h | 1d | verdict |
|---|---|---|---|---|
| derivatives | funding rate, 8h-normalized, xs | -0.013 (t -5.6) | -0.025 (t -2.1, folds unstable) | weak contrarian; 1h only |
| derivatives | funding z-score (90 settlements) xs / ts | -0.010 (t -4.2) / -0.018 (t -2.1) | n.s. | weak, 1h only |
| derivatives | cumulative 7d funding | n.s. | n.s. | dead |
| derivatives | perp-spot basis z (168h) xs | -0.032 (t -14.6) | n.s. | 1h only; likely perp-premium mean reversion plus bounce |
| derivatives | OI change 24h, long/short account ratio, taker long/short (fixed), liquidation proxy | n.s. | n.s. | dead |
| derivatives | OI / 30d quote volume xs | n.s. | +0.028 (t +2.1) | marginal |
| derivatives | top-trader long/short position xs | -0.007 (t -3.0) | n.s. | marginal |
| microstructure | VWAP(24h) deviation xs / ts | -0.033 (t -12.5) / -0.032 (t -3.7) | xs -0.037 (t -2.8) | strongest single feature: short-term reversal |
| microstructure | kline taker-flow OFI 1h / 24h | xs -0.005 / ts -0.026 | n.s. | weak reversal, not continuation |
| microstructure | trade-count z | xs -0.008 (t -3.3) | n.s. | marginal |
| microstructure | 1-minute trade OFI -> next minute | +0.001 | - | dead (+0.658 contemporaneous only) |
| cross-sectional | 24h momentum xs | -0.027 (t -10.1) | -0.028 (t -2.1) | reversal, not momentum |
| cross-sectional | 7d / 30d-skip-1d momentum | -0.008 / n.s. | n.s. | no momentum premium in 2025-26 |
| cross-asset | BTC residual 24h xs | -0.024 (t -9.1) | n.s. | idiosyncratic reversal |
| cross-asset | 30d BTC beta xs | -0.012 (t -3.9) | -0.043 (t -2.6) | high-beta alts underperformed (low-beta anomaly), sample-specific |
| cross-asset | dominance proxies (BTC-minus-alts 7d, BTC volume share) | - | -0.074 (t -1.8) / n.s. | not significant |
| quarterly basis | BTC annualized basis -> 1d / 7d | - | +0.042 (t +1.1) / +0.049 (t +0.5) | not significant; mean 5.97%/yr, corr 0.80 with funding |
| on-chain | 7d change in active addresses, lagged 1d | - | +0.032 (se 0.032) | not significant |
| on-chain | 30d change in hash rate: same-day vs lagged 1d | - | +0.068 vs +0.022 | the "signal" existed only without the publication lag |
Combined model. A purged walk-forward ridge on 18 cross-sectionally ranked features scored OOS xs rank IC +0.0195 (t 7.4) at 1h and +0.038 (t 2.7) at 1d. Everything that survived is a 1h reversal or crowding effect, and 1h close-to-close reversal is partly bid-ask bounce. Pass it through xgboost walk-forward with fees before calling it alpha. No trend, momentum, OI, liquidation, or on-chain feature survived at 1d.
Volatility is where features clearly work. Target: log next-day RV from 5-minute BTC spot returns over 966 days; the estimator is fit on the first half and scored on the second (OOS R2):
| estimator | 1 day | 7-day mean |
|---|---|---|
| close-to-close | 0.135 | 0.382 |
| Parkinson | 0.347 | 0.452 |
| Garman-Klass | 0.342 | 0.433 |
| Yang-Zhang | - | 0.429 |
| RV 5-minute | 0.403 | 0.440 |
Range estimators beat close-to-close by 2.6x on a single day. A 7-day Parkinson estimate was the best forecaster, even beating 5-minute RV.
Bars (AFML). 7 days of BTCUSDT perp aggTrades (5.31M trades), sampled to equal bar counts (2,016 bars each):
| bar type | excess kurtosis | Jarque-Bera |
|---|---|---|
| time (5m) | 6.93 | 4,305 |
| tick | 0.13 | 14 |
| volume | 0.20 | 5 |
| dollar | -0.08 | 2 |
Activity-clocked returns are close to normal, as López de Prado claims. Lag-1 autocorrelation stayed below 0.02 for every bar type.
crypto gotchas
metrics.create_timeis the window START. Measured: taker long/short atcreate_timeT correlates 0.999 with the kline covering [T, T+5m) and 0.121 with [T-5m, T). Joining oncreate_timeproduced a fake taker_ls IC of +0.178 at 1h (t 20). With availability set to create_time + 5m it fell to -0.003. (examples/crypto-features/check_metrics_time.py). Treat every vendor timestamp as unknown until tested.- Delisted perps keep publishing klines with volume 0 and a frozen close. FTMUSDT has 14,462 flat hours after 2025-01-06 and MKRUSDT has 8,583 after 2025-09-08. Cut each symbol at its last bar with volume > 0, or it enters the cross-section as 0% returns. Its funding file ran until 2025-06-19, months after trading stopped.
- Spot timestamps are µs from 2025-01-01; futures stay ms (the binance-public-data repository). Normalize by magnitude.
- Funding
calc_timecarries +1..+5 ms jitter (...000001), and the RESTfundingTimedoes too. Truncate to the hour before joining; an exact-equality join drops rows. - Funding intervals differ by symbol and change over time.
fapi/v1/fundingInfolists 4h symbols (e.g. LPTUSDT). Normalize to an 8h-equivalent rate (rate * 8 / interval_h) before comparing or z-scoring, using thefunding_interval_hourscolumn in the file. All 18 names in this universe stayed at 8h in 2024-26. - Survivorship. Delisted symbols disappear from
exchangeInfo(or showSETTLING), so a "current top-N" universe is survivor-biased. Here the as-of vs survivors-only difference in 30d momentum was small: IC -0.0168 vs -0.0177, and top-minus-bottom -4.4 vs -4.6 bp/day. The window held only 2 delistings. Choosing today's large caps is itself the bigger bias. - Yang-Zhang degenerates on 24/7 markets. Open equals the previous close, so the overnight term is about 0 and it collapses toward Rogers-Satchell.
- Quarterly basis must be rolled. Use the nearest contract with >= 14 days to expiry, and annualize by days to expiry (expiry 08:00 UTC per the symbol suffix). Near expiry the basis converges to 0 and fakes a regime change.
- On-chain publication lag. A CoinMetrics day-D value exists only after D closes. Hash rate is an estimate from block times and gets revised. The same-day value showed +0.068 IC, but only +0.022 once lagged. Revision history was not observable from the community API (unverified), so store fetch-time snapshots.
- Contemporaneous flow is not a feature. 1m OFI vs the same-minute return had rank corr +0.658; against the next minute it was +0.001.
- The data.binance.vision CDN resets bursts of parallel connections. Use about 6 workers with exponential backoff, since 16 workers failed with ECONNRESET.
crypto sources
- The binance-public-data repository documents the archive layouts and the µs switch.
- Binance USD-M futures API: funding rate history and funding info.
- CoinMetrics community API.
- López de Prado, AFML ch.2 (bars), ch.3, ch.5, ch.7.
- The community tradermonty crypto-regime-analyzer guide (2K installs) is a descriptive regime score, not a validated signal, as its own validation notes state. The funding and dominance ICs measured here agree.
ecosystem
- sklearn owns
ColumnTransformer,TargetEncoder,SplineTransformer,KBinsDiscretizer,permutation_importance, and themutual_info_*helpers. - polars and pandas each have their own guide;
kdense-polarscovers the pandas migration. DuckDB and ClickHouse belong to olap; dbt feature models to dbt. - Feature stores: Feast.
- Explainability: the
kdense-shapguide; XGBoostpred_contribsin xgboost. - Finance: López de Prado AFML (triple barrier, fracdiff, purged cv). Model training and walk-forward evaluation are in xgboost and ml.
- Learning: Kaggle course, FES book,
awesome-feature-engineering.