ml trained

Every measured claim here used Python 3.12, scikit-learn 1.9.1, LightGBM 4.7.0, XGBoost 3.4.1, CatBoost 1.2.10 (installed, not benchmarked), Optuna 5.0.0, MLflow 3.16.1, statsforecast 2.1.1, pandas 2.3.3, onnxruntime 1.30.0, onnxmltools (converter opset cap 15), and skl2onnx 1.20.0.

mental model

decision + prediction time
-> target and horizon (what is known at t? what is predicted for t+h?)
-> split that mirrors production (iid: stratified k-fold | entities: group k-fold | time: walk-forward, gap >= h)
-> baseline (prior / naive / seasonal naive / persistence)
-> model ladder (linear -> forest -> boosting; stop when gain < fold noise)
-> tuning nested inside training folds, one untouched holdout scored once
-> calibration + intervals (fit on data the model never saw; measure coverage)
-> track (MLflow) -> export -> reload parity -> serve -> monitor drift vs baseline

Leakage and overfitting both show up as a better score. The splitter decides what "out of sample" means, so choose it before choosing the model. See the official common pitfalls page and the cross-validation page.

examples

Three worked examples, each writing only to the working directory:

  • tabular-classification: model ladder, PR-AUC trap, and calibration.
  • timeseries-forecasting: baselines vs statsforecast vs LightGBM, interval coverage, and shuffled-KFold inflation.
  • reproducible-pipeline: Optuna inside walk-forward, MLflow sqlite, and export parity.

Honest time-series CV with a purge:

from sklearn.model_selection import TimeSeriesSplit, cross_validate
cv = TimeSeriesSplit(n_splits=5, gap=H) # gap >= horizon purges overlapping targets
cross_validate(model, X, y, cv=cv, scoring="neg_mean_absolute_error")

Calibration without refitting the model: wrap it in FrozenEstimator instead of the removed cv="prefit":

from sklearn.calibration import CalibratedClassifierCV
from sklearn.frozen import FrozenEstimator
cal = CalibratedClassifierCV(FrozenEstimator(fitted), method="sigmoid").fit(X_cal, y_cal)

The scikit-learn calibration guide and scikit-learn#29893 cover the change.

Conformal intervals in statsforecast:

from statsforecast.utils import ConformalIntervals
sf.cross_validation(df=df, h=24, step_size=24, n_windows=30, level=[80, 90],
prediction_intervals=ConformalIntervals(h=24, n_windows=5))

See the conformal prediction guide.

best practices

  • Put the baseline in the same table. A model is reported only as a ratio to its baseline. Use MAE/SNaive (MASE-style) for forecasts and the lift over prevalence for PR-AUC. In the practiced cases the baselines gave the needed context: PR-AUC 0.117 for the prior, and SNaive24 MAE 0.64.
  • Keep every fitted transform inside the CV loop (Pipeline, ColumnTransformer). Scalers, encoders, imputers, and feature selection fitted on all rows leak information (the common pitfalls page). Feature-level leakage audits belong in feature-engineering.
  • Match the splitter to production.
    • Use StratifiedKFold(shuffle=True) for iid rows with a rare class.
    • Use GroupKFold or StratifiedGroupKFold when a customer, patient, or ticker repeats.
    • Use TimeSeriesSplit(gap=h) or a rolling-origin loop for anything temporal.
    • scikit-learn has no purged/embargoed splitter. gap covers the purge. Add an embargo by widening gap past the feature lookback when labels also feed features. Rolling-origin min_train_size/step options are only proposed so far (scikit-learn PR #33424).
  • Nest tuning. Optuna and GridSearchCV score on inner folds. The holdout is touched once, after the choice is made (the grid search tutorial).
  • Seed the sampler with TPESampler(seed=...) and run it sequentially if you need identical studies. Parallel trials are not reproducible (the Optuna faq).
  • Use boosting defaults first. For LightGBM, control complexity with num_leaves (keep it below 2^max_depth) and min_data_in_leaf (the parameter tuning guide). For XGBoost depth, early stopping, and GPU, see xgboost.
  • Handle imbalance with the metric and the threshold, not resampling. Report PR-AUC, log loss, and Brier. Pick the operating threshold with TunedThresholdClassifierCV (the classification threshold guide). Class weights move probabilities, so recalibrate after using them. SMOTE inside a time series pulls future neighbours into synthetic rows.
  • Choose the tracking backend deliberately. Use MLflow with sqlite:///mlflow.db, and log params, seed, data hash, metrics, and the model with input_example, which infers the signature.
  • Prove the exported artifact. Reload it and compare predictions against the in-memory model before calling it shipped.

strengths

  • On tabular data, gradient-boosted trees are the default winner at little tuning cost. They were best or tied in both measured tables. LightGBM and XGBoost scored within fold noise of each other (PR-AUC 0.453 vs 0.453, SD about 0.005-0.011).
  • statsforecast MSTL is a strong, fast-to-write forecaster for multi-seasonal series. It was the best point forecast here (MAE 0.667x of SNaive24), with close-to-nominal conformal coverage.
  • Conformal intervals are model-agnostic and only need held-out residuals.
  • Optuna with the ask/tell or objective API wraps any validation loop, including walk-forward.

weaknesses / pain points

  • Tuning rarely pays as much as framing. 30 TPE trials improved walk-forward MAE from 0.4449 to 0.4387 (1.4%). Adding lag features over seasonal naive improved it 29% (0.633 to 0.448 holdout).
  • Rolling-origin evaluation with classical stats models is slow. AutoETS and MSTL with 5 conformal windows took 1,027 s for 30 cutoffs on 2,880-hour windows (single process, contending with another job).
  • deterministic=True in LightGBM does not cover every source of nondeterminism (LightGBM#6731). Pin the thread count, force_row_wise, and the seed as well.
  • Prophet is not benchmarked here. Treat it as one candidate that must beat seasonal naive and MSTL on rolling-origin CV before adoption, never as the default. It was not run here, so this is unverified.
  • Neural nets (darts, sktime deep models, TabNet-style models) are worth trying only with many related series, rich exogenous signals, or raw text/image/sequence inputs. None were measured here.

gotchas

  1. Accuracy on imbalance hides everything. The majority-class dummy scored 0.883 accuracy against 0.894 for the best model, an apparent gain of 1.1 points. PR-AUC for the same pair was 0.117 vs 0.453, a 3.9x lift. A class-weighted LightGBM scored below the dummy on accuracy@0.5 (0.840).
  2. ROC-AUC flatters imbalanced problems. The best model reached ROC-AUC 0.80 but PR-AUC only 0.45. Read PR-AUC against the prevalence baseline.
  3. Class weights destroy calibration. is_unbalance=True gave a mean predicted probability of 0.317 against a true rate of 0.117, and Brier went from 0.1354 to 0.0820 after sigmoid calibration. Calibration does not change ranking: ROC-AUC was identical at 0.7891. Isotonic calibration slightly lowered PR-AUC (0.4425 to 0.4217) because it creates ties.
  4. Shuffled KFold on overlapping time targets invents skill. On the forward 24h BTC return, shuffled KFold gave R² 0.347, a 69.6% hit rate, and a gross Sharpe of 9.7. Walk-forward with gap=24 on the same model gave R² -0.221, a 50.3% hit rate, and a Sharpe of -0.97.
  5. Unshuffled KFold is also wrong for time. It trains on the future. Here it gave R² -0.11 and a Sharpe of 0.77, spuriously positive.
  6. Interval coverage is not what the label says. Nominal 90% intervals covered 75-76% for naive-family conformal intervals, 79.7% for AutoETS, 88.5% for MSTL, and 83.2% for LightGBM quantile intervals, measured over 720 held-out hours. Always measure coverage. Quantile GBMs under-cover because each quantile is overfit separately.
  7. MAPE explodes near zero and is asymmetric. Prefer MAE/MASE, or model the log target (as in the volume case) where the error is relative. Also, conformal scores computed in log space become multiplicative after expm1 and can blow up (statsforecast#1119).
  8. A leak can come from a column. The bank-marketing V12 column is call duration, known only after the outcome. Ask for every column: is it known at prediction time?
  9. Binance spot kline timestamps are µs from 2025-01-01 onward and ms before that. Detect the unit per file (open_time > 1e14 means µs), or 2025+ bars land in 1970.
  10. LightGBM lives at lightgbm-org/LightGBM, and issue links under microsoft/ redirect. Categorical columns must stay pandas category dtype at predict time (LightGBM#5244).

known bugs

Each entry lists the version, the defect, and the workaround. No patches are used.

  • MLflow 3.16.1: set_tracking_uri("file:./mlruns") raises MlflowException: filesystem tracking backend ... is in maintenance mode. The official tracking docs still say "mlruns by default", which no longer holds. The fix is sqlite:///mlflow.db (the local database guide), or mlflow migrate-filestore for old data. MLFLOW_ALLOW_FILE_STORE=true opts out but is not recommended. Discussion: mlflow#23525.
  • onnxmltools with LightGBM: target_opset=17 raises "higher than ... converter support (15)". Pass target_opset=15 or omit it. Categorical LightGBM splits have a long-open conversion issue (onnxmltools#309). Convert numeric/ordinal-encoded models, or verify parity per model.
  • XGBoost 3.4.1 / LightGBM 4.7.0: fit(..., early_stopping_rounds=...) raises TypeError. XGBoost takes it in the constructor. LightGBM uses callbacks=[lgb.early_stopping(n)]. The community xgboost-lightgbm snippets use the removed form.
  • XGBoost 3.x: use_label_encoder=False only warns ("Parameters not used"). Remove it from the community signal-classification snippets.
  • statsforecast conformal intervals and log transforms: the intervals blow up after inverse transform (statsforecast#1119). Report intervals in the modelled space, or conformalize on the original scale.

troubleshooting

symptomroot causefix
the CV score looks great and the live score is a coin flipshuffled or future-including splits on temporal data, or overlapping horizon targetsuse TimeSeriesSplit(gap=h) or a rolling-origin loop, and compare against the walk-forward number only
PR-AUC or AUC is high but precision at 0.5 is uselessthe threshold stayed at 0.5 on imbalanced or reweighted scorescalibrate, then tune TunedThresholdClassifierCV on the business cost
predicted probabilities average 3x the base rateis_unbalance, scale_pos_weight, or resamplingrecalibrate on a split the model never saw: sigmoid when calibration rows are few, isotonic above about 1,000 rows
a 90% interval covers 75%the intervals come from in-sample residuals or overfit quantile modelsconformalize on held-out windows and measure coverage per horizon
the tuned model is worse on the holdout than the defaultthe tuner overfit a single validation split, or the holdout leaked into the searchrun a nested walk-forward search and score the holdout once
mlflow refuses ./mlrunsFileStore is in maintenance mode (MLflow 3.x)track to sqlite:///mlflow.db
ONNX conversion fails on the opsetthe converter caps the supported opsetpass target_opset=15
the reloaded model gives different numbersfloat32 vs float64 inputs, a lost category dtype, or library version driftpin versions, store the category mapping, and assert parity at export
mlflow.log_model warns "Failed to resolve installed pip version"the uv venv has no piprun uv pip install pip, or ignore it: only the conda.yaml pin is affected

practiced cases

tabular classification, OpenML bank-marketing (id 1461)

The run used 45,211 rows with 11.7% positive. V12 (call duration) was dropped as a leak. Stratified 5-fold, mean scores:

modelaccuracyROC-AUCPR-AUCBrierlog loss
dummy prior0.88300.50000.11700.10330.3609
logistic (OHE+scale)0.89250.76540.40590.08600.3014
random forest0.89410.79420.44110.08280.2892
LightGBM0.89340.80020.45310.08120.2839
XGBoost0.89420.79880.45250.08160.2855

The calibration check used a held-out 25% test set. The model was fitted on 70% of train, and the calibrator on the remaining 30%:

modelBrierlog lossmean pacc@0.5
LightGBM is_unbalance raw0.13540.43800.3170.8399
+ sigmoid0.08200.28920.1140.8943
+ isotonic0.08170.29260.1140.8945

BTCUSDT 1h klines, data.binance.vision, 2024-01 to 2026-08 (23,376 bars, 0 gaps)

Forecasting log quote volume, 24h ahead, 30 daily cutoffs (720 hours). Stats models were fitted on a trailing 120 days. LightGBM used all history before each cutoff.

modelMAEvs SNaive24cov80cov90width90
Naive0.59150.9190.6610.7531.921
SeasonalNaive 240.64351.0000.6810.7602.165
SeasonalNaive 1680.71441.1100.6620.7582.201
AutoETS(24)0.53190.8270.7060.7971.760
MSTL(24,168)0.42900.6670.7830.8851.746
LightGBM lags>=24 (quantile q)0.45230.7030.7250.8321.569

Seasonal naive lost to plain naive because volume regimes shift faster than the daily cycle repeats. MSTL beat LightGBM on this single series.

Forward 24h log return, same LightGBM, 23,184 rows:

splitterR²hit rateICgross Sharpe (24h holds)
KFold(shuffle=True)0.34720.69550.5699.69
KFold(shuffle=False)-0.11250.49770.0180.77
TimeSeriesSplit(gap=24)-0.22070.50340.005-0.97

The base rate for "up" was 0.521. The honest model does not beat always-long, and no costs were applied. The trading-specific path lives in xgboost.

reproducible pipeline (Optuna 5.0.0 + MLflow 3.16.1)

The 30-trial run (sqlite backend) tuned and logged: walk-forward MAE 0.4449 by default and 0.4387 at best; holdout MAE 0.4484, against 0.6325 for SNaive24.

The pipeline with target_opset=15 in the ONNX export and LightGBM n_jobs=4 was run twice with --trials 12:

  • Both runs were identical. They produced the same best params (num_leaves=9, min_child_samples=252, lr=0.0167, n_estimators=353), the same data hash (04c5b2ac16f4fbbc), walk-forward MAE 0.43927, and holdout MAE 0.4539. The seeded TPESampler(seed) with sequential trials and LightGBM deterministic=True held here.
  • Export parity against in-memory predictions over 1,440 holdout rows:
    • joblib: exact.
    • LightGBM text model.txt: exact.
    • MLflow pyfunc reload: exact.
    • ONNX (float32): max |diff| 1.47e-05. ONNX is not bit-identical, so assert it with a tolerance.
  • Wall time was 36-45 s per run. The 30-trial run with default LightGBM threads took 1,184 s (sys 446 s) while another job was running. The cause was thread oversubscription. Pin n_jobs in tuning loops.
  • More trials did not win on holdout. 12 trials scored holdout MAE 0.4539, and 30 trials scored 0.4484. Both were about 0.7x of SNaive24. The search budget moved the result less than the baseline gap did.

ecosystem

  • Official docs (local text snapshots, pinned to a recent commit of each project):
    • scikit-learn reference and guides
    • LightGBM docs (repo at lightgbm-org/LightGBM)
    • Optuna docs and tutorials
    • MLflow classic ML docs
    • statsforecast README and notebooks (converted to markdown, outputs stripped)
  • Community skills (from skills.sh): mle-workflow (affaan-m/ecc, 5.6K installs), scikit-learn (k-dense-ai/scientific-agent-skills, 1.9K), ml-pipeline (jeffallan/claude-skills, 3.5K), xgboost-lightgbm (tondevrel/scientific-agent-skills), and signal-classification (agiprolabs/claude-trading-skills).
  • Where a community skill conflicts with the official docs, the official side wins:
    • Early stopping is passed in the constructor or through callbacks, not fit.
    • use_label_encoder is gone.
    • A bare TimeSeriesSplit(n_splits=5) without gap in the community scikit-learn evaluation guide is unsafe for horizon > 1.
    • SMOTE examples need a time-series caveat.
  • Awesome lists (trimmed): josephmisiti/awesome-machine-learning (Python section), EthicalML/awesome-production-machine-learning (serving, monitoring, and tracking sections), and cuge1995/awesome-time-series.
  • Libraries by job:
    • Models: scikit-learn, LightGBM, XGBoost (xgboost), CatBoost.
    • Classical forecasting: statsforecast.
    • Multi-model and deep forecasting: sktime, darts, neuralforecast (not practiced).
    • Tuning: Optuna.
    • Tracking and registry: MLflow.
    • Export: ONNX (skl2onnx, onnxmltools), native model files, joblib (same versions only).
    • Drift: Evidently, NannyML, and the others listed in awesome-production-ml (not practiced).
    • Upstream data: features (feature-engineering), warehouse extracts (olap), and modelled training tables (dbt).

Read the ml skill.

search pages

go to any page