ml

the careful experimenter

a model is only as good as its honest out-of-sample number. leakage and overfitting produce the best-looking scores, so every result is suspect until it beats a dumb baseline on data the model, the features, and the tuner never touched, split the way the model will meet production (by time, by group).

baseline firstvalidation that matches realitytracks every run

Use when building or evaluating machine learning models -- framing the problem, train/validation splits, baselines and model ladders, metrics from the decision, calibration and intervals, experiment tracking with MLflow, scikit-learn, LightGBM, Optuna, StatsForecast, and time-series forecasting.

methodology

  1. frame first: the decision, prediction time, target, horizon, unit of evaluation, cost of each error, and which data exist at prediction time.
  2. build the split before the model: stratified k-fold for iid data, group k-fold per entity, walk-forward with a gap at least the horizon; tune nested; keep one final holdout.
  3. baselines first (prior/majority, naive, seasonal naive, persistence), then the ladder linear/logistic -> random forest -> gradient boosting; adopt a bigger model only if it beats by more than fold noise.
  4. the metric comes from the decision: pr-auc/log loss/brier under imbalance, mae/mase against seasonal naive, ic/hit rate/net-of-cost sharpe; report the fold spread.
  5. calibrate on data the model did not fit; use conformal/quantile intervals and measure coverage.
  6. log params, data hash, seed, metrics, and the model with mlflow on sqlite; export-reload-assert identical; monitor drift in production.
  7. gradient boosting depth -> xgboost; feature and leakage questions -> feature-engineering.

contents

related gurus: xgboost, feature-engineering, olap, dbt.

search pages

go to any page