xgboost

the tree whisperer

XGBoost will fit leaked future information as readily as real signal, and on returns it often only ties a zero baseline. Before tuning, define the decision time, the target horizon, which features are available at that time, and where evaluation starts. Trust measured behavior on the installed version over memory: several popular snippets are written for the 1.x API.

early stopping alwaystuned on validationexplains with shap

Use when training, tuning, evaluating, or shipping XGBoost models -- XGBClassifier/XGBRegressor/XGBRanker sklearn API, native DMatrix/QuantileDMatrix + xgb.train, objectives and eval metrics, early stopping, categorical features, missing values, monotone/interaction constraints, custom objectives, multi-output, quantile regression, learning-to-rank, GPU (`device="cuda"`), SHAP (pred_contribs) and importance, model save/load (JSON/UBJ vs pickle), determinism, Dask/Spark, 1.x->2.x->3.x migrations, LightGBM/CatBoost choice, and especially financial price/return prediction on OHLCV candles (Binance klines) with leak-free walk-forward validation. TRIGGER when code imports `xgboost`, the user mentions xgboost, gradient boosting, GBDT, or asks to predict price, returns, or direction from market data.

methodology

  1. Check xgb.__version__. This guru targets 3.x. Map any old code through the migration table before use: constructor-only early_stopping_rounds, device=, iteration_range.
  2. Pick the front door. Use the sklearn estimator by default. Use native xgb.train + QuantileDMatrix(ref=) when you need memory savings, custom loops, or continuation. On the native path, slice predictions to best_iteration + 1 yourself.
  3. For trading questions: choose a horizon return, direction, or quantile target; build features from bars at or before t only; run walk-forward folds with a gap of at least h; print zero-return and persistence baselines on every fold.
  4. Choose the objective and constraints deliberately: quantile for risk bands, monotone where domain logic demands it, native categoricals with a policy for unseen levels.
  5. Explain with pred_contribs, which sums to the margin, and check with permutation importance.
  6. Ship with save_model("*.json" | "*.ubj"). Pin the version, and load with the same or a newer xgboost. Pin random_state and n_jobs.
  7. Verify any claim by running the matching script from the api-checks or walkforward guide. Record the result in trained with the date and versions.

contents

Related gurus: ml, feature-engineering, and olap.

search pages

go to any page