xgboost
the tree whisperer
XGBoost will fit leaked future information as readily as real signal, and on returns it often only ties a zero baseline. Before tuning, define the decision time, the target horizon, which features are available at that time, and where evaluation starts. Trust measured behavior on the installed version over memory: several popular snippets are written for the 1.x API.
Use when training, tuning, evaluating, or shipping XGBoost models -- XGBClassifier/XGBRegressor/XGBRanker sklearn API, native DMatrix/QuantileDMatrix + xgb.train, objectives and eval metrics, early stopping, categorical features, missing values, monotone/interaction constraints, custom objectives, multi-output, quantile regression, learning-to-rank, GPU (`device="cuda"`), SHAP (pred_contribs) and importance, model save/load (JSON/UBJ vs pickle), determinism, Dask/Spark, 1.x->2.x->3.x migrations, LightGBM/CatBoost choice, and especially financial price/return prediction on OHLCV candles (Binance klines) with leak-free walk-forward validation. TRIGGER when code imports `xgboost`, the user mentions xgboost, gradient boosting, GBDT, or asks to predict price, returns, or direction from market data.
methodology
- Check
xgb.__version__. This guru targets 3.x. Map any old code through the migration table before use: constructor-onlyearly_stopping_rounds,device=,iteration_range. - Pick the front door. Use the sklearn estimator by default. Use native
xgb.train+QuantileDMatrix(ref=)when you need memory savings, custom loops, or continuation. On the native path, slice predictions tobest_iteration + 1yourself. - For trading questions: choose a horizon return, direction, or quantile target; build features from bars at or before
tonly; run walk-forward folds with a gap of at leasth; print zero-return and persistence baselines on every fold. - Choose the objective and constraints deliberately: quantile for risk bands, monotone where domain logic demands it, native categoricals with a policy for unseen levels.
- Explain with
pred_contribs, which sums to the margin, and check with permutation importance. - Ship with
save_model("*.json" | "*.ubj"). Pin the version, and load with the same or a newer xgboost. Pinrandom_stateandn_jobs. - Verify any claim by running the matching script from the api-checks or walkforward guide. Record the result in trained with the date and versions.
contents
- trainedlearned layer: gotchas, known bugs, fixes, measured practiced cases (3.4.1). Read first
- xgboost-docsofficial docs snapshot: parameters, prediction, sklearn, categorical, constraints, custom objectives, multi-output, ranking, saving, dask/spark, release notes 2.0-3.4
- xgboostcore API cheat sheet: two APIs, objectives for market targets, parameter order, persistence, version migration table
- price-predictionmarket playbook: targets, features, leakage checklist, walk-forward, baselines, Binance data
- walkforwardrunnable leak-free walk-forward on real Binance klines, with measured results
- api-checksrunnable checks: parity, categorical, quantile, monotone, SHAP, model IO, cross-version, determinism, objectives
- xgboost-lightgbmcommunity: XGBoost vs LightGBM patterns. Uses the 1.x fit(early_stopping_rounds) and gpu_hist
- signal-classificationcommunity: trading signal classification with walk-forward scripts
- shapcommunity: SHAP explainers, plots, and workflows
Related gurus: ml, feature-engineering, and olap.