xgboost trained
Official sources win over memory. This page flags community snippets that the official documentation contradicts. Every measurement below was taken on xgboost 3.4.1 on macOS arm64, CPU only.
mental model
- XGBoost fits trees additively on the gradient and hessian of a loss. Each round adds one tree per output, or one vector-leaf tree when
multi_strategy="multi_output_tree". The prediction isbase_score(the intercept) plus the sum of the leaf values, in margin space. For classifiers,predict_probaapplies the link function to that margin, per the official model and prediction documentation. - There are two front doors to one engine. The sklearn estimators (
XGBRegressor,XGBClassifier,XGBRanker) take every training knob in the constructor; after early stopping,predict()automatically usesbest_iteration. - The native API is
DMatrixorQuantileDMatrixplusxgb.train. It returns the last-round model, andpredictuses every tree unless you passiteration_range. histhas been the defaulttree_methodsince 2.0. Features are binned into quantile histograms (max_bin=256).QuantileDMatrixbuilds those bins once. An evaluationQuantileDMatrixmust reuse the training bins throughref=.- For market data, the hard contract is what was known at decision time, not the model. Read the price-prediction playbook before any fit on candles.
examples
These snippets ran verbatim against 3.4.1. The api-checks and walkforward guides hold the full runnable scripts.
# sklearn: every knob goes in the constructor (2.1+). predict() uses best_iterationm = xgb.XGBRegressor(n_estimators=2000, learning_rate=0.05, max_depth=6,early_stopping_rounds=50, eval_metric="rmse", random_state=7, n_jobs=4)m.fit(x_tr, y_tr, eval_set=[(x_va, y_va)], verbose=False)# native: the same model. Slice to the best iteration yourselfdtr = xgb.QuantileDMatrix(x_tr, y_tr)dva = xgb.QuantileDMatrix(x_va, y_va, ref=dtr) # ref is required (3.0+)bst = xgb.train({"learning_rate": 0.05, "max_depth": 6, "seed": 7}, dtr, 2000,evals=[(dva, "val")], early_stopping_rounds=50, verbose_eval=False)p = bst.predict(xgb.DMatrix(x_te), iteration_range=(0, bst.best_iteration + 1))# quantile bands in one model (3.0+ accepts a vector alpha)xgb.train({"objective": "reg:quantileerror", "quantile_alpha": [0.05, 0.5, 0.95]}, dtr, 2000,evals=[(dva, "val")], early_stopping_rounds=100) # predict -> shape (n, 3)# monotone: tuple by column position, or dict by name (dict needs a DataFrame)xgb.XGBRegressor(monotone_constraints={"MedInc": 1})# SHAP: contributions sum to the MARGIN (logit for classifiers). The last column is the biasc = bst.predict(xgb.DMatrix(x), pred_contribs=True) # c.sum(1) == predict(output_margin=True)# save: extension picks the format. Always load with the same or a newer xgboostm.save_model("model.json") # or model.ubj
best practices
- Use the sklearn API by default. Switch to the native API for
QuantileDMatrixmemory savings, multiple eval sets with custom callbacks, or training continuation. - Set
early_stopping_roundsandeval_metricin the constructor. The eval set must sit after train and before test in time.predictthen uses the best iteration. - On any native path, always pass
iteration_range=(0, best_iteration + 1). In the practiced parity check, the full model and the best-iteration model differed by up to 0.21. - Pin
random_stateandn_jobs. Save models as.jsonfor diffable and portable artifacts or.ubjfor smaller ones (UBJ was 29% smaller in the practiced case). Store the column list, the training window, and the xgboost version next to the model. - For categoricals, use the pandas
categorydtype and native support. Set a policy for unseen categories before inference: retrain, add an explicit "other" level, or map during ETL. The official categorical guide covers the details. - Choose importance with care.
gainis the sklearn default.weightranks features differently. In the practiced case,weightandgainshared only 1 of their top 3 features. For honest attribution, use SHAP or permutation importance on a held-out fold. - Market data:
- Target a horizon return, not a price level.
- Use walk-forward folds with a gap of at least
hbefore both validation and test. - Print baselines beside the model.
- Treat a win against zero-return RMSE as necessary but not sufficient.
- The price-prediction playbook works through the full leakage checklist.
strengths
- Strong accuracy with little preprocessing on tabular, nonlinear data.
NaNis handled natively through a learned default direction. - A rich objective set: squared, pseudo-Huber, quantile (vector alpha), expectile (3.3+), logistic, softprob, ranking (
rank:ndcg,rank:pairwise), survival (AFT, Cox), Tweedie, and gamma, per the official parameter reference. - Exact TreeSHAP is built in (
pred_contribs,pred_interactions). Noshapdependency is needed, and it matchedshap.TreeExplainerto 0.0 in the practiced check. - Monotone and interaction constraints are first class.
- The same model can run on GPU with
device="cuda". - Distributed training is available on Dask and Spark.
- JSON is a stable, documented model schema.
weaknesses / pain points
- Trees do not extrapolate past the leaf values they learned. Level targets such as price fail in new regimes.
- Low-signal targets: on BTC 1h returns, XGBoost matched zero-return RMSE and never beat it on average. Boosting has no protection against fitting noise, so early stopping and the baselines carry the weight.
- Hyperparameter surface:
learning_rate,max_depth/max_leaves,min_child_weight,subsample,colsample_*,lambda/alpha,gamma, andmax_binall interact. - Categorical support is newer than CatBoost's. Unseen categories are a hard error at inference.
gblinearis deprecated (3.3). Random-forest wrappers are deprecated (3.4); usenum_parallel_treeinstead.- Comparison with alternatives (from the vendor docs; this skill has not benchmarked them):
- LightGBM grows trees leaf-wise by default and is often faster on large, wide data.
- CatBoost uses ordered target statistics for high-cardinality categoricals and symmetric trees, and its defaults work well untuned.
- XGBoost is the choice when you need the most mature tooling for GPU, distributed training, constraints, SHAP, and model schemas.
- The xgboost-lightgbm guide collects side-by-side community patterns.
gotchas
- Native
predictignores early stopping.xgb.trainreturns every round. Passiteration_range=(0, best_iteration + 1). ABoosterrebuilt fromsave_raw()also uses all trees: in the practiced case its predictions differed from the sklearn model's by up to 0.48. The sklearnload_modelpath restoresbest_iteration. pred_contribssums to the margin, not the probability. For a classifier the sum is the logit. In the practiced case the sum and the probability differed by up to 10.8.- A newer model loaded into an older xgboost can predict silently wrong.
- 3.x writes
base_scoreas a vector string ("[-4.2386947E0]"). - xgboost 2.1.4 loads JSON, UBJ, or pickle from 3.4.1 without error but reads
base_score=0.5, so every prediction shifts by a constant 4.74. - Always load with the writer's version or a newer one.
- 3.x writes
- The sklearn
fit(early_stopping_rounds=...)is gone (2.1+). It raisesTypeError. The xgboost-lightgbm community guide still uses it, along withtree_method="gpu_hist"andgpu_id(both removed in 3.1; the error lists{'approx','auto','exact','hist'}). - Old keyword arguments fail silently.
use_label_encoderis ignored with only a warning ("Parameters ... are not used"). The signal-classification community guide still passes it. - Monotone constraint dicts are matched by name. A dict with numpy input raises
Constrained features are not a subset of training data feature names. Use a DataFrame, or use a positional tuple. - Categorical support is on by default in 3.3+. A pandas
categorycolumn is split as a categorical withoutenable_categorical=True. On older versions it raised. Frames with reordered categories are re-coded (the diff was 0.0), but an unseen category raisesFound a category not in the training set. - Custom objectives skip intercept estimation.
base_scorestays 0.5 (built-in objectives fit it). Setbase_scoreexplicitly. Withbase_score=0.0a custom squared-error objective matched the built-in exactly. The sklearnpredictwith a callable objective returns the raw margin. - Quantile coverage is not guaranteed. A 90% band covered 0.857 on California housing when early stopping never triggered (1,999 rounds). On BTC 4h returns it covered 0.907. 3.4 changed quantile and MAE leaf estimation to a smooth approximation, so retrained quantile models differ from 3.3.
- Binance spot kline timestamps changed units. From 2025 they are microseconds (
1735689600000000); earlier files use milliseconds. Normalize before any time index or join. Also drop the last, still-open bar from live API pulls. XGBRankerrequires sortedqidvalues (non-decreasing, grouped). Unsortedqidraisesqid must be sorted in non-decreasing order.save_modelwith an unknown extension (.bin,.model) writes UBJSON with a warning. The legacy binary format was removed in 3.1.
known bugs
| issue | versions | effect | workaround |
|---|---|---|---|
| #12527 (open) | 3.2.0 | training continuation from a model with an empty category container silently drops feature_types. GPU trains a wrong model; CPU raises an error | pass feature_types and categories explicitly per categorical.rst "Auto-recoding"; retrain from scratch |
| #12177 (open) | 3.2.0, Windows | sparse polars categorical codes over-allocate, and the process dies silently (access violation) | use pandas category, or dense codes |
| #11826 (open) | 3.1+ | an unseen category at inference is a hard error with the re-coder (reproduced here on 3.4.1) | map unseen values to NaN or an "other" level in ETL |
| #12248 (open) | 3.x hist | recomputed hessian sums in tree nodes do not exactly match the cover stored in trees | do not rely on cover as an exact row count |
| #11210 (open) | all | reg:squaredlogerror produces NaN when a prediction is <= -1 | clip the labels, or use a log-target with squared error |
| #12545 (open) | 2.1.4–3.4.0, JVM/Linux | native memory grows for every DMatrix during predict | reuse DMatrix objects, or run inference from Python or C |
| #11982 (closed) | 3.x | resuming from a checkpoint occasionally dropped accuracy | 3.3 serializes RNG state (v3.3.0); upgrade |
problems -> fixes
| symptom | cause | fix |
|---|---|---|
TypeError: fit() got an unexpected keyword argument 'early_stopping_rounds' | removed from fit in 2.1 | move it to the constructor |
Invalid Input: 'gpu_hist' | removed in 3.1 | device="cuda", tree_method="hist" |
unexpected keyword argument 'ntree_limit' | removed in 2.0 | iteration_range=(0, n) |
Training dataset should be used as a reference... | eval QuantileDMatrix without ref (3.0+) | QuantileDMatrix(x_va, y_va, ref=dtr) |
| predictions shifted by a constant after a deploy | a newer model loaded by an older runtime (base_score parse) | align versions; load with a version >= the writer |
| native predictions worse than sklearn's for the same params | all trees used after early stopping | iteration_range |
Constrained features are not a subset... | dict constraint on numpy input | pass a DataFrame or a tuple |
TypeError: _estimator_type undefined from xgboost 2.1.x | scikit-learn >= 1.6 with an old xgboost | scikit-learn<1.6 for 2.1.x, or upgrade xgboost |
| walk-forward h=1 on 32k bars took 6.7 min of wall clock at 12% CPU | host load from unrelated work at the time (not xgboost) | set n_jobs explicitly; time on an idle host |
| backtest looks strong | leakage: target shift, fold overlap, test reuse | walk the leakage checklist in the price-prediction playbook |
practiced cases
Environment: xgboost 3.4.1, scikit-learn 1.9.1, pandas 3.0.6, numpy 2.5.3, shap 0.52.0, Python 3.12, macOS arm64 CPU. The api-checks and walkforward guides hold the scripts.
- walkforward on real Binance data (BTCUSDT 1h, 2023-01..2026-08, 32,135 bars, 59 folds). Numbers are fold means.
- h=1: xgb RMSE 0.00472 vs zero 0.00471. Hit rate 0.508. IC 0.021. XGBoost beat zero in 30.5% of folds.
- h=4: 0.00924 vs 0.00923. Hit rate 0.510. IC 0.000. Beat zero in 40.7% of folds.
- Synthetic random walk (null case): hit rate 0.500, IC -0.023.
- Conclusion: no predictive edge from these OHLCV features. The script's leak-free fold layout held.
- (a) sklearn vs native parity (California housing, early stopping 50): both reached
best_iteration=220; the native model kept 271 rounds. Sklearn vs native with the best-iteration range: max difference 0.0. With the full native model: max difference 0.214. - (b) categorical (OpenML adult, 48,842 rows, 8 categorical columns, 99 levels, same random split, early stopping 100):
variant columns AUC logloss iterations native 14 0.9246 0.2839 275 default (no flag, identical to native) 14 0.9246 0.2839 275 one-hot 113 0.9255 0.2821 415 integer codes as numbers 14 0.9242 0.2839 411 Native support was not more accurate here. It gave a smaller design matrix and needed fewer rounds. Fit times were not comparable because of host load.
- (c) quantile [0.05, 0.5, 0.95], one model, no crossing (0.0000):
data 90% coverage share below each quantile stopped at California (random split) 0.857 0.073 / 0.500 / 0.931 round 1,999 (did not stop) BTC forward 4h, time split with a gap 0.907 0.049 / 0.504 / 0.956 round 259 - (d) monotone (MedInc +1, 500 test rows, ICE grid of 60 values):
- Unconstrained: 500/500 rows had a decreasing ICE curve. Worst step -0.95. RMSE 0.4559.
- Constrained: 0/500. RMSE 0.4837, 6% worse.
- A dict with numpy input raised; a dict with a DataFrame worked.
- (e) SHAP:
sum(pred_contribs) - marginmax 6.2e-6.- Interactions summed back to the contributions (6.4e-6).
- The bias column (1.98594) equaled the config
base_score(1.98596). shap.TreeExplainermatched to 0.0.- For the classifier, the contributions summed to the logit (9.5e-6), not the probability.
- (f) model IO:
.json(355,791 B),.ubj,.bin, and.model(251,855 B each; the last two are UBJ with a warning) all reloaded with a max difference of 0.0,best_iterationrestored, and categorical feature names kept.- Pickle round trip: 0.0.
- Cross-version: 2.1.4 → 3.4.1 loads were exact for JSON, UBJ, and pickle. 3.4.1 → 2.1.4 loaded without error and was wrong by a constant 4.7388 (
base_scoreread as 0.5).
- (g) determinism:
- Identical prediction and model hashes across repeated runs and across
n_jobs1, 4, and 8 (CPU hist,subsample/colsampleat 0.7). random_state=Noneequaled seed 0. Seed 1 changed the hash.- GPU determinism is not verified here.
- Identical prediction and model hashes across repeated runs and across
- objectives:
- A custom squared objective with
base_score=0matched the built-in (0.0). - Multi-output:
one_output_per_treebuilt 200 trees (RMSE 6.915);multi_output_treebuilt 100 trees (RMSE 7.756). XGBRankerrank:ndcgreached train ndcg@5 0.9518. Unsortedqidraised.
- A custom squared objective with
- Not verified: CUDA (
device="cuda"), Dask and Spark, external memory, LightGBM and CatBoost benchmarks, and 1.x models loading in 3.x.
ecosystem
- Official documentation:
- parameters, prediction, tree methods, GPU
- sklearn, saving models, categorical, multi-output, ranking, custom objectives
- Dask, Spark
- release notes 2.0 through 3.4
- awesome-xgboost
- Distributed at a glance:
from xgboost import dask as dxgb(3.0+) givesdxgb.DaskXGBRegressor, andxgboost.spark.SparkXGBRegressorcovers Spark (PySpark 3.4+). The official parameter reference notes that the distributed AUC is an approximation. - Community guides, the top skills.sh results for "xgboost":
- xgboost-lightgbm (tondevrel/scientific-agent-skills, 450 installs) uses the 1.x-era API; see gotcha 4.
- signal-classification (agiprolabs/claude-trading-skills, 465 installs).
- shap (davila7/claude-code-templates, 367 installs).
- Siblings:
- ml: model selection and evaluation.
- feature-engineering: features.
- olap: storing and querying klines at scale.