xgboost trained

Official sources win over memory. This page flags community snippets that the official documentation contradicts. Every measurement below was taken on xgboost 3.4.1 on macOS arm64, CPU only.

mental model

  • XGBoost fits trees additively on the gradient and hessian of a loss. Each round adds one tree per output, or one vector-leaf tree when multi_strategy="multi_output_tree". The prediction is base_score (the intercept) plus the sum of the leaf values, in margin space. For classifiers, predict_proba applies the link function to that margin, per the official model and prediction documentation.
  • There are two front doors to one engine. The sklearn estimators (XGBRegressor, XGBClassifier, XGBRanker) take every training knob in the constructor; after early stopping, predict() automatically uses best_iteration.
  • The native API is DMatrix or QuantileDMatrix plus xgb.train. It returns the last-round model, and predict uses every tree unless you pass iteration_range.
  • hist has been the default tree_method since 2.0. Features are binned into quantile histograms (max_bin=256). QuantileDMatrix builds those bins once. An evaluation QuantileDMatrix must reuse the training bins through ref=.
  • For market data, the hard contract is what was known at decision time, not the model. Read the price-prediction playbook before any fit on candles.

examples

These snippets ran verbatim against 3.4.1. The api-checks and walkforward guides hold the full runnable scripts.

# sklearn: every knob goes in the constructor (2.1+). predict() uses best_iteration
m = xgb.XGBRegressor(n_estimators=2000, learning_rate=0.05, max_depth=6,
early_stopping_rounds=50, eval_metric="rmse", random_state=7, n_jobs=4)
m.fit(x_tr, y_tr, eval_set=[(x_va, y_va)], verbose=False)
# native: the same model. Slice to the best iteration yourself
dtr = xgb.QuantileDMatrix(x_tr, y_tr)
dva = xgb.QuantileDMatrix(x_va, y_va, ref=dtr) # ref is required (3.0+)
bst = xgb.train({"learning_rate": 0.05, "max_depth": 6, "seed": 7}, dtr, 2000,
evals=[(dva, "val")], early_stopping_rounds=50, verbose_eval=False)
p = bst.predict(xgb.DMatrix(x_te), iteration_range=(0, bst.best_iteration + 1))
# quantile bands in one model (3.0+ accepts a vector alpha)
xgb.train({"objective": "reg:quantileerror", "quantile_alpha": [0.05, 0.5, 0.95]}, dtr, 2000,
evals=[(dva, "val")], early_stopping_rounds=100) # predict -> shape (n, 3)
# monotone: tuple by column position, or dict by name (dict needs a DataFrame)
xgb.XGBRegressor(monotone_constraints={"MedInc": 1})
# SHAP: contributions sum to the MARGIN (logit for classifiers). The last column is the bias
c = bst.predict(xgb.DMatrix(x), pred_contribs=True) # c.sum(1) == predict(output_margin=True)
# save: extension picks the format. Always load with the same or a newer xgboost
m.save_model("model.json") # or model.ubj

best practices

  • Use the sklearn API by default. Switch to the native API for QuantileDMatrix memory savings, multiple eval sets with custom callbacks, or training continuation.
  • Set early_stopping_rounds and eval_metric in the constructor. The eval set must sit after train and before test in time. predict then uses the best iteration.
  • On any native path, always pass iteration_range=(0, best_iteration + 1). In the practiced parity check, the full model and the best-iteration model differed by up to 0.21.
  • Pin random_state and n_jobs. Save models as .json for diffable and portable artifacts or .ubj for smaller ones (UBJ was 29% smaller in the practiced case). Store the column list, the training window, and the xgboost version next to the model.
  • For categoricals, use the pandas category dtype and native support. Set a policy for unseen categories before inference: retrain, add an explicit "other" level, or map during ETL. The official categorical guide covers the details.
  • Choose importance with care. gain is the sklearn default. weight ranks features differently. In the practiced case, weight and gain shared only 1 of their top 3 features. For honest attribution, use SHAP or permutation importance on a held-out fold.
  • Market data:
    • Target a horizon return, not a price level.
    • Use walk-forward folds with a gap of at least h before both validation and test.
    • Print baselines beside the model.
    • Treat a win against zero-return RMSE as necessary but not sufficient.
    • The price-prediction playbook works through the full leakage checklist.

strengths

  • Strong accuracy with little preprocessing on tabular, nonlinear data. NaN is handled natively through a learned default direction.
  • A rich objective set: squared, pseudo-Huber, quantile (vector alpha), expectile (3.3+), logistic, softprob, ranking (rank:ndcg, rank:pairwise), survival (AFT, Cox), Tweedie, and gamma, per the official parameter reference.
  • Exact TreeSHAP is built in (pred_contribs, pred_interactions). No shap dependency is needed, and it matched shap.TreeExplainer to 0.0 in the practiced check.
  • Monotone and interaction constraints are first class.
  • The same model can run on GPU with device="cuda".
  • Distributed training is available on Dask and Spark.
  • JSON is a stable, documented model schema.

weaknesses / pain points

  • Trees do not extrapolate past the leaf values they learned. Level targets such as price fail in new regimes.
  • Low-signal targets: on BTC 1h returns, XGBoost matched zero-return RMSE and never beat it on average. Boosting has no protection against fitting noise, so early stopping and the baselines carry the weight.
  • Hyperparameter surface: learning_rate, max_depth/max_leaves, min_child_weight, subsample, colsample_*, lambda/alpha, gamma, and max_bin all interact.
  • Categorical support is newer than CatBoost's. Unseen categories are a hard error at inference.
  • gblinear is deprecated (3.3). Random-forest wrappers are deprecated (3.4); use num_parallel_tree instead.
  • Comparison with alternatives (from the vendor docs; this skill has not benchmarked them):
    • LightGBM grows trees leaf-wise by default and is often faster on large, wide data.
    • CatBoost uses ordered target statistics for high-cardinality categoricals and symmetric trees, and its defaults work well untuned.
    • XGBoost is the choice when you need the most mature tooling for GPU, distributed training, constraints, SHAP, and model schemas.
    • The xgboost-lightgbm guide collects side-by-side community patterns.

gotchas

  1. Native predict ignores early stopping. xgb.train returns every round. Pass iteration_range=(0, best_iteration + 1). A Booster rebuilt from save_raw() also uses all trees: in the practiced case its predictions differed from the sklearn model's by up to 0.48. The sklearn load_model path restores best_iteration.
  2. pred_contribs sums to the margin, not the probability. For a classifier the sum is the logit. In the practiced case the sum and the probability differed by up to 10.8.
  3. A newer model loaded into an older xgboost can predict silently wrong.
    • 3.x writes base_score as a vector string ("[-4.2386947E0]").
    • xgboost 2.1.4 loads JSON, UBJ, or pickle from 3.4.1 without error but reads base_score=0.5, so every prediction shifts by a constant 4.74.
    • Always load with the writer's version or a newer one.
  4. The sklearn fit(early_stopping_rounds=...) is gone (2.1+). It raises TypeError. The xgboost-lightgbm community guide still uses it, along with tree_method="gpu_hist" and gpu_id (both removed in 3.1; the error lists {'approx','auto','exact','hist'}).
  5. Old keyword arguments fail silently. use_label_encoder is ignored with only a warning ("Parameters ... are not used"). The signal-classification community guide still passes it.
  6. Monotone constraint dicts are matched by name. A dict with numpy input raises Constrained features are not a subset of training data feature names. Use a DataFrame, or use a positional tuple.
  7. Categorical support is on by default in 3.3+. A pandas category column is split as a categorical without enable_categorical=True. On older versions it raised. Frames with reordered categories are re-coded (the diff was 0.0), but an unseen category raises Found a category not in the training set.
  8. Custom objectives skip intercept estimation. base_score stays 0.5 (built-in objectives fit it). Set base_score explicitly. With base_score=0.0 a custom squared-error objective matched the built-in exactly. The sklearn predict with a callable objective returns the raw margin.
  9. Quantile coverage is not guaranteed. A 90% band covered 0.857 on California housing when early stopping never triggered (1,999 rounds). On BTC 4h returns it covered 0.907. 3.4 changed quantile and MAE leaf estimation to a smooth approximation, so retrained quantile models differ from 3.3.
  10. Binance spot kline timestamps changed units. From 2025 they are microseconds (1735689600000000); earlier files use milliseconds. Normalize before any time index or join. Also drop the last, still-open bar from live API pulls.
  11. XGBRanker requires sorted qid values (non-decreasing, grouped). Unsorted qid raises qid must be sorted in non-decreasing order.
  12. save_model with an unknown extension (.bin, .model) writes UBJSON with a warning. The legacy binary format was removed in 3.1.

known bugs

issueversionseffectworkaround
#12527 (open)3.2.0training continuation from a model with an empty category container silently drops feature_types. GPU trains a wrong model; CPU raises an errorpass feature_types and categories explicitly per categorical.rst "Auto-recoding"; retrain from scratch
#12177 (open)3.2.0, Windowssparse polars categorical codes over-allocate, and the process dies silently (access violation)use pandas category, or dense codes
#11826 (open)3.1+an unseen category at inference is a hard error with the re-coder (reproduced here on 3.4.1)map unseen values to NaN or an "other" level in ETL
#12248 (open)3.x histrecomputed hessian sums in tree nodes do not exactly match the cover stored in treesdo not rely on cover as an exact row count
#11210 (open)allreg:squaredlogerror produces NaN when a prediction is <= -1clip the labels, or use a log-target with squared error
#12545 (open)2.1.4–3.4.0, JVM/Linuxnative memory grows for every DMatrix during predictreuse DMatrix objects, or run inference from Python or C
#11982 (closed)3.xresuming from a checkpoint occasionally dropped accuracy3.3 serializes RNG state (v3.3.0); upgrade

problems -> fixes

symptomcausefix
TypeError: fit() got an unexpected keyword argument 'early_stopping_rounds'removed from fit in 2.1move it to the constructor
Invalid Input: 'gpu_hist'removed in 3.1device="cuda", tree_method="hist"
unexpected keyword argument 'ntree_limit'removed in 2.0iteration_range=(0, n)
Training dataset should be used as a reference...eval QuantileDMatrix without ref (3.0+)QuantileDMatrix(x_va, y_va, ref=dtr)
predictions shifted by a constant after a deploya newer model loaded by an older runtime (base_score parse)align versions; load with a version >= the writer
native predictions worse than sklearn's for the same paramsall trees used after early stoppingiteration_range
Constrained features are not a subset...dict constraint on numpy inputpass a DataFrame or a tuple
TypeError: _estimator_type undefined from xgboost 2.1.xscikit-learn >= 1.6 with an old xgboostscikit-learn<1.6 for 2.1.x, or upgrade xgboost
walk-forward h=1 on 32k bars took 6.7 min of wall clock at 12% CPUhost load from unrelated work at the time (not xgboost)set n_jobs explicitly; time on an idle host
backtest looks strongleakage: target shift, fold overlap, test reusewalk the leakage checklist in the price-prediction playbook

practiced cases

Environment: xgboost 3.4.1, scikit-learn 1.9.1, pandas 3.0.6, numpy 2.5.3, shap 0.52.0, Python 3.12, macOS arm64 CPU. The api-checks and walkforward guides hold the scripts.

  • walkforward on real Binance data (BTCUSDT 1h, 2023-01..2026-08, 32,135 bars, 59 folds). Numbers are fold means.
    • h=1: xgb RMSE 0.00472 vs zero 0.00471. Hit rate 0.508. IC 0.021. XGBoost beat zero in 30.5% of folds.
    • h=4: 0.00924 vs 0.00923. Hit rate 0.510. IC 0.000. Beat zero in 40.7% of folds.
    • Synthetic random walk (null case): hit rate 0.500, IC -0.023.
    • Conclusion: no predictive edge from these OHLCV features. The script's leak-free fold layout held.
  • (a) sklearn vs native parity (California housing, early stopping 50): both reached best_iteration=220; the native model kept 271 rounds. Sklearn vs native with the best-iteration range: max difference 0.0. With the full native model: max difference 0.214.
  • (b) categorical (OpenML adult, 48,842 rows, 8 categorical columns, 99 levels, same random split, early stopping 100):
    variantcolumnsAUCloglossiterations
    native140.92460.2839275
    default (no flag, identical to native)140.92460.2839275
    one-hot1130.92550.2821415
    integer codes as numbers140.92420.2839411

    Native support was not more accurate here. It gave a smaller design matrix and needed fewer rounds. Fit times were not comparable because of host load.

  • (c) quantile [0.05, 0.5, 0.95], one model, no crossing (0.0000):
    data90% coverageshare below each quantilestopped at
    California (random split)0.8570.073 / 0.500 / 0.931round 1,999 (did not stop)
    BTC forward 4h, time split with a gap0.9070.049 / 0.504 / 0.956round 259
  • (d) monotone (MedInc +1, 500 test rows, ICE grid of 60 values):
    • Unconstrained: 500/500 rows had a decreasing ICE curve. Worst step -0.95. RMSE 0.4559.
    • Constrained: 0/500. RMSE 0.4837, 6% worse.
    • A dict with numpy input raised; a dict with a DataFrame worked.
  • (e) SHAP:
    • sum(pred_contribs) - margin max 6.2e-6.
    • Interactions summed back to the contributions (6.4e-6).
    • The bias column (1.98594) equaled the config base_score (1.98596).
    • shap.TreeExplainer matched to 0.0.
    • For the classifier, the contributions summed to the logit (9.5e-6), not the probability.
  • (f) model IO:
    • .json (355,791 B), .ubj, .bin, and .model (251,855 B each; the last two are UBJ with a warning) all reloaded with a max difference of 0.0,best_iteration restored, and categorical feature names kept.
    • Pickle round trip: 0.0.
    • Cross-version: 2.1.4 → 3.4.1 loads were exact for JSON, UBJ, and pickle. 3.4.1 → 2.1.4 loaded without error and was wrong by a constant 4.7388 (base_score read as 0.5).
  • (g) determinism:
    • Identical prediction and model hashes across repeated runs and across n_jobs 1, 4, and 8 (CPU hist, subsample/colsample at 0.7).
    • random_state=None equaled seed 0. Seed 1 changed the hash.
    • GPU determinism is not verified here.
  • objectives:
    • A custom squared objective with base_score=0 matched the built-in (0.0).
    • Multi-output: one_output_per_tree built 200 trees (RMSE 6.915); multi_output_tree built 100 trees (RMSE 7.756).
    • XGBRanker rank:ndcg reached train ndcg@5 0.9518. Unsorted qid raised.
  • Not verified: CUDA (device="cuda"), Dask and Spark, external memory, LightGBM and CatBoost benchmarks, and 1.x models loading in 3.x.

ecosystem

  • Official documentation:
    • parameters, prediction, tree methods, GPU
    • sklearn, saving models, categorical, multi-output, ranking, custom objectives
    • Dask, Spark
    • release notes 2.0 through 3.4
    • awesome-xgboost
  • Distributed at a glance: from xgboost import dask as dxgb (3.0+) gives dxgb.DaskXGBRegressor, and xgboost.spark.SparkXGBRegressor covers Spark (PySpark 3.4+). The official parameter reference notes that the distributed AUC is an approximation.
  • Community guides, the top skills.sh results for "xgboost":
    • xgboost-lightgbm (tondevrel/scientific-agent-skills, 450 installs) uses the 1.x-era API; see gotcha 4.
    • signal-classification (agiprolabs/claude-trading-skills, 465 installs).
    • shap (davila7/claude-code-templates, 367 installs).
  • Siblings:

search pages

go to any page