Model Performance
Training metrics for every model in the registry · last trained 2026-09-24 02:05
21
Models
0
Training fights
0
Upcoming fights
Strategy Backtest — vs. Main
Recency-Heavy
CatBoost (ordered boosting) on the full engineered feature set (strategy: Recent form dominates career-long averages much more (365-day recency half-life vs main's 730) — tests whether a slump reflects something main's longer memory is underweighting.)
5 tied
With Odds
CatBoost (ordered boosting) on the full engineered feature set (strategy: Main's exact feature config plus de-vigged market-implied-probability features from BestFightOdds (opening line, closing range, and how much the line moved) — isolates whether the betting market's own signal improves on our features alone. Needs the historical odds backfill run at least once (ufcscraper_scrape_bestfightodds_data) for meaningful training coverage; works before that too, just with has_market_odds=0 everywhere.)
4 ahead
1 tied
No Age
CatBoost (ordered boosting) on the full engineered feature set (strategy: Main's exact feature config with every age-derived feature removed (calendar age, fighting age/years since debut, age-vs-prime, age-gap buckets, age tiers) — isolates how much the model leans on a fighter's age specifically, and whether it's still competitive without it.)
5 behind
Model Comparison (temporal holdout)
| Model | Algorithm | Features | Holdout AUC (95% CI) | Holdout Acc | Log Loss | Brier | vs. Champion | vs. Main |
|---|---|---|---|---|---|---|---|---|
|
XGBoost
PRIMARY
Gradient-boosted trees on the full engineered feature set
|
xgboost | 122 (full) |
0.6886
[0.657, 0.716]
|
63.6% | 0.6367 | 0.2228 | Statistically tied | — |
|
LightGBM
LightGBM on the full engineered feature set
|
lightgbm | 122 (full) |
0.6840
[0.653, 0.712]
|
63.0% | 0.6408 | 0.2247 | Significantly behind | — |
|
Logistic (Core)
Regularized logistic regression on curated differential features
|
logistic | 18 (core) |
0.6850
[0.654, 0.713]
|
63.7% | 0.6406 | 0.2247 | Statistically tied | — |
|
CatBoost
CatBoost (ordered boosting) on the full engineered feature set
|
catboost | 122 (full) |
0.6868
[0.656, 0.714]
|
65.2% | 0.6425 | 0.2253 | Significantly behind | — |
|
Ensemble
Soft-vote of XGBoost, LightGBM and logistic regression
|
ensemble | 122 (full) |
0.6942
[0.664, 0.722]
|
65.6% | 0.6333 | 0.2211 | CHAMPION | — |
|
Market Margin · XGBoost
Gradient-boosted trees on the full engineered feature set (strategy: Starts from the betting market's OPENING line and learns only what the market gets wrong, instead of predicting the winner from scratch. Because the market baseline already encodes 'favorites usually win', the model cannot score points by re-learning that, so it spends its capacity on genuine mispricings — which makes it both the most accurate configuration measured (66.3% vs Main's 63.0%, +3.3pts, McNemar p=0.0001, better in all 5 walk-forward windows) and much better at calling upsets (its picks against the market land 55.7% of the time, vs 42.9% for Main). Only predicts fights that have an opening line — typically the next card or two — so it complements Main rather than replacing it.)
|
xgboost | 271 (full) |
0.7310
[0.702, 0.756]
|
67.6% | 0.6062 | 0.2089 | — |
Ahead
+3.8% AUC
|
|
No Age · CatBoost
CatBoost (ordered boosting) on the full engineered feature set (strategy: Main's exact feature config with every age-derived feature removed (calendar age, fighting age/years since debut, age-vs-prime, age-gap buckets, age tiers) — isolates how much the model leans on a fighter's age specifically, and whether it's still competitive without it.)
|
catboost | 117 (full) |
0.6709
[0.640, 0.699]
|
62.7% | 0.6509 | 0.2293 | Significantly behind |
Behind
-1.9% AUC
|
|
No Age · Ensemble
Soft-vote of XGBoost, LightGBM and logistic regression (strategy: Main's exact feature config with every age-derived feature removed (calendar age, fighting age/years since debut, age-vs-prime, age-gap buckets, age tiers) — isolates how much the model leans on a fighter's age specifically, and whether it's still competitive without it.)
|
ensemble | 117 (full) |
0.6829
[0.652, 0.710]
|
63.8% | 0.6408 | 0.2246 | CHAMPION |
Behind
-1.7% AUC
|
|
No Age · LightGBM
LightGBM on the full engineered feature set (strategy: Main's exact feature config with every age-derived feature removed (calendar age, fighting age/years since debut, age-vs-prime, age-gap buckets, age tiers) — isolates how much the model leans on a fighter's age specifically, and whether it's still competitive without it.)
|
lightgbm | 117 (full) |
0.6722
[0.642, 0.701]
|
62.6% | 0.6480 | 0.2280 | Significantly behind |
Behind
-1.6% AUC
|
|
No Age · Logistic (Core)
Regularized logistic regression on curated differential features (strategy: Main's exact feature config with every age-derived feature removed (calendar age, fighting age/years since debut, age-vs-prime, age-gap buckets, age tiers) — isolates how much the model leans on a fighter's age specifically, and whether it's still competitive without it.)
|
logistic | 15 (core) |
0.6608
[0.629, 0.688]
|
60.7% | 0.6527 | 0.2304 | Significantly behind |
Behind
-2.6% AUC
|
|
No Age · XGBoost
Gradient-boosted trees on the full engineered feature set (strategy: Main's exact feature config with every age-derived feature removed (calendar age, fighting age/years since debut, age-vs-prime, age-gap buckets, age tiers) — isolates how much the model leans on a fighter's age specifically, and whether it's still competitive without it.)
|
xgboost | 117 (full) |
0.6788
[0.647, 0.708]
|
62.8% | 0.6437 | 0.2259 | Significantly behind |
Behind
-1.8% AUC
|
|
Recency-Heavy · CatBoost
CatBoost (ordered boosting) on the full engineered feature set (strategy: Recent form dominates career-long averages much more (365-day recency half-life vs main's 730) — tests whether a slump reflects something main's longer memory is underweighting.)
|
catboost | 120 (full) |
0.6839
[0.654, 0.711]
|
64.5% | 0.6434 | 0.2258 | Significantly behind |
Tied
-0.1% AUC
|
|
Recency-Heavy · Ensemble
Soft-vote of XGBoost, LightGBM and logistic regression (strategy: Recent form dominates career-long averages much more (365-day recency half-life vs main's 730) — tests whether a slump reflects something main's longer memory is underweighting.)
|
ensemble | 120 (full) |
0.6959
[0.665, 0.723]
|
66.0% | 0.6325 | 0.2208 | CHAMPION |
Tied
-0.0% AUC
|
|
Recency-Heavy · LightGBM
LightGBM on the full engineered feature set (strategy: Recent form dominates career-long averages much more (365-day recency half-life vs main's 730) — tests whether a slump reflects something main's longer memory is underweighting.)
|
lightgbm | 120 (full) |
0.6842
[0.652, 0.713]
|
63.6% | 0.6403 | 0.2244 | Significantly behind |
Tied
-0.4% AUC
|
|
Recency-Heavy · Logistic (Core)
Regularized logistic regression on curated differential features (strategy: Recent form dominates career-long averages much more (365-day recency half-life vs main's 730) — tests whether a slump reflects something main's longer memory is underweighting.)
|
logistic | 18 (core) |
0.6852
[0.654, 0.713]
|
63.5% | 0.6409 | 0.2248 | Statistically tied |
Tied
-0.1% AUC
|
|
Recency-Heavy · XGBoost
Gradient-boosted trees on the full engineered feature set (strategy: Recent form dominates career-long averages much more (365-day recency half-life vs main's 730) — tests whether a slump reflects something main's longer memory is underweighting.)
|
xgboost | 120 (full) |
0.6920
[0.660, 0.720]
|
65.9% | 0.6347 | 0.2218 | Statistically tied |
Tied
+0.2% AUC
|
|
With Odds · CatBoost
CatBoost (ordered boosting) on the full engineered feature set (strategy: Main's exact feature config plus de-vigged market-implied-probability features from BestFightOdds (opening line, closing range, and how much the line moved) — isolates whether the betting market's own signal improves on our features alone. Needs the historical odds backfill run at least once (ufcscraper_scrape_bestfightodds_data) for meaningful training coverage; works before that too, just with has_market_odds=0 everywhere.)
|
catboost | 125 (full) |
0.7425
[0.713, 0.768]
|
68.7% | 0.5986 | 0.2054 | Statistically tied |
Ahead
+5.5% AUC
|
|
With Odds · Ensemble
Soft-vote of XGBoost, LightGBM and logistic regression (strategy: Main's exact feature config plus de-vigged market-implied-probability features from BestFightOdds (opening line, closing range, and how much the line moved) — isolates whether the betting market's own signal improves on our features alone. Needs the historical odds backfill run at least once (ufcscraper_scrape_bestfightodds_data) for meaningful training coverage; works before that too, just with has_market_odds=0 everywhere.)
|
ensemble | 125 (full) |
0.7433
[0.714, 0.770]
|
68.5% | 0.5958 | 0.2044 | CHAMPION |
Ahead
+4.8% AUC
|
|
With Odds · LightGBM
LightGBM on the full engineered feature set (strategy: Main's exact feature config plus de-vigged market-implied-probability features from BestFightOdds (opening line, closing range, and how much the line moved) — isolates whether the betting market's own signal improves on our features alone. Needs the historical odds backfill run at least once (ufcscraper_scrape_bestfightodds_data) for meaningful training coverage; works before that too, just with has_market_odds=0 everywhere.)
|
lightgbm | 125 (full) |
0.7418
[0.712, 0.768]
|
69.0% | 0.5988 | 0.2057 | Significantly behind |
Ahead
+5.1% AUC
|
|
With Odds · Logistic (Core)
Regularized logistic regression on curated differential features (strategy: Main's exact feature config plus de-vigged market-implied-probability features from BestFightOdds (opening line, closing range, and how much the line moved) — isolates whether the betting market's own signal improves on our features alone. Needs the historical odds backfill run at least once (ufcscraper_scrape_bestfightodds_data) for meaningful training coverage; works before that too, just with has_market_odds=0 everywhere.)
|
logistic | 18 (core) |
0.6850
[0.654, 0.713]
|
63.7% | 0.6406 | 0.2247 | Significantly behind |
Tied
+0.0% AUC
|
|
With Odds · XGBoost
Gradient-boosted trees on the full engineered feature set (strategy: Main's exact feature config plus de-vigged market-implied-probability features from BestFightOdds (opening line, closing range, and how much the line moved) — isolates whether the betting market's own signal improves on our features alone. Needs the historical odds backfill run at least once (ufcscraper_scrape_bestfightodds_data) for meaningful training coverage; works before that too, just with has_market_odds=0 everywhere.)
|
xgboost | 125 (full) |
0.7417
[0.713, 0.769]
|
68.3% | 0.5974 | 0.2051 | Statistically tied |
Ahead
+5.1% AUC
|
Holdout metrics are computed on the most recent ~20% of fights, which the models never saw
during training or tuning — an honest estimate of forward, real-world performance.
AUC and accuracy: higher is better. Log loss and Brier: lower is better (they punish
confidently wrong probabilities). The AUC range in brackets is a 95% bootstrap confidence
interval — how much that number could plausibly move on a different set of holdout fights.
"vs. Champion" is a paired significance test against the model with the best point-estimate
AUC: "Statistically tied" means the difference could easily be noise, not a real gap.
"vs. Main" is a separate comparison for alternate strategies: this exact algorithm's holdout
AUC against Main's version of the same algorithm, on the same held-out fights — this is what
tells you whether the strategy is actually worth keeping (vs. Champion only compares
algorithms within the same strategy).
Outcome Model — Method & Round
| Sub-model | Classes | Holdout Accuracy | Holdout Log Loss | Holdout Size |
|---|---|---|---|---|
| Method (Decision / KO-TKO / Submission) | 3 | 54.3% | 0.9588 | 1280 |
| Round bucket (finishes only) | 3 | 50.1% | 1.0021 | 647 |
A separate pair of models answering "how" a fight ends rather than "who" wins — shown on every
fight's detail page. Method accuracy is measured against 3 classes (chance ≈ 33-48% depending on
class balance); round accuracy is measured only on fights that were predicted correctly to be finishes.
Feature Importance — CatBoost
Net striking dominance
12.329
Age edge (significant gap)
11.542
Elo rating edge
7.765
Experience edge (total fights)
6.962
Fighter 2 · Strike rate delta r1 to r3
5.831
Fighter 1 · Strike rate delta r1 to r3
5.687
Defensive skill edge
3.824
Fighter 2 · Age
3.164
Fighter 1 · Age
3.132
Fighter 2 · Avg takedowns attempted
2.731
Fighter 1 · Avg takedowns attempted
2.171
Head-strike accuracy edge
1.716
Fighter 2 · Strikes absorbed/min
1.649
Fighter 1 · Strikes absorbed/min
1.234
Fighter 2 · Avg control time
1.182
Fighter 2 · Distance strike defense
1.130
Fighter 1 · Fights per year
1.061
Fighter 1 · Elo rating
1.045
Fighter 1 · Avg control time
0.926
Fighter 1 · Distance strike defense
0.884
Global importance: how much the model relies on each feature across all fights.
For a single fight's explanation, open that fight's detail page.