The finding
FACEIT Elo is already a near-complete predictor. Recent form adds no statistically significant ranking value on top of it.
| Predictor | Accuracy | AUC | ECE |
|---|---|---|---|
| My model (stats + Elo, isotonic-calibrated) | 69% | 0.753 | 0.032 |
| FACEIT Elo formula (baseline) | 70% | 0.745 | 0.072 |
The AUC gain is +0.007, 95% CI [−0.001, +0.016], DeLong p = 0.09: not significant. That held even after testing four more orthogonal features, including a leakage-safe teammate "stacking" signal that tells a coordinated 5-stack from five strangers at the same rating.
Where the model does win is calibration: when it says 70%, the team wins about 70% of the time, which the raw Elo formula can't claim. A walk-forward backtest (8 expanding folds, 4,440 out-of-sample predictions) shows the same picture.

How the number is kept honest
No feature sees the future
A match's stat features use only each player's matches that finished before it started.
Always scored on newer matches
A chronological split: the model is only tested on matches newer than everything it trained on.
A real baseline
Measured against FACEIT's own Elo win-probability formula, which turns out to be very hard to beat.
"Better" means statistically better
AUC differences go through a paired bootstrap confidence interval and DeLong's test.
Cross-checked against a replayed Elo
FACEIT only exposes current Elo, so the comparison is re-run against an Elo rebuilt from scratch, in order, with no look-ahead.
Architecture
The ML library knows nothing about the web, so a feature vector means the same thing at training time and serving time, by construction.