Doosra win probability (T20 and ODI)
The batting side's chance of winning a limited-overs cricket match, after any ball. It powers Match Replay in Doosra, a cricket analytics app, and you can try it in the demo Space.
Use
from predict import WinProbability # predict.py in this repo; needs only NumPy
wp = WinProbability("winprob-t20.json")
# 2024 T20 World Cup final: South Africa 151/4 after 16 overs, chasing 177 at Kensington Oval
wp.predict(innings=2, score=151, wickets=4, balls_left=24, target=177, venue_par=159, elo_diff=-123) # 0.925
wickets is wickets fallen; venue_par the ground's typical first-innings total (leave it out if unknown);
elo_diff the batting side's Elo rating minus the bowling side's (0 = evenly matched). Use
winprob-odi.json for one-day matches.
How it works
One model per format (T20, ODI) and per innings. Each is a blend, averaged on the log-odds scale, of:
- gradient-boosted trees (LightGBM) with monotone constraints, encoding cricket common sense: more wickets in hand, more balls left, a higher score or a stronger side can only help the batting side, and more runs needed or a higher required rate can only hurt;
- a logistic regression on the same features, which keeps the curve smooth from ball to ball.
Features: score, wickets in hand and balls left; in the chase also the target, runs needed and required rate; the ground's par (the average of its last 20 first-innings totals before the match); gender; and each side's Elo rating built from earlier results only, so nothing from the match or later leaks in.
Data: every ball of 17,000+ T20 and ODI matches (men's and women's, internationals and leagues) from the Doosra dataset (Cricsheet). Training uses matches with 6-ball overs and a decided result, with no rain rule and a full-length chase.
Evaluation
Split by time: trained on matches up to 2022, early-stopped on 2023-24, and tested on everything from 2025 on, so the test is a genuine forecast of matches the model never saw.
| Test (2025+) | Matches | Log loss | Brier | AUC | Calibration error | Accuracy | Logistic alone | Trees alone | Par heuristic |
|---|---|---|---|---|---|---|---|---|---|
| T20 | 2,889 | 0.455 | 0.152 | 0.861 | 0.8% | 76.7% | 0.461 | 0.458 | 0.658 |
| ODI | 638 | 0.512 | 0.174 | 0.818 | 2.3% | 72.8% | 0.514 | 0.519 | 0.712 |
The blend beats both of its halves and a hand-built par-score heuristic. It is well calibrated: when it says
70%, the batting side wins about 70% of the time (plots in reports/). On an ordinary ball (no wicket or
boundary) the prediction moves 0.6 points on average in T20
(0.3 in ODIs), so the worm follows the game rather than noise.
Limitations
- Limited-overs only. Tests have draws and aren't covered.
- Rain rules. Rain-affected matches are left out of training. Scored with their revised target, they are approximate.
- No player or pitch information beyond the ground's par and team Elo. A side with its best batter still in is treated like any other side at that score.
- Coverage. Cricsheet coverage is uneven; some teams have few or no matches (for example, there are no Afghanistan matches), so their Elo ratings are weak.
Licence
The data is from Cricsheet under the Open Data Commons Attribution License (ODC-By 1.0); the model is released under the same licence. Credit Cricsheet and Doosra if you use it.

