Doosra win probability (T20 and ODI)

The batting side's chance of winning a limited-overs cricket match, after any ball. It powers Match Replay in Doosra, a cricket analytics app, and you can try it in the demo Space.

Calibration on 2025+ T20 matches

Use

from predict import WinProbability          # predict.py in this repo; needs only NumPy

wp = WinProbability("winprob-t20.json")
# 2024 T20 World Cup final: South Africa 151/4 after 16 overs, chasing 177 at Kensington Oval
wp.predict(innings=2, score=151, wickets=4, balls_left=24, target=177, venue_par=159, elo_diff=-123)   # 0.925

wickets is wickets fallen; venue_par the ground's typical first-innings total (leave it out if unknown); elo_diff the batting side's Elo rating minus the bowling side's (0 = evenly matched). Use winprob-odi.json for one-day matches.

How it works

One model per format (T20, ODI) and per innings. Each is a blend, averaged on the log-odds scale, of:

  • gradient-boosted trees (LightGBM) with monotone constraints, encoding cricket common sense: more wickets in hand, more balls left, a higher score or a stronger side can only help the batting side, and more runs needed or a higher required rate can only hurt;
  • a logistic regression on the same features, which keeps the curve smooth from ball to ball.

Features: score, wickets in hand and balls left; in the chase also the target, runs needed and required rate; the ground's par (the average of its last 20 first-innings totals before the match); gender; and each side's Elo rating built from earlier results only, so nothing from the match or later leaks in.

Data: every ball of 17,000+ T20 and ODI matches (men's and women's, internationals and leagues) from the Doosra dataset (Cricsheet). Training uses matches with 6-ball overs and a decided result, with no rain rule and a full-length chase.

Evaluation

Split by time: trained on matches up to 2022, early-stopped on 2023-24, and tested on everything from 2025 on, so the test is a genuine forecast of matches the model never saw.

Test (2025+) Matches Log loss Brier AUC Calibration error Accuracy Logistic alone Trees alone Par heuristic
T20 2,889 0.455 0.152 0.861 0.8% 76.7% 0.461 0.458 0.658
ODI 638 0.512 0.174 0.818 2.3% 72.8% 0.514 0.519 0.712

The blend beats both of its halves and a hand-built par-score heuristic. It is well calibrated: when it says 70%, the batting side wins about 70% of the time (plots in reports/). On an ordinary ball (no wicket or boundary) the prediction moves 0.6 points on average in T20 (0.3 in ODIs), so the worm follows the game rather than noise.

Accuracy through the match

Limitations

  • Limited-overs only. Tests have draws and aren't covered.
  • Rain rules. Rain-affected matches are left out of training. Scored with their revised target, they are approximate.
  • No player or pitch information beyond the ground's par and team Elo. A side with its best batter still in is treated like any other side at that score.
  • Coverage. Cricsheet coverage is uneven; some teams have few or no matches (for example, there are no Afghanistan matches), so their Elo ratings are weak.

Licence

The data is from Cricsheet under the Open Data Commons Attribution License (ODC-By 1.0); the model is released under the same licence. Credit Cricsheet and Doosra if you use it.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train Sarthak213/doosra-win-probability

Space using Sarthak213/doosra-win-probability 1