# Mega-dice-o-meter calculations The Mega-dice-o-meter started as a simple improvement to the statistical side of MrMesmer’s calculation, but it quickly turned into a much bigger rewrite once it became clear how messy “luck” in Blood Bowl really is. One thing I found quite early on is that any attempt to measure luck in Blood Bowl inevitably involves some subjectivity, and can be viewed as defining a model of “good” and “bad” luck. The probabilities themselves are straightforward enough, but deciding whether a result counts as good or bad luck depends heavily on context, including what the coach is trying to do. For example, there’s no way an automated system can know that you wanted a push instead of a pow because it was part of a five-player chain-push sequence leading to a 2D uphill on the ball. In practice, the dice-o-meter should get most things right most of the time. The main thing is not to over-interpret the exact numbers — the difference between a 1-in-10 “dicing” and 1-in-10.1 one isn’t very meaningful; but the gap between 1-in-10 and 1-in-100 clearly is. ## Measuring luck We want an additive measure of luck, so that we can sum the luck from individual rolls to get a total for the game. For quantifying the amount of "surprise" associated with an event that occurs with probability $p$, the only additive choice (up to scaling) is the negative-log-likelihood, $-\ln (p)$. This has ensures that the surprise of successfully rolling 3$+$, 5$+$ is the same as rolling a single 6$+$. Other choices such as $1/p$ [Raspel] or $\ln(1/(1+p)$ [MrMesmer] do not do this. ## Good, bad and indifferent luck We don’t just want a measure of how unlikely a series of rolls is (if we did, we could just sum the negative log likelihoods). We also want to include a notion of whether outcomes are good or bad. A perfectly mixed sequence of extremely unlikely good and bad rolls would still be extremely unlikely, but it would be hard to describe it as “dicing”. So the basic term in the luck calculation is $w \ln (1/p)$, where $w$ encodes whether the outcome is good or bad from the perspective of the player rolling the die. The natural choice is to take $w$ positive for good luck, negative for bad luck, and zero for neutral outcomes. At first glance, the classification into good, neutral, and bad looks straightforward: failing a rush is bad; defender down on a block is good; and a simple push is neutral. In practice, it is not that simple. Sometimes a push is good (surf a saurus); sometimes a defender-down is bad (one turn attempt); and even a failed rush may not be bad if it is part of an anti-stalling plan (3rd season, thanks GW). The point is that whether a roll is good, neutral, or bad is context dependent, and a large part of that context is coach intent, which is not recorded anywhere (except perhaps as some vague electrical pulses in the coach’s brain). Even if we agree that two outcomes are both good, there is still the question of how good? Aside from differences in probability, is a blitz kick-off result better or worse than a double skull on a LOS block? It is not obvious how to answer this consistently across all situations in Blood Bowl, so we keep things simple. We set $w \in \left\{+1,0,-1 \right\}$ for good, neutral, and bad outcomes, and then use as much context as we can from the replay file to assign $w$ based on the most likely interpretation of the play. This gives a sensible account of most dice rolls in a match, within the finite resolution of a particular model of luck. --- ## Luck difference The game can be described as two sequences of dice rolls, one for each coach, with outcomes classified as good ($g$), neutral ($n$), or bad ($b$). For coach $i=1,2$, we have realised sequence $$\{x_\mathrm{realised}^{(i,n)}\}_{n=1}^{N}; \quad x_i^{(n)} \in X = \{g, n, b\}.$$ The luck of coach $i$ is then $$ \ell_i = - \sum_n w^{(i,n)}\left(x^{(i,n)}_\mathrm{realised}\right) \, \ln \left[ p^{(i,n)}\left(x^{(i,n)}_\mathrm{realised}\right) \right]$$ in which $w^{(i,n)}(x)$ and $p^{(i,n)}(x)$ are the weight and probability functions for the $n$th roll of coach $i$. The summation runs over all dice rolls made by that coach. The expected luck is then $$ \langle \ell_i \rangle = - \sum_n \sum_{x^{(i,n)} \in X} w^{(i,n)}\left(x^{(i,n)}\right) \,p^{(i,n)}\left(x^{(i,n)}\right) \,\ln \left[ p^{(i,n)}\left(x^{i,(n)} \right) \right],$$ which, without the weights, would just be the entropy of the underlying dice process. There is no particular reason for this expectation to be zero, so when comparing how lucky the dice rolls have been, it makes sense to consider the centred luck $$ L_i = \ell_i - \langle \ell_i \rangle.$$ These are the quantities plotted in the Turn-by-Turn tab. The dice-o-meter is then built around the difference in luck between the two coaches $$ d = L_1 - L_2 $$ Values $d \gt 0$ correspond to coach 1 having had the better of the dice, and $d \lt 0$ coach 2. --- ## Dice-o-meter statistic Knowing $d$ alone is not enough to interpret how unusual a game is, since we lack a sense of scale — how large does $d$ need to be before we would call it a “dicing”? To answer this, we introduce a random variable $D$ built from the same underlying dice-roll processes that generate $d$. Since the structure of the dice is known, we can in principle determine the distribution $P(D)$. By comparing the realised value of $d$ with this distribution, we can assess how likely such a result is, and therefore whether the game shows evidence of extreme luck imbalance. We quantify this using a tail probability, defined as $$\rho = P(D \gt d) ~~\text{for} ~~ d\gt 0$$ and $$\rho = P(D \lt d) ~~\text{for} ~~ d\lt 0.$$ For the coach who had the worse luck (determined by the sign of $d$), the tail probability $\rho$ represents the probability of seeing dice at least as bad as those realised in the game. So $\rho = 0.1$ corresponds to something expected roughly once every 10 games, while $\rho = 0.01$ corresponds to once every 100 games. Whether a 1-in-10 or 1-in-1000 result counts as a “dicing” is ultimately subjective. The advantage of this formulation is that it puts that discussion on a consistent numerical footing. How we compute the tail probability is discussed in the appendix below. --- ## Rerolls and counterfactual dice Imagine a gutter runner and a wood elf lino, each in a tackle zone, both vanilla and with no Team rerolls available. Both attempt to dodge out into the open to score, each rolling a d6 and succeeding on a 2+. Of the two players, who was luckier? Even though the replay shows the same dice rolls for both, it has to be the lino who was luckier because the gutter runner has a built-in dodge reroll up its sleeve and would therefore fail that dodge 1 in 36 times, compared to the 1 in 6 for the skill-less lino. This highlights a general issue: once rerolls exist, it is not only the dice that were rolled that matter, but also the dice that *could have been rolled but weren’t*. These counterfactual rolls do not appear in the replay of course. For straightforward skill use (as in the gutter runner vs lino example), this can be handled reasonably cleanly. When a dice roll is failed and a reroll is taken, nothing extra needs to be done beyond what is already in the replay. When a roll succeeds, however, built-in rerolls need to be considered. In these cases, we assume that if a player is attempting an action, it intends to succeed, and would therefore use an available built-in reroll if required. The probabilities are therefore adjusted to reflect success and failure including that reroll. Team rerolls are more difficult. When a roll fails and a team reroll is used, this is captured directly in the replay and can be handled normally. However, modelling counterfactual use of team rerolls is not really feasible. In principle, a coach can choose to use rerolls at many different points, but there is no reliable way to infer which failed rolls would have triggered them. As a result, counterfactual team rerolls are ignored. The critical cases — snakes and double skulls — are still fully captured. Rerolls associated with Pro and Brawler are discussed below. --- ## Handling of various roll types & skills A game of Blood Bowl consists of many different types of dice rolls: block dice, action dice, weather dice, and so on. Some skills also change the interpretation of outcomes. What follows is a set of notes on how different roll types and skills are handled by the dice-o-meter. Many edge cases are not handled, either because the necessary context is missing, because the logic becomes too complex, or because extracting the required information from the replay is too difficult. Future versions may improve on these latter two issues, but some limitations are unavoidable. One clear omission in the current processing is that player and ball positions are not available. As a result, spatial effects such as lucky bounces or chain push → surf → outcome sequences are not explicitly captured. - **Actions.** Most non-block actions involving dice (dodge, rush, catch, etc.) are treated as pass/fail with no neutral outcome. Skill modifiers are taken from the replay. Built-in rerolls are handled as described above. - **Argue the call, bribe use.** Successful argue = good; otherwise bad. - **Armour.** Breaks are good for the attacker; non-breaks are bad. - **Block dice.** The default interpretation is defender down = positive, attacker down = negative, and all other results neutral. If the defender is the ball carrier, any result that causes the ball to come loose is treated as positive; otherwise it is negative. Skills considered include: block, brawler, dodge, juggernaut, strip ball, sure hands, wrestle. Whether the player is blitzing is included for brawler and juggernaut decisions. Brawler is handled under the assumption that it is used whenever the initial result is attacker down, though there are edge cases where a team reroll would be preferred instead. - **Breathe Fire.** 1 = bad; 2–3 = neutral; 4+ = good. - **Casualty & APO.** This is tricky because interpretation differs between resurrection and non-resurrection formats. We classify BH as neutral and anything worse as bad. This reflects typical league interpretation and aligns with potential APO usage. APO rerolls are included when present in the replay. - **Injury.** Stunned is neutral; anything more serious is good. Stunty injury table taken into account. - **KO recovery.** Recovery = good; otherwise bad. - **Kick-off events.** These are handled via a fixed interpretation table, from the kicking team’s perspective. | Table | Event | Success / Neutral / Fail rule (kicking team perspective) | |---|---|---| | 2 | Get the Ref | **Context-dependent.** If only one team has secret-weapon-eligible players, that team gets **Success** and the other gets **Fail**. If both or neither do, **Neutral**. | | 3 | Time-out | Usually **Neutral**. **Fail** only when the receiving team is on game turn 8 or 16. | | 4 | Solid Defence | Always **Success**. | | 5 | High Kick | Always **Fail**. | | 6 | Cheering Fans | Roll-off event: winner = **Success**, loser = **Fail**, tie = **Neutral**. | | 7 | Brilliant Coaching | Roll-off event: winner = **Success**, loser = **Fail**, tie = **Neutral**. | | 8 | Changing Weather | Depends on weather band shift: perfect -> adverse = **Success**, adverse -> perfect = **Fail**, otherwise **Neutral**. | | 9 | Quick Snap | Always **Fail**. | | 10 | Blitz | Always **Success**. | | 11 | Officious Ref | Roll-off event: winner = **Success**, loser = **Fail**, tie = **Neutral**. | | 12 | Pitch Invasion | Roll-off event: winner = **Success**, loser = **Fail**, tie = **Neutral**. | Notes: - Event 8 uses weather bands where perfect = totals 4-10 and adverse = totals 3/11/12. - Events 6/7/11/12 use opposed roll-off logic (with team-specific kickoff modifiers). - Events 4/5/9/10 are fixed one-hot outcomes. - **Kick-off landing square.** Not currently handled. - **Pick-me-up.** 1–4 neutral; 5+ good. - **Pro.** Assumed to be used when an action fails and a reroll is available. For blocks, assumed to trigger on attacker-down results where appropriate. Built-in re-rolls take precendence over Pro. - **Safe Pass.** Fumble is still treated as bad, even though the ball is not dropped. The reasoning is that the intended action still fails. Passing remains a pass/fail structure. - **Stab, Chainsaw.** No separate roll is recorded; these feed into armour resolution. ## OTTD Mode When a coach is going for an OTTD with pushes, the usual good/neutral/bad classification is no longer appropriate. In these situations, the primary signal of success is the generation of pushes from block dice, rather than knockdowns. It is difficult to define a fully general OTTD model that applies across all possible configurations, so the current implementation uses a heuristic “OTTD mode” flag that is activated when an OTTD is deemed likely, and deactivated when it is no longer considered feasible. Whilst OTTD mode is active, pushes on blocks are treated as success outcomes and all other block results are treated as failures. All other dice types continue to be processed under the standard rules. For a turn to be eligible as an OTTD turn, it must be the final turn of the half, and a touchdown must have just been scored. In addition, no Throw Team-Mate (TTM) action may take place during the turn, since the block structure in that case is broadly similar but not identical in its optimisation space. If a touchdown is scored in such a turn, this is taken as direct evidence of an OTTD attempt, and OTTD mode is activated for the entire turn. Otherwise, a heuristic is used to determine whether the receiving team setup is consistent with an OTTD attempt. The following requirements are used: * Fast runner (MA7+) on half way line * Enough supporting players , as per this table (these are overly generous, but I am reluctant to rule out too many combinations just because I can't think of them) | Pushes Needed | Min Players (has Sidestep) | Min Players (no Sidestep) | |---|---:|---:| | 1 | 4 | 6 | | 2 | 6 | 7 | | 3 | 6 | 8 | | 4 | 8 | 9 | | More | unavailable | unavailable | * At least one player back to retrieve the ball. * Defense not fully anti-push locked In evaluating the latter conditions, skill interactions such as Stand Firm and Side Step are explicitly accounted for. In particular, Grab may neutralise the effect of Side Step, and Juggernaut-enabled blocks may neutralise Stand Firm where applicable. If these criteria are satisfied, OTTD mode is activated. It remains active until one of the following conditions is met: * A non-push block result is selected by the receiving side (exception: blitz with Juggernaut selecting Both Down) * A monitored fast player reaches required depth before activating * All monitored fast starters have activated A failed non-block action also terminates OTTD mode, although in this case it does so via a standard turnover condition. ## Possible Future Improvements A number of Blood Bowl mechanics are not currently handled but could likely be added with relatively little effort. - Add mapping for the following skills: * pile driver * Cannoneer * Cloud Burster * Fumblerooskie * Dirty Player +2 These skills are currently unrecognised. This does not affect the dice-o-meter calculations, but it does result in incomplete skill processing. - Extra time and penalties. In principle, these should already be supported, but this has not been tested. - Add partial handling of team rerolls in situations where their use can be inferred with high confidence. For example, if a turn 8 ball carrier requires a rush to score and a team reroll remains available, it is reasonable to assume that the reroll would have been used if required. A second group of missing features is more challenging. Apart from the OTTD detection logic, the current processor has no access to positional information from the replay file. Successfully extracting positional data would make it possible to address the following limitations: - Determine probabilities and classify favourable and unfavourable kick outcomes. - Improve Throw Team-Mate handling. At present, this is based solely on success or failure of the PA test; a more accurate approach would evaluate whether the thrown player lands approximately where intended. - Adjust block calculations in situations where a surf requires only a push result. - Quantify the luck associated with ball bounces (for example, bounces leading directly into catch attempts). - Quantify the luck associated with Ball & Chain movement resulting in a block. --- ## Appendix: Tail Probability Estimation for a Sum of Independent Three-Point Random Variables ### What is being calculated? The code computes approximations to tail probabilities $P(X \gt x)$ and $P(X \lt x),$ of a random variable $$X = \sum_{i=1}^n Y_i,$$ where the $Y_i$ are independent and each takes one of three values: $$Y_i = \begin{cases} -\log p_i & \text{with probability } p_i,\\[4pt] \log q_i & \text{with probability } q_i,\\[4pt] 0 & \text{with probability } 1-p_i-q_i. \end{cases}$$ Monte Carlo estimation becomes inefficient in the tails, so instead the code uses a Lugannani–Rice (LR) approximation, which provides high accuracy in both the bulk and the tails whilst remaining computationally efficient and deterministic. ### The cumulant generating function (CGF) and its derivatives The method is built around the cumulant generating function $$K(t) = \log \mathbb{E}[e^{tX}].$$ Because the variables are independent, $$K(t)=\sum_i K_i(t),$$ with $$K_i(t)=\log \mathbb{E}[e^{tY_i}].$$ For the three-point distribution above, $$e^{tY_i}= \begin{cases} p_i^{-t} & \text{with probability } p_i,\\ q_i^{t} & \text{with probability } q_i,\\ 1 & \text{otherwise}. \end{cases}$$ Therefore $$\mathbb{E}[e^{tY_i}] = (1-p_i-q_i) + p_i^{1-t} + q_i^{1+t},$$ and hence $$\boxed{ K(t) = \sum_i \log\!\left( (1-p_i-q_i)+p_i^{1-t}+q_i^{1+t} \right) }$$ The derivatives of $K$ encode the cumulants of the exponentially tilted distribution, and these can be computed analytically. For example, the first derivative is $$K'(t) = \sum_i \frac{ -p_i^{1-t}\log p_i + q_i^{1+t}\log q_i }{ (1-p_i-q_i)+p_i^{1-t}+q_i^{1+t} }.$$ ### Saddlepoint and Lugannani–Rice To estimate the probability of observing $X\approx x$, saddlepoint theory introduces a tilting parameter $t^\star$ defined implicitly by $K'(t^\star) = x$. This identifies the exponential tilting under which $x$ becomes typical of the transformed distribution. Intuitively: $t^\star\gt 0$ corresponds to a right-tail event, and $t^\star\lt 0$ corresponds to a left-tail event. The code solves this equation for $t^\star$ numerically using Brent’s root-finding method with an initial guess $t_0 \approx \frac{x-\mu}{K''(0)}$ from a local Gaussian approximation. Once $t^\star$ is known, the key large-deviation quantity is $I(x)=t^\star x -K(t^\star)$. The leading asymptotic behaviour of the tail probability is $P(X\gt x)\sim e^{-I(x)}$, but this alone is not sufficiently accurate in moderate deviations or near the bulk. The Lugannani–Rice formula corrects this. Define $$w = \operatorname{sign}(t^\star) \sqrt{2\bigl(t^\star x-K(t^\star)\bigr)},$$ and $$u = t^\star\sqrt{K''(t^\star)}.$$ The LR approximation is then $$\boxed{ P(X\gt x) \approx 1-\Phi(w) + \phi(w)\left(\frac1w-\frac1u\right), }$$ with $\Phi$ is the standard normal CDF and $\phi$ is the standard normal density. For left tails, $$P(X\lt x) \approx \Phi(w) + \phi(w)\left(\frac1w-\frac1u\right).$$ The code detects which tail is being computed from the sign of $w$. Near the mean, $t^\star\to0$, so $w\to0, \qquad u\to0$, making the correction unstable. The code therefore switches to a Gaussian approximation when $w^2 \lt \varepsilon$. The fallback uses $Z=\frac{x-\mu}{\sigma}$ with $\mu=K'(0), \qquad \sigma^2=K''(0)$, and evaluates $P(X \gt x)\approx 1-\Phi(Z)$ or $P(X \lt x )\approx \Phi(Z)$ as appropriate. This stabilises the method in the central region. ### Analytic error estimate To sanity-check the approximation, the code estimates the asymptotic error. Define $$\gamma_1 = \frac{K'''(t^\star)}{K''(t^\star)^{3/2}},$$ $$\gamma_2 = \frac{K''''(t^\star)}{K''(t^\star)^{2}}.$$ The relative LR error scales as $\sim \frac18\gamma_1^2 + \frac1{24}\gamma_2$, so the implementation uses $$\varepsilon_{\rm rel} = \frac18\gamma_1^2 + \frac1{24}\gamma_2,$$ and reports $$\boxed{ \varepsilon_{\rm abs} = P_{\rm LR}\,\varepsilon_{\rm rel}. }$$ In practice, this estimate is typically far smaller than the probability itself, and comparisons with Monte Carlo support its reliability. ### Summary The code builds tail probability estimates using exact knowledge of the cumulant generating function and a saddlepoint approximation (Lugannani–Rice) rather than simulation. In practice, this gives accurate results for both typical and tail events with very low computational cost, which is sufficient for the intended application.