I Just Wanted to Roll D20

·25 min read·Cic7e
I Just Wanted to Roll D20

The first version of this bot shipped in 2018. It had dice notation and some flavor messaging, decent at the time and better than most Discord bots. Then some things happened, I moved, left the internet for a while, and lost the source code. When I came back, I rewrote the whole thing, and this time I just kept going.

This post is what eight years of kept going looks like.

Before the 2018 version, before notation, before flavor, before any of it, there was the line anyone writing a Discord bot starts with:

await ctx.send(f"You rolled {random.randint(1, 20)}!")

Bots shipped with hundreds of mediocre commands typically have some form of that line. Mine did. That random.randint is still in the code, technically, buried in roll_dice() under everything that grew on top of it. The last living descendant of the original one-liner.

A single /roll today parses dice notation, dice inside dice, evaluates through an AST-allowlisted arithmetic parser (no eval anywhere, we’ll get to that), compiles optional success conditions into reusable predicates, computes probability distributions in arbitrary-precision rationals, scores luck as roll percentiles, then logs the whole affair to SQLite on a separate thread. The surrounding cog has saved macros, autocomplete, reroll-from-history, and per-server feature toggles.

Mind you, I didn’t plan this! I justified forty small decisions, each one defensible the day I made it. Then I looked up, and I was maintaining an arbitrary-precision probability engine with a JavaScript twin. Whoops!

Here’s how it went, decision by decision.

Note

If you’d rather poke at this than read about it: the engine runs live, right now, on the dashboard! The source for everything quoted here is on GitHub. (lite version: dashboard and databases stripped, engine intact). All outputs in this post are real, pulled from a seeded run. Your dice might differ, your probabilities won’t.

A regex that learned recursion

Dice notation looks innocent: 2d6+3. Some dice, a d, some sides, arithmetic around it. The parser is still, at heart, a regex and a callback, and that part has barely changed since 2018:

DICE_PATTERN_STR = r'(?:[\d.]|\([^)]+\))*[dD](?:[\d.]|\([^)]+\))+(?:(?:k[hl]?|[hl])(?:[\d.]|\([^)]+\))*)?'

The callback is simple, it matches a die then rolls it:

def roll_callback(match):
    ...
    _, rolls = roll_dice(num_dice, num_sides)
    ...
    return str(total)

sanitized = re.sub(DICE_PATTERN_STR, roll_callback, dice_string, flags=re.IGNORECASE)

2d6+3 goes in, 7+3 comes out, and whatever remains is plain arithmetic for the next stage. One pass handles the dice, another the math, and the two concerns never really touch:

You rolled: `[6 + 1 = 7] + 3` = **10!**

Then I wanted dice inside the dice. 2d(5+1d5): roll a d5, add five, and that’s your number of sides, so now roll two of those. The sides expression contains another die, so the callback hands it straight back to the parser:

if any(x in raw_sides for x in 'dD'):
    sides_resolved, sides_breakdown = parse_and_roll(raw_sides, sort)

parse_and_roll calling parse_and_roll. Dice are allowed in the count, the sides, even the keep clause, and every nested layer resolves before the layer above it rolls.

The keep syntax is where breakdown strings earn their keep. 4d6kh3 rolls four dice and keeps the highest three, with braces marking the dropped one:

You rolled: `kh3[3 + 2 + 2 + {2} = 7]` = **7!**

And 4d6kl1, keep the lowest, cheerfully discards three sixes to save a one:

You rolled: `kl1[{6} + {6} + {6} + 1 = 1]` = **1!**

There’s a limit, of course. The pattern matches parenthesized sub-expressions with [^)]+, which cannot span nested parentheses, so deeply nested constructs fall over. When a roll fails and the input looks like that, you’ll reach this fun little easter egg:

await ctx.followup.send("My abacus just filed a restraining order. "
                        "Try something like 2d(5+1d5) instead")

I personally love when error messages end up doubling as documentation for exactly when the parser gives up.

Math you can’t eval()

Once every die has been substituted with its rolled total, what’s left is arithmetic: 7+3, 2*(3+4), maybe a caret for exponents. The lazy option is eval(). The lazy option is also what allows a dice bot to acquire a remote code execution vulnerability, so instead there’s an allowlist:

ALLOWED_OPERATORS = {ast.Add: op.add, ast.Sub: op.sub, ast.Mult: op.mul,
                     ast.Div: op.truediv, ast.Pow: op.pow,
                     ast.USub: op.neg, ast.UAdd: op.pos}
ALLOWED_NODES = [ast.Expression, ast.BinOp, ast.UnaryOp, ast.Constant,
                 *ALLOWED_OPERATORS.keys()]

Parse the expression into a syntax tree, walk it, and reject anything not on the guest list. No function calls, no names, no attribute access, nothing but numeric constants and seven operators:

tree = ast.parse(expression, mode='eval')
for node in ast.walk(tree):
    if type(node) not in ALLOWED_NODES:
        raise ValueError(f"Invalid expression: Disallowed node {type(node).__name__}")

A recursive-descent walk over the surviving nodes computes the result. Two dialect quirks get smoothed over first; users write ^ where Python wants **, and humans write 2(3+4) where Python wants 2*(3+4):

expression = str(expression).replace('^', '**')
sanitized_string = re.sub(r'(?<=[\d)])\(', '*(', sanitized_string)

The security story ends up being “allowlist the syntax tree,” which beats regex-filtering raw input on every axis I care about. And because the dice pass already replaced every die with a number, the arithmetic parser never has to know dice exist!

How lucky was that, exactly?

Flavor messaging was in the 2018 bot. “Nat 20!” and so on, calls out a notable result. Cute! The rewrite’s escalation was smaller than it sounds: I wanted the bot to know how lucky a roll was, not just how notable.

The natural definition is a percentile. Given the distribution of what you rolled, your luck score is P(X ≤ result), the probability of doing as badly or worse. Zero is the worst possible roll, 0.5 is the median, 1.0 means you topped out:

pmf = await asyncio.to_thread(expression_to_pmf, user_input)
p_better = p_at_least(pmf, int_total)
p_worse = p_at_most(pmf, int_total)
luck_score = float(p_worse)  # P(X <= result): 0 = worst, 0.5 = median, 1 = best

That trailing comment is quietly the most expensive line in the codebase. Computing P(X ≤ result) means knowing the probability of every outcome; the full distribution of the exact expression. One UX nicety later, and suddenly the dice roller needs to become a probability engine.

The lack of cache anywhere in the codebase means it pays for that engine on every single roll. Every /roll, including the four-thousandth 1d20 of a busy Friday night, spins up a thread, enumerates the entire distribution as fresh Fractions, sums two tail probabilities, throws the whole table away, and prints one italic line about it. The bot will try to pay full price in every transaction, so the user gets a flavored subtext.

Worth it, would do again. Roll the mode of 4d6kh3 and you get a 13, the single most likely outcome, and even that sits at the 64.5th percentile.

And the flavor tiers came along for free, because they’re tail probabilities with jokes attached. Roll the maximum: an 18, exactly 7/432, a 1-in-62 event, the engine picks a one-liner:

The maximum possible outcome - only a 1.62% chance!
Legendary. 1 in 62 rolls go this well (1.62%)
Couldn't have rolled higher if you tried (1.62%)
You're welcome... 1.62% odds

One in 62, exactly, because everything underneath is exact. Which brings me to the next decision.

PMFs and the tyranny of floats

A probability mass function is a lookup table: every possible outcome of an expression, mapped to its probability. For a d20 that’s twenty rows of 1/20. For 2d6+3 it’s eleven rows shifted up by three. The engine’s whole job is deriving that table for whatever a user typed.

The first structural decision: those probabilities are Fractions, arbitrary-precision rationals, not floats.

PMF = dict[int, Fraction]  # outcome -> probability

def uniform_die(sides: int) -> PMF:
    p = Fraction(1, sides)
    return {face: p for face in range(1, sides + 1)}

Because floats lie. 1/3 in floating point is not a third, and the lies compound with every operation. A luck score is a claim about rarity, and when someone rolls the 1-in-2,000 outcome, “approximately 0.00049999” is not the energy the moment deserves. Fraction keeps an exact numerator and denominator at any size, so the 1-in-62 above is actually 7/432, displayed rounded. That choice has to happen at the data structure; there’s no retrofit to floats later.

You do deserve some truth, exactness has a documented exit. mean, variance, and stdev cast to float the moment they leave the engine, and a target probability is only printed as a fraction when the denominator is convenient to be read:

frac_str = f" = {p_hit.numerator}/{p_hit.denominator}" if p_hit.denominator <= 10000 else ""

So the real position isn’t “floats lie.” It’s rationals all the way through the math, floats at the display boundary, exactness shown only when it’s legible.

The second decision: distributions combine by convolution. If X and Y are independent, the PMF of X+Y is every pair of outcomes, multiplied and summed at their combined value:

def convolve(a: PMF, b: PMF) -> PMF:
    result: PMF = defaultdict(Fraction)
    for av, ap in a.items():
        for bv, bp in b.items():
            result[av + bv] += ap * bp
    return dict(result)

Conceptually, that’s the whole engine. NdM is a uniform die convolved with itself N−1 times. Every + in an expression is a convolution; every - is a convolution with a negated distribution; every constant is a distribution concentrated on one value. One pairwise multiply-and-sum covers all the arithmetic the parser can produce. P(2d6 = 7) comes out as exactly 1/6, which is the textbook answer.

Keep-highest, exactly

Here’s where the textbook runs out, and scope creep gets fun!

D&D players roll 4d6kh3 for stats: four dice, keep the best three. You can’t compute that with convolutions, because keeping the best three couples the dice, AKA they’re no longer independent, and the beautiful pairwise math stops applying. Brute force would enumerate all 6^4 = 1,296 raw outcomes. Fine for 4d6. The same people also roll 20d6, and 6^20 ≈ 3.7 × 10^15 raw outcomes, which is… not so fine.

Dice kept or dropped together are only distinguished by their sorted multiset, meaning which values came up and how many times, order irrelevant. The number of raw sequences behind each multiset is exactly the multinomial coefficient N! / (c₁! · c₂! · …), where cᵢ counts how many dice landed on face i. Python hands you both pieces:

denom = Fraction(1, sides ** num_dice)
for combo in combinations_with_replacement(range(1, sides + 1), num_dice):
    ways = math.factorial(num_dice)
    for face in set(combo):
        ways //= math.factorial(combo.count(face))
    kept_sum = sum(sorted(combo, reverse=highest)[:keep])
    result[kept_sum] += ways * denom

Enumerate sorted outcomes, weight each by how many of the M^N raw sequences produce it, sum the dice being kept. Exact probabilities for keep-highest, same Fraction arithmetic as everything else. And the collapse is the fun part:

Expression Raw outcomes Sorted multisets
4d6kh3 6^4 = 1,296 C(9,4) = 126
10d6 6^10 ≈ 6.0 × 10^7 C(15,10) = 3,003
20d6 6^20 ≈ 3.7 × 10^15 C(25,20) = 53,130

For 20d6 that’s a factor of about seventy billion. P(4d6kh3 = 18) still comes out as 7/432, i.e. 21/1296, the number printed in every D&D stat table since 1974.

Discord can’t figure out visuals

So you have the exact distribution. Now show it, in Discord. Which means we get to deal with monospace text, 2,000-character message limits, and people reading on phones. /odds 4d6kh3 target >=15 produces (real output, unabridged; this is the whole summary view):

If you got this far, congrats! Have a cookie (づ•ᴗ•)づ 🍪

Range:     3 … 18
Mean:     12.24
Median:   12
Mode:      13  (13.27%)
Quartiles: 10 | 12 | 14
Std dev:  ±2.847

 3 |   0.08%
 4 | ▍  0.31%
 5 | █▏  0.77%
 6 | ██▍  1.62%
 7 | ████▍  2.93%
 8 | ███████▏  4.78%
 9 | ██████████▌  7.02%
10 | ██████████████▏  9.41%
11 | █████████████████▏ 11.42%
12 | ███████████████████▍ 12.89%
13 | ████████████████████ 13.27%
14 | ██████████████████▌ 12.35%
15 | ▒▒▒▒▒▒▒▒▒▒▒▒▒▒▒░ 10.11%
16 | ▒▒▒▒▒▒▒▒▒▒▒░  7.25%
17 | ▒▒▒▒▒▒░  4.17%
18 | ▒▒░  1.62%

P(`>=15`) = 23.15% = 25/108
Odds: 1 in 4.32

The bars are block characters at eighth-block resolution, and the ▒ shading marks the outcomes your target accepts. The bar renderer lives in the shared engine module, imported by both the exact view and the simulation view. /odds 4d6kh3 and /simulate 4d6kh3 x100000 draw with the same function, so the bars mean the same thing when you eyeball one against the other.

Distributions with too many rows get bucketed into ranges first; the CDF view handles them differently, sampling evenly across the support so the table stays readable:

if len(items) > max_rows:
    step = len(items) // max_rows
    items = items[::step][:max_rows]

The CDF table itself is exactly / at-most / at-least, with ◄ marking rows your target hits (abridged here; the real one carries all sixteen rows):

Roll | Exactly | At most | At least
─────┼─────────┼─────────┼─────────
  13 |  13.27% |  64.51% |  48.77%
  14 |  12.35% |  76.85% |  35.49%
  15 |  10.11% |  86.96% |  23.15% ◄
  16 |   7.25% |  94.21% |  13.04% ◄
  17 |   4.17% |  98.38% |   5.79% ◄
  18 |   1.62% | 100.00% |   1.62% ◄

(The CDF is built from a running total kept in a dict named cum_lookup, commented in the source: “it’s cumulative but I am immature”. I’m keeping the comment, trust me there’s worse.)

Because everything underneath is exact, the target line is a fraction: 25/108, odds 1 in 4.32.

Both views live in the same message. Summary and CDF use first-class Discord components, so switching between them is one tap, editing the message in place. A green Roll button below fires an actual roll, scores it against the distribution you’re currently looking at, and logs it. Five-minute timeout, buttons grey themselves out.

The views are packed like sardines, deliberately. Six summary stats, sixteen histogram rows at eighth-block resolution, sixteen more rows of CDF, and the target as an exact fraction, all inside a 2,000-character budget you can drop into a channel mid-session without pushing everyone’s phone screen off the rails. You can even choose to hide it from others with whisper: True, sending an ephemeral message instead.

Everything target-shaped in this section, the ▒ shading, the ◄ markers, the verdict lines, runs off one compiled predicate. >=15, 8-12, !even, in {18, 19, 20} each compile once into a closure and a human-readable string:

if m := re.fullmatch(r'(?:between\s+)?(-?\d+(?:\.\d+)?)\s*-\s*(-?\d+(?:\.\d+)?)', t):
    lo, hi = sorted((float(m.group(1)), float(m.group(2))))
    return (lambda x: lo <= x <= hi), f"{lo:g}-{hi:g}"

Ranges sort their own bounds, so 12-8 and 8-12 are the same target, and ! wraps anything in negation. The whole DSL is about twenty lines, and every consumer just calls the closure.

Ceilings? It’s a feature

Exactness has a price, and convolution is where you pay it. The double loop is O(|a|·|b|) in outcomes, and something like 9999d999999999 would like a word. So the engine refuses before allocating anything, by estimating its total cost up front. For NdM the convolution work is a sum of arithmetic-series terms and the support grows predictably, so the estimate is exact:

def _estimate_sum_dice_cost(num_dice: int, sides: int) -> tuple[int, int]:
    if num_dice < 1:
        return 1, 0
    final_size = num_dice * (sides - 1) + 1
    work = 0
    current_size = sides
    for _ in range(num_dice - 1):
        work += current_size * sides
        current_size += sides - 1
        if work > MAX_CONVOLUTION_WORK:
            return final_size, work  # bail early, caller will reject
    return final_size, work

Too big and you’re told, politely, with actual numbers, then pointed at the right tool:

raise ValueError(f"{num_dice}d{sides} has {final_size:,} possible outcomes! "
                 f"(limit: 100,000) Try /simulate")

The ceilings are like a gradient with their own geography. A size check in the uniform-die constructor, a pre-flight cost estimate in sum_dice, a work check and a post-hoc size check inside the convolution, a multiset count in keep_dice, hard caps on dice and sides way down in the roller. Even the input limits descend in step with how expensive the command is: /roll accepts 1,024 characters, /simulate 512, /odds 256.

At the far end of /roll’s allowance sits the shortest refusal in the codebase:

await ctx.followup.send("I'm....not rolling this", ephemeral=True)

The amount of work it takes to keep everything together becomes too difficult to maintain solo, and error messages are starting to become load-bearing UX. Once computationally heavy expressions point you to /simulate, the routing layer becomes obvious:

Command Promise When the math is too expensive
/roll always rolls rolls anyway, luck scoring silently skipped
/odds exact or nothing refuses, with a cost estimate
/simulate always answers approximates but never refuses

9999d999999999: /roll shrugs and rolls 9,999 dice, /odds declines with a number in the trillions, /simulate grinds until its time budget and returns partial results under a warning banner.

There’s a second reason /odds declines, and it’s less dignified than cost. The roller and the exact engine don’t actually share a parser, there is a weaker one.

The roller’s parser is the regex-and-callback from the top of this post, recursive, rolls dice as it reads. The exact engine can’t use any of that, because substituting a die with its rolled total is precisely the move a probability engine cannot make. It needs the distribution, not a roll. So the odds module grew its own reader:

_TERM_RE = re.compile(r'(\d*)[dD](\d+)(?:(k[hl]?|[hl])(\d+))?')

Plus, a loop that walks the expression in pieces, accumulating dice terms and integer constants with a pending-sign state variable. It handles NdM, one keep clause, and +/- between terms. So nesting, implicit multiplication, or 2d(5+1d5) becomes unsupported, as the exact engine needs to speak a strict dialect of a language the roller is fluent in.

And /roll’s graceful degradation is a try/except with a confession in it:

except (ValueError, NotImplementedError, TypeError, SyntaxError,
        ZeroDivisionError, KeyError):
    pass  # PMF too complex for exact computation - skip luck/flavor

The luck feature became expensive enough to make itself optional at runtime. I’d love to claim this was foresight. It wasn’t; it’s a comment that grew a command economy around it.

It’s also doing double duty. Half the time that pass fires, it isn’t “the math was too expensive”, it’s NotImplementedError from the second parser, but that’s less fun to declare. That’s the twin looking at your expression and not being able to read it. Luck scoring degrades silently, and the comment lets you believe it was about complexity. After the JavaScript port, the count is four parsers for one notation, two of them deliberately dumber than the others.

While reviewing my own code, I noticed MAX_DICE_FOR_KEEP = 30 sits defined and unused in the odds module. A cost cap from an earlier era, superseded by the multiset threshold in keep_dice(). If you look hard enough you can find more, maybe I’ll comment it out one day.

Monte Carlo? Where!

/simulate is where everything too big for exact math gets routed: roll the expression up to a million times and report what actually happened.

A Discord bot is a single event loop with response deadlines, so the simulation runs on a thread under a hard eight-second budget. The clock is checked once per percent of trials rather than once per trial, because a million time.monotonic() calls are their own denial of service:

check_every = max(1, trials // 100)
for i in range(trials):
    if i % check_every == 0 and time.monotonic() > deadline:
        return results, True   # partial results, flagged, still delivered

Hit the budget mid-run, you’ll still get results! Just from however many trials finished, with a warning banner. I want to apologize in advance if you were checking whether your one-million-ogre army survived the battle, to find 567,000 magically vanished.

The modifier parameter is advantage, generalized: simulate 1d20 x100000 modifier 2 rolls each trial twice and keeps the higher. Elven accuracy, Lucky feat, and every ‘roll twice, keep best’ house rule all collapse into one integer.

When you estimate a hit rate from N trials, the honest move is a confidence interval, and the honest interval is Wilson’s, not that Wald formula everyone memorizes. Wald misbehaves at the extremes: intervals that dip below zero for rare events, wobble where they shouldn’t. Wilson stays bounded and sane:

def wilson_ci(successes: int, trials: int, z: float = 1.96) -> tuple[float, float]:
    p = successes / trials
    denom = 1 + z * z / trials
    center = (p + z * z / (2 * trials)) / denom
    margin = (z * math.sqrt(p * (1 - p) / trials + z * z / (4 * trials * trials))) / denom
    return max(0.0, center - margin), min(1.0, center + margin)

Result of a seeded run of a hundred thousand simulated d20s against target >=15. The true answer is 30%:

hits: 30,157 / 100,000  (30.16%)
Wilson 95% CI: 29.87% – 30.44%
longest hit streak: 8    (theory says ~10)
longest miss streak: 31  (theory says ~32)

Because 100k trials also answers the question people actually ask: “Am I cursed, or just unlucky?” There’s streak analysis: run lengths of hits and misses, plus the theoretical ceiling. For n trials with hit probability p, the expected longest streak of hits is about log(n)/log(1/p); a million fair coin flips should top out around 20 in a row. Above, the observed streaks sit right where theory says they should. Spoiler: The dice are not out to get you. Sorry.

BATTLE READY

At this point I needed insurance, as the question changes from “can the engine do this” and becomes “how do I know it isn’t lying.” Exact probabilities are a strong claim, and strong claims need a test suite. I wrote one, and it says more about this project than any feature in it.

The fuzzers generate random expressions, nested dice, keep clauses, exponents, implicit multiplication, thousands of seeds, and just fires them at the engine. Here’s what passing looks like:

EXPECTED = (ValueError, NotImplementedError, ZeroDivisionError, SyntaxError)
...
except EXPECTED:
    continue  # a clean refusal is a pass

A refusal is a pass, the suite encodes the thesis of the ceilings section as a constant: the engine is allowed to say no, and saying no correctly is success.

You don’t need to ask, of course there’s a time limit on refusals:

HOSTILE_TIME_LIMIT = 2.0  # seconds, refusals must be effectively instant

Not “eventually refuse.” Refuse fast, measured and asserted. Nobody writes that constant unless they’ve watched a bot sit there thinking about 9^9^9. The hostile-input roster reads like a YouTube griefing tutorial: 9^9^9, (9^9)^9, 2^999999999, 2d6; print('hi'), import os.

Every hostile-input test since is a thank-you note, the test names are tracked, safe_eval: hostile input is refused (issue #1), breakdown display regressions (issue #2). Somewhere out there are the people those numbers belong to. Issue #1, judging by the test that carries it, is somebody who figured out they could exponentiate the bot into a coma.

Then the suite proves the two halves of the engine describe the same universe:

pmf = expression_to_pmf(expr)
for _ in range(AGREEMENT_ROLLS):
    total, _ = roll_expression(expr)
    assert int(total) in pmf

25,000 rolls, every one landing inside its own exact distribution. The keep-highest math is checked against brute force on all 54 small combinations. The 7/432 from earlier is pinned in an assert.

My favorite is the structural one, to verify that breakdown strings don’t lie (3+2d6 ➜ 3+[6 + 1 = 7]) the suite strips brackets and parentheses in a fixpoint loop until nothing changes, then asserts the operator sequence survived:

def op_sequence(text: str) -> list[str]:
    prev = None
    while prev != text:
        prev = text
        text = re.sub(r'\[[^\[\]]*\]', '', text)
        text = re.sub(r'\([^()]*\)', '', text)
    ...

None of this runs on pytest, it’s a hand-rolled harness with a scoreboard that prints BATTLE READY when everything passes.

All of this rigor lives on one side of the project. The suite proves the roller and the odds engine agree. It says nothing about anything else in the system, and it does not cross borders.

The bot becomes a platform

I just kept yearning for more, more buttons, interactions, data, so I built them. Saved macros with Discord-native autocomplete, random tables with weighted entries, reroll-from-history, and per-server feature toggles. When commands became too cumbersome for these features, the next step was a website.

Every roll is logged (expression, breakdown, result, luck score), and that last field is the tell: the data was the point all along. One SQLite file per server, stored in a folder named after the server’s ID. The dashboard serves the log back as interactable data containing most rolls, luckiest, unluckiest, top rolls, most-used expressions:

self.cursor.execute("CREATE INDEX IF NOT EXISTS idx_roll_luck ON roll_history(luck_score)")

Logging exact percentiles means the dashboard can chart whether a server runs hot over time, and you can frame the 2d10000 ➜ 2 roll on the wall.

The exactness doesn’t extend to everything, though. The roll-history endpoint paginates, and its count function knows about fewer filters than the query does:

# Note: get_roll_count only supports user/date filters, so `total`
# can overestimate when other filters are active. Fine for now.

With the backend setup it was time to start the frontend. Quickly realized Flask isn’t async, so the dashboard is a Quart app on stilts, running as a coroutine sibling on the bot’s event loop:

await asyncio.gather(serve(app, config), bot.start(get_token()))

One process, the web server and the Discord bot as two arguments to gather.

The engine that computes 7/432 ships a page counter that can overestimate, and the code knows it. There’s also a macro-edit endpoint that doesn’t pretend:

return {"error": "Editing is not wired up yet"}, 501

A present route that declines is that same ‘half-feature’ species as the orphaned functions in the odds module, you’ll meet them later. Less importantly, there are the standard per-command cooldowns: rolls at three per five seconds, simulations at two per ten.

Then comes the part that finally scared me. I’m not even talking about self-hosting, docker, nginx, cloudflare, or the oauth full stack. An odds engine in a browser meant allowing anything running chrome to run arbitrary expressions directly on bare metal: the same expressions whose cost I’d spent this whole post refusing in /odds.

So the engine got ported to JavaScript. All of it. Your client crunching away numbers while my server serves static files and staying out of the way. BigInt rationals where Python had Fraction. The convolver, the multiset enumeration, the cost ceilings, the three-regime routing, the entire exact-plus-empirical apparatus, built twice. Maybe with some nicer visuals. While we’re opening doors, let’s throw in moderation too because why not? So I made starboard, reaction roles, welcome messages, etc., can’t even blame that on the dice.

I’d call you a fool if you asked me to double my workload. I did it anyway. Every feature from now on costs two implementations, a parity tax, forever. My problem, because the downside is worse: the Python side and the JS side have to agree, feature for feature, or the dashboard quietly lies to people.

Parity isn’t fair

I said the engine got built twice; what I meant is twice is costly. Python’s refusals emigrated with the code, for better and worse:

if (sides > MAX_OUTCOMES) throw new Error(`A d${sides.toLocaleString()} has more outcomes than the engine can handle! (limit: ${MAX_OUTCOMES.toLocaleString()}) Try /simulate`);

There aren’t even slash commands on a web page. The ceilings, Discord time budgets, and other server anxieties were ported into an environment where the routing they describe doesn’t exist. In a Web Worker, on the user’s own machine, nothing is waiting.

“Math you can’t eval()” was a twelve-line story because ast.parse is in Python’s standard library. JavaScript has no ast and I’m still avoiding eval, so the parity tax on that one section was a hand-rolled precedence climber:

function parsePower() {
    let base = parseFactor();
    if (peek() === '*' && expr[pos + 1] === '*') {
        pos += 2;
        return Math.pow(base, parsePower()); // right-associative
    }
    return base;
}

The same safety guarantee and five times the code becomes my personal responsibility.

The term regex exists twice, anchored and global, because JavaScript regexes carry lastIndex as mutable state. Which resulted in this beauty:

// Reset lastIndex in case (replace doesn't use it, but split might on some engines)
DICE_PATTERN.lastIndex = 0;

The renderers disagree on fundamentals, too. Discord’s 2,000-character budget caps the bot’s tables at 20 rows; the browser uses 120, 150, 40. Same expression, different bucketing, different bars. The constraint that shaped every rendering decision in this post simply doesn’t exist on the other side.

Now the specimen. Every command’s target supports odd, the implementations look the same line across both languages:

lambda x: int(x) % 2 == 1        # Python
x => Math.trunc(x) % 2 === 1     // JavaScript

Python’s % returns the sign of the divisor, so -3 is odd. JavaScript’s returns the sign of the dividend, so -3 is not odd. I wrote earlier that the two sides have to agree “or the dashboard quietly lies to people.”

I do want to bring up one more structural problem. The exact engine’s keep-highest weights come from a factorial built as a Number loop, and from Math.pow(sides, numDice) for the denominator. factorial(20) is about 2.4 × 10¹⁸, past the point where integers survive floating point. So for anything like 20d6kh3, which is well inside the engine’s own limits, the browser’s “exact” answer isn’t. The twin is exact right up until it silently isn’t, and I found out writing this section.

But the gap that matters is this one. The test suite proves the Python roller and the Python odds engine describe the same universe: 25,000 rolls, all landing inside their own support.

There is currently no equivalent on the JavaScript side. There is no cross-language test at all. Parity, that thing I said I’d be paying for forever? It’s verified the old-fashioned way: using eyes and aligning bar charts. It just looks right, the careful got me -3 anyway.

The seed incident

I want to close with the scope creep that happened while writing this post, because it’s the purest specimen I have.

While drafting the section about random.randint, I thought: reproducible rolls would be nice. A seed parameter; replay a sequence, settle “the bot rigged my nat 1” disputes, deterministic replays for testing. Reasonable, right? Here’s the chain it started:

  1. Python: random.seed(), Mersenne Twister. Done in ten minutes.
  2. JavaScript: Math.random() is not seedable, by spec. Fine… grab a seeded PRNG library.
  3. …which implements a different algorithm than Mersenne Twister, so the same seed produces different roll sequences on the bot and the dashboard.
  4. Ah, I know! To solve parity, I’ll just implement one shared PRNG in both languages, plus a test suite proving the two implementations produce identical streams.
  5. Which means ripping out the random.randint call in roll_dice(), the last living descendant of the one-liner this post opened with.

The newest feature would have killed the oldest fossil. I stopped at step five and wrote this section instead, but I want to be honest: I know exactly how to do it. I could, and that’s the scary part.

For further evidence that this project intends to continue without my consent, the odds module already contains p_greater(), p_equal(), and margin_pmf() — opposed-roll support, implemented, called by nothing. Half a feature, waiting for a slow weekend.

Though, none of this is a complaint. The one-liner still works; you can see it in there, buried in sediment, with eight years of somebody’s reasonable decisions layered on top. I haven’t decided it’s done, and I’m not sure I’m going to. There’s a non-zero chance this post is already the newest layer, and clearly there’s still work to be done.

Again, if you want to poke at the strata: the bot and its dashboard are live, add it to your server and experiment with Dice Playground. Or the engine’s source is on GitHub, lite version, everything in this post intact.

Thanks for reading, see you around! ✧ ദ്ദി