Enigma × Jev

Jev performance: how the analyst works, how it compares, and every result.

Research · Jev performance

A probabilistic analyst on trial

Evaluating Jev as crib selector and plaintext judge in a software Bletchley pipeline

Abstract

Breaking Enigma in 1940 needed two kinds of judgement that no machine supplied. Someone had to decide which probable plaintext (the crib) to set on a Bombe, and someone had to decide whether a Bombe stop's trial decryption was German.

We study whether Jev, a hosted model that answers only typed probabilistic questions, can take both roles. We reconstruct judge calls and crib-ranking calls from an audit log and label each against the known plaintext. Jev's probability that a candidate is correct separates right from wrong decryptions almost perfectly (AUC , candidates from messages), but it is systematically underconfident (ECE ). A monotone recalibration, fitted leave-one-message-out, lowers the Brier score from to . That beats the project's trigram judge calibrated the same way ().

Against the standard alternatives run on the same candidates, the picture splits. Quadgram and Kneser–Ney fitness, logistic regression and XGBoost match recalibrated Jev when trained and tested on the same kind of traffic. Trained on synthetic traffic and tested on real intercepts, XGBoost's decisions fall to , while zero-shot Jev keeps . A two-condition decision rule makes the right call in of cases. The answers are stable on repetition and unaffected by candidate order.

As a crib selector, however, Jev shows no skill over a fixed list. We discuss why, and what the division of labour between computation and judgement implies.

Research note · Enigma × Jev

AUC: Jev separating correct from wrong decryptionsn-gram judge
/right accept/reject decisions under the ruleraw top choice
Brier score, Jev recalibratedraw · n-gram
same decision when asked again with candidates reversed
mean rank of the true crib in Jev's orderfixed list · no skill
msmedian latency of a judge call~ tokens in, ~ out

1. Introduction

The electromechanical Bombe answered a narrow question: which rotor settings are consistent with this crib? Everything around that question was human judgement. Hut 6 chose which crib to run from traffic analysis and habit. Once a stop came back, an analyst decided whether the trial decryption was genuinely German before a key was declared broken (Welchman 1982). Turing and Good formalised that judgement as the accumulation of evidence in logarithmic units, bans and decibans (Good 1979). It was the first sustained use of sequential Bayesian inference.

The Enigma × Jev system rebuilds the computational parts faithfully (see the research paper) and hands the two judgements to Jev. Jev is a hosted model reached through TypeSafe's System One interface. It returns probabilities and choices, never prose. This paper asks how good those judgements are, how they compare with the obvious statistical alternative, and whether they are trustworthy enough to act on. The contributions:

  1. A labelled evaluation of Jev's behaviour, reconstructed from the complete audit log and scored against known plaintexts. The quantities are discrimination (AUC), accuracy of probabilities (Brier, log loss and Murphy's decomposition), calibration (reliability diagrams, ECE) and decision quality, each with cluster-bootstrap confidence intervals.
  2. A like-for-like comparison with an n-gram judge. The n-gram judge is given every advantage: it is the search's own language model, calibrated on labelled data. Jev receives no labels.
  3. Two controlled experiments on the instrument itself: test–retest reliability, and sensitivity to the order in which candidates are presented.
  4. An account of the one role Jev did not fill, crib selection, and why.

2. How Jev is used

2.1 The interface: a state and typed questions

Each call sends a plain-text state and a set of named questions, each of one of three types. A noul question returns one probability in [0, 1]. A choice question returns a probability for every named option, the argmax and a confidence. A score question returns a distribution over an ordered scale. Before anything downstream sees an answer, the client (src/jev/client.ts) checks five things:

  • every question is answered, with the requested type;
  • each choice distribution covers exactly the offered options and sums to one within 10⁻⁵;
  • the chosen option has the highest probability;
  • the reported model revision matches the one requested (jev-1.13.0);
  • token usage is reported.

A malformed answer is rejected rather than repaired. Jev cannot produce a key or a decryption, so it cannot hallucinate one either; every quantity it influences is a probability over options the software enumerated.

State (sent)

Questions (sent)

Answers (returned, )

Figure 1. A real judge call from the main run (), reproduced from the audit log with candidate texts shortened. The first candidate is the correct decryption. Its first letters were the crib, and the state says so.

2.2 Two roles: a prior and a posterior

Crib selection (prior). After crib dragging removes every listed phrase that would place a letter over itself, Jev receives the traffic type, date, message length, the surviving and eliminated phrases, and the first forty cipher letters. It answers one choice over the survivors plus “none of these”. The Bombe then runs the cribs in Jev's order.

Plaintext judgement (posterior). After a Bombe run, Jev receives the ciphertext and up to three distinct candidate decryptions. It answers a choice (which candidate is correct, or none) and one noul per candidate, P(correct). Two design decisions make this an independent judgement rather than an echo of the search. First, the search's own n-gram scores are withheld. Second, letters that were assumed as a crib are marked, so a wrong key cannot borrow credibility from German that was typed in by construction.

2.3 The decision rule

A candidate is accepted only when Jev's choice falls on it with probability at least 0.5 and its own P(correct) is at least 0.5: more likely than not on both questions. The rule was adopted after the pilot run (§4.4), which showed that Jev's argmax over sets of wrong candidates often lands on one of them at low confidence. The main and holdout runs evaluate the rule after it was fixed.

2.4 What is not claimed

We treat Jev as a black box. Its architecture, training data and inference procedure are not examined here, and none of our conclusions depends on them. We evaluate behaviour at one pinned revision, on one task family, through the interface above.

3. Evaluation design

3.1 Data and labels

Every request and answer is appended to an audit log (~/.enigma-jev/jev-calls.jsonl). Each judge state contains the ciphertext, and each crib state contains the ciphertext's opening, so every call can be matched to a message whose plaintext is known. The messages are nine historical intercepts with published keys and sixteen synthetic messages under two independent sets of random keys. A candidate is labelled correct when it matches at least 90% of the known plaintext over the letters shown. A crib is true when the plaintext begins with it.

The machine slept during the main run, which invalidates the report's elapsed-time field. Run boundaries are therefore taken from the burst of crib calls with which every run begins. Calls from the interactive page that match no run's messages are excluded.

Table 1. Calls by phase. The evaluation set pools main and holdout.

3.2 Baselines

  • n-gram, fixed bar. The German-ness statistic the search itself uses: normalised trigram log-probability, 0 for random letters and 1 for typical German, computed outside any crib. It accepts the most German candidate at ≥ 0.4.
  • n-gram, calibrated. The same statistic mapped to a probability by logistic regression (Platt 1999). The regression is fitted leave-one-message-out, so each message is scored by a model that never saw it. It accepts at ≥ 0.5. This baseline uses labels; Jev does not.
  • Accept the search's top candidate. What the pipeline would do with no judge at all.
  • For cribs: the fixed list order, a uniform distribution over the options, and a random order.

3.3 Metrics

Discrimination. The area under the ROC curve: the probability that a randomly chosen correct candidate is scored above a randomly chosen wrong one (Hanley and McNeil 1982). Accuracy of probabilities. The Brier score (Brier 1950) and log loss, both strictly proper scoring rules (Gneiting and Raftery 2007). The Brier score is decomposed as reliability − resolution + uncertainty (Murphy 1973). Calibration. Reliability diagrams (DeGroot and Fienberg 1983), and the expected calibration error over ten equal-width bins (Naeini, Cooper and Hauskrecht 2015). Decisions. Accuracy, precision and recall of each accept/reject rule.

Intervals are 95% percentile bootstraps with 2,000 resamples (Efron and Tibshirani 1993). They resample calls, not candidates, because the candidates in one call share a message and are not independent.

4. Results

4.1 Discrimination

Over the evaluation set of candidates, of which were correct (base rate ), Jev's P(correct) achieves an AUC of (95% CI ). The calibrated n-gram judge reaches (). Figure 2 shows why the difference, though small in AUC, matters. Apart from one near-miss, every correct candidate receives a higher P(correct) from Jev than every wrong one, with a clear gap around 0.5. The exception is the Scharnhorst decryption: it is 89% right, so the 90% rule labels it wrong, and Jev scores it 0.58. The n-gram statistic's two populations overlap near its threshold, because trigram hill-climbing can manufacture German-looking fragments from a wrong key.

Figure 2. Every evaluated candidate, by judge and by truth. Left: Jev's P(correct). Right: the n-gram German-ness statistic. Dashed lines mark each judge's acceptance bar. Hover a point for its value.
Data table
Figure 3. ROC curves. Both judges hug the top-left corner. Jev reaches full recall at a single false positive, the Scharnhorst near-miss. The n-gram judge needs several more false positives to reach full recall.

4.2 Calibration

Discrimination is not calibration. Jev's probabilities are underconfident. No candidate it scored below 0.5 was correct, and every candidate it scored above 0.6 was correct, yet its probabilities for correct decryptions average about , where 1 would be warranted. Its expected calibration error is , against for the calibrated n-gram judge, and its Brier score is () against . Murphy's decomposition (Table 2) locates the difference. Jev has the higher resolution, against , so it carries more information about which candidate is right. It loses entirely on the reliability term, the penalty for miscalibration.

Because Jev's errors are monotone, a one-parameter-pair logistic map, fitted leave-one-message-out exactly as for the n-gram judge, repairs them. The recalibrated Jev scores a Brier of () and an ECE of . It is better calibrated than the trigram judge, and its ranking is untouched; §4.8 shows that trained classifiers given enough labels can match it. This mirrors findings on modern neural models, whose confidence is informative but miscalibrated in a correctable way (Guo et al. 2017; Kadavath et al. 2022).

Figure 4. Reliability diagram. Each point is a bin of candidates: mean predicted P(correct) against the observed share correct. Point area shows the bin's size. A calibrated judge lies on the diagonal. Raw Jev bends away from it at both ends. Recalibrated Jev and the n-gram judge lie close to it where their bins hold most candidates; the sparse middle bins are noisy (see the data table).
Data table
Table 2. Proper scores and the Murphy decomposition (Brier = reliability − resolution + uncertainty). Lower is better except resolution.

4.3 Decisions

The rule of §2.3 made the right call in of judged calls (, CI ). It never rejected a correct decryption. Its single error was an acceptance: the Scharnhorst message, whose accepted decryption matched of the published plaintext, just below the 90% bar. That is a near miss rather than a wrong key. The alternatives in Figure 5 fare worse:

  • Jev's raw top choice accepts a wrong candidate times.
  • The search's fixed n-gram bar accepts one times.
  • The calibrated n-gram judge is precise but misses correct decryptions.
  • With no judge at all, the pipeline would be right of the time.
Figure 5. Accuracy of each accept/reject rule on the evaluation set, with 95% bootstrap intervals. Colour follows the judge: blue for Jev, orange for the n-gram statistic, aqua for no judge.
Confusion counts

Jev's two answers agree with each other. P(pick) alone separates right from wrong picks with AUC . The probability of “none” separates calls with no correct candidate from the rest with AUC , though, like P(correct), it is too cautious in magnitude (Brier ).

4.4 The question matters: a prompt ablation

The pilot run used the first version of the judge state, which did not mark crib letters, and applied no decision rule. On the pilot calls, Jev's top choice was right of the time, against after the change. Its mean P(correct) on wrong candidates fell from to , and on correct ones rose from to . Applied after the fact, the decision rule would have raised the pilot's decision accuracy from to . The comparison is observational: the pilot covered an earlier and smaller subset of messages. It nonetheless indicates that telling the judge what was assumed, and not only what was found, sharpens its judgement.

4.5 Crib selection: no skill

Jev's ranking of cribs was evaluated on the calls whose message opens with a listed phrase. The results:

  • Mean rank. The true crib's mean rank was in Jev's order, in the fixed list, and in expectation for a random order.
  • Top three. Both Jev and the list put the true crib in the top three times.
  • Log score. Jev's log score on the true crib, , was worse than a uniform guess, . Its distribution is concentrated, with a top option averaging , and usually on generic phrases such as KEINEBESONDERENEREIGNISSE or TAGESMELDUNG.
  • “None of these”. Its probability was indistinguishable whether or not a listed crib was present ( against ; AUC ).
Figure 6. Rank of the true opening crib for each call, in Jev's order and in the fixed list (lower is better). The grey track spans all the phrases on offer.

The contrast with §4.1 is instructive. Judging a decryption is a problem of recognition, and all the evidence is in the state. Choosing a crib is a problem of prediction from context, and the state supplied little: traffic type, date, length. Hut 6's crib-writers worked from call signs, time of transmission, routing, the previous day's traffic and the habits of individual operators (Welchman 1982). A prior can only be as informative as the context it conditions on. The negative result is a statement about the question asked as much as about the model.

4.6 Reliability of the instrument

Two experiments re-asked logged judge questions, drawn half from calls in which Jev had accepted a candidate and half from calls in which it had accepted none.

  • Test–retest. The identical state and questions, sent again, gave the same decision of the time. P(correct) moved by on average, with a correlation of between the two runs.
  • Position. The same candidates in reversed order gave the same decision of the time, and the picks moved with the content. Across the evaluation set Jev's top choice fell on the first slot times, where the correct candidate sat times. The experiment shows that surplus is deference to the search's best-scoring candidate, which is listed first, and not a preference for the slot.

Position bias is a documented failure of language models used as judges (Zheng et al. 2023; Wang et al. 2024). Its absence here matters for trusting the decisions. A third test re-asked crib rankings: Jev returned the same top crib of the time, with a rank correlation of . Its crib priors are stable, merely uninformative.

Figure 7. P(correct) for the same candidate on two askings. Left: identical prompt. Right: candidates reversed. Points on the diagonal are perfectly consistent; the dashed lines mark the 0.5 acceptance bar.

4.7 Cost and latency

A judge call took a median of ms (90th percentile ms), and a crib call ms. Calls averaged about input and output tokens. A full break on the page costs one crib call plus one judge call per Bombe run: for a successful break, typically two to five calls, well under a second of the analyst's time. For comparison, the Bombe itself takes seconds per run even on eight cores.

Figure 8. Latency of every logged call, by question type. The tick marks the median.

4.8 Against the state of the art: n-gram fitness and XGBoost

What a practitioner would use instead. Since Gillogly, Enigma breakers have scored trial decryptions with the index of coincidence in early stages and with n-gram log-likelihood, trigrams up to hexagrams, in the final ones. Lasry's thesis made that the standard toolkit for classical ciphers, and modern breaks still use it (Gillogly 1995; Weierud and Sullivan 2005; Ostwald and Weierud 2017; Lasry 2018). In machine learning, the default for a small tabular problem is gradient-boosted trees, with XGBoost as the reference implementation (Chen and Guestrin 2016). The research frontier for decipherment replaces n-gram models with neural language models (Kambhatla, Mansouri Bigvand and Sarkar 2018). Jev is itself a pretrained model used zero-shot; training a neural character model would need a far larger German corpus than this project has, so that variant is left to future work.

Protocol. Nine judges score the same candidates against the same 90% label:

  • the index of coincidence;
  • this project's trigram German-ness;
  • quadgram fitness, the usual hill-climbing score;
  • an interpolated Kneser–Ney character 5-gram model (Kneser and Ney 1995);
  • logistic regression and XGBoost on twelve text features;
  • Jev, raw and recalibrated.

The language models learn from the project's German corpus ( training letters). The features are:

  • the three n-gram scores;
  • coincidence, χ² against German letter frequencies, entropy, and vowel and common-letter shares;
  • unseen trigrams and the X rate;
  • dictionary coverage and the longest covered run, against corpus words.

All are computed, as for the n-gram judge, on the letters outside the assumed crib. Every judge but raw Jev turns its score into a probability by learning from labels. Evaluation is leave-one-text-out: all runs of a plaintext, which recurs in main and holdout under new keys, are held out together, so no trained judge ever sees the text it is scored on. Intervals come from a cluster bootstrap over calls. The code is analysis/judges.py, on scikit-learn and XGBoost (Pedregosa et al. 2011). XGBoost is deterministic on one platform but not across them, because arm64 and x86-64 builds break near-tied splits differently. The figures here were computed on . Rerun on x86-64 Linux, XGBoost's AUC is 0.988 rather than 0.993, and 0.965 rather than 0.993 under shift. Every decision count, including the 53% under shift, is identical, and so is every other judge's result.

Table 4. Candidate-level scores (AUC, Brier, log loss, ECE) and call-level decisions: accept the top-scoring candidate when its probability reaches 0.5, otherwise accept none. The last row is Jev's two-condition rule (§2.3). Best values in bold.

Given labels from the same traffic, the field closes. XGBoost reaches a Brier score of and plain quadgram fitness , against for recalibrated Jev. The paired difference between XGBoost and recalibrated Jev is , with a 95% interval from to , which is no difference at all. Zero-shot Jev still ranks best: AUC , against for XGBoost, whose trees misorder a few candidates that logistic regression on the same features does not (). What XGBoost leans on most is , then and . In effect it learns to be a dictionary-and-n-gram judge. On decisions, Jev's own rule is the only judge with a single error ().

Figure 9. What labels buy. Each judge is trained on k randomly chosen texts and scored on the rest, 40 draws per k; lines are mean Brier scores, plotted to k = 20 (the table adds 24, where one held-out text decides the score). The dashed line is zero-shot Jev, which uses no labels.
Table view

Labels are the price. Given one labelled text (about candidates), XGBoost's Brier score is , twice as bad as uncalibrated Jev's . It overtakes raw Jev at two texts ( labelled candidates, Brier ). Only at about twelve ( candidates, ) does it approach recalibrated Jev. Recalibrated Jev reaches from a single text, because it fits two numbers and inherits everything else. The Kneser–Ney model behaves similarly for the same reason: its knowledge of German comes from its corpus, not from the labels.

Table 5. Distribution shift. Every trained judge learns from the synthetic traffic alone ( candidates) and is scored on the real intercepts ( calls, candidates, correct).

The decisive test is traffic the judge has not seen. Taught only on the synthetic messages and asked about the real ones, XGBoost still ranks well, but its probabilities no longer mean what they meant. It accepts nothing wrong, yet it rejects of correct decryptions, and its decisions fall to . Logistic regression falls to and the Kneser–Ney judge to . Zero-shot Jev decides correctly, as it does everywhere else. Real naval and army traffic, with its numbers, abbreviations and garbles, lies outside anything learned from one author's clean synthetic German. A model that has read vastly more German than letters does not have that problem. Note too that Jev recalibrated on synthetic traffic alone decides worse than raw Jev (). Calibration learned on one kind of traffic carries that traffic's bias.

Cost. The classical judges win by orders of magnitude. The features take ms a candidate and a tree prediction µs, against ms and a network call for Jev. In a pipeline that asks for two to five verdicts per break, behind a Bombe that runs for seconds, the difference does not matter.

5. Discussion

What Jev contributes. Three results stand out. First, as a reader of trial decryptions Jev ranks candidates better than the search's own statistic, and does so without labels. The n-gram judge that rivals it had to be calibrated on labelled messages, and it shares its model with the search whose output it judges, so it is not an independent check. Second, Jev's miscalibration is monotone and cheap to repair. With a single labelled text its probabilities become about as well calibrated as a gradient-boosted classifier's, which needs roughly ten times the labels to get there (§4.8). Third, its capacity to abstain is what makes the pipeline safe. Without a judge, the system accepts a wrong key in more than half of all calls; with Jev under the rule, it did so once, on a near-correct decryption.

How it compares. Against the project's own trigram judge, Jev is better at ranking (AUC against ). Out of the box it is worse calibrated; recalibrated, it is better. Against the state of the art (§4.8) the honest summary has two halves. Where labelled traffic of the same kind is plentiful, quadgram fitness or XGBoost is as good as recalibrated Jev, about three hundred times faster, and needs no network. Where it is not, and in real codebreaking it never is, the trained judges' probabilities break: on real intercepts, after training on synthetic ones, XGBoost rejected of correct decryptions and Jev none. Against the historical analyst we can say less: Hut 6 left no forecast records to score. The division of labour, though, is the same. The machine enumerates and the judge decides, and the judge's value lies in knowing when not to accept.

The checking problem, then and now. A Bombe stop was an answer, not the answer. Each had to be tried on an Enigma analogue at something like a quarter of an hour apiece, so a hundred stops were more than a day's work, and the key changed at midnight (Veritasium n.d.). Welchman's diagonal board mattered because it cut the stops to a handful (Welchman 1982). The live Bombe in the research paper reproduces the effect on a real menu: a 14-letter crib gives 85 stops with the board and 724 without, and the full 23 letters give 1 against 2. In this pipeline Jev is the checker. At a median of ms a verdict, a hundred stops are checked in –, against more than a day in 1940. The step that once throttled the whole operation now costs less than the search it follows. The diagonal board, the consecutive-stecker switch and the Jumbo Bombes' “machine gun” all removed stops before anyone had to read them (Weinbaum 2025). Jev reads the stops that remain, and says “none of these” when none is German. In Turing's units the acceptance bar of 0.5 is 0 decibans, even odds, and a verdict of 0.99 is +20: 40 hubdubs in the half-decibans of the Banburists' scoring sheets. The machine page now prints each verdict in decibans.

Where it fails. Crib selection. The fault is at least partly in the question: a prior conditioned on so little context cannot be sharp. The obvious next experiment gives Jev richer traffic context and scores the change with the same log score.

6. Limitations

  1. Small, correlated sample. message-runs from 25 distinct texts. The synthetic texts reappear in the holdout under new keys, so candidates are not independent across runs. The call-level bootstrap addresses correlation within a message, not this reuse.
  2. Easy negatives. Most wrong candidates are clearly wrong. The hard cases, near-correct keys with one wrong plug or ring, are rare in this set, and they are where a judge is most tested.
  3. Post hoc rule. The decision rule was chosen after the pilot. Main and holdout evaluate it prospectively, but on messages from the same distribution.
  4. In-sample recalibration. Recalibration is leave-one-message-out, yet fitted within the same study. Its advantage should be confirmed on fresh traffic.
  5. One revision, one domain. The results describe jev-1.13.0 on this task, through this state wording. §4.4 shows the wording matters.
  6. A small shift test, and one home-grown corpus. The distribution-shift result (Table 5) rests on calls from nine intercepts. The classical judges' language models and dictionary come from the same author as the synthetic messages, which flatters them in-distribution and may exaggerate their fall on real traffic. A judge trained on a large, independent German corpus would narrow the gap; a neural character model in the style of Kambhatla et al. was not tried.
  7. Labels at 90%. The threshold makes an 89% decryption “wrong”. A different threshold would move the single wrong acceptance into the correct column.

7. Conclusion: can it decrypt Enigma?

The evidence divides cleanly by what the codebreaker knows in advance:

Table 3. What the workflow achieves, by prior knowledge (main backtest; historical and synthetic messages).
The codebreaker has…OutcomeHistoricalSynthetic
the day's key sheet (rotors, rings, plugs)✓ always broken: the message key falls in a single scan//
a crib: the message's first 14 letters✓ broken: the full key, in seconds to minutes//
only Jev's guess at a crib, top three tried✗ not broken: the guess is rarely right//
nothing but the ciphertext✗ not broken: ten plugs hide the statistical signal//

The machine decrypts; Jev decides. The key is recovered by computation: the software Bombe and the plugboard hill-climb. Jev only answers questions of probability, so it cannot produce a key or a plaintext, and it cannot invent one either. Its contribution is the verdict, and there it is superb. It separates real decryptions from impostors with an AUC of . Under a simple more-likely-than-not rule it made of accept/reject decisions correctly. Its single error was the Scharnhorst decryption, right against a 90% bar: a near-miss, not a wrong key. It gave the same verdict of the time when asked again, and when the candidates were shuffled.

Why Jev and not a classifier. Standard tools can do Jev's job where they have been taught. Quadgram fitness and XGBoost, given labelled breaks of the same kind of traffic, match recalibrated Jev (Brier against ). But codebreaking is the business of traffic nobody has labelled yet. Taught on synthetic messages and shown real intercepts, XGBoost's decisions fell to , while Jev, taught nothing, stayed at .

Why the judge matters. Remove Jev and trust the search's best guess, and the pipeline accepts a wrong answer in of calls. Wrong keys still produce German-looking fragments, because the plugboard search tunes them towards German. A judge that can say “none of these” is the difference between a codebreaker and a generator of plausible nonsense. Jev's one weakness here is modesty. It says 0.7 when it could say 0.99, and a two-parameter recalibration fixes that (Brier → ).

Where it stalls: the crib. Every break above began with a probable phrase. Asked to choose that phrase from the traffic type, date and length, Jev did no better than a fixed list: the true crib's mean rank was in Jev's order against in the list. A message whose opening is on the list still breaks, because the page tries every listed crib, typically within 5–25 seconds. A message whose opening is not on the list, and for which the user supplies no guess, usually does not break. The bottleneck of 1940 remains the bottleneck: not computation, but knowing what the enemy is likely to have said.

Let the machine enumerate, let the judge decide, and give the judge the evidence it needs. That division of labour broke Enigma in 1940, and it breaks it here.

Reproduce: bun run src/analysis/jev-eval.ts rebuilds every Jev figure on this page from the audit log; bun run src/analysis/jev-experiments.ts reruns §4.6 (about 58 calls). Table 3 is read from the newest backtest report. The 5–25 second range is from timed runs on the codebreaker page. Analysis generated .

References

  • Brier, G. W. (1950) ‘Verification of forecasts expressed in terms of probability’, Monthly Weather Review, 78(1), pp. 1–3.
  • Chen, T. and Guestrin, C. (2016) ‘XGBoost: A scalable tree boosting system’, Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 785–794.
  • DeGroot, M. H. and Fienberg, S. E. (1983) ‘The comparison and evaluation of forecasters’, Journal of the Royal Statistical Society, Series D (The Statistician), 32(1/2), pp. 12–22.
  • Efron, B. and Tibshirani, R. J. (1993) An Introduction to the Bootstrap. New York: Chapman & Hall.
  • Gillogly, J. J. (1995) ‘Ciphertext-only cryptanalysis of Enigma’, Cryptologia, 19(4), pp. 405–413.
  • Gneiting, T. and Raftery, A. E. (2007) ‘Strictly proper scoring rules, prediction, and estimation’, Journal of the American Statistical Association, 102(477), pp. 359–378.
  • Good, I. J. (1979) ‘Studies in the history of probability and statistics XXXVII: A. M. Turing's statistical work in World War II’, Biometrika, 66(2), pp. 393–396.
  • Guo, C., Pleiss, G., Sun, Y. and Weinberger, K. Q. (2017) ‘On calibration of modern neural networks’, Proceedings of the 34th International Conference on Machine Learning, PMLR 70, pp. 1321–1330.
  • Hanley, J. A. and McNeil, B. J. (1982) ‘The meaning and use of the area under a receiver operating characteristic (ROC) curve’, Radiology, 143(1), pp. 29–36.
  • Kadavath, S. et al. (2022) ‘Language models (mostly) know what they know’, arXiv:2207.05221.
  • Kambhatla, N., Mansouri Bigvand, A. and Sarkar, A. (2018) ‘Decipherment of substitution ciphers with neural language models’, Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pp. 869–874.
  • Kneser, R. and Ney, H. (1995) ‘Improved backing-off for M-gram language modeling’, Proceedings of ICASSP-95, vol. 1, pp. 181–184.
  • Lasry, G. (2018) A Methodology for the Cryptanalysis of Classical Ciphers with Search Metaheuristics. Kassel: Kassel University Press.
  • Murphy, A. H. (1973) ‘A new vector partition of the probability score’, Journal of Applied Meteorology, 12(4), pp. 595–600.
  • Naeini, M. P., Cooper, G. F. and Hauskrecht, M. (2015) ‘Obtaining well calibrated probabilities using Bayesian binning’, Proceedings of the AAAI Conference on Artificial Intelligence, 29(1), pp. 2901–2907.
  • Ostwald, O. and Weierud, F. (2017) ‘Modern breaking of Enigma ciphertexts’, Cryptologia, 41(5), pp. 395–421.
  • Pedregosa, F. et al. (2011) ‘Scikit-learn: Machine learning in Python’, Journal of Machine Learning Research, 12, pp. 2825–2830.
  • Platt, J. C. (1999) ‘Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods’, in Smola, A. J. et al. (eds) Advances in Large Margin Classifiers. Cambridge, MA: MIT Press, pp. 61–74.
  • Veritasium (n.d.) [Documentary on the breaking of Enigma, with Sir Dermot Turing and Jonah Weinbaum, filmed at Bletchley Park and The National Museum of Computing]. YouTube. Transcript consulted 27 September 2026.
  • Wang, P. et al. (2024) ‘Large language models are not fair evaluators’, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics.
  • Weierud, F. and Sullivan, G. (2005) ‘Breaking German Army ciphers’, Cryptologia, 29(3), pp. 193–232.
  • Weinbaum, J. (2025) Action This Day: The Mathematics and Machinations that Bested the German Enigma. MS thesis. Dartmouth College. Available at: https://digitalcommons.dartmouth.edu/masters_theses/247.
  • Welchman, G. (1982) The Hut Six Story: Breaking the Enigma Codes. New York: McGraw-Hill.
  • Zheng, L. et al. (2023) ‘Judging LLM-as-a-judge with MT-Bench and Chatbot Arena’, Advances in Neural Information Processing Systems, 36.