Research
Enigma, the Bombe, and a probabilistic analyst
Historical context, method of construction, and evaluation of the Enigma × Jev system
Abstract
The German Enigma was not broken by brute force. It was broken by structural observations about the machine, by exploiting how operators actually used it, and by pairing electromechanical search with human judgement about probable plaintext.
This paper reviews that history from the Polish Cipher Bureau (1932–39) through Bletchley Park and the naval four-rotor problem, and summarises the modern computational literature. It then documents how the present system rebuilds the codebreakers' pipeline in software: a verified machine simulator; a Turing–Welchman Bombe with the diagonal board; a Gillogly and Weierud–Sullivan ciphertext-only hill-climber; and German n-gram statistics. Here a language model with typed probabilistic outputs (Jev) makes the two judgements a Hut 6 analyst made: which crib to try, and whether a trial decryption is German.
We describe two algorithmic refinements that proved necessary, a ring-independent ranking of Bombe stops and a ring-aware reconstruction of message settings. Its test register reproduces, within their confidence intervals, the stop counts Weinbaum (2025) published for one- and two-loop menus. We report a tiered backtest on nine historical intercepts with published keys and on synthetic traffic, and state the threats to its validity.
A. Historical and technical context
This section draws in places on a documentary filmed at Bletchley Park and The National Museum of Computing, with Sir Dermot Turing and the running rebuild of the Bombe (Veritasium n.d.). Where the documentary simplifies, the text says so; its numbers are recomputed on this project's simulator.
- 1918machine
Scherbius patents a rotor cipher machine and sells it to banks and businesses.
- 1926–30machine
Navy (1926) and Army (1928) adopt Enigma with rewired rotors; the Army's Enigma I adds the plugboard (1930).
- 1931Poland
Hans-Thilo Schmidt sells documents to French intelligence, which passes them to Warsaw.
- 1932Poland
Rejewski recovers the rotor wiring with permutation theory.
- 1938Poland
Bomba and Zygalski sheets; in December rotors IV and V make 60 rotor orders.
- Jul 1939Poland
At Pyry the Poles hand methods and replicas to Britain and France.
- 1940Bletchley
Turing's Bombe (March); indicators no longer doubled (May); Herivel tip and cillies; the diagonal board (August).
- 1941naval
Captures from München and U-110; Banburismus reads naval traffic.
- 1942naval
U-boats move to the four-rotor M4 (February); Shark is read again from December.
- 1943–44naval
US Navy four-rotor Bombes take on Shark; Colossus attacks the Lorenz cipher, which was not Enigma.
- 1995–2017modern
Gillogly; Weierud and Sullivan; the M4 Project; Ostwald and Weierud.
A.1 The machine
Arthur Scherbius patented a rotor cipher machine in 1918 and sold it openly to banks and businesses; Poland and Britain, among others, simply bought one. A commercial Enigma was of little use against the military version, however. The German Navy adopted the machine in 1926 and the Army in 1928, each with rotors rewired in secret. In 1930 the Army introduced Enigma I with a plugboard (Steckerbrett), the component on which most of the machine's combinatorial strength rests (Kahn 1991; Bauer 2007). Without the new wirings no amount of ingenuity about settings could help, which is why Britain, which read unplugged commercial variants, had made almost no progress against the military machine by the summer of 1939 (Veritasium n.d.).
A key press sends current through the plugboard, the three rotors from right to left, a fixed reflector (Umkehrwalze), the rotors back again and the plugboard once more. At any rotor position the whole path is therefore a product P·S·P. Here P is the plugboard involution and S = R−1U R is a fixed-point-free involution. Two consequences followed, and both mattered to the codebreakers. First, the machine is self-reciprocal: if E enciphers to K, then at the same setting K enciphers to E, so the same setting enciphers and deciphers. Second, no letter is ever enciphered to itself. Hold down L and every lamp but L can light. That lets an analyst rule out a guessed plaintext at any alignment where a letter would coincide with its cipher letter, the “crib dragging” of Figure 6 and of the Break step on the machine page.
What set Enigma apart from the Caesar shift and every fixed substitution before it is that the substitution changes with every letter. In the video's demonstration, L lights Y, and L pressed again lights H. Behind the rotors sit three spring-loaded pawls, one per rotor, and all three push on every key press. Only the right rotor's pawl always engages. The others drop into a notch on the rotor to their right only when that rotor reaches its turnover letter, so the middle rotor advances once per 26 letters and the left rotor very rarely. When the middle rotor reaches its own notch, the left pawl drops in and pushes the left rotor and, through that notch, the middle rotor as well, so the middle rotor steps twice in succession (the double step) and the period is 26·25·26 = 16,900 rather than 26³.
Undo a small clip and the lettered ring turns against the wiring core; this is the ring setting (Ringstellung). If the core carries contact 1 to contact 4 (A to D), moving the ring by one makes the same wire carry B to E. The substitution depends only on the difference between window letter and ring, while the moment of turnover depends on the window letter alone. Many ring/position pairs are therefore cryptographically equivalent over a message, which is why the recovered key on the machine page can differ from the sender's letters yet decrypt identically.
A.2 The size of the key space, and why it was not the security
| Component | Choices | ≈ bits |
|---|---|---|
| Rotor order, 3 of 5 (Enigma I from Dec 1938) | 60 | 5.9 |
| Start positions, 26³ | 17,576 | 14.1 |
| Effective ring settings (middle and right), 26² | 676 | 9.4 |
| Plugboard, 10 pairs: 26! / (6!·10!·2¹⁰) | 150,738,274,937,250 | 47.1 |
| Enigma I, 3 of 5 rotors, 10 plugs (7–10 pairs from 1 January 1939; ten in wartime) | 1.07 × 10²³ | 76.5 |
| Enigma I, 3 of 5 rotors, 6 plugs (December 1938) | 7.16 × 10¹⁹ | 66.0 |
| Enigma I, 3 rotors, 6 plugs (the 1930 configuration) | 7.16 × 10¹⁸ | 62.6 |
| Commercial Enigma: rotor order and positions only | 105,456 | 16.7 |
| Naval M3, 3 of 8 rotors (336 orders; see below) | 6.02 × 10²³ | 79.0 |
| Naval M4, + greek wheel (2 × 26) and 2 thin reflectors | 6.26 × 10²⁵ | 85.7 |
An earlier version of this table gave the 1930 configuration as 7.16 × 10¹⁹. That figure silently assumed the five-rotor box of 1938. The 1930 machine had three rotors and six orders, so the correct figure is 7.16 × 10¹⁸, agreeing with the “over seven times ten to the eighteen” of the popular account (Veritasium n.d.). Two further rotors (×10) and four more plug pairs (×1,501.5) then multiplied it by 15,015 by the outbreak of war.
The plugboard contributes most of the bits, but it barely hindered the attacks that succeeded. The plugboard acts on the scrambler by conjugation, so it leaves invariant any property that depends only on cycle structure. Rejewski exploited exactly that (Figure 4), and a Bombe tests plugboard hypotheses by implication, not enumeration (Figure 7). Neither approach pays for the 10¹⁴ plugboards. The search that mattered to a Bombe was the 60 × 17,576 = 1,054,560 rotor settings, and for the naval M3 5,905,536. The number of plugboards is largest at eleven pairs; the Germans used ten.
Weinbaum (2025), following Hosgood, counts only 276 naval wheel orders, on a procedural rule that one of the three wheels must be a naval wheel (VI–VIII). The rule did not always hold. The three M4 Project messages of November 1942 in this project's ground truth use wheels II, IV and I. The Bombe and the backtest therefore search all 336.
Shannon's unicity distance U = H(K)/D gives the ciphertext length beyond which the key is, in principle, uniquely determined (Shannon 1949). Take H(K) ≈ 76.5 bits, and a redundancy of German of D ≈ 3.2–3.7 bits per letter (log₂26 ≈ 4.70 minus an entropy rate of roughly 1.0–1.5 bits; Shannon 1951). That gives U ≈ 21–24 letters. Almost every intercept therefore contained enough information to determine its key. The obstacle was always computational, and the history below is the history of closing that gap. It is also why the machine page will not claim to have broken a message much shorter than this.
A.3 Poland, 1932–1939: permutation theory
On 8 November 1931 Hans-Thilo Schmidt (“Asché”), an employee of the German Army's cipher office, let French intelligence photograph the Enigma operating instructions and keying procedure for 10,000 marks. The French could not use them and passed them to their Polish allies (Kozaczuk 1984). Through Schmidt the head of the Polish Cipher Bureau, Gwido Langer, eventually held 38 months of daily keys. Expecting that source to vanish in a war, he withheld all but two months of them from his own cryptanalysts, so that they would have to learn to recover keys without them (Weinbaum 2025). The Polish Cipher Bureau (Biuro Szyfrów) had recruited mathematics students from Poznań University, among them Marian Rejewski, Jerzy Różycki and Henryk Zygalski.
The weakness Rejewski attacked lay in the procedure, not the machine. Before each message the operator chose a three-letter message key, say GEX. He enciphered it twice at the day's common setting, a precaution against radio garbles, and sent the six cipher letters first. Letters 1 and 4 therefore encipher the same unknown letter, as do 2 and 5, and 3 and 6. Let A, …, F be the enciphering permutations at the six positions. Then the products AD, BE and CF can be read off a day's traffic once about eighty indicators are collected. Their cycle structure (the day's “characteristic”) is independent of the plugboard, because conjugate permutations share a cycle type (Rejewski 1980). With Schmidt's keys, Rejewski set up and solved systems of permutation equations and recovered the rotor wirings by the end of 1932, without ever having seen a military machine (Rejewski 1981).
Tools followed. The grill method recovered a day's right rotor by sliding precomputed permutations under the day's six, and worked while only six letters were plugged. The cyclometer produced the catalogue of characteristics, 6 × 26³ = 105,456 entries, and cut the recovery of a day's setting to about fifteen minutes (Rejewski 1981). On 15 September 1938 the operator began choosing and sending his own indicator setting, which made the catalogue useless. Two attacks then used females, the same letter at positions 1 and 4, 2 and 5, or 3 and 6 of an indicator.
- The bomba. Each bomba (1938) held three pairs of Enigma equivalents set three steps apart, and stopped where all three pairs produced their female. Six were built, one per rotor order. They relied on the repeated letter being unplugged, a fair bet with five to eight plugs but only 6 in 26 once ten were used, and a day's key took up to about two hours (Kozaczuk 1984).
- Zygalski sheets. A setting can produce a 1–4 female exactly when AD has a fixed point, a property the plugboard cannot change. So a perforated sheet per left-rotor position and order recorded, hole by hole, where a female was possible, and stacking the sheets for a day's females left light through only the consistent ring settings. Weinbaum (2025) puts the chance that a random setting can produce a female at 0.405. Measured on this simulator over all 60 orders and 17,576 positions, it is 0.406, from 0.375 to 0.433 across orders (Table B1). On his own formula about seven sheets leave a single hole among the 676 on a sheet, not the twelve or thirteen his text states. Bletchley cut the full set for 60 orders and sent it to Warsaw (Rejewski 1981).
Meanwhile the Germans turned the screw. Rotor orders had changed quarterly before 1936, then monthly, daily, and eventually every eight hours. From 1 January 1939 seven to ten plugs were used, and in December 1938 two further rotors multiplied the rotor orders tenfold, from 6 to 60, beyond Polish resources. In late July 1939, at Pyry near Warsaw, the Poles gave their methods and replica machines to French and British cryptanalysts. It was one of the decisive transfers of the war (Hinsley and Stripp 1993; Sebag-Montefiore 2000). The name bomba is disputed: a noisy machine, an ice-cream dessert eaten while it was conceived, or, in Rejewski's account, nothing at all (Weinbaum 2025). Five weeks later Germany invaded Poland; the cryptanalysts escaped, the bomby were destroyed, and the work passed to Britain.
A.4 The operator as the weak link: indicators, cillies and the Herivel tip
From September 1938 each operator chose the starting position for his indicator himself and sent it in clear. On 1 May 1940 the Army and Air Force stopped doubling the message key, which removed Rejewski's and Zygalski's footholds. The procedure then became the two-stage exchange of Figure 5: a clear Grundstellung (VER), and a message key (ITA) enciphered once from it. A receiver undoes the two checkpoints in turn. The procedure was sound. Its operators were not, because people are poor sources of randomness (Welchman 1982; Veritasium n.d.).
- Cillies. The habit was older than Bletchley. The Poles had exploited stations that reused keys such as AAA or keyboard runs such as ASD and PYX, some of which were forbidden from 1933 (Kozaczuk 1984). Operators chose keys they could type without thinking: three letters of a girlfriend's name (a “Cillie” gave the habit its name), keyboard runs such as QWE or diagonals down the QWERTZU layout, or the continuation of the clear setting. Clear BER invites the key LIN. An analyst who guesses the key knows three plaintext letters at a known setting: a three-letter crib.
- The Herivel tip. John Herivel reasoned that an operator setting the day's rings with the rotors already in the machine would leave the ring letters showing in the windows, and would choose his first clear setting of the day within a letter or two of them. Plotted across the network, the first clear settings of the morning therefore cluster around the day's ring settings. That removed the 676 ring choices that made the Bombe's problem hard. The tip was noted in February 1940 and produced breaks from May 1940, in time for the fall of France, chiefly on the Luftwaffe's general key, the “Red” network (Welchman 1982).
- Parkerismus. Reg Parker noticed that the officials who compiled key sheets sometimes reused whole columns from earlier months. Once one setting was recognised in a filed key, the rest of its column could be predicted; Welchman recalled a month of the keys Rommel's army would use in Africa worked out this way (Welchman 1982).
A.5 Bletchley Park: the crib, the menu and the diagonal board
The Government Code and Cypher School recruited from Oxford and Cambridge, from the chess and crossword worlds, and in its thousands from the Women's Royal Naval Service, whose members ran the Bombes. Perhaps 150 people worked there in 1939. At the peak there were some 9,000, bound by rules so strict that staff were told not to talk even by their own fireside. Army and Air Force Enigma went to Hut 6 under Gordon Welchman and naval Enigma to Hut 8 under Alan Turing. Their cribbing rooms kept files of what each station habitually said.
Turing's Bombe did not depend on indicator repetition, which the Germans could abolish and did. It depended on the crib, a guess at plaintext and its position (Turing c.1940; Copeland 2004). A weather ship in the Bay of Biscay sending at six every morning will, sooner or later, spell out WETTERVORHERSAGEBISKAYA. Because no letter enciphers to itself, the analyst slides that guess along the intercept and discards every alignment at which a letter lands on itself (Figure 6). The example is historical. On D-Day a Biscay weather intercept of 28 letters admitted the crib at only one of its six alignments (Weinbaum 2025), and Figure 6 can switch to that intercept.
Each surviving alignment pairs crib and cipher letters through the unknown scrambler at that position. The pairs form a graph, the menu, and closed loops in the menu are consistency conditions that no plugboard can hide. Write the machine at position i as P Si P. If the loop runs x₁ → x₂ → … → x₁ through positions i₁, …, ik, then each plugboard on the way out meets the next on the way in and cancels, leaving
P(x₁) = Sik ⋯ Si₂ Si₁ P(x₁):
the plugboard partner of x₁ is a fixed point of a product of unplugged scramblers. Turing turned that into a circuit. Assume a partner for one menu letter, the test register, and set that wire live. Current flows around every loop and sets live each wire the assumption implies. If all 26 wires light, every possible partner has been contradicted and the rotor position is impossible. If 25 light, the dark wire is the only partner left. Either way one guess, right or wrong, disposes of a whole rotor position at the speed of electricity (Turing c.1940). A position that is not refuted is a stop.
Turing's design needed rich menus, for a reason Weinbaum (2025) makes exact. Each scrambler is a product of 13 transpositions, an odd permutation, so a loop of even length composes to an even permutation, and a 26-cycle is odd. A single even loop therefore never lights all 26 wires: without the diagonal board it stops at every one of the 17,576 positions. An odd loop is refuted at only about 1 position in 13. Turing estimated 264−c stops for a menu with c closures, and before the diagonal board a menu typically needed three. Weinbaum's simulations put three closures at about 30 stops, against Turing's 26, and one closure at about 17,000 against 17,576. This project's register reproduces his one- and two-loop figures (B.5, Table B1). Welchman's diagonal board (1940) added the fact that a plugboard is symmetric: if A is plugged to B then B is plugged to A. That multiplied the implications of every hypothesis and made short or loop-poor menus usable (Welchman 1982). Figure 7 reproduces the effect quantitatively. At a stop the Bombe offers an answer, not the answer, and each stop had to be tried on an Enigma analogue. Checking took of the order of a quarter of an hour, so a hundred stops were more than a day's work, longer than the key would live. By cutting stops to a handful the diagonal board turned a demonstration into a production line (Veritasium n.d.). Two refinements cut them further. Air Force key sheets never plugged a letter to its alphabetical neighbour, and a consecutive stecker knock-out switch let the Bombe discard any stop that implied one. On the “Jumbo” Bombes a machine gun of uniselectors swept the diagonal board at each stop for two deductions that shared a letter (Weinbaum 2025).
The first Bombe, Victory, arrived in March 1940, six months into the war, engineered by Harold “Doc” Keen at the British Tabulating Machine Company in Letchworth. Agnus, with the diagonal board, followed in August. A Bombe carried 36 Enigma equivalents in three banks of twelve, and its indicator drums counted backwards from ZZZ so that, at a stop, they showed the ring setting directly. It ran through the 17,576 positions of one wheel order in about a quarter of an hour, twenty minutes in the best case (Weinbaum 2025), so a full 60-order run was a matter of Bombe-hours and wheel orders were rationed by rule and intuition. About two hundred were built by the end of the war. This is precisely the stop → verify → accept pipeline the machine page animates, with Jev in the role of the checker.
Cribs came from routine and from operator habit. Weather reports, keine besonderen Ereignisse (“nothing to report”), standard addresses and signatures, and messages re-sent in a broken key (“kisses”) all supplied them. So did minelaying to provoke predictable warnings (“gardening”), and the cillies and Herivel tips of A.4. Procedure helped too. A continuation began FORT and repeated the time of the first part in top-row letters (Q = 1, W = 2, … P = 0), so a follow-up to a message of 23:30 opened FORTYWEEPYWEEPY (Weinbaum 2025). The crib list used by this system (src/jev/cribs.ts) is drawn from these categories.
A.6 Naval Enigma, Banburismus, and the weight of evidence
The Kriegsmarine chose three rotors from eight, including three with two notches, and enciphered indicators through bigram tables. It resisted until material was captured. These were the “pinches”: the trawler Polares off Narvik on 26 April 1940; the armed trawler Krebs on 4 March 1941; the weather ship München on 7 May; U-110 on 9 May, with a machine and bigram tables; and the weather ship Lauenburg on 28 June (Kahn 1991; Weinbaum 2025). Turing's Banburismus set messages side by side in depth and scored their coincidences sequentially in bans and decibans, logarithmic units of odds. It narrowed the wheel orders a Bombe had to try. Two naval messages whose message keys differ only in the last letter are in depth at one unknown offset. There, a letter repeats with probability about 1/17, against 1/26 at random, and each observed repeat count was scored from printed sheets in half-decibans, called hubdubs. Six repeats in 52 letters, for example, is worth about 13.6 hubdubs, odds of roughly five to one for depth. A bigram counted for about three single repeats. A typical day meant some 400 messages and 6,000 comparisons, and Joan Clarke was among its best practitioners (Weinbaum 2025). Nearly nine naval messages in ten contained EINS, so an EINS catalogue then fixed the rings. It is one of the first sustained applications of Bayesian sequential inference, later documented by his assistant I. J. Good (Good 1979).
On 1 February 1942 the U-boat network moved to the four-rotor M4 (“Shark”). Its extra, non-stepping greek wheel occupied the space of a thinner reflector, so that with the greek wheel at A an M4 behaved exactly like a three-rotor machine and could still talk to one. That backward compatibility is one of this project's own regression tests (B.2), and it limited how far the Germans could change the machine without replacing it. As one historian in the video puts it, Turing's attack was so fundamental that only a new machine, or no traffic at all, would have defeated it. Shark was unreadable for most of 1942. Then in October 1942 Lieutenant Anthony Fasson and Able Seaman Colin Grazier boarded the sinking U-559 and passed out the short-signal and weather code books before drowning with the boat. With them Hut 8 read Shark again from December. Four-rotor Bombes built for the US Navy by the National Cash Register Company in Dayton, more than a hundred of them, carried much of the load from 1943 (Budiansky 2000; Erskine and Smith 2011).
Two threads from this history shape the present design. Evidence should be expressed as calibrated probability, and it should be scored with a proper scoring rule: Jev's judgements are probabilities, evaluated here with the Brier score. Brier devised it in 1950 for weather forecasts, fittingly the richest source of Enigma cribs (Brier 1950).
A.7 What it changed, and the machine that was not Enigma
By the end of 1941 Army and Air Force Enigma was read routinely. Decrypts showed Rommel's army critically short of fuel before the Second Battle of El Alamein (1942). When Shark was mastered in 1943, convoy losses fell and U-boat losses rose; Dermot Turing, interviewed in the video, calls it the turning point after which Germany began to lose the war, while stressing that battles are won by the people at the sharp end. F. H. Hinsley, the official historian of British intelligence and a Bletchley veteran, judged that Ultra shortened the war by perhaps two years (Hinsley and Stripp 1993). Such counterfactuals are contested, and the claimed influence on the Battle of Britain in particular is weaker than that on North Africa and the Atlantic.
The German high command's most secret teleprinter traffic did not go by Enigma at all but by the Lorenz SZ40/42 (“Tunny”). Bill Tutte deduced its structure in 1942 from a single depth, without ever seeing the machine. To read it in time, Tommy Flowers built Colossus, an electronic, programmable machine, in service from early 1944. Colossus decrypts helped confirm, before D-Day, that the deception about the landing site was believed (Copeland 2006). Colossus remained secret for decades after the Bombe was acknowledged.
Turing's life after the war is often read backwards from its end: his prosecution for homosexuality in 1952 and his death in 1954. His nephew, Sir Dermot Turing, argues against making him a martyr. From 1945 to 1950 Turing designed the ACE, programmed the Manchester computer and wrote on machine intelligence, often impatient that the machines were not built fast enough (Turing 2015). This project, in which a machine judges whether another machine's output is language, is a small descendant of that 1950 question.
A.8 Modern computational cryptanalysis
Friedman's index of coincidence measures the probability that two letters drawn from a text coincide (Friedman 1922). For German it is about 0.076, against 1/26 ≈ 0.0385 for random text, a normalised ratio near 2. The held-out German used in this project measured 1.98. Naval traffic, full of numbers and abbreviations, was leaner: Bletchley scored it at a repeat rate of 1/17, a normalised 1.53 (Weinbaum 2025). That is one reason the naval messages in this study read as the least German (C.4). Gillogly showed that Enigma can be attacked from ciphertext alone. He scored every rotor order and position by the index of coincidence of the trial decryption, then recovered the plugboard by trigram hill-climbing, which is effective for long messages (Gillogly 1995). Weierud and Sullivan refined hill-climbing and read several hundred German Army messages of 1941 without cribs (Weierud and Sullivan 2005). In 2006 Stefan Krah's distributed M4 Project broke naval M4 intercepts from November 1942 that had never been read. Ostwald and Weierud consolidated modern methods and their limits for short messages (Ostwald and Weierud 2017). The present system reuses several of these messages as ground truth. Some wartime intercepts remain unread today; they are short, naval, and without cribs, the very conditions under which this system also fails (C.4).
B. How this system was built
B.1 Division of labour
Jev is served through TypeSafe's System One endpoint. It is given a textual state and a set of typed questions. It returns a probability (noul), a distribution over named options (choice), or a position on an ordered scale (score), and never prose. Distributions are validated to cover every option and to sum to one within 10⁻⁵, and the model revision is pinned (jev-1.13.0). Jev therefore cannot emit a key or a plaintext. All cryptanalysis is algorithmic and locally verifiable. Jev is asked only the two questions that required an analyst's judgement in Hut 6:
- Prior: which crib? A
choiceover the listed phrases that survive crib dragging at the start of the message, plus “none”. The state carries the traffic type, date, length and the surviving and eliminated cribs. - Posterior: is it German? A
choiceamong up to three distinct candidate decryptions, plus “none”, and anoulP(correct) for each candidate. The state withholds the search's own n-gram scores and marks the letters that were assumed as a crib, so the judgement is independent of the search and is not flattered by the crib.
- 1Interceptciphertext and traffic type
- 2Crib draggingrule out phrases where a letter would meet itself
- 3Jev: which crib?
choiceover the survivors - 4Bombe60 rotor orders × 17,576 positions, 8 workers
- 5Rank stopstrigram score over all 26 right-ring choices at once
- 6Plugboard climbmenu pairs locked, then ring refinement
- 7Jev: is it German?
choice+ P(correct); accept at ≥ 0.5 on both
Not accepted → the next crib in Jev's order (step 4); then every crib again with the middle rotor stepping inside it; last, the ciphertext-only climb.
B.2 Machine simulator and its verification
src/enigma/ implements rotors I–VIII, the Beta and Gamma greek wheels, reflectors A, B, C, B-thin and C-thin, rings, plugboard and the double step. It is verified against:
- the textbook vector (I-II-III, UKW B, rings AAA, start AAA:
AAAAA → BDZGO); - tests of involution and no self-encipherment;
- an explicit double-step sequence (ADU → ADV → AEW → BFX);
- the identity between an M4 with Beta at A and B-thin, and an M3 with UKW B;
- the exact reproduction of all ten historical plaintexts from their published keys.
B.3 German language model
An interpolated trigram model (src/lang/) is trained on about 13,000 letters of German composed for the project: military reports and general prose. Each paragraph is rendered in four operator conventions (words run together; X between words; X with Q for CH; J between words), giving 53,728 training letters. The conditional probability is
P(c | ab) = w P₃(c | ab) + (1 − w) [0.85 P₂(c | b) + 0.15 P₁(c)], w = nab / (nab + 4),
so an unseen context backs off smoothly. A normalised German-ness g = (s − srand) / (sGerman − srand) maps mean log-probability to 0 for uniform random letters and 1 for typical German. Held-out German scores 0.78–0.93; random letters score about 0.
B.4 Ciphertext-only search, and why it is not enough
src/break/climb.ts follows Gillogly and Weierud–Sullivan. It scans 60 × 17,576 settings by index of coincidence, finds ring turnovers for the survivors, then hill-climbs the plugboard, first on coincidence and then on trigrams, using four move types. We measured why this fails on wartime keys. On a 532-letter message with ten plug pairs, the true rotor setting decrypted without the plugboard has a normalised coincidence of 1.04, against 1.08 for the best of the random settings: the signal is buried. With six pairs and the correct rings it rises to 1.27, and the method succeeds, recovering such a message completely in about eleven seconds. Ten-pair traffic needs cribs, as it did in 1940.
B.5 A software Turing–Welchman Bombe
src/break/bombe.ts implements the constraint P(Cj) = Sj(P(Pj)) for each crib position j, where Pj is the crib letter, Cj the cipher letter, P the plugboard and Sj the scrambler. For each rotor setting it assumes a partner for the busiest menu letter and propagates depth-first. Every implied pair is also imposed in reverse, as with the diagonal board, and a contradiction rejects the hypothesis. A consistent hypothesis is a stop and carries the plugboard pairs it implies. The test is stricter than the electrical Bombe's. That machine held one hypothesis and stopped whenever a wire stayed dark, Turing's “spider”. The software tries all 26 partners and rejects a partner at the first contradiction of any kind, including a letter implied to have two partners, the check the machine gun later automated. Its stop counts are therefore consistent hypotheses, Turing's “normal stops” summed over hypotheses and turnover variants, not positions. Four refinements proved necessary. Each was found by a failure in testing and is covered by a regression test.
- Turnover inside the crib. Rings are unknown, so the middle rotor may step within the crib. A first pass assumes it does not. A second pass lets it step before crib letter t for every t. For two-notch rotors (VI–VIII) steps recur every P = 13 letters, and a variant is admissible only if exactly one step falls inside the crib (t − P < 1 and t + P ≥ n). This is why the web page cuts cribs to 14 letters: a 25-letter crib almost surely contains a step and needs 24 variants.
- Ring-independent ranking of stops. A loop-poor menu yields thousands of stops, which must be ranked by how German their full decryption reads. That decryption depends on the unknown right ring. The ring enters only through the residue d (mod P) at which the middle rotor steps. For a letter at crib-relative index u = qP + r ≥ 0, the middle offset is sm + q + [1 ≤ d ≤ r], and symmetrically for u < 0. Each letter is decrypted under both offsets, and its trigram gain Δ is filed under its threshold R. Then score(d) = T₀ + ΣR ≥ d ΔR scores all 26 rings in O(N + P). Before this change the true stop of the 1930 manual message ranked 5,258th of 6,957.
- Ring-aware reconstruction. Turning a stop back into message-start settings must also search the middle ring. Taking it as A made the 1930 message's middle rotor meet its own notch and double-step where the true key does not, so no reconstruction existed.
- Completion and acceptance. The plugboard is completed by hill-climbing with the menu's pairs locked, followed by a ring refinement. Acceptance scores only the letters outside the crib: the crib is German by construction and would otherwise flatter a wrong key.
Validation of the test register. Figure 7 and the tests in test/register.test.ts use the spider logic itself (src/break/register.ts). To check it independently, src/analysis/bombe-stops.ts builds random menus of one or two closed loops of given lengths. It runs each over all 17,576 positions of a random Enigma I wheel order, without the diagonal board, and compares mean stops with the figures Weinbaum (2025) published from his and Bouchaudy's simulations. It also measures the share of settings that can produce a female (A.3).
Loading the stop-count validation…
B.6 Jev's decision rule
In the first backtest, Jev's distribution over all-garbage candidate sets was flat, and its most probable option often fell on a garbage candidate at low confidence. A candidate is therefore accepted only if Jev's pick has P(pick) ≥ 0.5 and P(correct) ≥ 0.5: more likely than not on both questions. This mirrors Hut 6's rule of confirming a stop before committing to it. The threshold was chosen after inspecting that first run, a limitation discussed in C.4.
B.7 The web system
A Bun server (src/web/server.ts) keeps the TypeSafe key server-side and streams the break to the page as server-sent events. Each Bombe or ciphertext-only run is split by wheel order across a pool of eight worker threads. The crib plan runs your own crib first, then Jev's ranking, each cut to 14 letters. Every crib gets a no-turnover pass, then an every-turnover pass. Jev reads each run's best candidates, and the search stops at the first acceptance. The ciphertext-only climb is the last resort. The plugboard climb of the accepted stop is replayed step by step, so the text is seen to resolve.
C. Evaluation
C.1 Data
Historical. Ten intercepts with published keys and plaintexts were collected, with sources and exclusions recorded in data/historical/messages.json:
- the 1930 instruction-manual example (UKW A);
- both parts of an Army message of 7 July 1941 from the Barbarossa campaign;
- Scharnhorst, December 1943 (M3);
- the three M4 Project messages of November 1942 (Looks, Schroeder, Rasch);
- a U-534 message of 1 May 1945 (M4, C-thin);
- a short Army message, excluded from scoring because no plaintext has been published;
- a textbook example, kept as a control.
Every entry decrypts exactly on this simulator. Synthetic. Sixteen German messages (67–493 letters) were written for the project and held out of the language model. Each is enciphered under a seeded random key: Enigma I for Heer and Luftwaffe traffic, M3 with rotors I–VIII for the Kriegsmarine, ten plug pairs, and six pairs for every fourth message.
C.2 Protocol
Every message passes through tiers of decreasing prior knowledge:
- verify: the full key is known;
- key: the daily key is known and the message key is sought;
- crib: the true first 14 letters are known plaintext, measuring the Bombe separately from crib choice;
- bombe: cribs are chosen in Jev's order from the fixed list;
- climb: ciphertext only.
For M4 messages both Bombe tiers are given the wheel order and greek wheel; a full four-rotor search is beyond one machine, as it was for Bletchley without the US Navy's Bombes. A tier breaks a message when its best candidate matches at least 90% of the published plaintext. Jev's verdict is right when it accepts such a candidate, or accepts none when none was shown. The n-gram baseline applies the same rule with German-ness ≥ 0.4. Jev's per-candidate probabilities are scored by the Brier score.
C.3 Results
Loading the latest backtest report…
C.4 Discussion and threats to validity
Findings.
- Machine and message keys. The simulator and the message-key search are exact on every case.
- The Bombe on real traffic. Given a 14-letter known crib, it recovers most historical and synthetic messages in seconds to minutes, including the M3 and M4 messages where the wheel order is given.
- Ciphertext only. The climb fails on ten-plug traffic of these lengths, consistent with the literature and with the measurement in B.4.
- Jev as judge. Under the decision rule, in the main run Jev never rejected a correct decryption. It accepted exactly one candidate below the 90% bar: the Scharnhorst decryption from the known-crib tier, with 89% of its letters right. Its raw top pick would have landed on twelve candidates below the bar. On the holdout run, with fresh keys, the rule was right in all 64 judged cases, against 57 for the raw top pick.
- Jev against the standard judges. On the same candidates, the Jev paper (§4.8) runs the codebreakers' n-gram fitness scores, a Kneser–Ney 5-gram model, logistic regression and XGBoost (Chen and Guestrin 2016), all cross-validated leave-one-text-out. Trained on the same kind of traffic, quadgram fitness and XGBoost match recalibrated Jev. Trained on the synthetic messages and tested on the real intercepts, XGBoost rejected 15 of 16 correct decryptions, while zero-shot Jev decided 31 of 32 calls correctly. The zero-shot judge is the one that transfers.
- Jev as crib selector. Its ranking beat the static list on the historical openers (mean rank 4.0 against 6.5, n = 4). It did not on the synthetic set (6.6 against 6.4, n = 5) or on the holdout (8.6 against 6.8, n = 5). With samples this small the evidence does not support an advantage. Asked only for the traffic type, date and length, Jev's prior over openers is weakly informative. That is the component most in need of richer context, such as the time of transmission, the call signs and the preceding traffic, which is what Hut 6's crib-writers actually used.
Threats to validity.
- Crib list. It was fixed before any backtest, but by an author who knew several historical openers. Its coverage of historical messages may be optimistic.
- Decision rule. Jev's acceptance threshold was set after inspecting the first run. The holdout run supports it out of sample in one respect only: it draws new keys but reuses the sixteen synthetic plaintexts, so it is not an independent test set for text.
- One author. The same author wrote the training corpus and the synthetic messages. Stylistic homogeneity may overstate the language model's fit relative to genuine traffic, whose abbreviations and garbles lowered German-ness to 0.44–0.54 for naval messages.
- Small historical set. Nine scored messages, several from one day or one source, give wide uncertainty on every rate.
- Two crib strategies. The
bombebacktest tier tries Jev's top three cribs at full length. The web page cuts cribs to 14 letters and tries every fitting crib with both passes, so the tier understates what the page achieves. - Dependence on the service. Jev's outputs are a model's judgements under a pinned revision. They may shift if the revision changes, and are recorded verbatim in
~/.enigma-jev/jev-calls.jsonlfor audit.
Reproducibility. bun test runs the unit and regression suite. bun run backtest all reproduces the main run with seed 1941, and … backtest synthetic --seed 2024 the holdout. Reports are written to reports/, and the table above is read from the newest of them.
D. References
- Bauer, F. L. (2007) Decrypted Secrets: Methods and Maxims of Cryptology. 4th edn. Berlin: Springer.
- Brier, G. W. (1950) ‘Verification of forecasts expressed in terms of probability’, Monthly Weather Review, 78(1), pp. 1–3.
- Budiansky, S. (2000) Battle of Wits: The Complete Story of Codebreaking in World War II. New York: Free Press.
- Chen, T. and Guestrin, C. (2016) ‘XGBoost: A scalable tree boosting system’, Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 785–794.
- Copeland, B. J. (ed.) (2004) The Essential Turing. Oxford: Oxford University Press.
- Copeland, B. J. (ed.) (2006) Colossus: The Secrets of Bletchley Park's Codebreaking Computers. Oxford: Oxford University Press.
- Erskine, R. and Smith, M. (eds) (2011) The Bletchley Park Codebreakers. London: Biteback.
- Friedman, W. F. (1922) The Index of Coincidence and Its Applications in Cryptography. Riverbank Publication No. 22. Geneva, IL: Riverbank Laboratories.
- Gillogly, J. J. (1995) ‘Ciphertext-only cryptanalysis of Enigma’, Cryptologia, 19(4), pp. 405–413.
- Good, I. J. (1979) ‘Studies in the history of probability and statistics XXXVII: A. M. Turing's statistical work in World War II’, Biometrika, 66(2), pp. 393–396.
- Hinsley, F. H. and Stripp, A. (eds) (1993) Codebreakers: The Inside Story of Bletchley Park. Oxford: Oxford University Press.
- Hodges, A. (1983) Alan Turing: The Enigma. London: Burnett Books.
- Kahn, D. (1991) Seizing the Enigma: The Race to Break the German U-Boat Codes, 1939–1943. Boston: Houghton Mifflin.
- Kozaczuk, W. (1984) Enigma: How the German Machine Cipher Was Broken, and How It Was Read by the Allies in World War Two. Ed. and trans. C. Kasparek. Frederick, MD: University Publications of America.
- Ostwald, O. and Weierud, F. (2017) ‘Modern breaking of Enigma ciphertexts’, Cryptologia, 41(5), pp. 395–421.
- Rejewski, M. (1980) ‘An application of the theory of permutations in breaking the Enigma cipher’, Applicationes Mathematicae, 16(4), pp. 543–559.
- Rejewski, M. (1981) ‘How Polish mathematicians deciphered the Enigma’, Annals of the History of Computing, 3(3), pp. 213–234.
- Sebag-Montefiore, H. (2000) Enigma: The Battle for the Code. London: Weidenfeld & Nicolson.
- Shannon, C. E. (1949) ‘Communication theory of secrecy systems’, Bell System Technical Journal, 28(4), pp. 656–715.
- Shannon, C. E. (1951) ‘Prediction and entropy of printed English’, Bell System Technical Journal, 30(1), pp. 50–64.
- Turing, A. M. (c.1940) Treatise on the Enigma (“Prof's Book”). Kew: The National Archives, HW 25/3.
- Turing, D. (2015) Prof: Alan Turing Decoded. Stroud: The History Press.
- Veritasium (n.d.) [Documentary on the breaking of Enigma, with Sir Dermot Turing and Jonah Weinbaum, filmed at Bletchley Park and The National Museum of Computing]. YouTube. Transcript consulted 27 September 2026.
- Weierud, F. and Sullivan, G. (2005) ‘Breaking German Army ciphers’, Cryptologia, 29(3), pp. 193–232.
- Weinbaum, J. (2025) Action This Day: The Mathematics and Machinations that Bested the German Enigma. MS thesis. Dartmouth College. Available at: https://digitalcommons.dartmouth.edu/masters_theses/247.
- Welchman, G. (1982) The Hut Six Story: Breaking the Enigma Codes. New York: McGraw-Hill.
Message sources: F. Weierud, CryptoCellar (cryptocellar.org); M4 Project pages, bytereef.org; Crypto Museum (cryptomuseum.com); Enigma-Hörenberg (enigma.hoerenberg.com); Franklin Heath Enigma wiki; Enigma (Maschine), German Wikipedia (textbook example).