Beinecke MS 408 · Cryptanalysis · September 2026

Deutsch

Trying to crack the Voynich Manuscript — with controls

I threw modern code-breaking at the world's most mysterious book: cipher attacks in eight languages, a published medieval-style cipher, self-citation tests, image cross-references. Every test ran against a text I had encrypted myself, and against a scrambled copy of the Voynich. Here is what survived.

The short version

  1. No decipherment. Substitution ciphers (letter-for-letter, word-level “Naibbe”-type, glyph-group, with memory, with null glyphs removed) failed in Latin, Italian, German, Middle High German, Ancient Greek, Old Occitan, Old French and Catalan — while the same attacks cracked my own encrypted control texts with 94–100 % accuracy.
  2. “It reads like Latin” is a trap. Scrambled Voynich produces equally Latin-looking output. Every attack was scored against that baseline.
  3. Voynich words copy their neighbours. A word is 1.36–1.38× more likely to be a one-glyph variant of a word in the line above than of a random line. Ordinary prose, plain or enciphered: 1.03–1.12×. Generated text and magic-book name lists: 1.58–1.68×.
  4. Position-dependent ciphers are ruled out. Vigenère, progressive, running-key (incl. Bible text) and per-page keys all flatten the text; Voynich is more repetitive and predictable than plain Latin.
  5. The pictures cross-reference, the text doesn't. Plants in the pharmaceutical section reappear in the herbal section — but their labels do not reappear on the matching herbal pages.
  6. No source text, no music, no rhyme, nothing through the page. The bathing section is not an encrypted De balneis Puteolanis, the recipe section not the Antidotarium Nicolai — tests that found the true source in 100 % of simulations. Read as notation, the glyphs don't move like 14th-century melodies, the lines don't rhyme, and text lying on top of each other on the two sides of a leaf is unrelated (try the see-through viewer).

1 · The manuscript in 60 seconds

The Voynich Manuscript (Yale, Beinecke MS 408) is a small illustrated codex of about 240 surviving pages written in an unknown script. Its parchment was radiocarbon-dated to 1404–1438. Sections show plants, astronomical and zodiac diagrams, bathing women connected by tubes, pharmaceutical jars with plant parts, and finally pages of short star-marked paragraphs. Researchers transcribe the script with the EVA alphabet (e.g. daiin, qokeedy, chedy). Prescott Currier noticed in the 1970s that the text comes in two statistical “dialects”, Currier A and B; I analyse them separately.

Words analysed (paragraph text)
33,088
B: 21,610 · A: 11,078 · unassigned: 400 (Takahashi EVA)
Most frequent word
daiin
863 times
Distinct word forms
8,021
in 37,840 tokens overall
Folio 9v of the Voynich Manuscript: a plant with blue and yellow five-petalled flowers resembling a wild pansy; the text lines stop at the flowers
f9v. The flowers resemble a wild pansy. The text lines stop where the blossoms begin — the drawing was made first and the text fitted around it, the reverse of normal medieval practice.

2 · How the tests work

The field is full of “decipherments” that read nicely and mean nothing. So every test here has two controls:

Positive controlA real text (e.g. Pliny's Latin) that I encrypt with the same kind of cipher. The attack must recover it — otherwise the attack is useless.
The Voynich textCurrier A and B separately, from the Takahashi EVA transcription.
Negative controlThe same Voynich text with its words shuffled. It contains no message. If Voynich scores no better than this, nothing was found.

The attack itself is a homophonic-substitution solver (simulated annealing with restarts) that searches for the key under which the text looks most like a given language, scored by a 5-gram character model of that language, trained on 0.08–8.6 million characters per language. The score is the average log-probability per character: higher is more language-like. A real crack shows up as a jump of about +0.5 or more over the shuffled baseline (that is the gap the positive controls showed) and as readable text.

The starting point: the Naibbe cipher

In 2025, Michael Greshko published the Naibbe cipher: a hand-executable verbose homophonic cipher (playing cards and dice choose between tables; each Voynich-like word encodes one or two plaintext letters). It reproduces many Voynich statistics strikingly well. It fails on one, though:

Voynich vs. Naibbe vs. real languages

h1/h2: character entropy and conditional entropy in bits, computed with word spaces; “type/token”: distinct words per token in the first 10,000 words; “boundary dependency”: mutual information between the last glyph of a word and the first glyph of the next (bits); shuffled baseline ≈ 0.004–0.018.

Texth1h2Ø word lengthType/tokenTop wordBoundary dependency
Voynich B3.872.005.160.2412.1 %0.238
Voynich A3.832.144.850.2994.6 %0.124
Naibbe cipher (Latin Pliny)3.882.085.230.2982.3 %0.006
Latin (Secretum secretorum)3.953.275.660.3169.8 %0.045
German (Luther 1545)3.883.124.420.1586.9 %0.089

3 · Cipher attacks in eight languages

Four cipher families were attacked, each in up to eight candidate languages and for Currier A and B. The chart shows the gain over the shuffled baseline (score of the real Voynich text minus score of shuffled Voynich; best of all runs on each side). Pick a cipher model:

Voynich vs. its own shuffled copy

Gain in log-probability per character. The grey band (±0.2) is where search noise lives; the line at +0.5 marks the typical gap when a cipher is actually broken (positive controls).

Voynich BVoynich Anoise range

The controls worked

Cipher modelControl texts (encrypted by me)Key recovered
Word-level verbose (Naibbe-type)Pliny (Latin), Parzival (MHG), Galen (Greek), Occitan, Old French, Catalan94–97 %
Glyph groups = letterssame six languages, 39–120 cipher symbols100 %
Glyph groups + memoryLatin, previous word ending selects 1 of 6 alphabets99 %
Word-level + memoryLatinnot decidable*

* With ~3,700 cipher symbols the Voynich text is too short: even on the control, wrong keys score better than the true key (−2.53 vs −2.61). This is a limit of the method, not a result about the manuscript.

The “sounds like Latin” trap

An optimiser always finds output that looks vaguely like the target language. Compare:

Only the control is real Latin (Pliny, Natural History 15: “pomiferae arbores quaeque mitioribus sucis voluptatem primae cibis attulerunt…”). The Voynich outputs are Latin-flavoured noise — and the scrambled Voynich does just as well. Many published “solutions” never run this comparison.

Apparent hits that did not survive. The three largest gains all came from the memory model on Voynich B: Middle High German +0.15, German +0.18, Greek +0.20. For German and Greek I ran a fairer control that keeps the word-ending rule but contains no message (“structured nonsense”): the gap shrank to +0.09 and +0.13. All outputs stayed unreadable. A +0.36 German result after removing null glyphs did not replicate on a second run (−3.08 vs −2.93/−3.00 shuffled).

4 · Words that talk to their neighbours

In the Voynich text, the end of a word predicts the start of the next — far more than in any natural language tested. A word ending in -y is followed by a word starting with q- 1.76× more often than chance; -r → a- 3.1×, -s → a- 4.9×. A cipher that encrypts each word independently (like Naibbe) cannot produce this. A cipher with memory can: if the previous word's ending selects the table for the next word, the effect even overshoots.

How strongly the end of a word predicts the start of the next

Mutual information between last glyph and next word's first glyph, in bits.

VoynichCipher with memory (constructed)Natural language / plain Naibbe

5 · Near-copies: the self-citation test

Torsten Timm and Andreas Schinner proposed that the text was generated by copying and slightly modifying earlier words on the page. I tested a simple, model-free prediction: how often is a word a one-glyph variant of a word in the line directly above, compared with a random line from another page of similar length? 1.0× means no preference.

One-glyph near-copies from the line above (ratio vs. a random line)

Means over 5 random seeds (±0.02–0.05). Reference line at 1.0 = no local copying.

VoynichGenerated text / constructed name listsProse (plain or enciphered)

The effect fades with distance

Near-copy ratio when comparing with the line 1, 2, 4 or 8 lines above. Hover for values.

The result holds within a single manuscript section (herbal, balneological, stars: 1.20–1.31×), so it is not just topic vocabulary. The most interesting comparison came from a magic book, the Sworn Book of Honorius (Liber Iuratus, Middle English version): its prose pages score 1.17×, but its pages with lists of invented angel names (“lafyel mazyel memyell paryel toupyel…”) score 1.58×. Voynich sits between prose and such lists. Both “generated nonsense” and “catalogues of systematically formed names or terms” remain compatible; ordinary enciphered prose does not.

6 · Position-dependent ciphers

What if the same glyph means different letters at different places — a Vigenère-type or running-key cipher, a new alphabet per page, or a key derived from the whole text (e.g. a Bible passage)? I encrypted the same Latin text in six such ways and compared their statistical fingerprints with Voynich.

Fingerprints of position-dependent ciphers

Each panel is one measure; Voynich highlighted. Position-dependent ciphers push every measure the opposite way from Voynich.

Voynich BVoynich ALatin plain text and its encryptions

There is one strong positional effect — but it is tied to the line, not to the position in the book:

First glyph of a word: line start vs. mid-line (Voynich B)

Share of words starting with each glyph. Line position explains 0.119 bits of the first glyph in Voynich B, 0.000 in Latin.

First word of a lineWords in mid-line

Each line behaves like a unit with its own opening rules. A cipher with a special alphabet at line start was part of the memory model above — also negative.

Is the line-initial glyph a key indicator?

Leon Battista Alberti (1467) inserted a capital letter to tell the reader that the cipher disk had been turned. Could the special line-initial glyphs play that role, switching the alphabet for the rest of the line? Then the first glyph of a line would predict the glyph mix of the rest of the line. It does — but only weakly: 0.0086 bits in Voynich B (z = 25; indicators shuffled within page and paragraph position), 0.0038 in A, compared with 0.002–0.003 when any other word of the line is used as “indicator”. A real key change is 35–65× stronger: 0.31 bits even if the alphabets differ by only three swapped letter pairs, 0.56 bits for unrelated alphabets. The effect is mostly one glyph: lines starting with the gallows p, t or k contain 1.8–3.4× more p — the known “gallows-rich top line” pattern, not a key switch.

7 · What the pictures tell us

The images are the only possible source of known plaintext (“cribs”), and also the source of most failed decipherments: a plant is “identified”, its label is “read”, and the reading fails on the next page. So I only used image clues that could be tested.

Zodiac labels are not day numbers

The 12 zodiac pages show 15–30 women each, 295 labelled figures in total. An obvious idea: 30 figures = 30 days or degrees of the sign, so the labels would be numbers. If they were numbers written in a Naibbe-type cipher, many labels would repeat across pages.

Labels found on more than one page
16 %
expected if day numbers: 57 % (never below 49 %)
Labels starting with ot- / ok-
53 %
running text: 11 %
Single-word labels
81 %
in Naibbe, one word = 1–2 letters only

The order of the labels around the ring shows no page-to-page alignment either (0.401 vs 0.403 ± 0.002 random). The common ot-/ok- start is the one real structure — a shared first word (“stella…”?) or a labelling convention. And the one-word labels are a problem for Naibbe: a star or plant name does not fit into one or two letters.

Pharmaceutical plants reappear in the herbal — their labels don't

I compared all 12 pharmaceutical pages with all 130 herbal pages by image only and fixed the list of matches before looking at any text. The clearest match:

Fragment from folio 99r: a hairy caterpillar-like root, a vine with arrow-shaped leaves and two hanging red clusters
f99r (pharma) — hairy “caterpillar” root, arrow-shaped leaves, hanging red clusters.
Folio 96v: plant with arrow-shaped leaves, hanging red-brown bead clusters and an elongated hairy root
f96v (herbal) — the same plant.

Eleven matches in total (1 high, 6 medium, 4 low confidence). The two strongest both link f99 to f96 — pages close together in the book. The illustrator reused plants on purpose. Then the test: do the labels of a pharmaceutical page occur on the matching herbal page more often than on comparable pages of the same Currier language and length?

Labels found (high + medium matches)
4 vs 6.3 expected
p = 0.94 · all 11 matches: 6 vs 8.1
Allowing one glyph difference
9 vs 19.6
fewer than chance
Label of the f99r fragment
tolsasy
nowhere in the herbal section, not even ±1 glyph

The “hits” were only everyday words (chol, cheol, char, chor, otor). Had all seven good matches carried their name on the herbal page, the test would have shown it (p = 0.016). This speaks against fixed-spelling systems (plain language, simple substitution, codebooks) in which the herbal text names the plant with the pharma label's word — while a cipher that spells the same word differently each time, labels that are not names, or meaningless text all remain possible.

The nine-rosette foldout: nine circular structures connected by causeways, with a castle and suns in the corners
The nine-rosette foldout: a cosmological diagram or a map — no close parallel is known.

8 · Two candidate source texts

If the Voynich encrypts a known medieval text, the lengths of that text's sections should reappear, in order, in the lengths of the Voynich paragraphs — whatever the cipher, as long as it roughly preserves length. Two candidates fit the pictures:

A dynamic-programming alignment maps each source section onto 1–4 consecutive Voynich paragraphs so that lengths match as well as possible (for the Antidotarium also allowing two recipes per group, a free start and end, and a gap for the lost bifolio f109–f110). The p-value compares the real order of the source with the same section lengths shuffled. As a positive control, I encrypted the real source in simulation — with paragraph splits, merges and length noise — and checked that the test finds it.

CandidateVoynich sectionVoynich pTest finds the true sourceVerdict
De balneis Puteolanis
34 sections, 14,103 letters; almost all 12 verses (±6 %)
Bathing, 82 paragraphs, 31,035 glyphs0.72
baths only: 0.86
100 % at ±5 % length noise✕Ruled out
Antidotarium Nicolai, Venice 1471
135 recipes, 16,929 words; 91–2,326 letters each
Recipes, 283–291 paragraphs, 10,457 words0.28–0.75
8 variants
100 % up to ±50 % change per recipe✕Ruled out
… as a selection
35 % of recipes left out at random
same0.36–0.8058 %!No support

Variants: order of the 1471 print or of the 1351 Middle Dutch manuscript; Voynich paragraphs from the transcription's paragraph marks or additionally split at lines starting with the gallows p/f (stars are not transcribed); lengths in letters/glyphs or words.

De balneis: 34 equal blocks?
±15 %
best grouping of the Voynich paragraphs; the encrypted source gives ±2–5 %
Recipe section vs. full Antidotarium
~72 %
of the space needed even at 1 glyph per letter or 1 word per word; Naibbe-type: ~21 %
Bathing section vs. De balneis (Naibbe-type)
64 %
31,035 glyphs available, ~48,700 needed

Neither text is hidden in these sections as a whole. What is not excluded: a heavily abridged selection, a cipher that inflates lengths at random, or paragraph breaks that ignore the source's structure.

9 · Is it music — or verse?

Two last ideas from outside cryptography. Music: around 1400, pitches were written with letters too (Guidonian letters, organ tablatures). If every glyph were a note, the text should move like a melody. Poetry: Voynich lines behave like units with their own opening rules (section 6) — just what one would expect if every line were a verse. Rhymed verse leaves a trace that survives many ciphers: line endings match their neighbours.

The melody test

For each text I searched for the ordering of its symbols — its “scale” — that makes the sequence move as stepwise as possible (seriation by simulated annealing), and compared it with the same symbols in random order. Validation: given only note sequences, the method rediscovers the exact scale — all 16 pitches of 103 Trecento upper voices (C3–D5, 14th-century Italy) and all 15 of 600 German folk songs (G3–G5), in correct order, without being told they are pitches.

How melodic is the sequence?

Melodic index = 1 − (average jump in the best “scale”) ÷ (average jump for the same symbols in random order). 0 = no order at all.

Real musicVoynichCipher / generated textPlain language (letters)

Voynich lies halfway: as ordered as a cipher or the self-citation generator, far from music. And it moves lopsidedly: in Voynich B, 69 % of all steps go the same way (melodies: 52–56 %, as they have to come back down), and repeated “notes” are rare (5 % vs 13–22 %). The words run through the glyph inventory in one direction (q → o → k → e → d → y) and jump back. That is the known “slot grammar” of Voynich words, not a tune.

The rhyme test

For every pair of neighbouring lines in a paragraph: do their last words share their last two glyphs — compared with lines at least four apart? The same is measured for the first, second and second-to-last word, so that general similarity between neighbouring lines cancels out. Rhyme index = excess at the line end ÷ excess elsewhere. The couplet test compares lines 3–4, 5–6 … with 2–3, 4–5 … (the first line of a paragraph is skipped, it has its own vocabulary).

Do line endings rhyme with the next line?

Rhyme index (1.0 = no rhyme). Parzival's bar is cut off at 2.

Rhymed verse (plain / encrypted)VoynichProse, generator, verse set as prose

No section of the Voynich rhymes — neither in couplets (z between −2.7 and −1.1) nor in alternating rhyme (ratio 0.88–1.09 for lines two apart). Rhymed verse written one verse per line would show even through a Naibbe-type cipher (z = 21). Two things this test cannot see: unrhymed verse (such as the hexameters of the popular verse herbal Macer floridus) and verse copied as running text — Parzival re-wrapped as prose scores 1.03, like Voynich.

What the two ideas add. Both fail in an informative way: the Voynich line is not a verse, and a Voynich word is not a musical phrase — it is a one-way walk through an ordered set of glyph slots. Any decipherment has to explain that walk.

10 · Through the page

One more idea: what if position in space matters? Which glyphs lie on top of each other when a leaf is held against the light, or which words hide behind a drawing on the other side? Stencil and grille ciphers that work like this are documented only from the 16th century (Cardano's grille, 1550), a century after the parchment. I tested it anyway.

Aligning the two sides

The parchment does the alignment itself: on the scan of a front page, the paint and ink of the back page show through faintly. I searched for the scale, rotation and shift at which these ghosts coincide with the mirrored back page (phase correlation). For 44 of 83 leaves the match is unambiguous: a correlation peak of 8–24 σ, while the back page of a wrong leaf never reaches more than 6.4 σ. The first version of this page aligned the leaves by their edges only and was off by 65 px on average (about 6 % of the page width, up to 10 %). The text-only recipe leaves hardly show through and cannot be aligned this way, so the test below uses herbal, bathing and pharmaceutical leaves.

Hold the leaf against the light

Front page on top, back page mirrored underneath, aligned by the show-through. Fade the front page, zoom in (+/−, Ctrl + scroll or double-click; drag to move), or fine-tune the alignment yourself.

Fine-tune the alignment
100 %

The see-through test

On these 44 leaves I placed every transcribed word on the scan (text lines followed along their curves, words matched to the ink), mapped the back page onto the front page and asked two questions. The control is the same leaf with the back page shifted by 2–8 lines: same scribe, same vocabulary, same left/right structure of the lines — only the physical overlap is gone. A planted relation tells how strong a real one would have to be to show up.

QuestionVoynichWould be found if …
Do words lying on top of each other resemble each other?
1,385 pairs · same word, one-glyph variant, same or dependent first/last glyph
p = 0.06–0.98
nothing significant
5 % of back-page words copied the word behind them, or 20 % took their first glyph from it (found every time)
Are words behind a drawing different?
461 words behind drawings vs 1,458 behind bare parchment · length, unique words, frequent “filler” words
length −0.18 glyphs, p = 0.04
others p = 0.57–0.69
30 % of the words behind drawings were fillers (found every time)

Nothing links the two sides. The one nominal hit — words behind drawings are 0.18 glyphs shorter — does not survive correction for testing three measures, and words next to a drawing are shorter anyway (see below). Limits: word positions come from automatic line tracking and are off by about a word in places, which would dilute a real relation; the planted tests assume the placement is right.

The wrong control would have produced a false positive. Compared with the back page of a different leaf, overlapping words looked related (same last glyph: p = 0.005). The reason is proximity, not position: the two sides of a leaf are consecutive pages by the same scribe with a shared vocabulary. Shifting the back page on the same leaf keeps that — and the effect disappears.

Where text meets a drawing

In 588 lines a drawing interrupts the text (1,270 words touch a drawing). Do these words behave like line starts and line ends — or like ordinary mid-line words?

Word positionstarts with y/sstarts with gallows p/tends in -mword occurs only once
Start of line15.5 %17.1 %0.5 %21.9 %
Right after a drawing24.4 %0.2 %0.5 %15.8 %
Mid-line5.9 %3.2 %1.4 %10.5 %
Right before a drawing––2.8 %16.0 %
End of line––16.0 %22.1 %
Three rules, not one. The start of every text segment, the opening of a line or paragraph, and the end of a line each follow their own rule. Any decipherment has to reproduce all three.

11 · History, owners, missing leaves

No second text in this script is known — only later copies by owners (Baresch's samples, Marci's columns).

Top of folio 116v with later marginalia in Latin script and a small drawing of an animal
f116v — later marginalia in Latin script, partly read as German (“…nim gasmich…”, “take goat's milk”?), next to a small animal drawing.

12 · Scoreboard of hypotheses

HypothesisStatusWhy
Letter or glyph-group substitution of Latin, Italian, German, MHG, Greek, Occitan, Old French, Catalan✕Ruled outControls cracked at 94–100 %; Voynich never beats shuffled Voynich beyond noise.
Periodic or position-dependent key (Vigenère, Trithemius, running key incl. Bible, per-page key, Alberti-style line indicator)✕Ruled outOpposite fingerprint; no period 2–40; words recur unchanged across pages and A/B.
Naibbe-type verbose cipher of ordinary prose!WeakenedNo word-boundary dependency, no near-copies, one-word labels too short.
Fixed-spelling system in which plant names recur!WeakenedPharma labels absent from the matching herbal pages.
Cipher with memory (word ending selects the next table)?OpenExplains the boundary effect; letter-level version negative, word-level version not testable at this text length.
Lists of systematically constructed words (names, terms, recipes)✓ConsistentMatches near-copy pattern, list-like layout, line-level rules.
Generated text without meaning (self-citation)✓ConsistentMatches near-copy pattern; doesn't by itself explain the picture cross-references.
De balneis Puteolanis as the plaintext of the bathing section✕Ruled outSection lengths don't align (p = 0.72; the test finds the true source 100 % of the time); too short for a Naibbe-type cipher.
Antidotarium Nicolai as the plaintext of the recipe section✕Ruled outNo alignment in 8 variants (p ≥ 0.28; power 100 % up to ±50 % length change per recipe). Only a heavy selection of recipes is not excluded (power 58 %).
Musical notation (one glyph = one note)✕Ruled outFar less stepwise than real melodies, and strongly one-directional like the slot grammar.
Rhymed verse, one verse per line✕Ruled outLine endings don't rhyme (couplet test z ≤ −1.1; rhymed verse through a Naibbe cipher: z = 21). Unrhymed verse remains untested.
Stencil or see-through relation between the two sides of a leaf✕Ruled outAligned by show-through on 44 leaves: overlapping words unrelated (p ≥ 0.06), words behind drawings not different after correction; a copy relation for 5 % of the words would have shown.
Nomenclator / codebook, abbreviated Latin, book cipher, other languages (Hebrew, Czech …)?UntestedNeed cribs, special corpora or the key book.

13 · Reproduce it

All scripts (Python + a small C solver) are in voynich-analysis-code.zip, with a README. Data sources used:

Limits of this study

Stochastic search can miss a solution (success rate per restart on controls: roughly 15–90 %, hence 8–24 restarts per run). Language models are built from edited modern texts, not from abbreviated medieval handwriting. Plant matching is subjective (hence the pre-registered list and graded confidence). All transcription-based results inherit the uncertainties of EVA, especially word spacing. The source-text alignments assume that Voynich paragraph breaks coincide with section breaks of the source and that the cipher roughly preserves relative length; the rhyme test assumes one verse per line.

References