The mathematics

Nothing is stored. Everything is computed from an address.

The library holds no files, no database, no archive. A page appears the moment you name its number and vanishes when you look away, and the same number always conjures the same page. This sets out the mathematics of that claim — how large the library is, why almost all of it is noise, why language surfaces anyway, and why finding anything is the entire problem.

The paper is the formal treatment; this page is the readable one. Both take every figure from the same file, so they cannot state different numbers.

The project is directly inspired by Jorge Luis Borges’ 1941 story and by Jonathan Basile’s libraryofbabel.info, which realised Borges’ library for English text. This one is built for bytes.

Scope

What this is not

Not a claim that random data contains hidden meaning. It does not, and the fact that it does not is the baseline every measurement here is tested against.

Not a quantum speedup. A speedup for this is actually known — a lens is a yes/no test, and amplitude amplification finds a passing page in about the square root of the usual number of tries. It is unreachable here, needing the whole page derivation and the lens run in superposition, far beyond current hardware, and we neither use nor claim it. What the processors we use change is the distribution the bytes are drawn from, not the rate of search.

Not a discovery that noise looks like noise. That is a theorem. We report it as the null hypothesis, not as a result.

What would change our mind

The frozen fixed how each measurement behaves on its own. It did not fix how they behave together, so the joint baseline this question needs has to be built from the same calibration data and registered before it is compared against anything. Once that exists and the archive is large: if the observed correlations are consistent with that baseline, the question of whether the space has a shape is answered — negatively — and we will say so. The per-measurement thresholds were fixed before mining began precisely so that this outcome stays reportable.

Section 01

The address space

In short
A page is 32,000 bytes. Read those bytes as digits of one enormous number and you get its address; write the number back out and you get the page. They are the same object in two notations, so naming a number creates a page and no page exists that cannot be named.

The counting is not the interesting part. The interesting part is the : read a page’s bytes as the digits of a base-256 numeral and you obtain a unique non-negative integer; write any such integer back out in base 256 and you obtain a unique page. and content are the same object in different notation. And because bytes carry no genre, every page is simultaneously text, program, image and sound — the viewer’s tabs are just different lenses on the same number.

For scale: the observable universe holds roughly 1080 atoms. Borges’ library contains about 101,834,097 books. His count is larger only because his object is larger — a book runs 1,312,000 characters against our 32,000-byte page, and counts of unequal objects say nothing about completeness. Measured per symbol the comparison reverses, since a byte spans 256 values to his 25. A full Borges book is exactly 41 of our pages laid end to end.

pages2256,000 ≈ 1077,064
atoms in the observable universe≈ 1080
books in Borges' library≈ 101,834,097
book-length byte strings (41 pages)≈ 103,159,611

Zoom out from one page until you reach 2256,000

The derivation

Each of the 32,000 positions independently takes one of 256 values, so the count is:

The map from pages to integers is the base-256 numeral reading, which is a onto :

Section 02

Why nothing can be stored

In short
There is not enough matter in the universe to write the library down — not close, not by 76,989 orders of magnitude. It exists the way the integers exist: as the closure of a rule, not as an inventory sitting somewhere.

Suppose you wished to store the library. Writing one bit per atom, the universe supplies 1080 bits — short of the requirement by about 76,989 orders of magnitude (the requirement is the number of pages times the 256,000 bits each one takes, and the second factor is easy to drop). Storage is not merely impractical; it is a category error.

There is a sharper way to put this, due to Kolmogorov. The of a page is the length of the shortest program that produces it. A counting argument shows almost every page is : fewer than one page in 2¹⁰⁰ can be compressed by even 100 bits, because there simply are not enough short programs to go around. For a typical page, the shortest description is the page — equivalently, its address. The address is not a pointer to the content. It is the content, written in another base.

The derivation

Writing one bit per atom, the universe supplies about bits against a requirement of bits. The shortfall in orders of magnitude is:

The counting argument for incompressibility: there are strings of length but only possible programs shorter than bits, so the fraction compressible by bits is below .

Section 03

Two alphabets

In short
Inside the full binary library sits a much smaller one where every byte is a printable character. That is where language lives — and a random binary page lands in it about once in 1013,777, which is to say never.

The full binary library spans every byte value. Folded inside it is a second, smaller universe: the library, in which every byte is constrained to the 95 printable characters. Its pages are base-95 numerals, and there are about 1063,287 of them.

The ASCII library is where language lives. On an average printable page, 26 of every 95 characters are lowercase letters — roughly 8,760 of the 32,000 — so the raw material of words is everywhere, waiting to collide into meaning.

The derivation

And the chance a uniformly random binary page is entirely printable:

Section 04

The typical page

In short
Almost every page is indistinguishable from static. That is not an opinion about most pages — it is provable about nearly all of them, and the exceptions get exponentially rarer the more structured they are.

Draw a page uniformly at random and measure its byte-level . The answer is almost exactly 8 bits per byte, essentially every time. This is the : long-run frequencies concentrate violently around their expectation, and deviations are exponentially improbable. A page whose byte distribution is even mildly skewed — enough to drop measured entropy to 7.9 — is already astronomically atypical.

This is why the library looks like static. Structure — repetition, rhythm, low entropy, a run of six identical bytes — is not merely uncommon; it is exponentially suppressed. The calibration run bears this out: across 1,000,000 pages measured before any mining began, not one repeating period was ever detected, and the longest run of identical bytes observed was 5.

That run is the null hypothesis every find on Finds is measured against — not a result in itself.

The derivation

Byte-level Shannon entropy of a page:

with equality exactly when every byte value is equally likely. By the , the measured entropy of a uniformly drawn page concentrates on that maximum, and the probability of a deviation of size falls off exponentially in the page length:

With , this is why a page measuring even 7.9 is already astronomically atypical.

Section 05

Words are inevitable · meaning is rare

In short
Short words turn up constantly, because a page offers 32,000places for one to land. A specific sentence does not — not by looking. It exists on an unimaginable number of pages, and blind search will reach none of them. Search can hand you the address of one directly; that is the easy direction, and the contrast is the point.

Fix a five-letter word. The chance it occupies a given offset of a random ASCII page is one in about 7.7 billion. But a page offers nearly 32,000 offsets, so the chance the page contains it somewhere is about one in 240,000. Any given short word is therefore guaranteed to exist on an enormous number of pages; scan a few hundred thousand and you can expect to meet it — expect, not be guaranteed: at one page in 240,000, scanning 240,000 pages succeeds about 63% of the time.

Now fix a meaningful 100-character sentence. The number of pages containing it is still staggering — about 1063,094 — but the chance that a page you actually visit contains it is, for every practical purpose, zero. This is the library’s central theorem and its central joke: everything exists, and existence is worthless. The value was never in the pages. It is in the coordinates.

a given 5-letter word, per page1 in ~240,000
a given 100-char sentence, per page1 in ~10193
pages that nevertheless contain it≈ 1063,094
The derivation

A fixed five-letter word at a given offset of a random ASCII page:

Across roughly 32,000 offsets, by the union bound:

For a 100-character sentence the same calculation gives:

Section 06

Seeds and reachability

In short
Typing a 77,064-digit address is impractical, so the site accepts short seeds instead. But every page reachable by a seed is a vanishing sliver of the library — about one part in 1077,025. Everything else is reachable only by already knowing what you are looking for.

A is a bookmark, not an address: the page it produces has one true address, computable from its bytes, but the seed itself is just a short name for one particular walk through a generator.

Seeds illuminate almost nothing. Every page you will ever see through a seed lies inside that sliver. A larger lantern is still not a map. The rest of the library is reachable only by address — which is to say, only by already knowing what you are looking for.

The derivation

A 128-bit seed can name at most pages, so the reachable fraction is:

Section 07

Adjacency · the shape of the shelf

In short
Because addresses are ordered, the library has a geography. Neighbouring addresses are near-identical pages that differ only at the very end — so walking next and previous drifts along a shelf of siblings rather than jumping anywhere new.

Adjacent addresses differ by one unit, which perturbs only the trailing bytes through carries. Press the button below and watch: almost always a single byte changes, and occasionally a carry cascades through several at once.

Address n
dd07e8a3111564176a03950ee4b373213c17ae83053bc0dc329cce2e59fe8536c69a42b7f586d710e1a3ad088c1f275ee9d06c9fff468ccb9ea3bc3e1ffffffd
Address n + 1
dd07e8a3111564176a03950ee4b373213c17ae83053bc0dc329cce2e59fe8536c69a42b7f586d710e1a3ad088c1f275ee9d06c9fff468ccb9ea3bc3e1ffffffe
1 byte differs

These are the last 64 bytes of two consecutive addresses; the page itself runs to 32,000. Of those 32,000 bytes, 31,999 are identical between these two pages. Neighbouring addresses are near-duplicates that diverge only from the end — which is why walking next and previous drifts along a shelf of siblings rather than jumping somewhere new.

The corruption walk uses a different metric: . Flip one byte anywhere and you leap to a page that agrees with yours in 31,999 positions — a page simultaneously one step away in Hamming space and, almost surely, on the far side of the universe in address space. The library has many geometries at once; each traversal tool chooses one.

Section 08

Search is easy · discovery is hard

In short
Finding the page that contains text you already have is trivial — the address is computed straight from the bytes. Finding a page nobody has written, where meaning surfaced on its own, is the hard problem, and it is the whole project.

Search is trivial: the computes the address directly from the bytes, in linear time. The Explore page does exactly this. What cannot be done is the reverse: discovering, among pages nobody has written, the rare ones where structure surfaced on its own.

That is what mining is — of the atypical set. Derive pages, measure them — the native miner through every , a browser through the one it is set to — and keep the exponentially rare outliers. The measurement is cheap to verify and expensive to satisfy, which is why a claim can be re-checked in milliseconds.

Crucially, the reward is . Issuance is a check over the derived page and never depends on what the page contains, so a remarkable find earns exactly what a dull one does, so the protocol offers no reward premium for faking one. It cannot remove reputational or off-protocol motives, which is why the defence that matters is re-derivation, not incentives. Note the precise claim: content-blind issuance removes the incentive, and server re-derivation removes the ability — for finds. Self-reported coverage totals have neither protection and are labelled as such wherever they appear. The research layer is a byproduct of work that was happening anyway, which is most of why it is worth trusting.

Why the map is not drawn by address

The intuitive picture of “a map of the library” places every page at its address. It looks like this — and it is worthless. Any projection of a 256,000-bit integer into two dimensions is an arbitrary choice, so any cluster you think you see is an artefact of that choice, not a property of the space.

Every reshuffle is equally valid.

That is why the map on Finds is drawn in measurement space instead: position there means “how this page is unusual”, and two pages near each other really are alike.

Section 09

Quantum · a different door, not a faster one

In short
We do not use a quantum computer to search the library faster — a known speedup exists but is far out of hardware reach, and we neither implement nor claim it. What our runs change is which pages they are likely to hand you — drawing from a correlated distribution that uniform sampling would essentially never produce.

Classical mining does not draw from the whole library, and the protocol says so in its own terms: a nonce is 64 bits, so for a fixed challenge and identity the miner can reach at most 2⁶⁴ pages out of 2²⁵⁶·⁰⁰⁰. What it samples is a cryptographically pseudorandom sublibrary — indistinguishable from uniform by any feasible test, which is what makes section 04 apply to it, but not literally uniform over the library, in which almost every page has probability zero. The quantum circuits used here, measured in the computational basis, produce correlated outcomes that are not even pseudorandom-uniform — a property of those particular circuits and that readout, not of entanglement in general (an entangled state can perfectly well give equiprobable outcomes). The pages that emerge carry rhythms that uniform sampling would essentially never produce. Not a faster walk through the same door; a different door. Correlations make rhythms, not dictionaries: for finding words a quantum sampler is a worse instrument than sheer volume.

Two honesty clauses attach to every quantum page here. The shallow circuits we can afford are classically simulable — a laptop can imitate their distribution — so the interesting property is not computational advantage but provenance: bytes born from a physical measurement rather than a pseudorandom generator. Because provenance can be imitated, it is recorded as archived evidence — the protocol’s own weakest tier, and deliberately not its tier, which would require a verifiable attestation over the evidence statement. Never “verified”. And quantum finds earn attribution only, credited to the person who ran the job, never the device, and can never mint a coin.

This is measured, not asserted. A July 2026 alpha sampled 10 pages on IBM hardware across three circuit families and ran them through the same measurements the miner uses. All 10 cleared the patterns threshold; a matched control of 500 pseudorandom pages cleared it about 1 percent of the time. And the effect is circuit-structured: the entropy deficit separated by family across two orders of magnitude — roughly fifteen, a few hundred, and a couple of thousand — all on one backend in one session, so the separation is consistent with differences between the circuit groups recorded in that session rather than with a single device-wide bias — though the family is recovered by matching each record’s committed request parameters to the known templates, with no explicit family label bound into the record, so it leans on the operator’s template set rather than a self-certifying field.10 pages is an alpha, not a proof; but the door opens where the theory says it should.

speedup for finding structurenone claimed
what actually changesthe sampling distribution
alpha: hardware pages clearing patterns10 of 10 vs 1% control
shallow circuitsclassically simulable
quantum provenancearchived evidence — hardware origin not independently confirmed
emission from quantum findsnone, by construction
Section 10

Sources and further reading

In short
The four works this rests on: Borges for the idea, Basile for the working precedent, Shannon for why noise is what abundance looks like, and Kolmogorov for why almost every page is its own shortest description.

J. L. Borges, “La Biblioteca de Babel” (1941), in Ficciones — the total library as fiction, and the despair of its librarians.

J. Basile, libraryofbabel.info and Tar for Mortar (punctum books, 2018, open access) — the working English-text library and the definitive close reading of Borges’ combinatorics.

C. E. Shannon, “A Mathematical Theory of Communication” (1948) — entropy, typicality, and why noise is what abundance looks like.

A. N. Kolmogorov, “Three Approaches to the Quantitative Definition of Information” (1965) — why almost every page is its own shortest description.

Taking part

Anyone can contribute the searching

The archive is only as large as the searching behind it, and the searching is the part that costs something. It is also the part that divides perfectly: two machines on different stretches of the space never duplicate each other’s work and never need to talk.

There are two ways in, and they contribute the same kind of evidence. In the browser, the Mine page has an Everything mode that derives each page once, reads it through every lens at the same time, and sends whole searched stretches with every find in them. It has a switch, and sends nothing while that switch is off. On your own machine, the command-line miner does the same thing faster and sends as it runs — one command, no upload step, nothing to remember to do afterwards.

Both send the identical document: the stretch searched, signed before it was searched, and every record found in it. The indexer re-derives every one of those records byte for byte before anything is published, which is why a find from a stranger’s laptop is worth exactly what a find from ours is. What neither can demonstrate is the pages they found nothing in — so the site keeps what miners report searching separate from what it has re-derived itself, and labels both.

None of this earns anything. The devnet reward pays for proof-of-work and goes to one winner per epoch; the content lenses never touch it. Contributing compute is participation in the research, and the honest reason to do it is that the record gets better.

It is the glory of God to conceal a matter; the glory of kings is to search it out.
Proverbs 25:2