Every file that could exist already does.
A page here is 32,000 bytes, and every possible arrangement of those bytes is a page. So every book, program, image and song that fits is already in the library — alongside an unimaginably larger amount of static. We are looking for the pages that aren’t static, and for whether they have a shape.
What is in there?
Words, rhythms, stretches that decode like a program, pages that sound like something. Every page a coordinate miner derives gets measured, and the strongest of what clears the bar is published — boards keep the top hundred, and a browser reports its best eight at a time, so what you see is the peak of the stream rather than all of it.
Does the space have a shape?
Not “is this page unusual” but “are the unusual pages unusual in the same ways.” Answering that needs volume, which is why this is a collective search rather than a solo one.
What you can realistically find
The odds below fall out of the mathematics rather than out of a model that was fitted to anything. Some have also been checked against the 1,000,000-page calibration corpus; others have not, because the corpus is not large enough to contain them — nothing as rare as a six-byte run turned up in 1,000,000 pages, which is exactly what the arithmetic predicts. The table says which alphabet each row is computed in, and the times are estimates.
| Find | Odds per page | Estimated time on one machine |
|---|---|---|
| Any word of 5 or more letters | 1 in 190 | A fifth of a second |
| Any word of 6 or more letters | 1 in 14,000 | 14 seconds |
| One particular 5-letter word | 1 in 240,000 | 4 minutes |
| A run of six identical bytes | 1 in ~34,499,869 | About 9½ hours |
| Any repeating period at all | never yet seen | Unknown — still open |
| One particular 100-character sentence | 1 in 10193 | 10¹⁵⁵ × the age of the universe |
Times are estimates, not measurements: they assume an unverified 1,000 pages a second on one machine, and every duration except the last scales directly off that assumption — the last row assumes something far larger, stated below. The odds are the calculated part; the durations are only there for a sense of scale.
These rows do not all use the same alphabet. The word rows are measured over raw bytes, the way the protocol’s own lens reads a page. The single-word and sentence rows are computed inside the printable-ASCII library, where every byte is one of 95 characters — a much smaller space, so those odds are not comparable with the rest of the column.
The last row assumes 10²⁰ pages a second — roughly an eighth of the Bitcoin network’s current hash rate — running since the Big Bang. Even then the chance of meeting that one sentence is about 10⁻¹⁵⁶. It is not that the sentence is absent: it is on an unimaginable number of pages. The number is simply small enough that “never” is the honest word for it.
What will your browser find?
Your browser mines with the same code the verifier runs, so anything it finds is real. Its strongest finds are sent on, and the server re-derives each one before publishing it; the rest stay counted but unnamed.
See what you can find →Page counts are self-reported by the mining browsers; every find shown was re-derived and verified by the server. Nothing here pays more for a remarkable page than a dull one, so there is no reward for faking — and the re-derivation is the part that makes a find itself impossible to invent.
“You’re not going to find anything.”
Mostly correct — and that is the design, not a disappointment.
A page derived from noise is supposed to look like noise. Structure is exponentially suppressed; the mathematics says so and every scan we have run agrees. Finding a five-letter word is not a discovery, it is arithmetic — at one page in 190, it would be strange not to. Any threshold set low enough to fire will fire.
So the question is not whether we found something unusual. It is whether the unusual pages hang together.
The native miner measures every page it derives through all five lenses — words, structure, decode density, execution, sound — and the structure lens alone takes 7 separate readings. A browser mines one lens at a time, whichever you have selected. 1,000,000 calibration pages fixed what each of those readings does on its own, and those thresholds were frozen before any mining began.
What that calibration does not yet contain is how the readings behave together. It measured each one separately, which is why the interesting question is still open rather than already answered: if pages turn out to be unusual on several lenses at once, more often than independent thresholds would produce by chance, that is the one result here that a generous threshold cannot explain away. Establishing that baseline is work still to do, and stating it plainly is the point — a joint claim made before the joint baseline exists would be exactly the kind of result this project is built to avoid.
That is Question B. It is also why it needs many people: a correlation between lenses only becomes measurable at volume.
Mining pays for work, not for what you find.
The reward is a proof-of-work check over the page a miner derived. It never looks at what the page contains, so a spectacular find earns exactly the same as a boring one — so the protocol pays no premium for faking one. That weakens the motive without removing every motive — reputation and off-protocol payment are untouched — what actually protects the record is that the server re-derives every published find. The research layer is a free byproduct of work that was happening anyway.
One machine is not enough.
Every page a coordinate miner derives is scanned by the same frozen code the verifier runs — the native client through every lens, a browser through the one it is set to — and what clears is recorded to a shared map. The seed explorer is deliberately different — it keeps its finds on your own machine and uploads nothing. The map grows with the number of people mining coordinates, and Question B does not become answerable until it is large.
Join the search →These numbers are small because the search has barely started. Pages looked at counts only signed, non-overlapping archive coverage — self-reported browser totals are shown separately, further down. Finds counts archived notable records plus quantum records.