Source Library’s whole promise is that the words you read on a page are the words printed on that page. So it was an unwelcome afternoon when we took a 1566 Vulgate, covered four lines of Genesis 1:2 with an opaque grey rectangle, and watched our production OCR return the verse complete — correct, seamless, in period orthography, with no gap, no warning, and a character-accuracy score of 99.3%.
Then we spent the rest of the day failing to fix it. Four different prompts, each asking the model in a different way to admit when it cannot read something. All four fabricated. This note is about that failure, because we think the negative result is more useful than a tidy success would have been.
What we did
The page is from a Louvain Bible of 1566. Verse 1 sits under a historiated initial in the left column; verses 3–8 run down the right. We masked the middle of verse 2 — about 27% of the reference passage — with a solid grey rectangle, then sent the image to the same model and prompt that transcribe our library.

is, át, &, a- — plus tur super aquas. below it. About ten characters of a hundred-character passage.What came back was the whole verse:
Terra autem erat inanis & vacua, & tenebræ erant super faciem abyssi: & spiritus Dei ferebatur super aquas.
No <unclear>. No <warning>. Nothing to tell a reader — or a downstream quotation API — that ninety of those characters were never visible. Ten runs at temperature 0 produced ten fabrications and, notably, one distinct output: this is not a rare stumble, it is a deterministic fixed point.
Ruling out the boring explanations
The obvious objection is that the model saw the text somehow. We checked three ways.
The mask is opaque. Sampling 203,304 interior pixels gives min 116, max 124, standard deviation 0.174 — an eight-level range, which is JPEG noise. A control patch of real text on the same page runs 45 to 235 with a standard deviation of 43.1. There is no signal under the box: the masked region carries 0.4% of the variance of actual text.
No unmasked image reached the model. Our harness passes a source URL only when the image is unmodified, and this run was resized, so no URL was sent. The Gemini runner base64-encodes the buffer and never consults a URL anyway.
The text is not elsewhere on the page. The left column holds verses 1–2, the right column starts at verse 3, and the Summarium at the top is a chapter précis, not verse text.
So the visible fragments were enough to identify which passage this was, and the model supplied the rest. A human scholar could do the same. The difference is that the scholar would say so.
Four ways to ask, four failures
Our prompt already offers the model a <warning> tag for quality problems and <unclear> for illegible readings. It used neither. So we tried asking harder.
| Prompt | Result |
|---|---|
| Production baseline | Fabricated |
| + open question: “is anything making this hard to read?” | Fabricated |
+ required <legibility> field, answer on every page | Fabricated, and reported clean on 4 of 5 masked pages |
| + inline marker, with explicit anti-fabrication language | Fabricated 6/6; marker never used |
The third is the worst outcome of the four. Making the field mandatory did make it appear — it was emitted on every page — but it appeared saying clean over a page the model was simultaneously inventing. Silence is at least ambiguous. A confident “clean” is a claim, and it is wrong.
The fourth included the sentence “reciting a remembered text as though you had read it is the worst error you can make.” It changed nothing.
Why the wording doesn’t matter
Every one of those four asks the model to report a state it cannot observe. To flag that it is reconstructing, it would need a signal distinguishing I read this from I completed this — and it doesn’t appear to have one. From the inside, completing a memorized verse from a few visible characters presumably feels exactly like reading. Instructions aimed at honesty cannot reach a failure the model cannot detect in itself.
The shape of the hole matters more than the prompt
One thing did change the behaviour, and it wasn’t anything we said. It was where we put the grey.

| Mask | Marked the gap | Fabricated |
|---|---|---|
| Rectangle | 0 / 4 | 4 / 4 |
| Diamond | 4 / 4 | 0 / 4 |
| Bowtie | 0 / 4 | 4 / 4 |
Under the diamond, the model wrote this, unprompted:
2 Terra [...] erat inanis [...] & spiritus [...] ferebatur super aquas.
Nothing asked for those brackets. The capability was there the whole time. The diamond leaves text at both edges of every line, so each line is visibly partial; the rectangle and the bowtie both leave some lines covered edge to edge, and a fully covered line apparently reads as continuous text to be completed rather than a hole to be marked. Whether the model admits a gap depends on whether the gap looks like one.
The honest caveat: we probed the rarest problem
A grey rectangle is not what damages a real book. We sampled 400 quality warnings the model wrote for itself across our corpus, and counted what it actually complains about:
| Condition | Share of warnings |
|---|---|
| Bleed-through from the reverse | 35% |
| Staining | 21% |
| Ink blots, saturation | 19% |
| Fading | 16% |
| Foxing | 14% |
| Gutter loss, tight binding | 12% |
| Cropping / occlusion | 2% |
We spent a day probing the 2% case. That is worth saying plainly, because it cuts both ways.
It makes the result a lower bound. A grey box is an unambiguous absence — and even then the model completed rather than reported. The conditions that actually dominate produce ambiguity rather than absence: bleed-through leaves a word half-legible, foxing leaves it smudged. Ambiguity is where a prior does its work invisibly, because resolving a blurred word toward the expected reading is indistinguishable, from the inside, from reading it. There is no stark grey rectangle to notice.
What we cannot yet tell you is how much of our corpus this touches. We have a mechanism, not a rate.
What this changed for us
The investigation started somewhere much more mundane. A reader turned off annotations on a Ming military atlas and the page went blank — because the model had wrapped the map’s cartouche labels in a tag our reader treated as AI commentary, and hiding commentary hid the entire page. That bug is fixed and shipped.
But it and the fabrication are the same defect wearing different clothes. Both are failures of span-level provenance: given a rendered page, can you tell which words came from the book and which came from the model? We have excellent artifact-level provenance — every page traces back to its prompt, model, job, and revision. Within a page, we had much less than we thought.
The working principle now is one line: every span of rendered text must be attributable — the page said it, or the model did. Grouping annotation tags by who wrote them rather than by whether they’re visible is what fixed the atlas. Making unreadable regions visible as unreadable is the unsolved half.
Where we think the fix has to come from
Not the prompt. Four attempts say the model cannot report what it cannot detect, and there is no obvious reason a fifth wording would land. The fix has to come from something that does know the pixels are missing — measuring the image directly for low-information regions, or checking the transcription against the region it claims to transcribe. Cheap, deterministic image statistics know a grey box is a grey box, and they know a bleed-through smear is low-contrast, without needing a model to introspect.
One more thing worth passing on, from the methodology rather than the result. Our first evidence for all of this was a striking number: in an earlier experiment, exactly one of twenty-eight runs mentioned a mask covering a quarter of the page. We built three issues on it before noticing that experiment had used a prompt reading “output only the raw text, no commentary.” The model had been told not to comment, and hadn’t. We had measured our own instructions and called it a finding. It is worth asking, of any number that surprises you, what would have to be true of the instrument to produce it.
The page is a Louvain Bible of 1566, held in our library. Masks, prompts, and raw model outputs are in the repository under scripts/eval/ so the runs can be reproduced or re-scored; the whole investigation cost about a dollar fifty in inference.
