I turn collections nobody can search (newspapers, manuscripts, records, film) into data you can question, at the scale of hundreds of thousands of pages. Every answer points back to the page it came from.
La Democracia, 1881, on the copy stand with an mm ruler and a 24-patch color target. Blocks are outlined in the color of the data class they hold; other blocks in grey.peopleplacesdatessourcesbox positions and ruler scale: illustration
capture logR031 · F0499
PaperLa Democracia
Year1881
Page IDR031 · F0499
MasterTIFF + JPEG copy
Checksumverified on every copy
Cited in“yerba mate” answer
106,393 pages captured · 290.6M words
How were public lands sold, and how did yerba mate grow after the war?
The papers show the state did not sell the yerbales on the open market. It kept them as public property and leased them. Yerba exports grew fivefold, concentrated in a few companies.
A picture of text becomes a record you can search.
One page, twice: as the scanner left it, and as the pipeline reads it, with every article, ad and table found.
Before · La Razón, R042 F0901. Text trapped in an image.raw scan · no text layer
After · The same page: every article, ad and table found and read.boxes: illustration
What the model returns
per page · validated JSON
segmentsevery article, ad and table on the page
typewhat each segment is
columnsthe column range it spans
textexactly as printed, old spelling kept
doubtuncertain words flagged, never guessed
sourcepaper, date and page, kept with every passage
The fields every page gets. A page that comes back short, empty or invalid moves up a tier.
How it’s done
Eight steps, from a fragile page to a cited answer. Pages enter once and are read by the cheapest model that can handle them; hard pages escalate. Every search runs keyword and meaning together, then answers only from cited pages.
1
Find and protect the pages
Track down the surviving copies. Give every page a stable ID, a master scan and a checksum.
stable ID · TIFF + JPEG master · checksum
2
Pick the right model
Six models raced on five of the hardest pages. No model won every page.
370 saved results across 34 runs
3
Turn the prompt into a spec
A one-line request became an editorial spec: keep the printed wording, flag doubtful words.
inspect · segment · transcribe · validate
4
Escalate only the hard pages
A fast model reads every page first; only the pages it fails move up a tier.
under 50 words, invalid JSON or empty: next tier
5
Make ideas findable
Pages become overlapping passages stored by meaning, so different words for one idea meet.
ferrocarril · camino de hierro · railway · transporte
6
Search six decades at once
Keyword search finds exact names; meaning search finds ideas. A second model reranks both.
merge · filter · balance · rerank
7
Answer with evidence
The best passages form an evidence pack. The answer may use only what it cites.
Quick 12 · Deep 30+ · Full: all passages
8
See the source, every time
Every claim links to its paper, date and page. The original scan is one click away.
La Democracia 1881 · R031 · F0499
The full build, by email
The overview is above. The full technical build, step by step, arrives as a one-time link sent to the address you enter.
305 Agentic · full build
Contents8 parts
Pipeline diagram · ingest once, search in seconds
Sources and preservation · stable IDs, masters, checksums
Model tests · 370 saved results across 34 runs
Prompt spec · layout first, exact wording, doubtful words flagged
Batch escalation · three tiers and a rescue route
Index and hybrid search · keyword and meaning, pgvector, reranking
Answer rules · evidence pack, citations, three depths
Live search screens · both archives
Sealed · opens with a one-time link
One email, no password. I see who reads the full build; nothing else is shared.
Full build unlocked for ·
The full build
Every step above, in full: what was tested, what was chosen and how the parts fit together.
Hybrid searchkeyword + vector · date and title filters
Rerankbalance decades · re-read top hits
Evidence pack12–30+ passages · paper · date · page
Answercites sources · flags uncertainty
Verifyopen the original scan
Pages enter once and are read by the cheapest model that can handle them; hard pages escalate. Every search runs keyword and meaning together, reranks, and answers only from cited pages.
01Find and protect the pages
source + preserve
The first problem was historical, not technical: finding where the papers still existed, then turning fragile pages into protected digital objects.
Rare copies located through historians, library holdings and private files such as the Cooney archive
High-resolution TIFF and JPEG masters
Stable IDs, metadata and checksums on every page
Every later step links back to this evidence layer
scattered sources
LibrariesPartial runs, fragile volumes
HistoriansKnow where copies survive
Private filesCooney archive, family papers
Microfilm~89,000 pages saved by the U.S. Embassy
one protected record
ID
R031 · F0499
Paper
La Democracia · 1881
Master
TIFF + JPEG copy
Checksum
Verified on every copy
Copies scattered across institutions become one record per page, with a stable ID that every later step points back to.
02Pick the right model
model selection
No two pages share a layout. Plain OCR reads lines and mixes columns. A multimodal model reads the page structure first. Six models went head to head on five of the hardest pages.
370 saved results across 34 documented runs
Tested resolution, reasoning level, JSON shape and prompt family
One advanced model was the most stable; no model won every page
A separate model writes answers; it never reads scans
Model benchmark: score on five hard pages, valid JSON outputs out of five, and average time per page
Model
Score, 5 hard pages
0102030405060708090100
Valid JSON
Avg time
Advanced model Achosen
93.6%
Valid JSON
avg122 s
Advanced model B
65.7%
Valid JSON
avg174 s
Fast model C
58.9%
Valid JSON
avg83 s
Fast model D
57.8%
Valid JSON
avg169 s
Small model E
56.4%
Valid JSON
avg231 s
Small model F
25.7%
Valid JSON
avg138 s
Model names withheld. Squares: valid JSON out of 5 runs. Time: average per page.370 saved results · 34 documented runs
Hard pages, with what makes each one hard.
Cabichuíwoodcuts + GuaraníR002 · F0078
La Autografiadaornate typeR007 · F0019
El Centinelablocks + engravingsR015 · F0019
Cacique LambarémultilingualR054 · F0009
Ads + tiny printdense advertisingR070 · F0013
El Puebloextreme densityR097 · F0677
03Turn the prompt into a spec
prompt testing
Small wording changes moved coverage. Dozens of experiments turned a one-line request into an editorial specification.
Archaic spelling and missing accents are evidence, not errors to fix
Uncertain words are flagged for review, never guessed
Every rule came from a regression on a real hard page
Output is validated JSON: segments, type, column range, source
“Extract the text”, then “Transcribe exactly”, then inspect · segment · transcribe · validate
La Razón · page segmented by the model
the four rules
1Identify the page layout before extracting text.
2Preserve the printed wording exactly.
3Mark doubtful words instead of inventing them.
4Test on real hard pages, not clean samples.
04Escalate only the hard pages
batch pipeline
Running the best model on every page is slow and expensive. A tiered factory sends each page through the cheapest model that can handle it.
Batches run unattended, page by page, with retries
Cheap fast model first; most pages stop here
Failures and known-hard titles move up a tier
Rescue route handles output limits and empty responses
1
Volume · Simple, fast model
Fast first read of every page, temperature zero, structured JSON.
escalate if < 50 words, invalid JSON or empty output
2
Complexity · Advanced model
Dense, damaged or sparse pages. Known-hard papers such as La Democracia go here directly.
escalate if < 50 words, invalid JSON or empty output
3
Rescue route
Output limits, recitation blocks, empty answers. Handled by observed pattern, page by page.
05Make ideas findable
embeddings + index
Words change over 160 years. Each page is split into overlapping passages and each passage becomes a vector stored in PostgreSQL with pgvector.
Overlapping chunks so no sentence is cut in half
Paper, date, page and collection stored with every vector
One index across every year and title
Coral Gables and Paraguay run on the same stack
overlapping passages
one idea, many words, 160 years
ferrocarril
camino de hierro
railway
transporte
All land near each other in vector space, so a search for one finds the others.
06Search six decades at once
hybrid search + rerank
Keyword search alone misses meaning. Vector search alone misses exact names. Running both, then reranking, finds the right passage across tens of thousands of pages.
Keyword and vector results merged into one list
Metadata filters narrow the search before ranking
Results balanced across decades and papers
Reranker re-reads the top candidates before anything is shown
keyword
Exact names, places, prices. Finds “Biltmore” and “R031”.
vector
Meaning. Finds “land sales” when the paper says “arriendo de yerbales”.
Filterdate range, title, collection
Balanceno single decade or paper crowds out the rest
Penalizeshort commercial fragments and ad noise
Reranka second model re-reads the top candidates
evidence packet · best passages with source links
Merge, filter, balance, rerank: hundreds of candidates narrow to a small packet of passages, each linked to its source.
07Answer with evidence
search prompt + answer
This is a research tool, not a chatbot. The best passages form an evidence pack, and the answering model writes under strict rules.
Three depths: quick, deep and full extraction
Streaming answers with follow-up questions
Each claim links to its paper, date and page
Limits stated: OCR errors and missed passages are possible
Question, then Evidence pack, then Answer model, constrained, then Answer
search prompt rules
Answer only from the passages retrieved
Cite paper, date and page for every claim
Compare sources when they disagree
Say what the evidence does not show
depth
Quick
12 passages
Deep
30+ passages
Full
all, full extraction
08See the source, every time
search UI + results
The technology disappears behind a place to ask, browse and check. The original scan is always one click away.
Chat with citations and a deep-research mode
Visual archive filtered by date, title and collection
Image search across photos and engravings
Export findings as infographics, slides and short bios
How were public lands sold, and how did yerba mate grow after the war?
The papers show the state did not sell the yerbales on the open market. It kept them as public property and leased them. Yerba exports grew fivefold, concentrated in a few companies.
paraguayhistory.net/appArchivo Histórico Paraguayo. Ask in plain words; every answer cites the original papers. Beside it: infographic, presentation, archive download and connection graph.
Profundidad: rápida · 12 pasajes · 1877–1897 · 5 periódicos · Todas las colecciones
¿Cómo se vendieron las tierras públicas y creció la yerba mate tras la guerra?
Los periódicos muestran que el Estado no vendió los yerbales en mercado abierto: los mantuvo como propiedad fiscal y los arrendó. La exportación de yerba mate se quintuplicó, concentrada en pocas empresas.
La Democracia · 1881
La Reforma · 1877
La República · 1891
El Tiempo
Real answer · every claim links to its paper, date and page
coralgableshistory.net/app?q=Venetian PoolCoral Gables Historical Archive. 6,825 pages, 1941–1977. The answer is written from the pages; each citation opens its page, and the pages it used are listed beside it.
Six models, five of the hardest pages
Every model read the same five pages. One advanced model was the most stable; no model won every page. Model names are withheld.
Model benchmark: score on five hard pages, valid JSON outputs out of five, and average time per page
Model
Score, 5 hard pages
0102030405060708090100
Valid JSON
Avg time
Advanced model Achosen
93.6%
Valid JSON
avg122 s
Advanced model B
65.7%
Valid JSON
avg174 s
Fast model C
58.9%
Valid JSON
avg83 s
Fast model D
57.8%
Valid JSON
avg169 s
Small model E
56.4%
Valid JSON
avg231 s
Small model F
25.7%
Valid JSON
avg138 s
Squares: valid JSON out of 5 runs. Time: average per page.370 saved results · 34 documented runs
The five pages every model had to read.
Cabichuí 1867woodcuts + GuaraníR002 · F0078
La Autografiadaornate typeR007 · F0019
El Centinelablocks + engravingsR015 · F0019
Cacique LambarémultilingualR054 · F0009
El Puebloextreme densityR097 · F0677
Two archives, one search engine
Both were built end to end and both are public. Search them yourself.
Two collections on one ruler. band height = average pages per year
“…recreation facilities including the municipal golf course, tennis courts and the Venetian Pool.”
Riviera · Apr 25, 1941 · p. 1
Venetian Pool · 1940sBiltmore · 1926 postcard
“I thank former Ambassador James Cason and his family for the invaluable project. This technology-driven historical archive will reach our schools as an innovative tool for students to explore our past interactively.”
Santiago Peña, President of Paraguay, via La Nación
“Through AI this has life and can be used. Had it stayed in an archive in a traditional format, no one would have seen it.”
Luis Fernando Ramírez, Minister of Education, Paraguay (La Tribuna)