top of page

The first fully decoded Herculaneum scroll shows why AI search and knowledge recovery are becoming practical work tools

Researchers used high-resolution X-ray tomography and machine learning to read an entire Herculaneum scroll without opening it. Scroll PHerc.1667 became the first papyrus roll decoded from start to finish in continuous sequence.

The text contains a Stoic treatise that mentions Aristocreon, nephew of the philosopher Chrysippus. A second scroll, PHerc.Paris4, received higher-resolution scans that made ink visible in three-dimensional data and confirmed earlier contest results. A third scroll, PHerc.139, revealed its title and author, Philodemus, and the heading "On the Gods, Book VIII."

All data and code from the work sit in public repositories on GitHub and Zenodo.

Scroll project reaches continuous reading milestone

The Vesuvius Challenge team achieved the breakthrough by combining CT scans with trained models that detect carbon ink on rolled layers. Earlier efforts recovered scattered passages. This round produced connected text across multiple meters of surface. See the team's publications on X-ray tomography and machine learning methods at Nature and arXiv.

The decoded content is modest in length yet complete in sequence. Stoic philosophy appears in clear paragraphs rather than fragments. The result shows that the same pipeline can now process additional scrolls without physical damage.

Same methods convert hidden archives into usable records

The technical steps follow a repeatable pattern. First, non-destructive scans capture every layer. Second, segmentation models separate sheets that remain tangled. Third, detection models locate ink patterns on each surface. Fourth, alignment tools stitch the surfaces into readable columns.

The output is a searchable text file. The process requires no manual unrolling and leaves the original artifact intact. The same workflow applies to any collection of stacked or rolled documents that cannot be opened safely.

Knowledge workers face identical problems with legacy files. Meeting notes, old reports, and scanned contracts sit in shared drives or email archives. Manual search fails when file names carry no context. AI pipelines that recover layered surfaces can also index these stores. Real-world examples include the Smithsonian Institution using similar X-ray and ML pipelines to recover text from sealed historical manuscripts, and law firms applying AI segmentation tools to index decades of scanned case files in legacy formats as reported by Reuters.

remio turns recovered context into daily task output

remio captures documents, conversation history, and meeting transcripts as they arrive. It maintains five memory layers that keep recent activity, past events, and long-term concepts in one place. When a user asks for a report or slide deck, the agent already holds the relevant sources.

The Herculaneum result demonstrates that inaccessible material can become readable. remio applies the same principle to current work files. A stack of quarterly notes or policy drafts no longer requires hours of manual review. The agent extracts decisions, action items, and metrics directly from the stored record.

Users can ask questions such as "What pricing points were discussed last quarter?" and receive answers drawn from their own files rather than generic templates.

Public data release lowers the barrier for new collections

The Vesuvius Challenge published both images and trained models. Any lab with access to CT equipment can now run similar pipelines on other holdings. Libraries and corporate archives hold comparable material that cannot be opened without risk.

The public code removes the need to rebuild detection models from scratch. Teams can adapt existing segmenters and ink detectors to their own scan formats. This mirrors how remio connectors import data from Notion, Linear, or email without custom engineering for every source.

Once the material is indexed, retrieval becomes a single natural-language query rather than repeated folder navigation.

Limits of current recovery methods remain visible

The decoded scrolls still contain gaps where ink contrast fell below detection thresholds. Higher-resolution scans and additional training data are required before every character becomes legible. The same constraint applies to modern document recovery. Poor scan quality or handwritten notes in unusual formats reduce accuracy until models receive targeted fine-tuning.

The project also shows that human review stays necessary for final verification of proper names and technical terms. Automated output supplies the bulk of the text, yet edge cases require cross-checks against known references.

Three signals to watch in the next quarter

Additional scrolls from the same collection will receive the same treatment. Published results will indicate whether ink detection rates continue to rise with each new batch.

Corporate and university archives will test open-source segmentation tools on their own holdings. Early case reports will show which document types reach usable accuracy first.

remio will report adoption metrics for its Deep Research and Report Writer skills. Higher usage rates will reflect whether recovered context from existing files is producing measurable time savings in day-to-day reporting.

These three developments together will determine how quickly the pattern established with ancient scrolls scales to contemporary knowledge work.

Get started for free

A local first AI Assistant w/ Personal Knowledge Management

For better AI experience,

remio only supports Windows 10+ (x64) and M-Chip Macs currently.

​Add Search Bar in Your Brain

Just Ask remio

Remember Everything

Organize Nothing

bottom of page