PROVENANCE
Methodology and provenance.
How a record gets from the official release to a page on this site, what the AI does to it, and what we know is imperfect. If you are going to cite this archive, read this page first.
Where the data comes from
The only source is the official release: the catalogue the U.S. Department of War publishes at war.gov/ufo, plus the video assets served through DVIDS. Nothing here is taken from UFO forums, private collections or secondary literature. If a record is not in the official release, it is not in this archive.
How each record is processed
- 01The official catalogue is downloaded and compared against what we already hold, so only new or changed records are reprocessed.
- 02Location text is geocoded where it names a real place; entries such as "Various" or "Low-Earth orbit" are deliberately left off the map.
- 03Documents are text-extracted, and scanned pages go through OCR. Pages that come back illegible are marked as such instead of being silently dropped.
- 04A language model writes a summary and a transcription for documents, and a visual description for video and imagery, so the material becomes searchable by what it actually contains.
- 05Summaries and descriptions are translated into the other five languages, the search index and the entity graph are rebuilt, and the whole site is regenerated as static pages.
What the AI does, and what it must not do
The model transcribes, summarises, describes imagery and translates. It is instructed not to interpret, conclude or fill gaps: if a page cannot be read, the correct output is to say so. Every AI-generated block on the site is labelled as such, and the original document is always linked next to it. The AI never decides whether a sighting is explained — that judgement, where it exists, belongs to the agency comment inside the file.
Known limitations
Illegible scans
Much of the corpus is typewritten carbon copy, microfilm or heavily redacted paper from the 1940s to the 1960s. Where OCR fails, the page counter tells you exactly how much was readable — "4 of 55 pages legible" — rather than implying full coverage.
Documents without a summary
Not every document carries an AI summary. Where OCR returns most of a document as illegible, no summary is published at all: 27 records currently show only the official description instead. A summary written over pages that cannot be read would be a guess, and this archive does not guess. Those records regain a summary only if the source is processed successfully later.
Machine translation
The five non-English versions of every summary are machine-translated, and the documents themselves are never translated — the transcription you search is always the original English. Where a translation and the original disagree, the original wins.
Automatic entity extraction
The knowledge graph is built by extracting people, places, agencies, craft descriptions and phenomena from the text. It is a navigation aid, not a finding: an entity is only kept when the term appears in legible text or in the official description, and near-duplicates are merged conservatively.
Corrections policy
If something here is wrong, it gets fixed and the fix is documented. Write to leonardo.sapuy@hotmail.com naming the record, the passage and, where possible, the page of the original PDF. Factual errors in transcriptions, summaries or translations are corrected at the source data and the site is rebuilt; where an error affected an AI summary, the summary is withdrawn rather than patched into a guess. This page is updated when a systemic problem — like the one described above — is found and corrected.
Last reviewed: 2026-07-26
About — an independent archive of declassified UAP records →