Episode 3: Ghosts in the Machine
4 parts
OCR hallucinations, misread names, and 2.38 million extracted entities — the machine-reading errors that shape what the corpus appears to say. Each part builds on the last — investigative analysis drawn from court filings, financial records, and released documents. Browse all episodes or the full investigation archive.
- Part 1 OCR Hallucinations Machine Intelligence
OCR engines (software that converts images of text into searchable characters) produce phantom entities from blank form labels and repeated document hea...
- Part 2 The Wrong Robert Machine Intelligence
A FedEx shipment from ".EFFERY EOSTIEN" to "ROBERT" at "ART CF WOMEN" in Hawaii was initially identified as linking Epstein to artist Robert Crumb. Exte...
- Part 3 VLM on a Consumer GPU Machine Intelligence
Qwen2.5-VL-7B (a vision-language model — an AI system that reads images the way a human would) running on a single NVIDIA RTX 4070 with 8 GB of video me...
- Part 4 2.38 Million NER Entities Machine Intelligence
Automated name detection software extracted 2,383,898 entities from 2.05 million documents — 57% organizations, 37% persons, 6% locations. Nearly one in...