Episode 1: The Corpus
4 parts
What 2.1 million released documents actually contain: the processing pipeline, the 42% still withheld, and statistical estimates of what remains unseen. Each part builds on the last — investigative analysis drawn from court filings, financial records, and released documents. Browse all episodes or the full investigation archive.
- Part 1 What 2.1 Million Documents Look Like Corpus & Data
The Jeffrey Epstein document corpus contains 2,100,266 files across 12 DOJ data sets and 6 source directories, totaling approximately 331 GB. It spans f...
- Part 2 The 16-Script Pipeline Corpus & Data
A pipeline of 27+ Python scripts transforms 2.1 million raw government documents into a searchable PostgreSQL database with 2.38 million extracted entit...
- Part 3 The 42% Gap Corpus & Data
The DOJ released approximately 3.5 million pages while acknowledging that more than 6 million pages were identified as potentially responsive — a 42% ga...
- Part 4 Chao1 Species Richness Corpus & Data
The Chao1 estimator — a statistical method originally designed to estimate the total number of species in an ecosystem, applied here to estimate total e...