Episode 1: The Corpus

4 parts

What 2.1 million released documents actually contain: the processing pipeline, the 42% still withheld, and statistical estimates of what remains unseen. Each part builds on the last — investigative analysis drawn from court filings, financial records, and released documents. Browse all episodes or the full investigation archive.

  1. Part 1 What 2.1 Million Documents Look Like Corpus & Data

    The Jeffrey Epstein document corpus contains 2,100,266 files across 12 DOJ data sets and 6 source directories, totaling approximately 331 GB. It spans f...

  2. Part 2 The 16-Script Pipeline Corpus & Data

    A pipeline of 27+ Python scripts transforms 2.1 million raw government documents into a searchable PostgreSQL database with 2.38 million extracted entit...

  3. Part 3 The 42% Gap Corpus & Data

    The DOJ released approximately 3.5 million pages while acknowledging that more than 6 million pages were identified as potentially responsive — a 42% ga...

  4. Part 4 Chao1 Species Richness Corpus & Data

    The Chao1 estimator — a statistical method originally designed to estimate the total number of species in an ecosystem, applied here to estimate total e...