Episode 17: Search and Statistics

4 parts

Sub-second search across 2.1 million documents, the German Tank Problem applied to missing records, and why nothing in the corpus is a finding by itself. Each part builds on the last — investigative analysis drawn from court filings, financial records, and released documents. Browse all episodes or the full investigation archive.

  1. Part 1 Sub-Second Search Across 2.1 Million Documents Technical Deep Dives

    A fast full-text search system built into the database (known technically as a GIN-indexed tsvector) enables sub-second queries across 2.1 million docum...

  2. Part 2 Detecting Missing Documents With a World War II Statistical Method Technical Deep Dives

    A World War II-era statistical method for estimating how many items exist based on the serial numbers you have seen (known as the German Tank Problem) i...

  3. Part 3 The Many Rodgers: How Entity Resolution Handles Scanning Chaos Technical Deep Dives

    A software tool that uses statistics to decide which database records refer to the same person (called Splink) merged eight or more scanning variants of...

  4. Part 4 Why Nothing in This Corpus Is a Finding Technical Deep Dives

    Every analytical result in the pipeline receives a confidence grade from a military-origin system for rating how reliable a source is and how credible t...