Methodology

The transcript talks. We check if it was right.

This site runs on a pipeline, not a vibe. Here's exactly what happens between someone saying something on the show and a verdict showing up here - and unlike some trackers, it grades guests as closely as it grades the four hosts.

  1. Capture the transcript

    Pull the episode's YouTube captions, timestamped line by line - no paid transcription service, no audio download.

  2. Pull out the calls

    Claude reads the transcript and extracts concrete, falsifiable, time-bound predictions - skipping jokes, hot takes, and vague futurism that can't ever be graded.

  3. Attribute the speaker

    Direct address, self-reference, and recurring topics assign each call to whoever said it - host or guest alike - with a stated confidence level, since there's no voice matching involved.

  4. Check the record

    A second pass searches the web for what actually happened and writes a sourced verdict, citing exactly what it found.

  5. Publish the receipts

    The verdict, the explanation, and every source cited go straight on the site, so any grade here can be checked by a visitor too.

The five verdicts

Right
Reality matched the call.
Wrong
Reality went the other way.
Partly Right
Some of the call came true, some didn't - a mixed or partial outcome.
Too Early
The forecast's timeframe hasn't elapsed yet, so there's nothing to grade yet.
Unvalidated
Extracted and attributed, but not yet checked against reality - still working through the archive.

How speaker attribution works

Because this project does not download or process audio, there is no voice-based speaker diarization. Instead, each prediction's speaker is inferred from the surrounding transcript text alone (direct address, self-reference, recurring topics). Each prediction carries a confidence level, and low-confidence or unattributed predictions are excluded from host and guest scorecards, though they remain visible on the relevant episode page for transparency.

Where this differs from other trackers

Some prediction trackers for this show score only the four permanent hosts and exclude guests entirely. This site tracks and scores guest predictions too (see any guest's page under Hosts), with a visible sample-size caveat when a guest's count is too small to be statistically meaningful.

About this project

This site tracks predictions made on the All-In Podcast and checks whether they came true, using a fully free toolchain: YouTube captions instead of paid transcription, and Claude Code (rather than a billed LLM API) for extraction, speaker attribution, topic tagging, and web-search validation.

It is a rewrite of an earlier version of this project that used paid transcription and OpenAI/xAI APIs. The rewrite trades a small amount of accuracy (see above) for a pipeline that costs nothing to run and can be regenerated by anyone with Claude Code and no API keys.

Disclaimer

This is an unofficial, best-effort project. It may contain transcription errors, misattributed speakers, or incorrect validation verdicts. It is not affiliated with the All-In Podcast or its hosts.