Find the Truth & Claim Evidence Matrix

Type a claim, get a probability. How Find the Truth researches, weighs, and scores a claim's truth — and how to read the claim evidence matrix behind it.

What It Does

Find the Truth takes a claim — or a question — and researches it from scratch. It gathers evidence for the claim, deliberately hunts for the strongest evidence against it, weighs everything by how trustworthy each source is, and returns a probability that the claim is true plus a cited report.

  • Find the Truth (every signed-in user, token-billed) — as of W1.44, reach it from the ✦ Deep button on the chat composer, under Interrogate. Type a claim, confirm a reshaped version if needed, pick a research depth, and run it. The verdict streams back as a probability meter, a plain-language finding, an evidence breakdown, and the strongest counter-evidence found.
  • Claim Evidence Matrix (every signed-in user, free to view) — a companion page at /research/[cafeId]/claims that shows every claim in your Cafe (from Find the Truth and from regular Verify claim runs) as a table of supporting / contradicting / silent sources, grouped by credibility tier, with independent-source counting.

The two features share data: every claim Find the Truth checks appears in the matrix, and the matrix works even if you've never run Find the Truth — it's populated by any Verify claim run too.


Find the Truth

How It Works

  1. Open a Research Cafe, click ✦ Deep on the chat composer, and choose Find the Truth under Interrogate (next to Claims). If you've already typed a claim or question in the chat box, it carries straight into the box below.
  2. Type your claim — or a question. Examples: "Coffee consumption reduces the risk of type 2 diabetes" or "Is the Great Wall of China visible from space?"
  3. A quick preview shows how the system understood your input:
    • Statements pass through mostly as-is.
    • Questions are reshaped into a checkable statement before anything runs — e.g. "Is the Great Wall of China visible from space?" becomes a statement you can accept or edit.
    • Claims that aren't checkable (opinions, nonsense, unfalsifiable statements) are rejected with a reason — nothing is charged.
  4. Pick a research depth: Quick answer, Solid research, or Leave no stone unturned. Each shows a rough cost estimate before you commit — no tokens are charged until you click Run. This estimate is capped at what that depth is actually allowed to spend, so it's never quoted higher than the run's own budget ceiling.
  5. Click Run. A progress panel streams through the phases: Gathering → Hunting counter-evidence → Weighing evidence → Building the matrix → Scoring → Writing report.
  6. When it finishes, the verdict card appears with the probability, the finding, the evidence breakdown, and a link to the full cited report.

Which model runs the check: Find the Truth runs on whichever model is selected in the Cafe's model picker (the same selector in the search bar's info row that drives Auto-Research and Ask Your Cafe — see Research Mode → Choosing a Model). Pick a different model there before you click Run if you want the check reasoned through a specific model; the run's token ledger reflects whatever model actually did the work.

Runs in the Background — Survives Closing the Tab

Find the Truth is a background job, same as Auto-Research and Verify claim:

  • Navigate away from the setup, switch tabs, or close the browser — the run continues server-side.
  • Reopen the Cafe and the run reattaches automatically if it's still going.
  • Every completed run is saved to a re-openable "Past truth checks" list, so you never lose access to a verdict or its cited report after leaving the setup.
  • If the server restarts while your run is mid-flight (rare — a deploy or crash), the job can't finish. Rather than leaving the run spinning forever, the next time the Cafe loads its truth-run list, any run that's been stuck for a couple of hours is automatically marked failed with a clear message ("The run was interrupted (server restart) and did not finish.") instead of showing "running" indefinitely. Just start the check again.

Waiting for a Free Slot

There are two caps on deep background runs — Find the Truth, Auto-Research, and Research a Person — and either one can make a run wait:

  • Your own cap — how many of your runs can be executing at the same time. This is the one you're most likely to meet: start a person research and a truth check back to back and the second waits for the first.
  • The platform cap — how many runs can be executing at once across everyone, so a busy moment degrades to waiting rather than slowing every run down.

If either cap is full when you click Run, a "Waiting for a free slot" panel appears in place of the usual progress panel:

  • It shows your place in line (e.g. "You're #3 in the queue").
  • Your run starts automatically the instant a slot frees up — nothing to click.
  • Nothing is charged while you wait — tokens are only spent once the run actually starts.
  • You can close the window; the run holds its place while it waits and reattaches when you come back, same as any other background run.
  • A Cancel button lets you give up your place at any time.
  • The wait is bounded. If no slot frees up in time the run is released rather than hanging forever, nothing is charged, and you can start it again.

When a run is released it tells you which cap held it up, because the two need different responses: "You already have a run going" means finish or cancel that one; "The platform is at capacity" or "Other runs were ahead of yours" means try again shortly.

What Sources It Uses

Find the Truth researches with web search, academic databases (OpenAlex, arXiv, PubMed), news search, Wikipedia, and newspaper archives (W1.12a). It deliberately does not use social media, Google Scholar, or manual URL scraping for a truth run — those are Research Mode tools for gathering material generally, but they're excluded here because a truth verdict should only draw on sources that can responsibly move a probability. (These same tools remain available for regular Cafe research via the search bar and Auto-Research.)

Newspaper results are filtered for topical relevance. A historic-archive match has to share a real word from your claim with the article's title or description, not just a matching year — an old periodical that happens to be from the right year but is about something unrelated is no longer admitted as evidence.

Newspaper evidence is weighed claim-type-aware. A digitised historic newspaper article is strong evidence for whether something happened or was reported at the time (e.g. "was the assassination widely covered in the press the next day?"), but on its own it's weak evidence for whether an underlying factual claim is actually true (e.g. a period newspaper repeating a contemporary rumor doesn't make the rumor true). Find the Truth's scoring reflects that split automatically — you don't need to do anything differently when a newspaper source shows up in the evidence breakdown. Snippet-only or reference-only newspaper sources (in-copyright items the archive could only partially return) are additionally capped so they can't single-handedly decide a verdict — see Newspaper Archives for how that tiering works.

Cost

Find the Truth is available to every signed-in user and costs tokens — it's one of the more expensive research actions because it runs two full research passes (gathering evidence for the claim, then separately hunting for the best evidence against it) plus scoring and report writing. The exact estimate for your chosen depth is shown before you click Run, and nothing is charged until you confirm; if your balance is too low to cover the estimate, you'll be prompted to buy tokens before the run starts.


Reading the Verdict Card

The Probability

The headline number is a percentage that the claim is true, always between 5% and 95% — it never shows 0% or 100%. This is intentional: Find the Truth produces an evidence-weighted estimate, not a certainty. Even overwhelming evidence is clamped short of absolute certainty, and even claims with no evidence never round down to a flat zero.

The card always shows the dominant side of the verdict. A claim the evidence supports reads as likely true with its percentage; a claim the evidence refutes reads as likely false with that percentage — so a claim scoring 8% true is shown as "likely false" at 92%, never as a confusing "true, 8%". Same underlying number, always displayed from the side the evidence actually lands on.

Below the percentage is a qualitative bandtrue, likely true, leans true, uncertain, leans false, likely false, false, or mixed — and a confidence read (low / medium / high) driven by how much evidence was found and how one-sided it is.

When the Band Says "Mixed"

Mixed means the model judged the claim genuinely two-sided and deliberately declined to call it true or false. It's its own band, not a fifth flavor of "leans true" or "leans false" — the probability meter still lands somewhere in the middle for a mixed verdict, but the headline word never picks a direction. It renders amber, the same tone as uncertain.

Before this existed, a genuinely mixed judgement was still forced through the same seven directional bands as everything else — so a two-sided verdict could print as "leans true" or "leans false" depending on which way the underlying number happened to fall, even while the "What we found" text described it as mixed. That's fixed going forward. It is not retroactive: a Find the Truth report you ran before this change keeps whatever label it showed at the time, even if the same claim would score as "mixed" if you re-ran it today.

This is a different feature from Verify claim's own "Mixed" badge (see the FAQ's verify verdicts table) — that's a separate, source-level cross-reference check with its own four-way result (Verified / Contradicted / Mixed / Couldn't verify). Find the Truth's "Mixed" is one of its probability bands. Both describe roughly the same idea — evidence on both sides — which is why the two tools share the word.

Confident Bands Have to Be Earned

A verdict can only show a confident bandtrue, likely true, likely false, or false — when it's backed by either 3 or more independent sources (any credibility mix), or 2 independent sources where at least 2 are medium/high credibility (for example, a solid news outlet plus a government source). A single low-quality source, or two thin web pages, is not enough to earn a confident reading.

When the evidence is thinner than that, the verdict is capped at "leans true" or "leans false" (or stays in the uncertain middle) no matter what the raw math would otherwise produce, and the small label under the percentage reads "insufficient-evidence estimate" instead of the usual "evidence-weighted estimate" — an honest flag that the number is a rough read on thin evidence, not a confident finding. Run the check at a deeper level, or wait for more sources to accumulate in the Cafe, if you need a firmer answer.

"What We Found"

Above the probability meter, a "What we found" block gives you the plain-language answer in a sentence or two — the actual finding, not just the number. This is often the fastest way to understand why a claim landed where it did (for example, a claim about a specific statistic might be technically false because the real figure only applies under different conditions — the finding will say so, even though the headline number alone wouldn't).

The Evidence Breakdown

Below the verdict, a breakdown table shows what the score is built from:

  • An "Overall reading" row at the top — this is the model's holistic read of all the gathered evidence together, and it's the dominant factor in the score. It's shown as its own row (with an info tooltip explaining what it means) specifically so the visible rows account for the number you see — nothing is hidden.
  • Per-source rows below it — each independent piece of evidence, its credibility tier, whether it supports or contradicts the claim, and its contribution to the score.

Independent-root counting applies here too (see Claim Evidence Matrix below) — several articles that all trace back to one wire story or press release count once, not once each.

Strongest Counter-Evidence

Find the Truth always runs a dedicated counter-research pass — it actively hunts for the best evidence against the claim, even when the claim looks obviously true. The verdict card surfaces the single strongest piece of counter-evidence it found, so you see the best case against the claim even on a claim that ultimately scores as true.

When the Card Says "We've Weighed the Sources as a Whole"

Occasionally the per-source breakdown is suppressed and replaced with a neutral note: "For this claim, we've weighed the sources as a whole rather than relying on individual for/against labels." This happens when the system detects that labeling each source as "for" or "against" would be unreliable for this particular claim (this is more likely on claims involving negation or absolute/universal wording, e.g. "always," "never," "everyone"). Rather than guess and risk mislabeling evidence, the card falls back to the honest "we can't confidently split these individually" framing — the headline probability is unaffected; it's still computed from the same weighed evidence, just displayed without the unreliable per-source split.

Copying and Downloading the Cited Report

Every Find the Truth run produces a full cited report — click through from the verdict card to open it in the reading pane or full-screen viewer. From there, the report viewer's Copy, Download (.md), Download PDF, and Save to Notebooks (W1.17 — saves the whole report as a note) buttons work the same as on any other research report (see Auto-Write Research Reports → Copying and downloading). A short AI-generated content notice appears at the bottom of every truth-verdict report, same as every other report type.

Every source the report discusses is in the References list. If a source is named in the "Supporting evidence" or "Contradicting evidence" prose, it always has a matching [N] marker and a corresponding entry in the References bibliography — the report never names more sources than it actually cites.

When a web search turns up several real articles at once, the report can cite them individually. Each article keeps its own real title and URL in the References list — not one merged reference standing in for a handful of different pages.

"Silent sources" only ever means sources this specific run gathered. If the verdict's report lists sources it didn't ultimately weigh in the finding, that list is scoped to what this run actually collected — it never pads the list with unrelated sources already sitting in the Cafe from earlier, unrelated research.

As of W1.17, a Find the Truth report is also searchable in Ask Your Cafe — a Grounded Q&A answer can cite it the same way it cites a source, and clicking the citation opens the report in the reading pane.

The Sources Rail and Reading Pane Refresh Automatically When a Run Settles (W1.17)

When a Find the Truth run finishes (verdict, error, or stopped), the sources rail and reading pane refresh themselves a couple of seconds later — you don't need to manually refresh the page to see the run's gathered sources or click through to its report. If you had a source open when a run you started elsewhere in the Cafe completes, this is why the rail updates on its own.

Download PDF (W54) is worth calling out specifically for Find the Truth, because this is the only report type with a fixed section order: the PDF opens with the verdict graphic (a green/amber/red square, the dominant-side percentage, and the band label), then the report body, then the claim/truth matrix for that report's claim (supporting / contradicting / silent, grouped by credibility tier), and all citations last. The verdict square and matrix are reproduced as server-rendered vector graphics — not a screenshot — so they stay crisp at any zoom or print size. It's generated entirely by code (no AI model call), free for every signed-in user with no token cost, and it's available on a shared report too — anyone you shared it with can download the same PDF. The References section (a collapsible list on the web viewer, W55.1) always starts on its own page in the PDF, same as every other report type.

Why Your Number Might Look Lower Than You Expect

Credibility tier drives how much weight each piece of evidence carries: high (peer-reviewed papers, government sources) > medium (encyclopedias like Wikipedia, established news, historical newspapers) > low (general web pages with little to assess) > unknown (nothing to assess at all — no recognisable source class, author, date, or domain signal). Academic papers with a DOI or PubMed link rate high even before their full details are fetched; Wikipedia rates medium — a solid consensus summary, deliberately not the equal of a primary or peer-reviewed source. The run also deliberately cites a variety of source classes where it can (an encyclopedia and an academic paper and a period newspaper beats three encyclopedia entries saying the same thing) — a claim corroborated across independent source classes is better verified than one corroborated by a single class. If a claim you know to be well-established lands lower than you'd expect, check the breakdown for the credibility tiers behind it — thin, hard-to-assess web sources are the usual reason.


Claim Evidence Matrix

The matrix at /research/[cafeId]/claims is the claim-centric view of everything your Cafe has checked — reached via Claims under Interrogate in the ✦ Deep picker (as of W1.44). It's free to view for every signed-in user — there's no additional research cost to look at claims you've already gathered evidence for.

What It Shows

For each claim, the matrix groups every source into three columns:

  • Supporting — sources that corroborate the claim
  • Contradicting — sources that dispute it
  • Silent — sources in your Cafe that don't address the claim at all

Each column is further grouped by credibility tier (high / medium / low / unknown), and a line under each claim reads something like "4 sources · 2 independent roots".

Independent Roots

"Independent roots" is a dedup count: if three news articles all quote the same original wire story, that's one independent root, not three. This stops corroboration from being inflated just because a story got picked up widely — a heuristic pass runs automatically (matching by DOI or normalized URL), and it's what independent-root counts on both the matrix and Find the Truth's evidence breakdown are based on.

Refine Independence (Optional, Costs Tokens)

On a claim with several same-domain or ambiguous sources, a "Refine independence" button lets you run a one-time AI pass that judges whether each ambiguous source reports original findings or just relays another one. It shows a cost estimate and requires your consent before running — nothing runs automatically. The result is saved, so re-running the same claim later doesn't re-charge you.

Where Claims Come From

The matrix populates from two paths:

  • Find the Truth runs — every claim you check with Find the Truth appears here automatically.
  • Verify claim runs — every claim you verify on a source card (see Verify Claim & Cited Report) also appears here. If you have an older Cafe with verify results from before this feature existed, the matrix builds itself from that history automatically the first time you open it — no action needed.

Deep Linking

Click a claim to focus it — the URL updates with ?claim=<id>, so you can copy the link, share it, or bookmark it, and it survives a page refresh. With more than one claim in a Cafe, a claim selector lets you pick which one to view instead of scrolling through all of them.


Tips

  • Read the finding, not just the number — "What we found" often explains an important nuance (e.g. "true, but only under different conditions than usually assumed") that the headline percentage alone can't convey.
  • Check the evidence breakdown before trusting a surprising result — if a confident-sounding claim scores lower than expected, look at the credibility tiers behind it. A basis of thin, hard-to-assess web sources is the most common reason.
  • A run takes a few minutes — that's normal. A deep evidence run usually takes 2–6 minutes: it searches, downloads and reads sources, then runs a second full pass hunting the best evidence against the claim. The progress panel names what each phase is doing, shows how long you've been waiting, and its token counter ticks up as the run spends — you can also navigate away and the run keeps going, its progress right where you left it when you come back.
  • Use Quick depth for a first pass — you can always re-run with more depth on a claim that matters.
  • The counter-evidence callout is worth reading even on a "true" verdict — it shows you the strongest case against, so you know what you'd need to rebut if challenged.
  • Use the matrix to see everything your Cafe has checked at a glance — viewing it never costs tokens, so it's worth checking any time you've been Verifying claims on source cards.
  • "Refine independence" is optional — the always-on heuristic dedup is usually good enough; reach for the AI refine pass only when a claim has several ambiguous same-domain sources and the root count looks off.

Availability

FeatureAvailabilityCost
Find the Truth (probability verdict)Every signed-in userCosts tokens (estimate shown before you run)
Claim Evidence Matrix (/research/[cafeId]/claims)Every signed-in userFree to view
Refine independence (AI provenance pass)Every signed-in userCosts tokens (estimate shown before you run)
Cited truth-verdict reportEvery signed-in user (view if shared, generate if you run it)Included in the Find the Truth run
Download report as PDF (fixed verdict → body → matrix → citations order, W54)Every signed-in user, incl. shared-report viewersFree, no tokens
Save to Notebooks — whole report (W1.17)Every signed-in user, owner onlyFree
Cited/searchable in Ask Your Cafe (W1.17)Every signed-in userFree

Find the Truth costs tokens because it's a relatively expensive background research run (two full research passes plus scoring). The Claim Evidence Matrix never costs tokens to view because it's a read view of evidence you (or your Cafe's Verify runs) already gathered.

See Also