How It Works
The Narrative Gap monitors news coverage from a curated set of sources across multiple regions and languages. Our system identifies what is being claimed, who is saying it, and how consistently those claims hold up across independent reporting.
Multi-source analysis
Rather than relying on any single outlet, we look at how a story is reported across different media ecosystems. Claims that appear independently in multiple regions carry more weight than those from a single source or media tradition. When coverage is heavily concentrated in one region, we flag that so readers can judge accordingly.
Competing explanations
For contested questions — where reasonable people disagree about what is happening or why — we track multiple explanations and assess how well each one holds up against the available evidence. This helps readers see the full landscape of possibilities rather than being funnelled toward a single narrative.
Confidence levels
Our confidence assessments reflect how strongly the available evidence supports a claim — how strong and consistent that evidence is, weighed for its quality and independence. Confidence is not a count of how many outlets carried a story: a single strong, uncontested source can warrant high confidence, while many copies of the same report add little. How widely a claim was reported is a separate signal, shown alongside the claim. We use plain language rather than misleading percentages:
- High confidence — the evidence strongly supports this claim
- Moderate confidence — the evidence leans toward this claim, but leaves real room for doubt
- Low confidence — the evidence is thin, weak, or mixed
- Very low confidence — the evidence we have weighs against this claim
- Not scored — we haven't evaluated the evidence for this claim yet. A claim can appear here while it's new, or while it's waiting in the analysis queue. Come back later, or check other claims in the same event for context.
Two ways we describe likelihood
Event pages describe competing explanations two ways at once, because they answer two different questions:
- Rank ("Leading", "Most likely", "Less likely", "Least likely") compares explanations against each other — which one the evidence currently favours relative to its rivals.
- The word in parentheses — one of the seven bands below — says how believable that explanation is on its own, independent of whether it's currently in the lead.
A field of weak explanations can have a "Leading" one that's still unlikely in absolute terms — both readings are honest at once. We never call an explanation "Most likely" unless it also clears a minimum bar on the absolute scale; a weak leader is labelled "Leading" instead.
How likely an explanation is, on its own
The seven bands, from most to least likely:
- Almost certain
- Very likely
- Likely
- Possible — roughly a coin flip
- Unlikely
- Very unlikely
- Almost certainly not
For events we treat as graver — where getting it wrong matters more — we hold competing explanations to a stricter standard before calling one "likely" or better.
How we phrase a question's verdict
When a question has enough signal to say something, we phrase it one of four ways:
- "Evidence suggests: …" — one explanation clearly leads the field.
- "Evidence is split — … leads slightly" — a leader exists, but the field remains genuinely competitive.
- "Evidence is contested — competing explanations remain open" — for graver events, the same split state without naming a lead when the case for doing so isn't strong enough.
- "No clear answer yet" — no explanation has enough support yet to say anything more specific; we say so plainly instead of guessing.
How clearly a source states something
Separately from confidence and likelihood, we also note how directly the underlying source material states a claim — "very clearly stated", "clearly stated", "less clearly stated", or "ambiguously stated" by the source. This is about how cleanly the claim was extracted from the reporting, not about how believable the claim itself is.
Intent and motive assessments
When we assess what an actor "likely" intends or what their motives may be, these are analytical inferences drawn from patterns in reported evidence — not allegations or statements of personal knowledge. Evidence may be incomplete or misleading, and our inferences may be wrong. They represent the best reading of available reporting, not established fact.
What we don't do
- Our assessments are evidence-based — when we say something is "likely", that reflects the weight of independent reporting, not opinion or speculation
- We don't take sides — we present the evidence landscape
- We don't rely on any single source or algorithm
- We don't make allegations — our assessments are analytical inferences, not accusations
How we measure confidence
Our confidence in a claim reflects how strongly the available evidence supports it. We weigh the quality of that evidence, whether it is independent and mutually corroborating, and whether any of it points the other way — and for higher-stakes claims we set a higher bar. Confidence is not a headcount of outlets: a claim carried by a single strong, uncontested source can reach high confidence, and a story echoed by many outlets that all trace back to one report is not thereby more certain. How widely a claim has been reported is shown separately, alongside the claim, so you can weigh reach and confidence as two distinct things.
Checking our predictions
When a claim we track makes a specific, checkable prediction, we follow up once it's due and record what actually happened, using three labels:
- Borne out — what was predicted happened.
- Did not hold — what was predicted did not happen.
- Partly borne out — what happened was close to, but not exactly, what was predicted.
Our track record page shows the current state of that checking — including if we've had to withdraw judged outcomes. Not every prediction can be checked before it expires, so the judged set is a subset of everything we track, and it skews toward the predictions we were most confident about, since that's the set we prioritise checking. An expired, unchecked prediction is never counted as wrong.
Limitations
- We can only analyse what gets published. If something isn't reported by any of our sources, we won't know about it.
- Our source coverage is uneven. Some regions and languages are better represented than others.
- Evidence quality varies. Official statements from governments aren't necessarily more truthful than investigative reporting — we weight by independence and corroboration, not by authority.
- Our system scores evidence automatically. It can miss nuance, context, and connections that a human analyst would catch.
- Predictions are probabilistic estimates, not certainties. Even high-confidence assessments can be wrong.
All analysis is advisory. Confidence assessments reflect the balance of available evidence at the time of analysis and may change as new information emerges.