Off-Nadir Delta
OSINTtradecraftevidenceAI analysisverification

Auditing What an AI Told You: Keeping a Claim Ledger

Kazushi MotomuraAugust 23, 20266 min read
Auditing What an AI Told You: Keeping a Claim Ledger

Quick Answer: The failure mode of AI-assisted analysis is not the wrong answer, it is the unre-examinable one: by the time it matters, nobody can reconstruct what was claimed or how well it was sourced. A claim ledger records each assertion with its evidence class (confirmed, reported, party claim, assessment), the sources behind it, and whether corroboration is genuinely independent or traces back to one claimant. It also links restatements, so when a later answer restates something with weaker evidence, that downgrade is visible rather than silent. Grading source reliability and information credibility separately — the two-axis approach in NATO's STANAG 2511 — is what keeps 'unreliable outlet, corroborated fact' from collapsing into the same bucket as 'reliable outlet, single report'.

Ask an AI system about an unfolding event and you get a fluent, plausible answer. Ask it again three days later and you get another one. Whether the second contradicts the first, and whether the confidence behind either was ever justified, is usually unrecoverable — nobody kept the first answer in a form that could be compared.

That gap is the real risk in AI-assisted analysis. Not hallucination, which people already watch for, but the ordinary case where an answer was reasonable, got restated, quietly weakened, and no one noticed.

What is a claim ledger?

A claim ledger records the individual assertions an analytic system made, separately from the prose it made them in. Each entry holds the claim text, when the event occurred and when the answer was generated, the sources cited, how many of those sources are genuinely independent, and an evidence class saying what kind of knowledge it is. The prose is the delivery; the ledger is what can be checked.

The separation matters because prose hides structure. A paragraph can assert six things at four different confidence levels and read as one confident statement. Split into rows, the weak claim stops borrowing credibility from the strong one beside it.

How should claims be classified by evidence?

By what kind of knowledge each one is, before any judgement about how likely it is to be true. Four classes cover almost everything in event analysis:

ClassMeaningWhat it implies
ConfirmedEstablished by independent reportingCan be relied on, with the sources named
ReportedStated by media, not independently establishedUsable, but not settled
Party claimAsserted by a party to the eventInterested testimony — treat as such
AssessmentThe system's own judgementNot evidence at all; inference

The class that gets skipped most often is party claim. A belligerent's casualty figure, a company's statement about its own incident, a ministry's account of its own operation — these are evidence that the claim was made, and the distinction disappears the moment they are folded into a summary as facts. Keeping them in their own class is the whole point of evidence-based event intelligence.

Assessment deserves the same discipline in reverse. An inference presented in the same voice as a sourced fact is the most common way an analytic product overstates what it knows.

Why grade the source and the information separately?

Because collapsing them loses the information you need to act. "An unreliable outlet reporting something that is independently corroborated" and "a reliable outlet carrying a single uncorroborated report" are very different situations with different next steps — chase the original reporting in the first case, find a second source in the second — and a single confidence label makes them look identical.

Intelligence practice solved this long ago with a two-axis grade. NATO's STANAG 2511, implemented in AJP-2.1, rates source reliability with a letter from A (completely reliable) to F (reliability cannot be judged), and information credibility with a number from 1 (confirmed by other sources) to 6 (truth cannot be judged), evaluated independently and combined as a pair such as B2. Ninety years on, the reason it survives is the reason a single score fails: the two axes tell you different things about what to do next.

What does "independently corroborated" actually require?

More than a count of links. Three outlets running the same wire story are one source wearing three hats, and three articles that all quote the same ministry spokesman are one claimant. Corroboration means the evidence chains reach different origins, not that the URLs differ.

This is worth checking mechanically rather than by eye, because the failure is systematic. Language like "multiple independent sources" or "independently confirmed" is exactly the wording that appears when a model has seen many articles, and it is often false in the specific sense that matters: every cited family traces back to one claimant's statement. An assertion of corroboration that the underlying sources do not support is worse than no assertion, because it manufactures confidence rather than merely lacking it.

The useful record therefore separates the number of citations from the number of independent evidence chains, and states which basis applies — independent reporting, a single claimant's statement, or no sources at all.

Why track restatements and downgrades?

Because the interesting failure is the one that happens over time. A claim asserted as confirmed on Monday and restated as merely reported on Thursday is the system contradicting itself, and unless the two are linked, the contradiction is invisible: both answers looked confident on the day they were given.

Linking each claim to what it restates makes three questions answerable that otherwise are not. Did the evidence get stronger, weaker, or just more detailed? Which of my earlier conclusions rested on a claim that has since been downgraded? And was any of this ever settled, or has it been circulating at the same weak evidence class for two weeks while sounding more certain each time?

The downgrade case is the one worth filtering for directly. It is the shortest path to "what did I tell someone that I should now correct" — a question with a real answer, and one that analytic products almost never let you ask. Analytic standards for key judgments cover the other half: expressing the confidence honestly in the first place.

What does this cost you?

Reading your own audit trail should be free, and it is here. Charging for access to the record of what you were already told would undermine the point of keeping it — auditability that is metered gets skipped precisely when budgets are tight, which is when it matters most.

The real cost is discipline at write time: the classification has to happen when the answer is produced, not reconstructed afterwards. A ledger assembled later is a second interpretation of the prose, with all the same problems.

What a ledger does not give you

It records what was asserted and how well it was sourced. It does not make the claim true. A well-sourced claim from reliable outlets can still be wrong, and the ledger will faithfully record that it was well-sourced.

Nor does it substitute for imagery or ground truth. Reported activity and observable change are different evidence, and the strongest analysis pairs them — which is why an assertion about damage carries more weight when satellite imagery over the area is consistent with it, and why an assertion nobody can observe should stay labelled as reported.

For the vocabulary used here, the glossary covers evidence class, corroboration, and confidence. The applied version of this discipline is described in sensor tasking tradecraft.


Evidence classes and source grades describe the state of the record at the time an answer was produced. They are a statement about sourcing, not a verdict on what happened.

Kazushi Motomura
Kazushi Motomura

Remote sensing specialist with 10+ years in satellite data processing and AI. Founder of Off-Nadir Lab. Master's in Earth System Science and Technology (Kyushu University). Co-author, Remote Sensing Encyclopedia. More about the author →

From headline to satellite evidence

One connected intelligence workflow across four surfaces — free to start, no GIS software or remote-sensing background required.