Which Satellite Observation Would Actually Refute Your Hypothesis?
Quick Answer: An observation is only worth collecting if its outcome could eliminate an explanation you currently hold. The CIA's own tradecraft primer says evidence 'consistent with only one hypothesis' is diagnostic, while evidence that 'could support more than one or indeed all hypotheses' is 'unimportant to determining which hypothesis is more likely correct.' Off-Nadir Delta now computes this before a collection: it decomposes each competing statement into the observables that would test it, then ranks those observables by how many statements their ABSENCE would refute. The inference runs one way only — an absent observable refutes every statement requiring it, while a present one refutes nothing. The answer separates what is available today, what must be bought elsewhere, what no overhead collection settles at all, and what every statement already predicts and therefore cannot separate.
An observation is worth collecting only if its outcome could eliminate an explanation you currently hold. That is a different question from when can this place be imaged, and it is the one that decides whether the imagery budget buys anything. Off-Nadir Delta now answers it directly: give it the competing explanations, and it returns the observation whose absence would refute the most of them — and, in the same answer, the observations that every explanation already predicts and that therefore cannot separate them.
This post explains the reasoning, the direction the inference runs in, and where it stops.
What makes an observation diagnostic?
An observation is diagnostic when its outcome is different under different hypotheses. If every explanation on the table predicts the same thing, seeing that thing eliminates nothing, however striking the image is. The CIA's Center for the Study of Intelligence states the rule directly in A Tradecraft Primer: Structured Analytic Techniques for Improving Intelligence Analysis (March 2009):
The "diagnostic value" of the evidence will emerge as analysts determine whether a piece of evidence is found to be consistent with only one hypothesis, or could support more than one or indeed all hypotheses. In the latter case, the evidence can be judged as unimportant to determining which hypothesis is more likely correct.
The same primer defines Analysis of Competing Hypotheses (ACH) as the "identification of alternative explanations (hypotheses) and evaluation of all evidence that will disconfirm rather than confirm hypotheses." The emphasis on disconfirmation is not stylistic. It is the whole mechanism, and it is what most collection planning quietly inverts.
Why does evidence that fits everything feel decisive?
Because fit is easy to see and diagnosticity is not. A satellite image that shows exactly what you expected is vivid, specific and immediately legible, and none of those properties are evidence. If four explanations for a disrupted shipping lane all predict that traffic moved off its usual corridor, then imaging the corridor and finding displaced traffic confirms all four at once — which is to say it ranks none of them.
This is the failure ACH exists to prevent, and it survives contact with satellite imagery unusually well, because imagery arrives with a strong impression of having settled something. The correction is mechanical rather than attitudinal: compute which observations differ across the hypotheses before choosing what to collect.
Which direction does the inference actually run?
Only one direction is sound, and it is the unglamorous one.
A hypothesis, decomposed, yields the observables that would have to be present for it to be true — a necessary condition. From that, exactly one inference follows:
| Outcome | What follows |
|---|---|
| The required observable is absent | Every hypothesis requiring it is refuted (denying the consequent) |
| The required observable is present | Nothing is refuted |
The second row is the one that gets written wrong. It is tempting to say that observing something rules out the explanations that did not predict it, but a hypothesis that does not require an observable is not thereby predicting its absence. It is simply silent about it. An observation that comes back positive leaves the field exactly as it was.
So the correct ranking metric is not "how evenly does this split the hypotheses" but "how many hypotheses die if this turns out not to be there." Off-Nadir Delta ranks by that number and states the asymmetry in every response: this measurement can eliminate, it cannot confirm.
How do you get from a statement to an observation?
Through an explicit middle term. A claim does not decompose into a sensor; it decomposes into the things that would have to be seen, and only then into instruments:
statement → required observables → data source → what it settles / what it leaves open
Skipping the middle term is how sensor-first reasoning produces confident nonsense. "We hold radar, radar sees vessels, therefore radar can settle a claim about vessels" is true at every step and false as a conclusion, because nothing in the chain asked what would have to be observed for the claim to be true or false. A claim about vessels being turned back over five weeks decomposes into a trajectory change, an event time, an event position, a vessel identity and an interaction — of which overhead imagery addresses part of one, and the five-week total none at all.
Each observable carries its own strength, and the strength of the whole test is the weakest required observable, never the strongest:
- Direct — the observation settles the claim (flood extent, a burn scar, a collapsed span)
- Indirect — it tests something consistent with the claim (a blockade shows as vessel clustering)
- Not observable — nothing overhead settles it (a boarding, an intent, a registry fact, a running total)
What does the answer look like in practice?
Consider two competing explanations for the same disruption: naval forces rerouted shipping away from the strait, versus shipping halted because the port was closed. Decomposed and ranked, the result separates into four kinds of answer:
| Verdict | Observable | What it means |
|---|---|---|
| Available here | How full a port is | If berths are not occupied as the closure account requires, that account is refuted |
| Obtainable elsewhere | A vessel changed course | Real and decisive, but it comes from transponder tracks, not pixels |
| Not settled by observation | When it happened | Imagery brackets an event between two passes; it does not time it |
| Separates nothing | Traffic moved off its usual lane | Both accounts predict it. Collecting it produces evidence consistent with each |
The last row is the point of the exercise. Route displacement is the most intuitive thing to go and image, it is well within the reach of free radar, and it would have told you nothing about which account was right.
What if nothing you can collect would settle it?
Then that is the answer, and it is worth more than a collection plan. Three distinct outcomes are kept apart rather than collapsed into "unavailable":
- Obtainable elsewhere. The right instrument exists and is simply not this product — vessel identity comes from transponder broadcasts or a registry, not from resolving a hull. Naming the correct instrument is part of an honest answer even when it is not on offer; presenting a coarser substitute silently would be claiming a capability that does not exist.
- Not settled by observation. A boarding lasting minutes, an intention, a five-week cumulative total. No acquisition contains a total: one image is one instant, and a figure accrued over weeks is arithmetic over events most of which no scene ever saw.
- Separates nothing. Collectable, cheap, and diagnostically empty.
Timing constrains the first outcome even when the observable is available. The Copernicus Sentinel-1 constellation has a 6-day exact repeat cycle with two satellites and a 12-day repeat cycle with one, and Sentinel-2 offers a 5-day revisit at the Equator at 10 m in its four highest-resolution bands. Those figures are the reason "when it happened" is not an imagery question: an event bracketed between two passes days apart is not timed by them.
How does this map to formal analytic standards?
It implements two requirements that IC analytic standards already state. ICD 203, Analytic Standards — signed 2 January 2015, with a technical amendment effective 21 January 2022 — lists nine Analytic Tradecraft Standards. Two are directly relevant.
Standard 3 requires that products "properly distinguish between underlying intelligence information and analysts' assumptions and judgments," and specifies that products "should, as appropriate, identify indicators that, if detected, would alter judgments."
Standard 4 requires that products "incorporate analysis of alternatives," which it defines as "the systematic evaluation of differing hypotheses to explain events or phenomena," adding that products "should identify and assess plausible alternative hypotheses."
Naming the indicator that would alter a judgment is normally a prose exercise. What is different here is that for the subset of indicators that are physically observable from orbit, the indicator, its instrument, its availability and its limits can be derived rather than asserted.
ICD 203 also fixes the vocabulary for expressing likelihood, mapping terms to probability bands — almost no chance at 01–05% through nearly certain at 95–99% — and instructs analysts "not to mix terms from different rows." Off-Nadir Delta uses that same vocabulary when a judgment is recorded against a watch, so a stated likelihood means the same thing here as it does in a finished intelligence product.
Where is this in Off-Nadir Delta?
In three places, all running the same computation:
- In the app, on a watched event, under the claim history. It takes no input: the statements already recorded about that event are used directly, and the answer names which observation could break the most of them at once.
- Over the REST API, at
GET /api/v1/discriminators. Pass two to eight competing statements, or an event id to test the claims already on your key. - Over MCP, as the
test_hypothesestool, so an AI assistant connected to Off-Nadir Delta can ask the question mid-conversation.
The response separates competing hypotheses (mutually exclusive, where the goal is to separate them) from a jointly held set of statements (not exclusive, where the goal is to falsify the most at once). The distinction matters because it inverts the answer: an observable required by every hypothesis has no diagnostic value in the first case and is the best available test in the second.
What this does not do
It does not tell you which hypothesis is true. It ranks observations by their power to refute, which is a claim about the questions, not about the world.
It does not confirm. A positive result leaves every hypothesis standing, and any reading of the output that treats a successful collection as support for a preferred explanation has inverted the method.
It can also fail to decompose a statement at all — some assertions yield no observable, and that is reported separately from "nothing would settle this." A decomposition failure is a limit of the method, not a finding about the world, and collapsing the two would let a gap in the analysis read as a fact about the event.
Related reading
- Why This Sensor? How the Delta Agent Explains Its Imagery-Tasking Choices — the sensor-selection reasoning that runs once you know what to observe
- Glossary — SAR, off-nadir angle, GRD, revisit and the rest of the vocabulary used above
- API and MCP documentation — the full request and response contract

Remote sensing specialist with 10+ years in satellite data processing and AI. Founder of Off-Nadir Lab. Master's in Earth System Science and Technology (Kyushu University). Co-author, Remote Sensing Encyclopedia. More about the author →