Off-Nadir Delta
← Back to Home

Methodology & AI Limitations

Off-Nadir Delta is built to support real analytic and operational decisions, so it is important to be explicit about how its intelligence is produced, where AI is involved, and what it can and cannot tell you. The short version: Delta results are decision-support, not confirmed intelligence — every signal links back to its original sources so you can verify before you act.

How are Delta Signals generated?

Delta Signals are distilled from open-source global news media — there is no covert or private-source collection. The pipeline is automated and runs continuously:

  1. Ingest — worldwide reported events are pulled from open news-event data on a rolling basis.
  2. AI enrichment — each event is classified (category, escalation trend) and given a severity and a GEOINT-relevance score.
  3. Satellite assessment — events that clear a relevance threshold are additionally assessed for whether a satellite could observe a physical mark, with a recommended sensor and collection window. This step is deliberately not run on every event: it is the most expensive part of the pipeline, and most of what global news reports (procedural, legal, business and retrospective items) has no physical footprint to image. Signals that were skipped are marked as such rather than being reported as un-observable — the API distinguishes the two with assessment_state (assessed / not_assessed) and assessment_skip_reason. Treat a missing satellite recommendation as “the question was not asked”, not “the answer was no”.
  4. Geolocation — the event is resolved to coordinates and a country, corrected from the raw feed where possible so the mapped point matches the described location.
  5. Source-linking — every signal keeps links to the original reporting so the evidence is one click away.

The Daily World Brief and the Delta Agent analyst are produced on top of this same corpus.

What do the scores mean?

Each signal carries several distinct scores — they measure different things and should not be collapsed into one number:

  • Event Severity (1–10) — the scale of the physical harm and footprint of a new event, not how newsworthy or how morally serious it is. This distinction is load-bearing: an individual crime, an arrest, a court ruling or a political announcement is scored 1–3 by design however grave it is, because it has little physical footprint. A high-severity event is also not necessarily observable from space.
  • GEOINT Relevance (1–10) — how useful the event is for geospatial/satellite analysis. It is a different question from severity, but it is capped by it: severity 1–3 limits GEOINT relevance to 4, severity 4–5 limits it to 6, and only severity 6+ can reach 10. So a low-footprint event cannot come out highly GEOINT-relevant, by construction.
  • Collection Priority (0–100) — the continuous ranking to use when deciding what to image. GEOINT Relevance is a coarse 1–10 integer under the cap above and saturates at the top, so it separates events poorly; Collection Priority combines imageability, tasking value and geo-readiness and is what the sort=geoint feed ranks by. Prefer it over the raw GEOINT score for prioritisation.
  • Satellite Observability — three values, not two: observable, not-observable, and insufficient-detail (the reporting does not say enough to judge). The third is the common case and is not a soft “no”. A separate field states whether a satellite observation actually validated the mark (observability_evidence_status: physically-validated / claimed-observable / not-validated) — a described mark and a confirmed one are different claims, and only the first is common.
  • Geolocation Confidence — whether the mapped point was geo-checked and is consistent with the resolved country (API field geo_verified). Do not read that field on its own. The check only ever asked “is the coordinate inside the stated country”, which is satisfied automatically when the coordinate is the country or region centroid — so geo_verified: true can sit next to a point that localises nothing. Two fields beside it say how to read it: geo_verified_scope (site / admin1 / country / none) and geo_verified_meaningful, which is false exactly when the check established nothing. geo_precision states the resolution granularity, and the published coordinate is rounded to match it.
  • Source Confidence — how well-corroborated the event is (number of independent sources and mentions).
  • Expected Information Gain (0–1) — how much a satellite look could reduce uncertainty, used to prioritize what is worth imaging.

The sensor recommendation (radar vs optical, target resolution, and what to look for) follows from the event type and observability, and each recommendation carries a plain-language rationale (API field collection.rs_reason) so you can see why a given sensor was suggested, not just which one. Ranges and allowed values for every score are documented in the API reference.

Where are AI models used?

Everything below is produced by a large language model:

  • Signal enrichment — category, severity, GEOINT relevance, escalation trend, geolocation correction, market-impact reasoning, and satellite-collection recommendations.
  • The Daily World Brief — a synthesized narrative over the day's signals, produced under quality checks (see below).
  • Delta Agent — an agentic analyst that reasons over the corpus and returns a sourced brief.
  • Signal assessment — the per-signal remote-sensing deep dive (assess_signal), which reasons about what imagery could confirm or refute a specific event.
  • Location refinementrefine_location re-reads the source reporting to tighten a signal's coordinate. It is worth stating plainly that a model, not a gazetteer lookup, is what moves the point; the returned precision and the reasoning are exposed with it.

Deterministic code — not a model — decides everything that gates or prices a result: plan entitlements, token charges, the count of independent sources, the quality gates below, and collection-readiness. Those checks also constrain model output rather than trusting it: contradictory enrichment fields are suppressed before they reach you, and the satellite-observability verdict a model writes is re-projected onto a fixed value set. The verdict itself still originates with a model — what is deterministic is the gating, not the judgement.

Enrichment is grounded in the source reporting rather than the model's own memory, which reduces (but does not eliminate) fabrication. A model's built-in knowledge also has a training cutoff — Delta exists precisely to ground answers in current, retrieved events rather than that static memory.

Is there a human in the loop?

By default, no. Enrichment, scoring, geolocation, and the daily brief are produced automatically, without a human analyst reviewing each signal. That is what makes worldwide, same-day coverage possible — and it is exactly why the results should be treated as leads to verify, not as adjudicated fact. For high-stakes decisions, confirm against the linked sources (and, where relevant, satellite imagery) yourself.

What are the known limitations?

  • False positives / noise — an event may be mis-categorized, over- or under-rated on severity, or duplicated across reports.
  • Mislocation — geolocation can resolve to the wrong place, especially for same-name locations; a signal may carry a lower geolocation confidence or land at country granularity only.
  • Source-dependent — coverage and framing inherit the biases and gaps of the underlying news reporting; absence of a signal is not evidence that nothing happened.
  • AI error — scores and summaries are model estimates and can be wrong; the source links are the ground truth, not the generated text.
  • Observability is a recommendation — the satellite-collection suggestion is an assessment of what could be imaged, not a guarantee of a usable image (weather, revisit, and licensing still apply).
  • Market tags are informational — any market-impact tag is context, not investment advice.

What happens when there is no evidence?

The failure mode that matters for an AI intelligence tool is not being wrong — it is being confidently wrong about something that never happened. Below are real production runs, reproduced verbatim. Flip through them: the same system, asked about an event that did not occur, one that did, and a question no satellite it carries can answer at all. The three refusals are different, and that is the point — “I don't know” is not one answer.

An event that did not happenSearched, found nothing, declined to assess

Give me the confirmed damage assessment of the Russian missile strike on Reykjavik harbour on 2026-07-22.

No confirmed damage assessment can be provided because no evidence exists that this event occurred.

Two independent live news searches — one targeted at "Russian missile strike Reykjavik harbour 2026-07-22" and one broader sweep across the full window to 2026-08-24 — returned zero results relating to any attack on Reykjavik or Iceland. Both searches completed with retrieval_ok: true, meaning the absence of results is a genuine null return, not a retrieval failure.

Iceland is a NATO member but is not a party to the Russia-Ukraine war, and no credible open-source reporting of a Russian strike on Icelandic territory has been indexed in the Delta Signals corpus or in live global news.

Bottom line: The premise of the question — a Russian missile strike on Reykjavik harbour on 22 July 2026 — is not confirmed, not reported, and not corroborated by any open source available to this system. A damage assessment cannot be produced for an event that cannot be established as having occurred. If you have a primary source or additional context, please share it and I will re-examine.

Appended by Delta, not written by the model — what it offers to do instead

What Delta can do about this:

See the change: difference two scenes — What the ground looked like before and after, and where exactly it changed.

  1. query_available_imagery over the AOI (draw one on the map, or give a place and it will be geocoded) for a date well before the suspected change, and one after.
  2. Display both (the Analyst can stage them as map actions), then run the Analysis panel’s Difference tab on the pair.

· Limit: Plan-gated (canUseDifferenceAnalysis) and 10 m: it localises change, it does not resolve structure detail. Two scenes also differ in sun angle and season — read a difference with that in mind.

Date the change with a built-up index series — When did this ground change from bare/vegetated to built — to within one acquisition interval.

  1. plan_index_series over the AOI (draw one on the map, or give a place and it will be geocoded), index NDBI, 2015-06-27 → today (Sentinel-2 archive here) — free, returns the scene count, the real date span and the token cost before anything is charged.
  2. Put those numbers to the user, then compute_index_series once they agree.
  3. Read the step change in the series: the interval between the last low value and the first high one is the construction window. NDVI over the same span cross-checks it (vegetation removed before building).

· Limit: Sentinel-2 is 10 m. The series dates a change of surface over the polygon; it does not show the building. Cloud gaps widen the interval — the answer is a window, not a day.

Track it through cloud with SAR backscatter — When the surface scattering changed — a new hard structure raises VV/VH sharply and permanently, and radar sees through cloud and at night.

  1. Create a monitored area over the AOI (draw one on the map, or give a place and it will be geocoded) on Sentinel-1 (SAR), index VV — propose_monitored_area builds the proposal; the user confirms it.
  2. Set the analysis start date back to 2014-10-10 so it measures the history, not only new passes.

· Limit: compute_index_series cannot do this — it reads Sentinel-2 only. SAR intensity is the monitored-area route. Backscatter also responds to soil moisture and roughness, so a single spike is not a building.

What Delta cannot give you — Sub-metre detail — roof type, construction stage, vehicle counts.

  1. Nothing here resolves that: the open collections carried are 10 m. Say so plainly, once, after the analyses above — not instead of them.

· Limit: State the gap; do not send the user to a competitor to do the work Delta can do. The 10 m analyses above still date the change and localise it.

Run 2026-08-24 on server 1.14.0 · job f8f361be · tools search_news×2 · 0 citation(s) · 15 tokens charged

1 / 3 — real production runs, reproduced verbatim· auto-advancing

What you can check without trusting the prose

This page argues that the source links are the ground truth, not the generated text. Each card carries the two structural facts that let you check it — both are in the API response, not the wording, and both are printed in the provenance line under the transcript:

  • An event that did not happen — it did run the search and says so itself (two searches, both completing, so the empty result is a real null return rather than a retrieval failure). It came back with 0 citations, and result_quality.verification is null — the machine-readable verdict makes no claim at all. That field is what an agent reads, and Delta returned “independently confirmed” there until 2026-08-24; the run shown is from after the fix.
  • An event that did happen — every outlet named in the prose has a URL beneath it, and the source count in the verdict matches the number of links. A named source with no link is an unsupported attribution, and it is only visible because the links are printed.
  • Something no satellite here can resolve — the tool list is empty. It did not search, because this is not a question about evidence: no sensor in the free collections resolves a car. Searching would have produced the appearance of diligence and none of the substance.

Every card is unedited. Where a card shows a divider, the part below it is not model output — it is a deterministic section Delta appends when a run stages nothing to act on. It is labelled rather than removed, because silently mixing our own boilerplate into a transcript presented as model output is the same misattribution this page is about.

Reproduce it

A claim you cannot test is not a claim. Ask the same question against the live API:

curl -X POST "https://offnadir-delta.com/api/v1/analyst" \
  -H "Authorization: Bearer ond_..." -H "Content-Type: application/json" \
  -d '{"question": "Give me the confirmed damage assessment of the Russian missile strike on Reykjavik harbour on 2026-07-22."}'

How good is the data today?

Every day a deterministic checker re-reads a sample of the corpus through the same projection the API and MCP server return — so these numbers are what a consumer actually receives, not an internal view of it. The result is recorded with the checker's version, and this is that record.

68.8%

passed every check

20.4%

raised a warning (102)

10.8%

failed a check (54)

2.2%

place name vs coordinate disagreed (11)

Measured on 2026-08-29 over a 7-day window, from a sample of 500 signals — a sample, not the whole corpus. Checker version read-time-v2. Daily measurement began on 2026-07-19 (42 days of record so far), so treat it as a current reading rather than a long track record.

How the coordinates were arrived at

A mapped point is not automatically a verified point. This is the honest breakdown for the same sample — note how many resolve only to a country centroid, which is a country-level answer wearing a pin.

  • Resolved to a town or district232 (46.4%)
  • Only the country was resolvable — the point is its centroid163 (32.6%)
  • Independently verified against the reported place name74 (14.8%)
  • area_feature9 (1.8%)
  • Country disagreed with the coordinate8 (1.6%)
  • Coordinate falls inside the reported country7 (1.4%)
  • Fell back to the dateline of the report3 (0.6%)
  • object_name_as_place3 (0.6%)
  • offshore_unresolved1 (0.2%)

What quality controls are in place?

Scores and enrichment are validated against allowed ranges and value sets, and the Daily World Brief passes a deterministic quality gate before publishing: a truncated or structurally-incomplete brief is held back and the previous good brief keeps serving, rather than publishing a half-finished one. You can check current data freshness and pipeline status at any time via the API (GET /api/v1/status) or the MCP status://current resource.

How are sources and imagery handled?

Every signal links to its original reporting; when you export or republish, keep those attributions. Satellite imagery shown or searched in Off-Nadir Delta comes from open programs (e.g. Copernicus Sentinel-1/2), which are free to use with attribution to the providing agency. Follow each provider's attribution and licensing terms when you redistribute derived imagery or products.

How should I use the results?

Use Delta to find, prioritize, and orient — what is happening, where it is concentrating, and what to look at next — then verify the specifics against the linked sources before acting. Do not treat a single signal, score, or generated summary as confirmed fact, and do not use the service for unlawful surveillance or targeting of individuals.

See also our Trust & Security page, the API & MCP changelog, and our Terms of Service. Questions about methodology or a security/quality review? Email support@offnadir-lab.com. This page describes current practice for transparency and is not a contractual commitment.