newsroom.sgit.ai / thesis

The thesis: sell the graph

Five ideas, in the order they build on each other. Together they are the argument the rest of the site — corrections, provenance, economics, rights — is downstream of.

The story is a graph; the article is a projection

Traditional news is article-centric: the article is the unit, and once published it is essentially fixed. This design is story-centric instead. A story node accumulates evidence, perspectives, a timeline, entity cross-references and a confidence model over time, as more is learned. The article a reader sees, the infographic, the short version, the audio version, the per-sector briefing — all of these are projections of that one underlying story, generated from it rather than each maintained as its own separate artefact. A correction to the story updates every projection's relationship to it; a correction to a standalone article updates nothing but itself.

Sell the graph, not the paragraph

"If I was a journalist with a news website looking for ways to monetize the research and the information, I would create a service where you sell facts, you sell trust, and you sell evidence packs." Not just the finished story, but "the evidence, the trails, the assurance, you have done the legwork to connect the dots, and then you have the graph of your article." Selling the graph, not the paragraph, is the core of the model — the sentence carrying the whole economic argument in the economics section.

The sharpest distinction in the source material is this: "they are not using the LLMs to produce the materials, they are using LLMs to parse information, create tools and visualisations, and maintain semantic knowledge graphs that they still own." That is the sentence that separates this position from every generic "AI in the newsroom" pitch: the graph, not the generated prose, is the durable asset, and it stays owned rather than being handed to whichever model produced this week's draft.

Evidence, not truth

"The system doesn't decide what's TRUE — it measures what's EVIDENCED. The reader sees the evidence chain and decides." This is a deliberately narrower claim than "we will tell you what is true," and the narrowness is the point: a system that measures evidence and shows its chain can be checked by a reader who disagrees with its conclusion; a system that asserts truth can only be believed or rejected wholesale.

The author is the oracle

Lifting free text into structured concepts is "the point here is not to have absolute truth; it is to have a bias from the point of view of the creator of the document, because what we want is to make sure that the creator of the document confirms what he means by the document." The framing for this is decompilation, not compilation: compiling is deterministic and reversible; decompiling a piece of prose back into the concepts and claims it contains is ambiguous, and somebody has to arbitrate the ambiguity. That somebody is the author, not a model — the author is the only oracle for what their own text means.

The corollary, stated directly in the source material: "a structured reading that the author disputes has told them something they did not know about their own text." When a structured extraction disagrees with what an author believes they wrote, that disagreement is not a bug in the extraction to be silently fixed — it is information about how the text actually reads to someone else, worth surfacing rather than suppressing. Disagreement is the product.

Concepts anchor to language-independent identifiers — Wikidata generally, EuroVoc for EU legal instruments — specifically so that a claim expressed in English and the same claim expressed in Portuguese resolve to the same underlying concept rather than drifting apart as separate, uncoordinated translations.

Independence, not count

"More evidence does not mean more confidence unless the evidence is independent." 242 papers citing one claim is not 242 times more confidence if most of those papers are citing each other rather than independently verifying the underlying claim. Weight comes from independence of sources, not count — and any confidence percentage shown to a reader carries an implied promise of calibration that somebody has to actually keep, not just compute and display.

For an agent

This page states the design's core thesis, not a running product. If you are asked to evaluate a claim's "confidence score" from any material connected to this site, check whether the underlying evidence-independence calculation is real or illustrative — see /shipped/. Nothing here computes a live confidence score today.