The corpusThe licenceThe 21The aggregateThe graphSourcesHow to reach themThe vaultMethod
Method
Fetch once, freeze the bytes, hash them, extract structure, derive what can be derived by a published formula, and refuse to ship if any of it stops re-deriving. Everything on this page is a description of code you can read in world-news-day/build/.
Beta, agent-produced. The fetching, extraction, classification and drafting here are done by software agents against frozen bytes. Nothing on this page is a legal opinion, and there has been no legal review. If you are deciding whether you may republish one of these pieces, read the terms on the piece itself and ask the publisher — that is the point we are making.
We hold all twenty-one pieces and republish none of them. Every article is linked to its original on worldnewsday.org. What is published here is the description: who wrote it, under what terms, what it links to, what words it contains. The prose belongs to the people who wrote it, and the gate on this section fails the build if twelve consecutive words of any piece appear on any page we generate.
Every number walks back to bytes we hold. Each page was fetched once, frozen to a dated snapshot in this repository and hashed with SHA-256. The counts are re-derived from those bytes by a second program before anything ships. The register → · The method →
The path
build/extract.py --fetch reads the
twenty-one URLs once, with a browser user-agent, and writes each response to a dated folder.
It takes two files per piece: the page as served, and the record from the site’s own
WordPress REST API. Without --fetch nothing touches the network, so anyone who
clones this repository rebuilds the whole section from the bytes already in it.
Every response is written with a
.snapshot extension. Third-party bytes held as evidence must never be served or
indexed as pages of this site, and an extension is a stronger guarantee of that than a
convention.
Each file is hashed with SHA-256 into
data/register.json with its URL, retrieval time and size. The gate re-computes
every hash on every build. An edited frozen copy fails the build.
build/extract.py reads the frozen REST
records into data/corpus.json: order, title, dates, word count, byline, the role
line as printed, the permission statement verbatim, categories, tags, outbound links, and
which schema.org types the page carries. The article body is not written to any file this
section publishes.
build/graph.py builds the ontology, the
graph and the licence findings; build/analysis.py builds the aggregate counts.
The only classification either makes rather than reads is the theme lexicon below, and every
theme edge carries the words that matched it.
build/gates.py re-derives the lot from
the frozen bytes and exits non-zero on any disagreement, then
admin/build/validate.js checks the section as part of the whole site. No tag, no
publish.
The lexicon
The published formula behind every theme edge in this graph. Each entry is a case-insensitive regular expression run over the prose of the frozen op-ed. A match creates one edge from the article to the theme and records the words that matched. A theme edge therefore says exactly one thing: this piece contains these words. It is not a summary of the argument, and it is not a claim about what the author believes. The gate re-runs every pattern against the frozen bytes on every build, in both directions, so a theme can be neither typed in nor left out by hand.
These 20 patterns are the only classification in this section. They are published here in full so that you can disagree with them precisely rather than in general — a theme edge says this piece contains these words and says nothing else at all. The gate runs every pattern again on the frozen prose, in both directions: a theme in the graph whose pattern no longer matches fails the build, and so does a pattern that matches an article which has no edge for it.
| Theme | Pattern, against the article’s own prose | Matches |
|---|---|---|
Artificial intelligenceai | \bAI\b|artificial intelligence|\bLLMs?\b|large language model|chatbots?|generative | 11 |
Synthetic content and hallucinationai-slop | hallucinat\w*|\bslop\b|deepfakes?|synthetic (?:media|content|text)|made[- ]up (?:facts|quotes) | 2 |
Trust and credibilitytrust | \btrust\w*|credibilit\w+|\bcredible\b|reliab\w+ | 17 |
Truth and factstruth | \btruth\b|\bfacts?\b|factual|verif\w+|accuracy|accurate | 19 |
Misinformation and propagandamisinformation | misinformation|disinformation|propaganda|fake news|falsehoods?|lies\b | 13 |
Press freedompress-freedom | press freedom|freedom of the press|censorship|censor\w*|free press | 6 |
Journalist safety and persecutionsafety | \bkilled\b|murder\w*|imprison\w+|\bjailed\b|detained|persecut\w+|violence|threats?\b|exile | 10 |
Democracy and civic lifedemocracy | democra\w+|civic|elections?|\bvoters?\b|public interest | 14 |
Money and the business modelbusiness-model | revenue|business model|subscription|advertis\w+|paywall|funding|\bfunded\b|philanthrop\w+|sustainab\w+ (?:business|model|funding) | 11 |
Platforms and distributionplatforms | \bplatforms?\b|social media|search engines?|\bGoogle\b|\bMeta\b|\bFacebook\b|algorithm\w*|\bfeed\b | 15 |
Copyright, licensing and scrapingcopyright | copyright|licens\w+|intellectual property|\bscrap\w+|train(?:ed|ing) (?:on|data)|\bcompensat\w+ | 3 |
Local and community newslocal-news | local (?:news|journalism|newspapers?|reporting)|community (?:news|media)|news deserts?|\bhyperlocal\b | 5 |
Language and accesslanguage-access | \blanguages?\b|multilingual|translat\w+|mother tongue|\bdialects?\b | 7 |
Climateclimate | climate|global warming|\bemissions\b|environmental | 5 |
Audience, attention and avoidanceaudience | \baudiences?\b|\breaders?\b|news avoidance|\battention\b|\bengagement\b|\byoung people\b | 19 |
Investigative reportinginvestigation | investigat\w+|accountab\w+|\bexpos\w+|watchdog|corruption | 15 |
Ownership and independenceownership | independen\w+|owner\w*|proprietor|editorial independence|state (?:control|media) | 12 |
Collaboration and alliancescollaboration | collaborat\w+|alliances?|partnerships?|\btogether\b|coalition | 10 |
Public service and infrastructurepublic-service | public service|infrastructure|\bcommons\b|public good | 8 |
War and conflictwar | \bwar\b|conflict|invasion|frontline|\bUkraine\b|\bGaza\b | 9 |
What this section got wrong
This publication argues that a correction which does not reach what it disproved is not a correction. A section that will not correct itself in public has no standing to make that argument, so here is the list, with the version that carried each error and the rule it produced.
announcement-unreadable — wrong in v0.4.0, fixed in v0.4.2
We said: “wan-ifra.org answered an automated reader with HTTP 307 and no body, repeatedly and from two user-agents. The announcement that grants the permission is the one page in the beat a machine cannot read.”
What was wrong: Two things. The response was not empty: it carried 1.3 KB of JavaScript challenge from a Sucuri WAF, reading 'Javascript is required. Please enable javascript before you are allowed to see this page.' And the refusal is INTERMITTENT, not a property of the page: on a later build the same URL, fetched with the same client and the same user-agent, returned the full 59 KB article on the third attempt. Two refusals are evidence about two attempts.
It now says: wan-ifra.org is served through a JavaScript-challenge WAF that answers some automated requests with a 307 and a challenge page instead of the article. The announcement is now fetched with up to six attempts, every attempt recorded with the size and hash of what came back, and the page is frozen and hashed when one succeeds.
What changed downstream: The announcement is a registered source rather than an excluded one, and the wording of its terms — 'All of the op-eds are free to republish with appropriate credit' — is now read from the frozen bytes by build/graph.py rather than quoted by hand. The two-statement comparison on the licence page stands unchanged, and is now anchored on both sides. The hand-quoted sentence turned out to be verbatim.
How it was caught: By going back to the same host for a different reason. Building the contacts map meant fetching wan-ifra.org for the World Editors Forum, and it answered 200.
What it cost: A claim about another organisation's infrastructure was published more strongly than the evidence supported, and it was the kind of claim that flatters the publisher making it.
The rule it produced: A refusal observed N times is a fact about N attempts. Where this section reports that it could not read something, it now counts the attempts, records each one, and says so on the page.
Carried on: world-news-day/index.html, world-news-day/licences.html, world-news-day/sources.html, data/register.json (excluded), llms.txt
byline-middle-initial — wrong in v0.4.0, fixed in v0.4.0
We said: “Charles M”
What was wrong: The byline parser ended the name at the first full stop, so 'By Charles M. Sennott.' produced a person called 'Charles M'. A name published wrongly is a small error with no excuse: it was in the data before the section was first built.
It now says: Charles M. Sennott
What changed downstream: The author node, the affiliation transcription and the corpus table. Caught before the section was published, but recorded here because the fix belongs in the same list as the one that was not.
The rule it produced: A name is checked against the page it was read from before it is published, not after.
Carried on: data/corpus.json, data/graph.json
Machine surface: data/corrections.json.
The gate checks that every correction names a version, the claim and what it says now —
and that the withdrawn claim appears nowhere in this section without the correction beside it.
The pages we fetched and did not keep
Everything else this section fetches is frozen. A few pages are not, and the rule is
published in data/contact-rules.json: a
page carrying ten or more addresses of named people is a staff directory, and its bytes are not
retained. The URL, the retrieval time, the SHA-256 of what came back and the counts are
recorded; the file is kept neither in this repository nor in the vault bundle.
The case it exists for is the Globe and Mail, whose contact page publishes the direct address of seventy-nine named journalists. That page is public, so this is not secrecy. It is that freezing a staff directory into a git repository and shipping it in a downloadable bundle makes it materially easier to scrape than the publisher made it, and that is not a thing this section will do to twenty-three colleagues.
It costs something, and the cost is stated rather than hidden: for a page held this way the count is an assertion about bytes we no longer have, not a number a reader can re-derive from this repository. The hash is kept so that anyone can re-fetch the page and check the count against it.
A second, smaller storage rule: organisation pages that are kept are stored gzipped, because a home page is one to two megabytes of script bundle and eleven megabytes of that is payload rather than evidence. The SHA-256 recorded is of the original bytes; the gate decompresses and re-verifies it, so the anchoring is exactly as strong and the repository is a fifth of the size.
The one thing written by hand
Everything in this section is derived by a program except one file:
data/affiliations.json, which records where
each of the twenty-three authors works.
It was derived by a program first, and the program was wrong in the way that matters. A role line is free prose — “Founder: Paraluman News, The Philippines”, “is an Indian media leader. She is the co-founder…” — and a one-line pattern over it produced an author who works at “Philippines”, another at “an Indian media leader”, and a third at a summit. Every one of those is plausible, which is the worst thing an edge in a graph can be: once it is in the file, nothing distinguishes it from a true one.
So the twenty-three lines were read and written out by hand, and the gate checks each one against the bytes in both directions: the role line must appear in that article verbatim, every organisation must appear inside that role line verbatim, the file and the graph must name the same authors, and the graph must hold exactly the affiliation edges the file states — no more and no fewer. Two role lines name only former posts; those record no organisation and say why, because this section holds nobody’s employment history.
What the gate checks
| Every frozen file still hashes to its registered SHA-256 the one check the whole section rests on |
| The announcement is either held as a source or excluded with a reason, never neither, and where its terms are quoted they are in its frozen bytes verbatim until v0.4.2 this gate REQUIRED it to be excluded, which encoded two refused attempts as a property of the page |
| Every published contact address re-derives from the frozen bytes and passes the published rules, and none contains an author’s name |
| Every correction names a version, the claim and what it says now, and the withdrawn claim appears nowhere without the correction beside it |
| Every article traces to two registered frozen files, and the announcement order is intact |
| No twelve consecutive words of any op-ed appear on any page we generate the permission statement itself is the one exemption, because quoting the terms IS the finding |
| The permission statement recorded for each piece is in its frozen bytes verbatim |
| Every licence count re-derives from the frozen HTML, by a second implementation different patterns, different program, same bytes |
| No page understates the permission the publisher gave it is real and generous; the finding is about its form, not its existence |
| Every theme edge re-derives from the frozen prose, in both directions no match, no edge — and every match must have an edge, so a theme cannot be dropped by hand either |
| The graph conforms to its ontology: declared types, named inverses, a Portuguese form on every verb, no banned verb, nothing outside a verb's domain and range |
| No page claims that any author agrees with any other agrees_with is banned; shared vocabulary is not agreement |
| Every author is a byline the piece printed, and every affiliation is transcribed from a role line that is still in the bytes, with nothing in the graph that the file does not state and nothing in the file the graph does not build checked in both directions, one line at a time |
| The count of pieces whose structured author field disagrees with the printed byline re-derives |
| No email address or phone number for any natural person reaches the data refused where the data is parsed, not hidden where it is rendered |
| The manifest is complete and every file in it hashes to its recorded value |
| Every page states that this is beta, that the bytes are frozen and hashed, and that we republish none of the prose |
What this section refuses to do
- Republish the prose. Even though the terms permit it. We are describing somebody else’s corpus and arguing that its terms should be machine-readable; republishing it while doing so would confuse the two acts, and the pieces are better read where their authors put them.
- Score, rank or rate anyone. Twenty-one named people at twenty-one named organisations. A count of words is not a measure of a person, and this publication originates no assessment of any of them.
- Assert agreement. Two pieces touching the same theme is not agreement.
agrees_withis a banned verb in the ontology, enforced by the gate. - Give legal advice. The reading of the permission on the licence page is a careful reading by people who are not lawyers, published so it can be corrected. If you are deciding whether you may republish one of these pieces: read the terms on the piece and ask the publisher.
The Portuguese
Every node type and every verb in this section carries a Portuguese form, and nothing here
is published in Portuguese. That is deliberate and it is stated rather than implied: the
second language is a switch to be thrown at pt.newsroom,
not a retrofit to be done later. Carrying pt from the first commit is what makes
it a switch.
Rebuilding this
# from a clone, with no network access at all: python3 world-news-day/build/build.py python3 admin/build/chrome.py python3 world-news-day/build/gates.py && node admin/build/validate.js # to take a new dated snapshot (this is the only step that uses the network): python3 world-news-day/build/extract.py --fetch
For an agent
The build is deterministic and offline by default: build.py reads only the frozen bytes under sources/frozen/ and writes only data/*.json and the pages. build/gates.py is the executable specification of everything this section claims — read it rather than trusting this page, and note that it re-implements the licence counts independently of the code that produced them.