ZEILX.AI — Verification: Method & Practice
← Verification
Method & Practice

Checking
the Work

Verification fails quietly — one missed check and a bad source travels all the way into a finished report. This is a working method for building redundancy into that process: four passes that fail differently, plus a notation for how much weight a source can carry.

Four verification layers

  1. Existence — is it real?
  2. Provenance — where did it originate?
  3. Corroboration — does anyone independent confirm it?
  4. Disconfirmation — what would prove it wrong?

Reliability tiers

T1Primary / authoritative T2Credible secondary T3Supporting / contextual T4Flagged / use with caution

§ 01What can go wrong

Research assisted by large language models introduces a failure mode that traditional citation practice never had to guard against: a source can be fluently described and entirely fabricated. A plausible author, a real-sounding journal, a DOI-shaped string — none of it existing.

That is the new problem. It sits on top of three older ones that predate AI and are just as corrosive to a research record: the single-source cascade, where a dozen articles repeat one unverified claim until repetition is mistaken for confirmation; authority drift, where a blog post and a court record are cited with the same visual weight; and confirmation bias, where evidence that fits the thesis is gathered and evidence against it is quietly never sought.

Each of those failures is invisible in a finished document. A clean reference list looks identical whether every entry was verified or none were. What follows is not a fix for that — no method eliminates it — but a way of arranging the checking so that a failure slipping past one pass is likely to be caught by another, and so a reader can see which checks were actually run.

§ 02Two working principles

The method rests on two ideas, neither of them original here. Both are borrowed from disciplines that have been dealing with unreliable information far longer than research writing has.

Redundancy through independent checks. Any single verification step has a blind spot. The response is not a better single check but several checks whose weaknesses do not overlap — so a failure that slips past one is caught by the next. This is the logic of layered defense: in safety engineering it is James Reason's "Swiss cheese" model, where multiple imperfect barriers are arranged so their holes rarely align. The four passes below are meant to fail differently on purpose. That redundancy is the whole point; the tiering and the appendix exist to support it.

Reliability as a graded scale, not a binary. A source is not simply "good" or "bad." Evidence-based medicine formalized this decades ago with levels of evidence — a systematic review outranks a case report, which outranks expert opinion — because the strength of a conclusion should track the strength of what supports it. Intelligence analysis reached the same conclusion independently: the Admiralty Code, developed by British naval intelligence and later standardized by NATO, has graded source reliability on a lettered scale since the Second World War. The T1–T4 scale applies that reasoning to general research, making the evidentiary weight of every citation explicit so a reader can see what is carrying a claim.

Neither idea is new, and nothing here claims to be. This is assembly, not invention: established practice from intelligence analysis, evidence-based medicine, archival scholarship, and journalism, arranged into one sequence and written down so it can be followed consistently and criticized openly. The value, if there is any, is in the arrangement and the disclosure — not in the parts.

§ 03The four passes

Applied in order. Each looks for a different kind of failure — which is what makes them redundant in the useful sense rather than simply repetitive.

01

Existence

Catches → fabrication

Every cited work is confirmed to exist as described — author, title, outlet, and year corroborated against a retrievable record, ideally a live URL that has actually been fetched. A citation that cannot be resolved to a real, locatable source is removed or flagged, never carried on faith.

Grounded in: basic bibliographic verification, elevated to a formal first step because model-generated references can be fully plausible and fully invented. A category of automated checkers now performs this against scholarly registries such as Crossref and OpenAlex; those tools are treated here as instruments for this layer, not substitutes for it — they confirm that a source exists and that its metadata is faithful, but not that it supports the claim it was cited for. That question belongs to the three layers below.

02

Provenance

Catches → the single-source cascade

Each significant claim is traced to its origin rather than to whoever last repeated it. If five outlets cite one study, the study is what gets cited and tiered — not the echo. The original artifact, its author, and the standing of its outlet are assessed directly.

Grounded in: the archival principle of provenance and the historian's primary/secondary/tertiary distinction; operationally close to the "trace claims to the original context" move in Mike Caulfield's SIFT method for digital source evaluation.

03

Corroboration

Catches → the isolated or misreported claim

Each load-bearing empirical claim is checked against at least one independent account that confirms it as reported. Independence is the operative word — two outlets drawing on the same wire story are one source, not two.

Grounded in: methodological triangulation (Denzin) — the practice of confirming a finding across independent lines of evidence — and journalism's "discipline of verification" (Kovach & Rosenstiel), which distinguishes reporting from assertion.

04

Disconfirmation

Catches → confirmation bias

Each claim is actively searched for what would undermine it — failed replications, retractions, author disavowals, methodological critiques, standing scholarly disputes. What is found is disclosed rather than omitted. A claim that survives a genuine attempt to break it is stronger for it; a claim that doesn't is flagged before it reaches the reader.

Grounded in: Popper's falsification principle — a claim earns confidence by surviving attempts to refute it, not by accumulating agreement — and the scientific norm of organized skepticism (Merton). It is a deliberate counter to the human tendency to seek only confirming evidence.

§ 04The reliability scale — T1 to T4

The tier is a judgment about a source's authority for the specific claim it supports. The same outlet can sit at different tiers for different claims — authority is contextual, not a fixed property of a masthead.

T1
Primary / authoritative

The thing itself, not a description of it: peer-reviewed studies, government and institutional data, court records, official statistics, an organization's own original statements.

EARNS T1 → direct primary status, formal review or official standing, and a resolvable original document.

T2
Credible secondary

Established reporting built on primary data: major outlets with editorial standards, reputable trade publications, expert analysis by a named, credentialed author.

EARNS T2 → editorial accountability, a traceable evidentiary basis, and a named author or institution answerable for accuracy.

T3
Supporting / contextual

Useful but weaker: opinion pieces, advocacy sources, single-author expert blogs, industry communications. Fit for color or context, or to corroborate a claim already carried elsewhere.

EARNS T3 → relevant perspective or expertise, but limited independence, review, or evidentiary transparency.

T4
Flagged / use with caution

Aggregators, sources with a clear agenda, unverified or contested material. Included only when explicitly marked — often to show what a claim's weak origin actually is.

EARNS T4 → a reason to be in the record despite unreliability, always paired with a disclosure of why it is flagged.

The load-bearing rule. No T3 or T4 source carries a key claim on its own. A weak source can support a strong claim only alongside independent corroboration at T1 or T2. This one rule is what keeps the notation from being decorative — without it, tiering is just labeling.

The tiering criteria draw on standard source-evaluation frameworks — the CRAAP test's authority and accuracy dimensions, the ACRL Framework for Information Literacy's principle that authority is constructed and contextual, and the SIFT method's emphasis on investigating a source before trusting it.

§ 05How a source is credentialed

The path from a raw citation to a tiered, verified entry — the sequence a person, or an assisting workflow, works through for each source before it goes into a report.

Resolve it (Existence)

Confirm the work exists as described against a retrievable record. Unresolvable → removed or explicitly flagged. This gate runs before any tiering.

Trace it (Provenance)

Follow the claim to its origin. Reassign the citation from the repeater to the original source. Assess the outlet's standing and the author's accountability.

Assign a provisional tier

Place the source at T1–T4 for the specific claim it supports, using the criteria in §04. Contextual, not automatic — the same outlet may tier differently elsewhere.

Corroborate load-bearing claims

For any claim the report leans on, confirm it against an independent account. Apply the load-bearing rule: no key claim rests on a lone T3/T4 source.

Attempt disconfirmation

Search for retractions, failed replications, disavowals, and live disputes. Record what is found. Downgrade the tier or flag the claim as contested where warranted.

Log it in the appendix

Write the final entry: identifier, tier, what was verified, the origin, and any flag. The citation identifier is then carried unchanged from appendix into the body text.

§ 06The verification appendix

Every report ends with an appendix that documents the process rather than just listing references. Each entry pairs a stable identifier with its tier, a short record of what was verified, the origin the claim was traced to, and any flag. The identifier appears unchanged next to the claim in the body, so a reader can move from any statement to its verification record and back.

This is what makes the work auditable: the reasoning behind each citation is visible, not buried. Contested claims are logged separately so a reader can weigh them, rather than being silently dropped or silently trusted.

KC-01 (T1)  verified: institutional report, figures cross-checked to source PDF  ·origin: primary report
USAT-02 (T2)  verified: outlet database, salary figures corroborated by two independent trade reports  ·origin: USA TODAY database
BIRD-01 (T4)  verified: attribution confirmed  ·flag: social-media artifact — object of study, not evidentiary support  ·origin: requester-supplied

Illustrative format. Identifiers are project-specific and human-readable so the appendix stays legible without a decoder.

§ 07Instruments

No single tool serves all four passes, and it would be a mistake to expect one to. The redundancy argument applies to instruments as much as to checks — each tool below is strong where the others are weak.

Finding sources. For academic literature, Elicit searches roughly 138 million papers through the Semantic Scholar corpus using semantic rather than keyword matching, and links every summary back to the source paper. Because it retrieves from a real corpus rather than generating from memory, anything it returns exists — which satisfies much of the first pass at the point of discovery. Its limits are worth stating plainly: the index thins out for government technical reports, small conference proceedings, and grey literature, and it does not cover journalism, court records, or regulatory filings. It is the right instrument for doctoral literature work and close to the wrong one for investigative reporting, where most sources are not papers.

Checking them. Crossref and OpenAlex resolve citations and carry retraction metadata, which makes the first pass mechanically checkable. Scite classifies citations as supporting, contrasting, or mentioning — the closest thing available to an automated fourth pass, since it surfaces where a finding has been challenged. Consensus filters by journal quality ranking, which is a useful sanity check against a tier assignment but not a substitute for making one.

General web search remains necessary for provenance and for corroboration across source types — tracing a claim past the outlets repeating it is a crawling problem, not a corpus problem.

The source ledger

Designed · in build

Verification done once and written nowhere does not accumulate. The ledger is a two-table record — one row per source, one row per use of that source — that a scheduled n8n workflow re-checks on a cycle: does the link still resolve, has a retraction notice appeared, has the version changed. When something moves, every claim in every report that depended on that source is flagged for review automatically.

This is the part that makes living document mean something checkable rather than aspirational. A source verified in March can be retracted in August, and nothing about the original verification would reveal it.

The boundary in the schema. Ledger columns are split into two classes. A workflow may write only mechanical facts — link status, retraction flag, version, date last checked — and may mark a source needs review. It may never assign a tier, never mark a source still-useful, and never clear a claim for publication. Detecting that something changed is not the same as knowing what the change means for a particular claim. That judgment stays with the researcher, and the schema is built so that a breach of the boundary would be visible.

Status, stated plainly: the four passes described above are current working practice and are applied by hand. The ledger and its monitoring workflow are designed and being built; they are not yet running. The full ledger design is published here, and this page will be updated when the workflow is live.

§ 08What this method is not

A verification method that oversold itself would fail its own disconfirmation pass. So, plainly:

not infallible

The tiers encode judgment, and judgment can be wrong. Redundancy lowers the rate at which errors survive; it does not eliminate them.

not automated clearance

Workflow nodes produce drafts and suggestions. No source is "cleared" by a machine — the human review step is load-bearing, not ceremonial.

not peer review

This is a working research practice for independent and coursework writing, not a substitute for formal peer review or editorial fact-checking.

not settled

It is one possible arrangement of well-established checks, not the correct one. It is published in order to be tested, criticized, and revised — and it changes when something better is found.

not a truth machine

It credentials sources and traces claims. It cannot certify that a well-sourced claim is ultimately correct — only that it is honestly and transparently supported.

single-axis, for now

T1–T4 grades the source. The Admiralty Code grades source reliability and information credibility on two independent scales, on the reasoning that a reliable source can still report something wrong. This method currently handles that by tiering per claim rather than per outlet. Whether to adopt a second axis is an open revision.

§ 09Foundations

Everything here traces to an existing tradition. Listed so the borrowing is visible and checkable:

  • Layered defenseJames Reason — the "Swiss cheese" model of layered, independently-failing barriers (Human Error, 1990).
  • DisconfirmationKarl Popper — falsification as the test of a claim (The Logic of Scientific Discovery); Robert K. Merton — organized skepticism as a norm of science.
  • CorroborationNorman K. Denzin — methodological triangulation (The Research Act); Bill Kovach & Tom Rosenstiel — the discipline of verification (The Elements of Journalism).
  • ProvenanceArchival provenance and the primary/secondary/tertiary source distinction; Mike Caulfield — the SIFT method for online source evaluation.
  • Tier scaleThe Admiralty Code / NATO intelligence grading system (AJP-2.1, STANAG 2511), which grades source reliability A–F and information credibility 1–6; levels-of-evidence hierarchies from evidence-based medicine (incl. GRADE); the CRAAP test (CSU Chico, Meriam Library); the ACRL Framework for Information Literacy for Higher Education.
  • Existence toolingAutomated citation-verification tools that check references against scholarly registries — Crossref, OpenAlex, Semantic Scholar — developed in response to documented fabrication rates in LLM-generated bibliographies.
  • Public disclosureThe Trust Project's Trust Indicators, particularly Citations and References and Methods — the precedent for publishing a verification standard rather than merely following one.

By this method's own logic, this list is provisional until each reference is run through the four passes and tiered in the appendix. A page about verification should be the first thing verified.