Technical Guides7 minOctober 8, 2026

How AI-Generated Content Gets Fact-Checked Before It Publishes

A technical breakdown of a guardrail layer that automatically cross-checks every claim in AI-written text against its source and live search, fixes what it can't verify, and keeps a correction log — before a reader ever sees it.

How AI-Generated Content Gets Fact-Checked Before It Publishes

Text you shouldn't trust at face value

An AI model can confidently write a date, a figure, a company name, or a comparative conclusion — and all of it can simply be wrong. No warning signs, no "I'm not sure" — just smooth, convincing text with an invented detail buried inside. At the scale of regular publishing, that's not one typo, it's a reputational and compliance risk.

Here's how the layer that catches this before publication — not after a client complaint — actually works.

Layer one: generation anchored to a source

A generation request never goes out into the void — there's always an explicit designated source (a topic anchor and/or a source URL), and the model gets a direct instruction: if you're not certain about a date, company name, researcher's name, or figure, don't invent it — state it more generally instead. This cuts down fabrication at the source, but doesn't guarantee zero — which is why there's a second, independent layer.

Layer two: an independent fact-verification pass

A separate model call reads the already-written text and looks for claims not backed by the source anchor: numbers, dates, company and product names, statistics, quotes, comparative conclusions ("wins"/"loses", "better"/"worse"). For each one, it cross-checks against a primary source via live search, and either corrects it with accurate data or, if it can't be verified at all, rewrites it more generally — never swapping one invented detail for another.

// Simplified from the real verification pass
const prompt = `
Check the text for claims not present in the source topic
and not confirmed by web_search.
For each one: find correct data via search, or rewrite it
generally, removing the invented specific.
`;

A separate check covers the direction of comparative conclusions on numeric metrics: if a metric is "lower is better" (time, cost, errors) but the text frames a rising number as a win, that's a misread even when the figures themselves are correct — and it gets fixed too.

Every run returns not just a corrected text, but a list of the specific corrections made — a log reviewable independently of the article itself:

const CORRECTION_TOOL = {
  name: 'submit_corrected_article',
  input_schema: {
    type: 'object',
    properties: {
      content: { type: 'string' },      // full text, corrected or unchanged
      corrections: { type: 'array', items: { type: 'string' } }, // what changed
    },
    required: ['content', 'corrections'],
  },
};

One engineering detail explains why this is a separate HTTP call rather than another step in the same generation request: three sequential Claude calls (write → verify facts → self-review) inside one Vercel serverless function hit the execution limit and failed with 504 FUNCTION_INVOCATION_TIMEOUT. The fix wasn't "raise the timeout" — it was splitting the steps into separate HTTP requests: the dashboard calls fact-verification as its own request, after the main generation has already returned a draft.

Layer three: links aren't trusted, they're validated by code

A prompt instruction to "use the right links" is a request, not a guarantee. So a separate, non-AI layer parses every link in the final text against the real, current list of published pages for that language, and against the UK↔EN slug map:

// If a link points to an article that only exists in the other language,
// don't just strip it — try to resolve the same-language equivalent first
if (otherKnownSlugs.has(slug)) {
  const mappedSlug = getAltLocaleSlug(slug, urlLocale);
  if (knownSlugs.has(mappedSlug)) {
    return `[${text}](/${locale}/blog/${mappedSlug})`; // swapped to the right locale
  }
}
// no equivalent exists — better to drop the link than leave it broken
return text;
Verification layer What it catches Type
Prompt instruction Baseline fabrication at write-time AI (request)
Independent fact-verification pass Unconfirmed dates/figures/names/comparisons AI (audit)
Link validator Broken or wrong-locale links Code (deterministic)
Self-review (below) Negation/logic/punctuation defects AI (cheap model)

This layer caught a real production issue: before it existed, 291 links in English-language articles pointed at the Ukrainian version of the page instead of the English one — the model "believed" it was doing the right thing, and only a code-level check, not another instruction to the model, could actually catch that.

Layer four: a cheap model that looks for formal defects only

A separate, deliberately cheap pass on a smaller model (Haiku, not the same Sonnet that writes and fact-checks) re-reads the final text and looks for exactly three defect types — not style, not content:

  1. An extra or missing negation word that flips the sentence's meaning
  2. A logical inconsistency within a sentence
  3. A punctuation defect — a question missing a "?", an unclosed mark
const REVIEW_TOOL = {
  name: 'submit_review',
  input_schema: {
    properties: {
      ok: { type: 'boolean' },          // true = no defects found
      fixedContent: { type: 'string' },  // full text, only if ok=false
    },
    required: ['ok'],
  },
};
// safety guard: a correction is only accepted if the fixed text is at
// least 70% of the original length — otherwise it's discarded and the
// original text ships without the self-review edit
if (ok === false && fixedContent.length >= content.length * 0.7) { /* ... */ }

This layer exists because of a specific incident, not a theory: one already-published article had an extra negation word before a conjunction in a localized sentence that flipped the sentence's meaning. None of the other three layers catches this class of error — fact-verification looks at claims, not at negation grammar.

Bilingual: each language runs all four layers independently

Translation isn't treated as a proxy for correctness. Each language version runs its own independent fact-verification pass and its own self-review — one language doesn't inherit "trust" from the other just because it's a translation of it.

Why this isn't just "another AI call on top"

The difference is in the architecture of responsibility. The model that writes and the model that verifies are separate calls with separate instructions — the verifier never sees the writer's "intent," only the finished result and the source. And the link validator isn't AI at all — it's plain code, which can't be talked into anything by a well-phrased sentence. More on combining content-matching with fabrication-proofing: Hybrid RAG: Precise Search and No AI Hallucinations.

Why this actually matters

AI hallucination risk isn't an abstraction for tech blogs. For real cases where an invented AI detail nearly caused serious consequences, see AI Hallucination: Military and Corporate Risk Lessons. A related approach — applied to a system auditing its own quality — is in another technical breakdown: Inside a Self-Analyzing AI Lead-Qualification Chatbot.

Result in practice

This layer already runs inside a live bilingual publishing system: 145+ articles in each of two languages, every one with its own fact-verification and link-correction log.

Have questions? Ask the AI agent right now

Responds in seconds, knows everything about our services and will help with your situation