How AI-Generated Content Gets Fact-Checked Before It Publishes
A technical breakdown of a guardrail layer that automatically cross-checks every claim in AI-written text against its source and live search, fixes what it can't verify, and keeps a correction log — before a reader ever sees it.

Text you shouldn't trust at face value
An AI model can confidently write a date, a figure, a company name, or a comparative conclusion — and all of it can simply be wrong. No warning signs, no "I'm not sure" — just smooth, convincing text with an invented detail buried inside. At the scale of regular publishing, that's not one typo, it's a reputational and compliance risk.
Here's how the layer that catches this before publication — not after a client complaint — actually works.
Layer one: generation anchored to a source
A generation request never goes out into the void — there's always an explicit designated source (a topic anchor and/or a source URL), and the model gets a direct instruction: if you're not certain about a date, company name, researcher's name, or figure, don't invent it — state it more generally instead. This cuts down fabrication at the source, but doesn't guarantee zero — which is why there's a second, independent layer.
Layer two: an independent fact-verification pass
A separate model call reads the already-written text and looks for claims not backed by the source anchor: numbers, dates, company and product names, statistics, quotes, comparative conclusions ("wins"/"loses", "better"/"worse"). For each one, it cross-checks against a primary source via live search, and either corrects it with accurate data or, if it can't be verified at all, rewrites it more generally — never swapping one invented detail for another.
// Simplified from the real verification pass
const prompt = `
Check the text for claims not present in the source topic
and not confirmed by web_search.
For each one: find correct data via search, or rewrite it
generally, removing the invented specific.
`;
A separate check covers the direction of comparative conclusions on numeric metrics: if a metric is "lower is better" (time, cost, errors) but the text frames a rising number as a win, that's a misread even when the figures themselves are correct — and it gets fixed too.
Every run returns not just a corrected text, but a list of the specific corrections made — a log reviewable independently of the article itself:
const CORRECTION_TOOL = {
name: 'submit_corrected_article',
input_schema: {
type: 'object',
properties: {
content: { type: 'string' }, // full text, corrected or unchanged
corrections: { type: 'array', items: { type: 'string' } }, // what changed
},
required: ['content', 'corrections'],
},
};
One engineering detail explains why this is a separate HTTP call rather than another step in the same generation request: three sequential Claude calls (write → verify facts → self-review) inside one Vercel serverless function hit the execution limit and failed with 504 FUNCTION_INVOCATION_TIMEOUT. The fix wasn't "raise the timeout" — it was splitting the steps into separate HTTP requests: the dashboard calls fact-verification as its own request, after the main generation has already returned a draft.
Layer three: links aren't trusted, they're validated by code
A prompt instruction to "use the right links" is a request, not a guarantee. So a separate, non-AI layer parses every link in the final text against the real, current list of published pages for that language, and against the UK↔EN slug map:
// If a link points to an article that only exists in the other language,
// don't just strip it — try to resolve the same-language equivalent first
if (otherKnownSlugs.has(slug)) {
const mappedSlug = getAltLocaleSlug(slug, urlLocale);
if (knownSlugs.has(mappedSlug)) {
return `[${text}](/${locale}/blog/${mappedSlug})`; // swapped to the right locale
}
}
// no equivalent exists — better to drop the link than leave it broken
return text;
| Verification layer | What it catches | Type |
|---|---|---|
| Prompt instruction | Baseline fabrication at write-time | AI (request) |
| Independent fact-verification pass | Unconfirmed dates/figures/names/comparisons | AI (audit) |
| Link validator | Broken or wrong-locale links | Code (deterministic) |
| Self-review (below) | Negation/logic/punctuation defects | AI (cheap model) |
This layer caught a real production issue: before it existed, 291 links in English-language articles pointed at the Ukrainian version of the page instead of the English one — the model "believed" it was doing the right thing, and only a code-level check, not another instruction to the model, could actually catch that.
Layer four: a cheap model that looks for formal defects only
A separate, deliberately cheap pass on a smaller model (Haiku, not the same Sonnet that writes and fact-checks) re-reads the final text and looks for exactly three defect types — not style, not content:
- An extra or missing negation word that flips the sentence's meaning
- A logical inconsistency within a sentence
- A punctuation defect — a question missing a "?", an unclosed mark
const REVIEW_TOOL = {
name: 'submit_review',
input_schema: {
properties: {
ok: { type: 'boolean' }, // true = no defects found
fixedContent: { type: 'string' }, // full text, only if ok=false
},
required: ['ok'],
},
};
// safety guard: a correction is only accepted if the fixed text is at
// least 70% of the original length — otherwise it's discarded and the
// original text ships without the self-review edit
if (ok === false && fixedContent.length >= content.length * 0.7) { /* ... */ }
This layer exists because of a specific incident, not a theory: one already-published article had an extra negation word before a conjunction in a localized sentence that flipped the sentence's meaning. None of the other three layers catches this class of error — fact-verification looks at claims, not at negation grammar.
Bilingual: each language runs all four layers independently
Translation isn't treated as a proxy for correctness. Each language version runs its own independent fact-verification pass and its own self-review — one language doesn't inherit "trust" from the other just because it's a translation of it.
Why this isn't just "another AI call on top"
The difference is in the architecture of responsibility. The model that writes and the model that verifies are separate calls with separate instructions — the verifier never sees the writer's "intent," only the finished result and the source. And the link validator isn't AI at all — it's plain code, which can't be talked into anything by a well-phrased sentence. More on combining content-matching with fabrication-proofing: Hybrid RAG: Precise Search and No AI Hallucinations.
Why this actually matters
AI hallucination risk isn't an abstraction for tech blogs. For real cases where an invented AI detail nearly caused serious consequences, see AI Hallucination: Military and Corporate Risk Lessons. A related approach — applied to a system auditing its own quality — is in another technical breakdown: Inside a Self-Analyzing AI Lead-Qualification Chatbot.
Result in practice
This layer already runs inside a live bilingual publishing system: 145+ articles in each of two languages, every one with its own fact-verification and link-correction log.
Have questions? Ask the AI agent right now
Responds in seconds, knows everything about our services and will help with your situation
You might also like
Inside a Self-Analyzing AI Lead-Qualification Chatbot
A technical breakdown of a production AI chatbot: deferred mounting for PageSpeed, input validation tuned on real conversations, UTM attribution, and a built-in quality-analytics layer — with real production numbers.
Technical GuidesRAG vs CAG: Which AI Architecture Fits Your Business
RAG vs CAG explained for business leaders: how each architecture works, when to use which, and a real case with numbers to guide your AI decision.
Technical GuidesQwen3Guard: Real-Time AI Content Safety
Alibaba's free open-source Qwen3Guard filters toxic tokens in real time. Here's how SaaS teams and CTOs can deploy it to protect LLM-powered products.
