ValuSignal

The Signal / AI in practice / October 4, 2026

AI hallucinations, three years on

Fewer inventions. Not zero.

Since 2023, AI has become far more reliable at working from facts you give it. Asked to answer from memory, it still invents comps, citations, and figures. Here is what changed, why the rest persists, and what to expect next.

Benchmarks, incidents and standards, 2023 to 2026Scroll to explore the charts9 chapters
01 / Two jobsThe distinction that explains the data

When ChatGPT arrived, appraisers learned it would invent a plausible comp out of nothing. That lesson is still half right.

Many appraisers decided in 2023 that generative AI was too risky for professional work. Three years later, the evidence supports both those who distrust AI and those who rely on it. Models have improved substantially at one kind of task and only slightly at another.

Grounded

You hand it the facts.

Paste in an engagement letter, a contract, or your own neighborhood data and ask for a summary or a draft. The model's job is to stay faithful to what you supplied.

<2%Best 2026 models, summary errors on Vectara's leaderboard
Recall

It answers from memory.

“What did homes on this street sell for?” The model must retrieve a fact it may never have reliably learned, and it rarely says so.

40%+GPT-5 wrong answers on OpenAI's SimpleQA, browsing off

The two figures come from different tests and are not a direct comparison. Both show the same pattern found across benchmarks: errors are far rarer when the facts are in front of the model. Vectara ↗ GPT-5 system card summary ↗

02 / What improvedGrounded summaries and citations

The good news

Given the facts, it mostly sticks to them.

Vectara's leaderboard measures how often a model adds unsupported information when summarizing a document it was given. At launch in November 2023, the spread between models was enormous. AI Business ↗

Summaries with unsupported content, November 2023Vectara HHEM leaderboard · lower is better
Google PaLM-ChatEarly Bard
27.2%
Claude 2Anthropic
8.5%
GPT-3.5 TurboOpenAI
3.5%
GPT-4OpenAI
3.0%
Best model, 2026Harder test set
1.8%

The 2026 figure uses a larger, harder set of articles, so it is not a like-for-like comparison. Many widely used 2026 models score between 3% and 4%. Vectara leaderboard ↗

Fabricated citations also fell sharply in one model generation. A 2023 study found that 55% of the references produced by GPT-3.5 did not exist, compared with 18% for GPT-4. Scientific Reports ↗ With browsing enabled, OpenAI reports that GPT-5 made factual errors in a much smaller share of its claims than GPT-4o or o3. GPT-5 system card summary ↗

03 / From memoryRecall without a source

The other half

Ask from memory, and the error rate barely moved.

With browsing switched off, OpenAI's own system card shows GPT-5 answering 40% to 47% of SimpleQA's short factual questions wrongly. Much of the reported improvement comes from looking facts up, not from a model that knows more. GPT-5 system card summary ↗

Artificial Analysis tested models on 6,000 difficult knowledge questions in November 2025. All but three were more likely to give a wrong answer than a correct one. Artificial Analysis ↗

Wrong answers, browsing offOpenAI SimpleQA
GPT-4o2024
52%
GPT-5 main2025
47%
GPT-5 thinking2025
40%
Fabricated referencesShare of AI-generated citations that do not exist
55%GPT-3.52023 study
18%GPT-42023 study
19.9%GPT-4o2025 study

A sharp drop, then a plateau. The figures come from two different studies, so treat the shape, not the exact gap. 2023 study ↗ 2025 study ↗

04 / Newer, not betterThe reasoning-model surprise

A reversal

The next model is not always more accurate.

OpenAI's own testing found that its o3 and o4-mini reasoning models gave false information about people more often than the earlier o1. OpenAI said the newer models make more claims overall, producing more correct and more incorrect ones, and that more research was needed to understand why. Summary of the o3 system card ↗

Answers about people containing false informationOpenAI PersonQA · lower is better
16%
o1Late 2024
33%
o3April 2025
48%
o4-miniApril 2025
05 / Why it persistsThree research findings

Not a simple bug

Some of the problem is built in.

01 / OpenAI, Sept 2025

Guessing gets rewarded.

A fact seen only once in training data cannot be learned reliably. Most benchmarks give no credit for “I don't know,” so confident guessing scores better, and training to be helpful made models more overconfident. OpenAI ↗

02 / Anthropic, 2025

Familiar is not the same as known.

Models default to declining on subjects they do not recognize. A familiar name can switch that default off even when the model knows nothing reliable, so it answers anyway. A local street, builder, or subdivision can feel familiar enough. Research summary ↗

03 / Mathematical proof, 2024

Zero is out of reach.

Researchers argued formally that no general-purpose language model can eliminate hallucination entirely, whatever its design or training data. arXiv ↗

It tends to agree with youGPT-4o catching a false claim · Stanford AI Index 2026
“Someone claims…”Attributed to a third party
98.2%
“I believe…”Presented as the user's view
64.4%

Higher is better here. State your conclusion first and then ask the AI to check it, and it is more likely to agree. Stanford HAI ↗

06 / The incidentsBetter benchmarks, more failures

In the real world

The mistakes are increasing, not shrinking.

2,145

Court decisions involving AI-fabricated legal citations, as of early October 2026. AI Hallucination Cases database ↗

~719
Jan 2026
~1,600
Jun 2026
2,145
Oct 2026

More people now use AI for professional work, and checking has not kept pace. The problem extends well beyond courtrooms.

2025

Deloitte Australia

Partially refunded a roughly AU$440,000 government report that contained a fabricated court quote and references to papers that do not exist.

AP ↗
2024

Air Canada

A tribunal held the airline responsible for a bereavement refund policy that its website chatbot invented.

AI Business ↗
2025

Chicago Sun-Times

Published a syndicated summer reading list in which 10 of the 15 recommended books did not exist.

Axios ↗
Specialist tools still errLegal research queries with false or misleading answers · Stanford, 2024
Lexis+ AI, Practical LawMore than
17%
Westlaw AIMore than
34%

Thomson Reuters disputed the methodology. In every incident above, a person used AI output without checking it. Stanford HAI ↗

07 / For appraisersStandards and safe uses

Your desk

Let it write. Make it cite.

No independent study has yet measured AI error rates on appraisal tasks. The examples available point the same way as the evidence from law. A commercial appraisal writer asked ChatGPT for ground-lease comparables involving national tenants and received entirely fictional sales that looked convincing. RealWired ↗

The Appraisal Standards Board has proposed Advisory Opinion 41 on the use of technology, including AI. The draft emphasizes that the appraiser remains responsible for the work and must evaluate the data rather than trust the software, consistent with USPAP's Ethics, Competency, and Record Keeping Rules. TALCB ↗

Lower risk

  • Drafting neighborhood or market commentary from your own data and notes
  • Summarizing an engagement letter, contract, or document you provide
  • Reorganizing or proofreading text you wrote
  • Checking your draft against a source document you supply

Higher risk

  • Asking for comparable sales, prices, or listing details from memory
  • Asking for zoning rules, tax figures, flood data, or market statistics without a source you can open
  • Asking for citations to guidelines, case law, or regulations
  • Asking AI to confirm a conclusion you have already stated

AI report-writing tools marketed to appraisers generally limit AI to drafting from the data the appraiser supplies, with the appraiser reviewing the result. That is the grounded job where the evidence of improvement is strongest.

08 / Three promptsTechniques with measured effects

Use today

Three prompts that cut the risk.

How you ask changes how often AI invents an answer. These three techniques have the strongest evidence behind them. They reduce errors substantially; they do not remove the need to check the result.

  1. 01

    Hand it the source, and let it say “I don’t know.”

    Answering from a document you supply is the job AI does best, and permission to decline removes the pressure to guess. Asking for word-for-word quotes first ties every point to text you can check. Anthropic recommends all three steps.

    Wrong answers, by job
    From memoryGPT-5, SimpleQA
    40%+
    From a supplied documentMany 2026 models, Vectara
    3–4%

    Different tests, so compare the scale, not the exact gap. Anthropic guidance ↗Vectara ↗GPT-5 system card summary ↗

    Prompt
    Use only the documents below. Do not use general knowledge.
    
    First, copy the word-for-word quotes that are relevant to the task and number them.
    Then [summarize the zoning restrictions / draft the neighborhood description], citing the quote number after each statement.
    
    If the documents do not contain the answer, write "Not in the documents." It is fine to say you don't know.
    
    <documents>
    [paste the zoning ordinance, contract, or your notes]
    </documents>
  2. 02

    Make it check its own answer, one claim at a time.

    In Meta’s chain-of-verification study, a model drafted an answer, wrote a verification question for each fact, answered those questions separately, and then revised. Hallucinated items fell by more than three-quarters.

    Hallucinated items per answer
    Plain answerFew-shot baseline
    2.95
    With verificationChain-of-verification
    0.68

    Wikidata list questions, Llama 65B, 2023 research setting. Chain-of-verification study ↗Anthropic guidance ↗

    Prompt
    Review your answer above.
    
    1. List every factual claim it makes: numbers, dates, names, addresses, sales, and rules.
    2. For each claim, write one question that would verify it.
    3. Answer each question separately, using only the documents I provided.
    4. Rewrite the answer. Remove any claim you could not confirm and mark where it was with [ ].
  3. 03

    Ask neutrally. Never lead with your conclusion.

    Models tend to agree with the person asking. Present a claim as your own belief and the model is far less likely to catch that it is false. Ask an open question about the evidence instead, and request the working.

    GPT-4o catching a false claim
    “Someone claims…”Neutral framing
    98.2%
    “I believe…”The user’s own view
    64.4%

    Stanford AI Index 2026. Higher is better. Stanford HAI ↗

    Instead of

    I think prices here fell about 5% this year. Can you confirm?

    Ask
    Using only the attached sales data, what was the change in median sale price between the first and third quarters?
    
    Show the calculation and list the rows you used. If the data is not enough to answer, say so.
09 / The next yearThe Signal's assessment

Our view

Where this goes by late 2027.

This is our reading of the evidence, not a reported finding. Gartner predicts that more than 40% of agentic AI projects will be canceled by the end of 2027, with reliability among the reasons. OODA Loop, citing Gartner ↗

  1. 01Likely

    Grounded work keeps improving.

    Summarizing and drafting from documents you supply should get more reliable, with the best models making errors in roughly 1% to 2% of summaries. Tools built for this job become more useful.

  2. 02Likely

    Recall is not solved by late 2027.

    Rare facts are where training cannot close the gap. A particular sale price or a county zoning rule is exactly that kind of fact. Search helps, but adds a new question: did the AI find the right source?

  3. 03Likely

    Safer products are narrower products.

    Professional tools will feel safer mainly because vendors restrict what the AI may do. An agent that picks comps, adjusts them, and drafts a value conclusion is the riskiest thing to adopt early.

  4. 04Our rule

    Verification beats model choice.

    The useful question is not which model hallucinates least. It is whether your process requires every AI-supplied fact to be checked against a source before it reaches a signed report.

A practical rule

Let AI help write from information you provide. Verify every fact it supplies from its own memory or search results. A source you can open and confirm is evidence. Without one, the fact is unsupported.

Sources checked October 4, 2026. Benchmark figures come from the cited vendors and researchers, use different methods, and should not be compared across charts. Some figures are drawn from published summaries of system cards.

The Signal / Monthly briefing

Track AI in appraisal, with sources.

One sourced briefing in the first week of each month.

Unsubscribe in one click. We don't sell your email or share it with sponsors. Privacy Policy

Found this useful? Share it