The Signal / AI in practice / October 4, 2026
AI hallucinations, three years on
Fewer inventions. Not zero.
Since 2023, AI has become far more reliable at working from facts you give it. Asked to answer from memory, it still invents comps, citations, and figures. Here is what changed, why the rest persists, and what to expect next.
When ChatGPT arrived, appraisers learned it would invent a plausible comp out of nothing. That lesson is still half right.
Many appraisers decided in 2023 that generative AI was too risky for professional work. Three years later, the evidence supports both those who distrust AI and those who rely on it. Models have improved substantially at one kind of task and only slightly at another.
You hand it the facts.
Paste in an engagement letter, a contract, or your own neighborhood data and ask for a summary or a draft. The model's job is to stay faithful to what you supplied.
It answers from memory.
“What did homes on this street sell for?” The model must retrieve a fact it may never have reliably learned, and it rarely says so.
The two figures come from different tests and are not a direct comparison. Both show the same pattern found across benchmarks: errors are far rarer when the facts are in front of the model. Vectara ↗ GPT-5 system card summary ↗
The good news
Given the facts, it mostly sticks to them.
Vectara's leaderboard measures how often a model adds unsupported information when summarizing a document it was given. At launch in November 2023, the spread between models was enormous. AI Business ↗
The 2026 figure uses a larger, harder set of articles, so it is not a like-for-like comparison. Many widely used 2026 models score between 3% and 4%. Vectara leaderboard ↗
Fabricated citations also fell sharply in one model generation. A 2023 study found that 55% of the references produced by GPT-3.5 did not exist, compared with 18% for GPT-4. Scientific Reports ↗ With browsing enabled, OpenAI reports that GPT-5 made factual errors in a much smaller share of its claims than GPT-4o or o3. GPT-5 system card summary ↗
The other half
Ask from memory, and the error rate barely moved.
With browsing switched off, OpenAI's own system card shows GPT-5 answering 40% to 47% of SimpleQA's short factual questions wrongly. Much of the reported improvement comes from looking facts up, not from a model that knows more. GPT-5 system card summary ↗
Artificial Analysis tested models on 6,000 difficult knowledge questions in November 2025. All but three were more likely to give a wrong answer than a correct one. Artificial Analysis ↗
A sharp drop, then a plateau. The figures come from two different studies, so treat the shape, not the exact gap. 2023 study ↗ 2025 study ↗
A reversal
The next model is not always more accurate.
OpenAI's own testing found that its o3 and o4-mini reasoning models gave false information about people more often than the earlier o1. OpenAI said the newer models make more claims overall, producing more correct and more incorrect ones, and that more research was needed to understand why. Summary of the o3 system card ↗
Not a simple bug
Some of the problem is built in.
Guessing gets rewarded.
A fact seen only once in training data cannot be learned reliably. Most benchmarks give no credit for “I don't know,” so confident guessing scores better, and training to be helpful made models more overconfident. OpenAI ↗
Familiar is not the same as known.
Models default to declining on subjects they do not recognize. A familiar name can switch that default off even when the model knows nothing reliable, so it answers anyway. A local street, builder, or subdivision can feel familiar enough. Research summary ↗
Zero is out of reach.
Researchers argued formally that no general-purpose language model can eliminate hallucination entirely, whatever its design or training data. arXiv ↗
Higher is better here. State your conclusion first and then ask the AI to check it, and it is more likely to agree. Stanford HAI ↗
In the real world
The mistakes are increasing, not shrinking.
Court decisions involving AI-fabricated legal citations, as of early October 2026. AI Hallucination Cases database ↗
More people now use AI for professional work, and checking has not kept pace. The problem extends well beyond courtrooms.
Deloitte Australia
Partially refunded a roughly AU$440,000 government report that contained a fabricated court quote and references to papers that do not exist.
AP ↗Air Canada
A tribunal held the airline responsible for a bereavement refund policy that its website chatbot invented.
AI Business ↗Chicago Sun-Times
Published a syndicated summer reading list in which 10 of the 15 recommended books did not exist.
Axios ↗Thomson Reuters disputed the methodology. In every incident above, a person used AI output without checking it. Stanford HAI ↗
Your desk
Let it write. Make it cite.
No independent study has yet measured AI error rates on appraisal tasks. The examples available point the same way as the evidence from law. A commercial appraisal writer asked ChatGPT for ground-lease comparables involving national tenants and received entirely fictional sales that looked convincing. RealWired ↗
The Appraisal Standards Board has proposed Advisory Opinion 41 on the use of technology, including AI. The draft emphasizes that the appraiser remains responsible for the work and must evaluate the data rather than trust the software, consistent with USPAP's Ethics, Competency, and Record Keeping Rules. TALCB ↗
Lower risk
- Drafting neighborhood or market commentary from your own data and notes
- Summarizing an engagement letter, contract, or document you provide
- Reorganizing or proofreading text you wrote
- Checking your draft against a source document you supply
Higher risk
- Asking for comparable sales, prices, or listing details from memory
- Asking for zoning rules, tax figures, flood data, or market statistics without a source you can open
- Asking for citations to guidelines, case law, or regulations
- Asking AI to confirm a conclusion you have already stated
AI report-writing tools marketed to appraisers generally limit AI to drafting from the data the appraiser supplies, with the appraiser reviewing the result. That is the grounded job where the evidence of improvement is strongest.
Use today
Three prompts that cut the risk.
How you ask changes how often AI invents an answer. These three techniques have the strongest evidence behind them. They reduce errors substantially; they do not remove the need to check the result.
- 01
Hand it the source, and let it say “I don’t know.”
Answering from a document you supply is the job AI does best, and permission to decline removes the pressure to guess. Asking for word-for-word quotes first ties every point to text you can check. Anthropic recommends all three steps.
Wrong answers, by job From memoryGPT-5, SimpleQA40%+From a supplied documentMany 2026 models, Vectara3–4%Different tests, so compare the scale, not the exact gap. Anthropic guidance ↗Vectara ↗GPT-5 system card summary ↗
PromptUse only the documents below. Do not use general knowledge. First, copy the word-for-word quotes that are relevant to the task and number them. Then [summarize the zoning restrictions / draft the neighborhood description], citing the quote number after each statement. If the documents do not contain the answer, write "Not in the documents." It is fine to say you don't know. <documents> [paste the zoning ordinance, contract, or your notes] </documents> - 02
Make it check its own answer, one claim at a time.
In Meta’s chain-of-verification study, a model drafted an answer, wrote a verification question for each fact, answered those questions separately, and then revised. Hallucinated items fell by more than three-quarters.
Hallucinated items per answer Plain answerFew-shot baseline2.95With verificationChain-of-verification0.68Wikidata list questions, Llama 65B, 2023 research setting. Chain-of-verification study ↗Anthropic guidance ↗
PromptReview your answer above. 1. List every factual claim it makes: numbers, dates, names, addresses, sales, and rules. 2. For each claim, write one question that would verify it. 3. Answer each question separately, using only the documents I provided. 4. Rewrite the answer. Remove any claim you could not confirm and mark where it was with [ ]. - 03
Ask neutrally. Never lead with your conclusion.
Models tend to agree with the person asking. Present a claim as your own belief and the model is far less likely to catch that it is false. Ask an open question about the evidence instead, and request the working.
GPT-4o catching a false claim “Someone claims…”Neutral framing98.2%“I believe…”The user’s own view64.4%Stanford AI Index 2026. Higher is better. Stanford HAI ↗
Instead ofI think prices here fell about 5% this year. Can you confirm?
AskUsing only the attached sales data, what was the change in median sale price between the first and third quarters? Show the calculation and list the rows you used. If the data is not enough to answer, say so.
Our view
Where this goes by late 2027.
This is our reading of the evidence, not a reported finding. Gartner predicts that more than 40% of agentic AI projects will be canceled by the end of 2027, with reliability among the reasons. OODA Loop, citing Gartner ↗
- 01Likely
Grounded work keeps improving.
Summarizing and drafting from documents you supply should get more reliable, with the best models making errors in roughly 1% to 2% of summaries. Tools built for this job become more useful.
- 02Likely
Recall is not solved by late 2027.
Rare facts are where training cannot close the gap. A particular sale price or a county zoning rule is exactly that kind of fact. Search helps, but adds a new question: did the AI find the right source?
- 03Likely
Safer products are narrower products.
Professional tools will feel safer mainly because vendors restrict what the AI may do. An agent that picks comps, adjusts them, and drafts a value conclusion is the riskiest thing to adopt early.
- 04Our rule
Verification beats model choice.
The useful question is not which model hallucinates least. It is whether your process requires every AI-supplied fact to be checked against a source before it reaches a signed report.
A practical rule
Let AI help write from information you provide. Verify every fact it supplies from its own memory or search results. A source you can open and confirm is evidence. Without one, the fact is unsupported.
Sources checked October 4, 2026. Benchmark figures come from the cited vendors and researchers, use different methods, and should not be compared across charts. Some figures are drawn from published summaries of system cards.
The Signal / Monthly briefing
Track AI in appraisal, with sources.
One sourced briefing in the first week of each month.
Unsubscribe in one click. We don't sell your email or share it with sponsors. Privacy Policy