What AI hallucinations actually look like on a live deal
A hallucination happens when a large language model (LLM) like ChatGPT, Claude, or Gemini generates output that sounds confident and plausible but is factually wrong. The model isn't lying — it's pattern-matching against its training data and filling gaps with statistically likely text. In real estate, "statistically likely" and "contractually correct" are not the same thing.
Stanford and Vectara hallucination leaderboard studies consistently show that even the best LLMs hallucinate on 2–5% of factual claims in controlled benchmarks. On unstructured real estate documents — scanned PDFs, photographed addenda, handwritten initials — the rate climbs because the model has less clean text to work with.
Most agents we talk to either over-trust AI output because it "looks right" or abandon the tools entirely after one bad experience. Both reactions cost time. The better move is understanding the specific failure patterns so you verify the right fields and let the rest flow.
Four hallucination types that matter in transactions
Not all AI errors are equal. A wrong square footage number in a listing draft and a wrong inspection deadline in a contract summary have completely different detection difficulty and downstream consequences. Here's a taxonomy built from the error patterns we've observed across agent workflows.
| Hallucination type | What it looks like | Detection difficulty | Consequence severity |
|---|---|---|---|
| Date hallucination | Closing date off by 3–7 days, inspection deadline pulled from wrong clause | Medium — numbers look plausible | High — missed contingencies, potential breach |
| Entity hallucination | Buyer name swapped with seller, wrong brokerage on amendment summary | Low — names are easy to spot-check | High — compliance violation, re-signing delays |
| Clause fabrication | AI invents a contingency term or misquotes addendum language | High — sounds legally plausible | Very high — E&O exposure, contract disputes |
| Numerical drift | Purchase price off by a digit, earnest money amount rounded incorrectly | Medium — requires source comparison | High — title issues, lender red flags |
Clause fabrication is the hardest to catch because it reads like real contract language. We've seen AI tools summarize an AS-IS addendum and insert a repair obligation that didn't exist in the original — the kind of error that wouldn't register unless you read the source document line by line.
Where in a transaction hallucinations cluster
Hallucination risk isn't evenly distributed across a deal. It concentrates at specific milestones where agents tend to feed AI complex, multi-page documents and ask for summaries or action items. Here's where we've seen the most failures.
- **Offer drafting and comparison** — When agents paste competing offers into ChatGPT for a side-by-side, the model frequently transposes terms between offers. Price from offer A ends up attributed to offer B.
- **Deadline tracking from executed contracts** — AI pulls dates from the document but sometimes grabs the wrong reference point (effective date vs. binding agreement date), throwing every downstream deadline off.
- **Amendment and addendum summaries** — Multi-page addenda with nested conditions are where clause fabrication peaks. The model fills logical gaps with plausible but invented language.
- **Disclosure prep and review** — AI listing tools like Restb.ai or Listing AI can misidentify property features from photos, generating disclosure-adjacent claims about conditions the agent hasn't verified.
- **CMA narrative generation** — Models fabricate comparable sales data when they don't have live MLS access. A made-up comp at a plausible address is dangerously hard to catch.
Input quality controls output reliability
This is the gap nobody talks about. The same AI tool given the same prompt produces dramatically different error rates depending on what you feed it. We've watched agents photograph a contract on their phone, upload it to ChatGPT, and trust the summary — not realizing the OCR layer misread half the dates before the model even started.
| Input type | Typical hallucination behavior | Risk level |
|---|---|---|
| Clean PDF (digitally signed) | Dates and names mostly accurate; clause interpretation still risky | Moderate |
| Photographed document | OCR errors compound with LLM pattern-matching; dates and numbers unreliable | High |
| Verbal summary or voice note | Model fills in details you didn't provide; invents specifics confidently | Very high |
| Copy-pasted text from email thread | Context confusion between parties; timestamps misattributed | High |
If you're using AI to process transaction documents, start with the cleanest digital source you have. A digitally executed PDF from your transaction platform is a fundamentally different input than a photo of a faxed addendum. Treat them accordingly.
The compounding error chain: how one bad date breaks a deal
A single hallucinated date doesn't stay contained. It cascades. Here's a scenario we've seen play out — not hypothetically, but in the kind of workflow agents run every week.
- Agent uploads an executed purchase agreement to an AI tool and asks for a deadline summary.
- The AI reads the effective date as July 8 instead of July 3 (a five-day OCR + hallucination error on a photographed contract).
- Every contingency deadline in the summary is now five days late — inspection, appraisal, financing.
- Agent sends the AI-generated summary to the buyer's lender and inspector as the reference timeline.
- The inspection period actually expires before the inspector's scheduled date. The buyer loses the right to request repairs under the contract.
- The seller's attorney flags the missed contingency. The deal either gets renegotiated under pressure or falls apart.
The error didn't happen at step 6. It happened at step 2. Everything after that was a reasonable person trusting a reasonable-looking document. That's what makes compounding hallucinations dangerous — they travel through email, through lender packets, through title instructions, before anyone checks the source.
This is also where E&O insurance gets complicated. NAR guidance and most state licensing boards hold the agent — not the AI tool — accountable for information accuracy in transaction documents. The CFPB and fair housing enforcement don't carve out exceptions for AI-generated errors. If a hallucinated detail creates a RESPA violation or a Fair Housing Act issue, the agent's license is on the line. As always, consult a licensed attorney for compliance questions specific to your state.
A 60-second verification protocol by document type
Generic advice to "always verify AI output" is useless because it doesn't tell you which fields to check first. Here's a lightweight protocol tied to the documents agents actually process with AI tools. It's designed to catch the high-severity hallucinations without requiring you to re-read every word.
| Document type | Check these fields first (60 seconds) | Why these fields |
|---|---|---|
| Purchase agreement summary | Effective date, closing date, all contingency deadlines, purchase price, earnest money amount | Date and numerical hallucinations cluster here; one wrong date cascades |
| Amendment / addendum summary | Which clause is being modified, exact language of the change, party names | Clause fabrication peaks in amendment interpretation |
| CMA / comp analysis | Property addresses of comps, sale prices, sale dates | AI invents plausible comps when it lacks live MLS data |
| Listing description | Square footage, bedroom/bath count, lot size, year built, HOA status | Numerical drift and feature hallucination from photo analysis |
| Disclosure review notes | Material condition claims, repair history references, permit mentions | AI fills knowledge gaps with plausible but unverified claims |
AI tools — whether it's ChatGPT Plus, Claude, a dedicated platform, or an [operational assistant built for real estate workflows](/blog/ai-operational-assistant-vs-another-app-to-manage) — save real time on real tasks. The goal isn't to stop using them. It's to verify the outputs that carry contractual weight and let the low-risk drafts flow. If you want a deeper look at which [AI tools are actually working for agents right now](/blog/real-estate-ai-agents-2026-what-works-what-doesnt), we broke that down separately.



