In April 2023 New York lawyer Steven Schwartz had ChatGPT draft a brief for a lawsuit brought against Avianca, the largest Colombian airline. Schwartz provided some background information to the system and used the two-page legal argument it produced, sending the text to a judge without proofreading. The judge responded that although the brief read well, none of the legal precedents cited by Schwartz actually existed.
The episode had serious consequences: Steven Schwartz was publicly embarrassed in New York and ultimately left his practice.
Why does generative AI "make things up"?
The explanation lies in how generative, statistical AI works. These systems are designed to produce the most probable responses to inputs: they are optimized for plausibility based on patterns learned from training data, not for verifying factual truth.
Probability, however, is not the same as truth. As a result, models can output entirely fictitious information that sounds credible in language and form — which makes such outputs misleading.
Two kinds of "hallucinations"
Errors produced by generative AI are usefully divided into two categories:
- Completely false outputs: plainly incorrect information that is usually easy to spot (for example, if the AI replies that the capital of France is Copenhagen).
- Convincingly realistic fabrications: statements rich in detail and formally plausible, which can convince users of their accuracy. In the legal case, generated docket numbers and judge names appeared authentic but were fabricated.
The second category is especially dangerous because users can be vulnerable if they do not verify concrete facts.
Practical experience: what happens on "expert" topics?
Researchers and practitioners often test AI systems on topics they know well. These tests typically reveal that generated texts mix correct and incorrect details: for instance, a birthdate might be right while graduation dates or employment entries are false or mixed with invented details.
One example described in the source: in a test the AI claimed the author was a professor at Aix-en-Provence University and the CEO of an unfamiliar French startup — assertions that were untrue.
Why these systems are not "truth-seekers"
Generative models do not verify sources the way a human fact-checker or researcher would. In generating an answer, the system computes an effective average of patterns present in its training data; asking the same question multiple times yields different responses because the averaging process is re-run on each iteration.
This behavior raises a broader question about intelligence: if these systems were truly intelligent in a human, epistemic sense, they would preferentially rely on verified sources (for example, authoritative encyclopedias) when available.
Source and context
The discussion above draws in part on an edited excerpt from Luc Julia’s book Robot, nem pilóta (published by HVG Könyvek). The Steven Schwartz case and related examples illustrate a fundamental limitation of generative AI: linguistic coherence does not guarantee factual accuracy, so human oversight remains essential, especially in legal, medical and other high-risk domains.



