Generative AI Lies

Examples of generative AI making stuff up

Posts

  • AI Hallucination Cases database

    ()

    That thing where lawyers (and others) use generative AI in court filings, and the AI makes stuff up? Now there’s a list of such situations: the AI Hallucination Cases database.

    “This database tracks legal decisions in cases where generative AI produced hallucinated content – typically fake citations, but also other types of arguments.”

    “While seeking to be exhaustive (201 cases identified so far), it is a work in progress and will expand as new examples emerge.”

    (Original Facebook post.)


  • Gell-Mann

    ()

    Mike Pope on the Gell-Mann Amnesia Effect/Knoll’s Law (“everything you read in the newspapers is absolutely true, except for the rare story of which you happen to have firsthand knowledge”) and ChatGPT.

    (Original Facebook post.)


  • DOGE

    ()

    We obtained records showing how a Department of Government Efficiency staffer with no medical experience used artificial intelligence to identify which VA contracts to kill. “AI is absolutely the wrong tool for this,” one expert said.”

    “Lavingia’s system also used AI to extract details like the contract number and “total contract value.” This led to avoidable errors, where AI returned the wrong dollar value when multiple were found in a contract. Experts said the correct information was readily available from public databases.”

    (Original Facebook post.)


  • Summarizing medical info

    ()

    About some of the problems with having generative AI summarize medical information.

    I summarize medical information for doctors, researchers, and patients every day for a living, and I can promise you that any summary you get from chatGPT will have at least one significant error. And how could you possibly know? If you don’t understand what your doctor is telling you, how could you effectively vet the summary for errors?

    (Original Facebook post.)


  • Summarizing research

    ()

    Generalization bias in large language model summarization of scientific research

    when summarizing scientific texts, LLMs may omit details that limit the scope of research conclusions, leading to generalizations of results broader than warranted by the original study. […] Even when explicitly prompted for accuracy, most LLMs produced broader generalizations of scientific results than those in the original texts[…] In a direct comparison of LLM-generated and human-authored science summaries, LLM summaries were nearly five times more likely to contain broad generalizations[…] Notably, newer models tended to perform worse in generalization accuracy than earlier ones. Our results indicate a strong bias in many widely used LLMs towards overgeneralizing scientific conclusions, posing a significant risk of large-scale misinterpretations of research findings.

    (Article from April.)

    (Indirectly via Aliette.)

    (Original Facebook post.)


  • Reading list

    ()

    Chicago Sun-Times prints summer reading list full of fake books

    “Reading list [created by generative AI] in advertorial supplement contains 66% made up books by real authors.”

    Apparently not created by the Sun-Times:

    “The reading list appeared in a 64-page supplement called ‘Heat Index,’ which was a promotional section not specific to Chicago. Buscaglia told 404 Media the content was meant to be ‘generic and national’ and would be inserted into newspapers around the country.”

    (Original Facebook post.)


  • Chain-of-Thought

    ()

    Generative AI company Anthropic tests its “Chain-of-Thought” “reasoning models” to see whether they’re “faithful”—that is, to see whether the models accurately report the steps that they’re following. Turns out that they don’t.

    “Reasoning models are more capable than previous models. But our research shows that we can’t always rely on what they tell us about their reasoning. If we want to be able to use their Chains-of-Thought to monitor their behaviors and make sure they’re aligned with our intentions, we’ll need to work out ways to increase faithfulness.”

    (Article from April.)

    (Original Facebook post.)


  • Hallucinations are getting worse

    ()

    A.I. Is Getting More Powerful, but Its Hallucinations Are Getting Worse

    “A new wave of ‘reasoning’ systems from companies like OpenAI is producing incorrect information more often. Even the companies don’t know why.”

    “[OpenAI] found that o3 — its most powerful system — hallucinated 33 percent of the time when running its PersonQA benchmark test, which involves answering questions about public figures. That is more than twice the hallucination rate of OpenAI’s previous reasoning system, called o1. The new o4-mini hallucinated at an even higher rate: 48 percent.”

    (Original Facebook post.)


  • Bio

    ()

    Yet another example of the kinds of falsehoods produced by LLMs: A post from early 2024 about Google Bard’s incorrect bio of Deirdre Saoirse Moen.

    (Original Facebook post.)


  • Naked emperor

    ()

    The little kid shouted, “The emperor has no clothes!”

    The other citizens all glared at the kid. “In the future, the emperor’s clothes will be awesome!” they said. “So awesome that they will solve all of our problems, including problems that have nothing to do with clothes. Besides, it’s really our fault—if we just looked at the emperor from the right angle, he would surely have clothes, so we need to get better at guessing how to look at him.”

    This is a post about generative AI.


    (This post was brought to you by the people who I’ve heard say things like “The answers I get are usually wrong, but that’s just because I haven’t learned how to write a good prompt yet.”)


    (Original Facebook post.)