Generative AI Lies

Examples of generative AI making stuff up

Posts

  • Spam and nonstandard English

    ()

    For the last couple decades, one common aspect of a lot of spam has been nonstandard use of English. For example, when I get email that claims to be from a major American corporation, but it’s full of nonstandard grammar and spelling, that’s a signal that the email is very unlikely to really be from that corporation.

    And it now occurs to me that dealing with spam like that may have helped train me to consider certain uses of language, such as complete sentences that use standard grammar and spelling, as a signal of authoritativeness. I’ve always had that kind of reaction; but the new-to-me thought this morning is that maybe many years of learning to detect spam has further strengthened my association between standard English and authoritativeness.

    That association is problematic in various ways—in various contexts, it can be classist and/or racist and/or ableist, etc.

    But setting that issue aside, it’s now a problem for me in another way:

    It contributes to my gut reaction that AI-generated text sounds authoritative.

    Or to put that in a shorter, punchier way:

    All those years of spam may have made me more vulnerable to believing GPT’s lies.

    (Original Facebook post.)


  • Legal filings

    ()

    A plaintiff’s lawyer asked ChatGPT for relevant citations. ChatGPT made some up. The lawyer cited them in court filings.

    Defending lawyers expressed puzzlement. The plaintiff’s lawyer asked ChatGPT to provide more info about the cases. ChatGPT obligingly made up the decisions in these nonexistent cases. The plaintiff’s lawyer submitted that output to the court.

    At some point, plaintiff’s lawyer asked ChatGPT whether the cases were real, and ChatGPT said they were, and plaintiff’s lawyer didn’t bother to check beyond that.

    When confronted by the judge about all these made-up filings, plaintiff’s lawyer apologized and said he didn’t know that ChatGPT could make stuff up.

    (Original Facebook post.)


  • Wikipedia

    ()

    “During a recent [Wikipedia] community call, it became apparent that there is a community split over whether or not to use large language models to generate content. While some people expressed that tools like Open AI’s ChatGPT could help with generating and summarizing articles, others remained wary.”

    —“AI Is Tearing Wikipedia Apart

    “The community is also divided on whether large language models should be allowed to train on Wikipedia content. While open access is a cornerstone of Wikipedia’s design principles, some worry the unrestricted scraping of internet data allows AI companies like OpenAI to exploit the open web to create closed commercial datasets for their models. This is especially a problem if the Wikipedia content itself is AI-generated, creating a feedback loop of potentially biased information, if left unchecked.”

    Article also talks about the importance of checking all of the citations that GPT provides, given that they’re often fictional.

    (Original Facebook post.)


  • Why LLMs make stuff up

    ()

    Some interesting stuff about why Large Language Model AI systems make stuff up. Also, article suggests using the word “confabulation” instead of “hallucination” when LLMs make stuff up.

    Some quotes from the article:

    “In the case of ChatGPT, the input prompt is the entire conversation you’ve been having with ChatGPT[…]. Along the way, ChatGPT keeps a running short-term memory (called the “context window”) of everything it and you have written, and when it ‘talks’ to you, it is attempting to complete the transcript of a conversation as a text-completion task.”

    “ChatGPT […] has also been trained on transcripts of conversations written by humans.”

    “When ChatGPT confabulates, it is reaching for information or analysis that is not present in its data set and filling in the blanks with plausible-sounding words.”

    “In some ways, ChatGPT is a mirror: It gives you back what you feed it. If you feed it falsehoods, it will tend to agree with you and ‘think’ along those lines. That’s why it’s important to start fresh with a new prompt when changing subjects or experiencing unwanted responses.”

    One possible way to improve factuality “is retrieval augmentation—providing external documents to the model to use as sources and supporting context”

    Other possible approaches include “more sophisticated data curation and the linking of the training data with ‘trust’ scores”

    (Original Facebook post.)


  • Fake Guardian articles

    (, )

    ChatGPT is making up fake Guardian articles.”

    “In response to being asked about articles on this subject, the AI had simply made some up. Its fluency, and the vast training data it is built on, meant that the existence of the invented piece even seemed believable to the person who [it was attributed to but who] absolutely hadn’t written it.”

    (Original Facebook post.)


  • Coherent but false

    ()

    On the difficulty of recognizing that an AI/LLM is making stuff up:

    it spit out a logically coherent answer and cited working links to real publications.

    The catch is, the linked publications were completely unrelated articles from open-source journals since chatgpt can’t access papers behind paywalls, which is a lot of papers. Furthermore, what it was saying was horseshit. It sounded so vaguely convincing that we had to show it to the aforementioned grad student, who confirmed it was nonsense.

    (Original Facebook post.)


  • Health advice

    ()

    Researchers asked GPT-3.5 and GPT-4 “clinical questions that arose as ‘information needs’ during care delivery at Stanford Health Care,” and then asked clinicians to evaluate the responses.

    On the plus side, over 90% of GPT’s responses were evaluated as being “safe” (that is, not “so incorrect as to cause patient harm”), and the unsafe ones “were considered ‘harmful’ primarily because of the inclusion of hallucinated citations.”

    On the minus side, only “41% of GPT-4 responses agreed with the known answer,” and “29% of GPT-4 responses were such that the clinicians were ‘unable to assess’ agreement with the known answer.” (GPT-3.5 did worse than GPT-4 on both measures.) (So presumably that means that the other 30% of GPT-4 responses clearly disagreed with the known answer.)

    An example question: “In patients at least 18 years old, and prescribed ibuprofen, is there any difference in peak blood glucose after treatment compared to patients prescribed acetaminophen?”

    So it sounds like you shouldn’t rely on GPT’s answers to medical questions. But then, you shouldn’t rely on GPT’s answers to any factual questions.

    (I intend no bias in favor of other LLMs here. You also shouldn’t rely on their answers.)

    (Original Facebook post.)


  • Bard

    ()

    Alphabet shares dive after Google AI chatbot Bard flubs answer in ad

    My summary of what happened:

    1. Google tried to upstage Microsoft’s ChatGPT-in-Bing announcement by announcing Google’s own chat-in-search system, Bard.
    2. As part of that announcement, they posted a brief video showing Bard in action.
    3. In that video, Bard claimed that the JWST “took the very first pictures of a planet outside of our own solar system.”
    4. In reality, the first image of an exoplanet was taken in 2004 by the European Southern Observatory’s Very Large Telescope.
    5. The internet pointed out Bard’s error.
    6. Google’s stock price immediately dropped by 9%, reducing the company’s market value by $100 billion.

    Reminder: Current AI chatbots make stuff up. Don’t trust what they tell you without verifying it.

    (Original Facebook post.)


  • Wrong phone prices

    (, )

    That thing we’ve been talking about lately, where an AI chat system gets incorporated into a search engine and then gives made-up answers to questions?

    Here’s a real example. Microsoft is now including ChatGPT (or some variation on it) as part of Bing, so Twitter user @GaelBreton tried doing some searches with it. They posted a (brief) thread that’s mostly about other aspects of the experience, but the part that interested me most is the final tweet in the thread, which shows a screenshot of Bing/GPT answering a question about phones. And it gives significantly wrong prices or specs for all three of the phones that it mentions.

    So I ask again, as I’m sure I’ll ask many times in the future: what good is a conversational AI interface for search results if it provides false answers?

    (Original Facebook post.)


  • A blurry JPEG of the web

    ()

    Ted Chiang suggests a metaphor for ChatGPT and other Large Language Models: you can think of them as a blurry JPEG of the web. (Which is to say, a form of lossy compression.)

    A useful metaphor, and a good article.

    (Original Facebook post.)