Generative AI Lies

Examples of generative AI making stuff up

Posts

  • Sum of consecutive cubes

    ()

    I try to avoid reading the AI Overviews when I do Google searches that show them. But this time I happened to glance at the AI Overview. And it was blatantly false.

    I searched for [sum of three “consecutive cubes”]. Here’s what the AI Overview told me:

    The sum of three consecutive cubes is always equal to the square of the sum of those three numbers; in other words, if you take any three consecutive integers, add them together, then square that sum, you will get the same result as adding the cubes of those three numbers.

    Example:

    • Consider the numbers 1, 2, and 3.
    • Adding them gives: 1 + 2 + 3 = 6
    • Squaring the sum gives: 6^2 = 36
    • Adding their cubes gives: 1^3 + 2^3 + 3^3 = 1 + 8 + 27 = 36

    Key points:

    • This property holds true for any set of three consecutive integers.
    • You can use this fact to quickly calculate the sum of three consecutive cubes without having to individually cube each number.

    It’s true that this pattern holds true if you happen to pick the numbers 1, 2, and 3. Super cool!

    So let’s look at the very next set of three consecutive integers: 2, 3, and 4.

    2 + 3 + 4 = 9, and 9^2 = 81.
    2^3 + 3^3 + 4^3 = (8 + 27 + 64) = 99.

    But wait, 81 doesn’t equal 99! What’s wrong???

    What’s wrong is that the AI Overview’s claim is false for all cases except for (1,2,3), (0,1,2), and (-1,0,1). There’s one other case where the sum of the cubes of three consecutive integers is a perfect square (of a different number), but the only cases where it’s specifically the square of the sum of those three consecutive integers are (1,2,3), (0,1,2), and (-1,0,1).

    (I’m reminded of a joke proof that all odd numbers are prime: 3 is prime; 5 is prime; 7 is prime; therefore, by induction, all odd numbers are prime.)

    Google’s AI Overview links to sources to support its statements. The sources that it linked to in this case are the following. Not one of them makes the claim that the AI Overview is making.

    As usual, the moral of this story is: Don’t believe anything that generative AI tells you.


    I often remember to add ” -ai” to the ends of searches (to tell Google not to give an AI Overview), but I often don’t.

    But I’ve been hearing reports that that isn’t working any more for some people. And, indeed, when I re-run this search with -ai, sometimes it gives me an AI Overview and sometimes it doesn’t.

    Interestingly, the specific contents of the AI Overview vary—if I do the search without -ai, I get the one that I posted about, but if I do the search with -ai, I get a different claim (that doesn’t mention squares) that I haven’t checked yet.

    Adding swear words does still seem to work to remove the AI Overview. It also provides search results that use those swear words, which may or may not improve your search results, depending on what kinds of search results you want. 

    🙂

    The reason I was doing this search was that an interesting fact about a particular number was mentioned in a recent TV show. I had played around with the relevant numbers a bit, and had reached the point where I was curious about what work had been done on this topic.

    If I hadn’t tried out the sums of cubes of three consecutive numbers on my own just prior, I might have been tempted to believe what the AI Overview said. But because I had just calculated several such answers myself, I knew immediately that the AI Overview was wrong.


    (Original Facebook post.)


  • Search engines

    ()

    AI search engines give incorrect answers at an alarming 60% rate, study says

    “A new study from Columbia Journalism Review’s Tow Center for Digital Journalism finds serious accuracy issues with generative AI models used for news searches. The research tested eight AI-driven search tools equipped with live search functionality and discovered that the AI models incorrectly answered more than 60 percent of queries about news content.”

    “Error rates varied notably among the tested platforms. Perplexity provided incorrect information in 37 percent of the queries tested, whereas ChatGPT Search incorrectly identified 67 percent (134 out of 200) of articles queried. Grok 3 demonstrated the highest error rate, at 94 percent.”

    Google’s Gemini appears to have also had an extremely high error rate.


    I should note that the study was only looking at one specific kind of queries. Here’s the methodology from the study:

    “We randomly selected ten articles from each publisher, then manually selected direct excerpts from those articles for use in our queries. After providing each chatbot with the selected excerpts, we asked it to identify the corresponding article’s headline, original publisher, publication date, and URL”

    “We deliberately chose excerpts that, if pasted into a traditional Google search, returned the original source within the first three results”

    Also: “More than half of responses from Gemini and Grok 3 cited fabricated or broken URLs that led to error pages.”

    Here’s the study (may be paywalled).


    In comments on my Facebook post, a friend indicated that exact text matching is a task that we wouldn’t expect LLMs to be good at. I replied:

    I don’t see this as an exact-text-matching task; I see it as a reference-finding task.

    Like asking the question: “Here’s a quote I found on the internet. Where does it come from?”

    When a search engine receives that question and responds with a made-up URL, that seems to me to be a problem.

    (But I agree that it doesn’t necessarily make sense to generalize from the results of studies that focus specifically on a particular kind of query. And I do feel like the Ars Technica article that I linked to should have said a little more about the specific focus of this study.)


    (Original Facebook post.)


  • AI in healthcare

    ()

    Artificial intelligence systems [in healthcare contexts] require consistent monitoring and staffing to put in place and to keep them working well.”

    “Evaluating whether these products work is challenging. Evaluating whether they continue to work — or have developed the software equivalent of a blown gasket or leaky engine — is even trickier.”

    “‘Even in the best case, the [LLMs] had a 35% error rate’”


    (To be clear: Some of this article is about LLMs, and some of it is about predictive algorithms that I assume are old-fashioned non-generative AI. So this is partly an LLM issue, but also partly a non-LLM issue.)


    (Original Facebook post.)


  • Expert testimony

    ()

    A federal court judge has thrown out expert testimony from a Stanford University artificial intelligence and misinformation professor[, Jeff Hancock], saying his submission of fake information made up by an AI chatbot ‘shatters’ his credibility.”

    “At Stanford, students can be suspended and ordered to do community service for using an AI chatbot to ‘substantially complete an assignment or exam’ without instructor permission. The school has repeatedly declined to respond to questions […] about whether Hancock would face disciplinary measures.”

    (Original Facebook post.)


  • Identifying sources

    ()

    Researchers asked ChatGPT’s search tool to identify the source of excerpts from a couple hundred online articles.

    The result: ChatGPT made up answers. (Not always, but often.)

    Gasp! Shock! Surprise!

    (Original Facebook post.)


  • We don’t like to talk about that

    (, )

    If your ChatGPT prompt includes certain not-uncommon names of humans, ChatGPT says “I’m unable to produce a response” and ends the session.

    Turns out that those names are names of some people who have prominently reported that ChatGPT was making up lies about them.

    So apparently, on learning that ChatGPT is lying about specific people, OpenAI has decided to prevent ChatGPT from responding to any prompt that mentions those people’s names.

    Of course, usually there’s more than one human who has a particular name, so OpenAI is also preventing ChatGPT from talking about anyone who has the same name as someone who ChatGPT has previously prominently lied about.

    (Original Facebook post.)


  • Album release date

    (, )

    Today I did a Google search for [“field of stars” mccutcheon] and I forgot to append “-AI” to leave out the AI Overview. When I forget to leave out the Overview, I normally try to not even look at the Overview; but this time the Overview caught my eye. It starts out:

    “Field of Stars is an album by American folk singer-songwriter John McCutcheon. The album was released on January 10, 2024.”

    McCutcheon had an online concert today to celebrate the release of the album, so I spent several seconds wondering why he waited 10+ months after its release to have the concert. And then I realized that of course the AI Overview is just wrong, once again. The album will be officially released on January 10, 2025. (But is available now in various pre-official-release contexts.)

    But this is one of the reasons that I usually try not to even look at the Overview, because they often read to me as so authoritative that even though I know they include false information, I still sometimes believe them.

    (Original Facebook post.)


  • Kosher bacon

    ()

    If you do a Google search for [salt pork substitute kosher], the AI Overview tells you to try pancetta or bacon as a kosher substitute for salt pork.

    Yet another example of why you should never believe anything that generative AI tells you.

    (Original Facebook post.)

    (Update: Sometime in the year after I posted this, Google stopped returning an AI Overview in response to that query.)


  • Transcription

    ()

    Researchers say an AI-powered transcription tool used in hospitals invents things no one ever said

    This is about Whisper, which I’ve heard praised in other contexts. 🙁

    “Whisper has a major flaw: It is prone to making up chunks of text or even entire sentences, according to interviews with more than a dozen software engineers, developers and academic researchers. Those experts said some of the invented text — known in the industry as hallucinations — can include racial commentary, violent rhetoric and even imagined medical treatments.”

    “medical centers [are using] Whisper-based tools to transcribe patients’ consultations with doctors”

    “While most developers assume that transcription tools misspell words or make other errors, engineers and researchers said they had never seen another AI-powered transcription tool hallucinate as much as Whisper.”

    (Original Facebook post.)