Sign up to see the future today. Innovations that cannot be missed from the cutting edge of science and technology. There are a litany of reasons to take the output of generative AI with a heavy dose of salt. Their answers are known to sound authoritative, convincing, and elaborate, and at the same time to be completely and utterly wrong. It’s a lesson that astrophysicist and science communicator Paul Sutter learned the hard way and was brave enough to share it with the world. In an article for the publication Nautilus earlier this month, Sutter retold a truly harrowing story that sounds like something straight out of a stress-induced nightmare. In February, he gave a presentation in front of a room full of “collaborators” and revealed a new update to his algorithm designed to detect voids, or empty regions, between galaxies. The new version was “ten times faster than the previous one,” had a “more sophisticated way of dealing with the unpleasant realities of a real data set,” and could “handle surveys a hundred times larger,” Sutter recalled. It was also partly the result of extensive consultation sessions with an AI to write the code for the update (“vibe coding,” in the parlance) that he recalled as “indispensable.” But just ten minutes after his presentation, a collaborator intervened saying that “something seemed strange to them.” It turns out that the way the new algorithm handled poll edges was “wrong,” Sutter admitted. “It was neither a typo, nor a missing quote, nor a factor of two,” he wrote. “It was subtle, but it was very wrong, and everything afterward was also wrong, and I had shared it all in a room full of people who trusted me.” Sutter’s chilling story perfectly illustrates how many users of AI tools are unknowingly putting themselves at risk, even at the most exclusive levels of academia. The world of science has been inundated with under-researched and often unedited AI detritus, a major reckoning that forces academics to be accountable for everything they publish, including potentially embarrassing hallucinations. The AI coding tool Sutter was using “sounded like it could be understood,” he wrote, despite being “little more than a sophisticated predictor of the next word.” “LLM fluency is not an accident and it is not an emerging mystery,” Sutter wrote. “It’s a trait we bred, the way we turn wolves into dogs that look at our faces when we open the bag of treats.” “So how do we implement a tool that is sometimes wrong but always nice? How do we trust AI?” he concluded. “The simple answer is: don’t do it.” Sutter compared our extensive use of artificial intelligence tools to alchemy, an ancient practice devised by what he calls “pre-scientists.” “We are in the prechemical era of AI,” said the astrophysicist. “The crucible is closed and like alchemists we are not going to stop using it.” To get to the truth, he says, we must carefully trace an AI’s chain of reasoning and audit each of its results. Following his embarrassing slip in February, Sutter promised that he now works “differently” and is willing to immediately “distrust” anything an AI says. It’s a warning that even some of the most talented thinkers can be easily tempted by the lure of AI. As such, when we pasted his latest article for Nautilus into the far-from-perfect Pangram AI detection tool, it informed us that 56 percent of the text appeared to be written by an AI. More on AI hallucinations: Academics in crisis now being held accountable for AI hallucinations in their research papers