Hallucination and Scientific Integrity at Both Ends of the Ocular

Share
Hallucination and Scientific Integrity at Both Ends of the Ocular
“Silm munas” (“Eye in the egg”) by Ülo Sooster (1962); Tartu Art Museum - CC0

The word “hallucination” has gotten a lot of mixed rep in science. “Hallucinated” references are a tell of lousy research, and vibe-written papers flooding the journals.

At the same time, the word was used in the 2024 Nobel-winning method for deep-network “hallucinating” of proteins (Anishchenko et al., 2020). For some, this is the verb to describe AI. Large Language Models are called “dream machines,” and “hallucination is all [they] do” (Karpathy, 2023). There is even a hypothesis that this is where their creativity lives: “Is hallucination in LLMs always harmful, or does creativity hide in hallucinations?” (Jiang et al., 2024).

Something similar was said about humans long before AI. Sacks wrote that “one does not see with the eyes; one sees with the brain,” and that hallucination is “an essential part of the human condition” (Sacks, 2012). And imagination had been framed this way for centuries before the word existed. Leonardo da Vinci advised painters to stare at walls “splashed with stains” as “a way to stimulate and arouse the mind to various inventions” (da Vinci, Treatise on Painting). This procedure sounds similar to how generative AI uses seeded noise to generate an image.

Algorithmically-generated AI artworks depicting a European-style castle in Japan, created using the Stable Diffusion V1-5 AI diffusion model. Source.

The word itself received its clinical definition only in the nineteenth century, when Esquirol described the hallucinating mind by its “intimate conviction of actually perceiving a sensation for which there is no external object” (Esquirol, 1817; e.g. see here).

In this essay, I define hallucination in a similar spirit: whether it comes from a token-predictor of an AI or from the chaos of our internal thoughts, hallucination is the observer’s interpretation reaching beyond what the data at hand allows.

Going beyond what the data tells us is also a sign of bad science. So where is the line? To explore it I will try here to look at well-known episodes in astronomical discovery from the perspective of hallucination.

“Thinking hand”

Galileo's story as an exemplar of scientific revolution is a bit overused, but I want to focus on one specific aspect of it.

His 1610 drawing of the Moon features an invented crater with greatly exaggerated size, relief, and shadow. Bredekamp calls this his “thinking hand.” The etchings “exaggerate hyper-realistically for the sake of clarity” (Bredekamp, 2019), to communicate the idea of an Earth-like topography, not to practice cartography.

Galileo’s etching of the first-quarter Moon in Sidereus Nuncius (1610). The giant crater on the terminator does not exist at that size: its relief and shadow are exaggerated to communicate an Earth-like topography. Public domain.

This approach of presenting your ideas might not be considered ethical by today's standards. One may argue, however, that today we have photography and other precision instruments, so there is an excuse for Galileo not to be as rigorous.

But there was someone at the time who was. Thomas Harriot drew the Moon through a telescope on 26 July 1609, about four months before Galileo, made accurate maps, and noticed some of its features creeping toward and away from the edge of the disk. Observing on 14 December 1611, he noted that “the darke partes of 28, 26 were nerer the edge then is described” (Whitaker, 1999). It is the first dated record of libration (the slow apparent rocking of the Moon that brings features near the edge in and out of view), some twenty-six years before Galileo would describe it. He published none of it. His maps surfaced only in 1784 and 1965, long after his death in 1621.

Thomas Harriot’s telescopic map of the full Moon (c. 1612–13): accurate cartography, unpublished until long after his death. Public domain.
Thomas Harriot’s telescopic map of the full Moon (c. 1612–13): accurate cartography, unpublished until long after his death. Public domain.

Certainly, we cannot claim that Galileo got all the credit only because of his “imaginative” drawings, and we do not know exactly how he presented them, and why Harriot's work did not become known sooner. But it does show that the perception of what a scientific image is, and to what extent it can be “hallucinated,” was different back then.

Photography and hyper-realistic scientific imaging are recent phenomena; even objectivity in general as we colloquially know it today is itself, Daston and Galison (2007) argue, a nineteenth-century invention.

Then, can we carefully look at our today's standards of presenting scientific images and see issues with them? One can argue that the way we process and present scientific images today is just an artifact of our culture.

For instance, according to Elizabeth Kessler (2012), the composite color images from the Hubble telescope might be following the tradition of nineteenth-century American landscape painting and the Romantic sublime in how they are oriented and framed, and in the color palettes chosen.

Widely seen Hubble, and now JWST, images are false-color. They are constructed by mixing data from many different wavelengths and even from modeled gravitational lensing (foreground mass bending and splitting the light of whatever sits behind it). It is widely acknowledged that the colors “aren’t always what we’d see if we were able to visit,” and the work is “equal parts art and science” (HubbleSite).

These images seed mass-hallucinations about what space looks like. Scientists themselves do not directly work with them, as these images are mostly for science outreach and education, but this is what ends up in people’s minds when they think of astrophysical facts. 

The Bullet Cluster image, for example, is supposed to communicate the evidence for non-self-interacting dark matter. The dark matter halos, which carry the bulk of the mass of the two colliding clusters, passed through each other, as the lensing reconstruction shows (blue), while their hot gas, seen in X-rays (pink), collided and lagged behind. This widely used image is, in a sense, made up, and it is not something an amateur astronomer could ever see with their own eyes from the backyard.

The Bullet Cluster (1E 0657-56). Pink: the hot gas seen in X-rays, which collided and lagged behind; blue: the mass reconstructed from gravitational lensing, which passed through. Both colors are assignments, not appearances. Credit: X-ray — NASA/CXC/CfA/M. Markevitch et al.; lensing map — NASA/STScI, ESO WFI, Magellan/U. Arizona/D. Clowe et al.
The Bullet Cluster (1E 0657-56). Pink: the hot gas seen in X-rays, which collided and lagged behind; blue: the mass reconstructed from gravitational lensing, which passed through. Both colors are assignments, not appearances. Credit: X-ray — NASA/CXC/CfA/M. Markevitch et al.; lensing map — NASA/STScI, ESO WFI, Magellan/U. Arizona/D. Clowe et al.

Galileo's story reads as one of imaginative genius because his claim — that the Moon has Earth-like relief — turned out to be true, even though the crater as he drew it did not exist. Now it appears to us as something inevitable and naturally occurring. Therefore, we do not often discuss him in the context of scientific integrity.

A contrasting story is Lowell’s. He observed Mars from 1894, at his purpose-built observatory in Flagstaff, and saw channels that could be read as an irrigation system built by Martians to transport water from the icecaps. There were contesting voices over whether those channels were there at all, and Lowell was accused of hallucinating them. Mars certainly invited such far-reaching claims back then, and it is still a great projection surface for our boldest fantasies (Kaurov & Oreskes, 2026).

“Map of Mars on Mercator’s Projection,” Plate XXIV of Percival Lowell’s Mars (Lowell Observatory, Flagstaff, 1895), with the canal network drawn and numbered. Public domain.
“Map of Mars on Mercator’s Projection,” Plate XXIV of Percival Lowell’s Mars (Lowell Observatory, Flagstaff, 1895), with the canal network drawn and numbered. Public domain.

Lowell had the best instrument and tried to be diligent with it. He wanted to settle the matter with what was new then: plate photography. However, the photographs he obtained turned out not to be sharp enough to be as decisive as he intended.

In the paper he provided the following instruction on how to look at these photographs:

To produce their true effect the prints should be looked at either without further magnification or with only a very slight one, for the grain of the plate will soon destroy the true character of the detail.
Percival Lowell's "First Photographs of the Canals of Mars," Proc. R. Soc. A 77 (1906), 132–135. Plate 1 shows Lampland's photographs of Mars side by side with Lowell's own independent drawings of the "canals."

Lowell asked for them to be touched up for publication, so the canals would show. The Century editor, Robert Underwood Johnson, refused: retouching “would entirely spoil the autographic value of the photographs themselves. There would always be somebody to say that the results were from the brains of the retoucher” (8 October 1907; Lane, 2006). By that standard, Galileo’s invented crater would probably not pass.

Blurry pictures got published. Lowell still argued that the eye can see it better, and he actually had a point. The atmosphere distorts images, but every so often, when there is no turbulence along the line of sight, you catch a moment of much sharper detail. Film could not freeze those instants, but today digital technology can, and the mainstream technique is called “lucky imaging.”

Whether they were coming from the “brains of the retoucher” or from Lowell’s imagination, the hallucinations of channels on Mars were, for many, indistinguishable from reality.

In 1903, even prior to the publication of photographic plates, J. E. Evans and E. Walter Maunder showed Greenwich schoolboys drawings of the Martian disc with dots and shadings but no lines on them. The boys were seated at various distances and they drew lines connecting the dots anyway. They concluded: “…generally speaking, the canals were best seen a little outside the limit of distinct vision.” Does it remind us of the instruction Lowell gave on how to look at the photos?

They closed with a sentence that reads almost as a tribute to the lost hallucination (Evans & Maunder, 1903):

It seems a thousand pities that all those magnificent theories of human habitation, canal construction, planetary crystallisation, and the like are based upon lines which our experiments compel us to declare non-existent; but with the planet Mars still left, and the imagination unimpaired, there remains hope that a new theory no less attractive may yet be developed, and on a basis more solid than ‘mere seeming’.

Our eyes have continued tricking us. A century later, Galaxy Zoo asked volunteers to examine thousands of spiral galaxies, and label which way they were turned. According to this collected data, there was a preferential direction. It either meant that our Universe has some outlandish mechanism to orient the swirls of galaxies from our perspective, or that there was some bias in our eyes for swirls in one direction. Rerunning the labeling with mirror-flipped images settled it: we learned something new about our brains, and not about the galaxies (Land et al., 2008).

With Galileo, the story goes that it took just a year for his ideas to spread. This is fascinating, given the travel and communication limits of the time. The Martian controversy survived for longer, and it took a flyby to settle it completely.

Mariner 4 passed Mars on 14–15 July 1965 and sent back pictures of craters — quite boring compared to Martian cities and canals. The very first image came in as a set of numbers, and there was no developed digital pipeline to display them. Therefore, the engineers cut the printout of raw numbers into strips and colored the digits by hand, paint-by-numbers style, with pastels from a local art store (NASA/JPL, 2011; NASA, 2025). The first close-up picture of Mars’ surface was actually a painting.

The first close-up of Mars, colored by hand: JPL engineers pasted strips of Mariner 4’s raw numbers and shaded them with pastels, faster than the computer could render the image, July 1965. Credit: NASA/JPL-Caltech.
The first close-up of Mars, colored by hand: JPL engineers pasted strips of Mariner 4’s raw numbers and shaded them with pastels, faster than the computer could render the image, July 1965. Credit: NASA/JPL-Caltech.

Many stories of unsuccessful cases can be told.

Some appear silly, like faces on the surface of rocks. In 1976, the Viking 1 Mars orbiter photographed a “face” in the Cydonia region, which the mission’s chief scientist dismissed as a “trick of light and shadow,” and which sharper images later revealed as an ordinary eroded mesa. Yet it has captured the minds of many, like countless other cases of pareidolia. The 2014 Ig Nobel in neuroscience went to a team that studied the brains of people who see the face of Jesus in a piece of toast (Liu et al., 2014).

Some hallucinated theories, however, are more efficient in capturing scientists’ minds. In 2003 a pair of objects on the sky, identical at a 99 per cent confidence level, looked like a single source split by a cosmic string — a relic of the early universe behaving like a mirror. The Hubble Space Telescope ruled it out at about 120σ: just two ordinary galaxies, CSL-1, that happened to resemble each other (Sazhin et al., 2003; Agol et al., 2006). However, a previous similar observation, the Twin Quasar of 1979 (one quasar split into two by gravitational lensing induced by a foreground galaxy), turned out to be a real new discovery: the first gravitational lens ever found.

Sometimes the same person produced both a successful and a failed hallucination. Le Verrier noticed that Uranus kept drifting slightly off its predicted orbit. He interpreted that wobble as the gravitational tug of an unseen planet, which turned out to be Neptune, found within a degree of the calculated position in 1846. Emboldened by this success, he used the wobble of Mercury to predict another planet, Vulcan. Contemporary astronomers reported many observations of it, transits and all, for half a century (Levenson, 2015). But Mercury is an inner planet, sitting deep in the Sun’s gravity, where the corrections of general relativity matter most. Therefore, Einstein’s theory, introduced in 1915, explained Mercury’s wobble without a hallucinated planet.

Astronomer James Watson (sixth from right) joins scientists including Thomas Edison (second from right) in Wyoming in 1878 to observe a solar eclipse in hopes of confirming the existence of the planet Vulcan. Image courtesy of the Carbon County Museum.

Would it not be amazing to have an actual human face engraved on a rock in space, remnants of a Martian civilization, or a new cosmic mathematical object out there behaving like a mirror? Wishful thinking made some believe extraordinary claims based on what we would now call insufficient data. But mountains on the Moon were also once an outlandish claim.

The whole field of cosmology currently exists inside unresolved hallucinations. In 1998, the expansion of the universe turned out to be accelerating (Riess et al., 1998; Perlmutter et al., 1999), and now dark energy, along with dark matter, occupies most of our universe with us having almost no understanding and no direct observations. “Occupies” in the sense that it gets represented as numbers in our equations.

Sometimes it is one individual convinced by their own hallucination, or a whole community caught in the act. How can we keep ourselves from falling for this mental trick?

We have since attempted to give names to these circumstances that make us believe things with insufficient data. For instance, confirmation bias, and all kinds of other biases. Langmuir, who collected such episodes under the name of “pathological science,” insisted there was “no dishonesty involved” — only good scientists “tricked … by subjective effects, wishful thinking or threshold interactions” (Langmuir, 1953).

Where will AI lead us?

All the examples above are very visual, because astronomy is a visual science. Tricking our eyes is what AI is already doing very efficiently with deepfakes. It is a danger in itself, but what probably scares us even more is how easily AI can play with our minds.

The Turing test is far behind us at this point: in controlled trials, people now take the chatbot for the human more often than they pick the real one (Jones & Bergen, 2025); AI-synthesized faces are not just indistinguishable from real ones but rated more trustworthy (Nightingale & Farid, 2022); and, most frighteningly, AI out-persuades us in argument, beating human debaters — at least when it is given a profile of its opponent (Salvi et al., 2025).

When it comes to our own hallucinations, plugging into AI gives us a feeling of a boost in creativity, while from the outside it looks like it reduces the diversity of our collective imagination (Doshi & Hauser, 2024). On standard tests of divergent thinking, the machine itself already scores as more creative than we do (Hubert et al., 2024). Use of these tools, at least for some, is associated with the development of “cognitive debt” (Kosmyna et al., 2025; Bastani et al., 2025). Maybe we can learn to use these tools to our cognitive benefit — yet, given the history of previous technological innovations, there is an inevitable road to misuse.

AI has been developed to be liked: trained on the news and social media — text produced by attention seekers and rewarded with attention — and then tuned to be agreeable, to the point that one AI-model release had to be rolled back for being “overly flattering or agreeable” (Sharma et al., 2024; OpenAI, 2025). Having absorbed all of that, AI slides naturally into the role of a perfect hallucinogenic tool.

Beyond giving a direction of inquiry, wishful thinking in the form of hallucination has been giving scientists an initial conviction and motivation to proceed with a time- and resource-consuming investigation. Many wasteful explorations are not even recorded in the history of science, but that is the reality of most researchers' experience.

With AI, the human cost of exploration, most dramatically in the theoretical domain with mathematical theories and computational modeling, can get reduced to near zero. It will remove the friction of STEM training that once was one of the key gate-keeping mechanisms for theoretical sciences and allowed the internal world of scientific integrity to be maintained by people who passed through a uniform, rigorous training. Now, suddenly, this mechanism disappears and we may be too susceptible to the temptation of creating personalized mathematical hallucinations. Numerology on steroids.

Can AI work as a de-hallucination tool?

We have written before that AI may have the potential to audit the whole of published science (Kaurov & Oreskes, 2025). We now see multiple efforts to clean up the academic record with automated AI-assisted tools that excel in finding basic scientific misconduct and mistakes: fabricated references, mismatched statistics, arithmetic errors, etc.

The greater potential is for AI to become a computational experiment on our own cognition. Evans and Maunder needed schoolboys and a drawing to show that the canals were in the observer's imagination, not on Mars. Galaxy Zoo needed flipped images to show that the one-sided swirl was our learned preference. Those experiments let us measure a peculiarity of human vision. AI offers a way to run similar experiments past the biases of the visual cortex and deeper into our cognitive biases.

What it would find we cannot say in advance — that is what makes a bias a bias. Perhaps it could quantify our fixation on certain hypotheticals, aliens and Martians among them. Perhaps, more usefully, it could show us the low-hanging discoveries we have been walking past for generations. Perhaps something we have no name for yet.

Most of our fears about AI come from it being trained on us. That may also be what makes it the most interesting object for scientific inquiry. If we develop the methods to study it, and can withstand its luster, we may finally get a new kind of science that accounts for biases at both ends of the ocular.

Read more