Do LLMs pass the mirror test?
The article argues that current mirror tests for Large Language Models are flawed because they rely on visual recognition rather than the model's primary sensory modality. It suggests that, similar to animal cognition studies, researchers should develop tests that account for the specific nature of AI processing.
Why it matters
As AI systems become more advanced, establishing valid benchmarks for machine self-awareness and cognition is critical for future development and safety.
The mirror test — Gallup's original, the one with the red dot on the chimp's forehead — has been adapted for LLMs several times in the past, but as far as I'm concerned, every adaptation gets it wrong in very similar ways: they build visual mirror tests translated into text. Show the model its own output and ask "is this yours?" or have it identify its responses among an (anonymized) lineup. Some models pass it, others fail, and I think neither outcome is particularly informative because I think they all test for the wrong thing ; which is, coincidentally, exactly the criticism that led Alexandra Horowitz to build a different kind of mirror test for dogs.
The article presents a philosophical and scientific argument regarding AI testing methodology without taking a political stance.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in