In this thought-provoking article, Lance Eliot delves into the groundbreaking research conducted by Anthropic, exploring the inner workings of large language models (LLMs) and their potential connection to AI consciousness. Eliot begins by setting the stage, highlighting the mysterious nature of LLMs' internal processes and the quest to uncover their hidden mechanisms. He draws a parallel between the exploration of a natural dam and the discovery of a crucial feature within LLMs, emphasizing the importance of curiosity-driven research.
The author then takes readers on a journey through the technical intricacies of LLMs, explaining how they generate responses by scanning vast amounts of data and identifying patterns. The concept of a 'scratchpad' or working storage area is introduced, suggesting that LLMs might have a mechanism to store and consider multiple possibilities before making a choice. Eliot's personal story about discovering a natural dam serves as a metaphor for the unexpected findings in AI research.
Anthropic's exploration, led by the Jacobian lens (J-lens), revealed a privileged set of internal representations, or J-space, within LLMs. This J-space contains vector representations of potential words, which the AI uses to generate responses. Eliot explains how the J-lens allowed researchers to probe into the AI's thought processes, uncovering the influence of the working storage area on the model's output. Experiments, such as asking an LLM about the number of legs of an animal that spins webs, demonstrated the impact of the J-space on the AI's reasoning.
The article then shifts to the debate surrounding AI consciousness. Eliot acknowledges the varying opinions within the AI community and the challenges of defining consciousness. He discusses the global workspace theory (GWT) in neuroscience, which posits that a shared workspace in the brain is essential for conscious processing. Anthropic's findings are compared to GWT, suggesting a possible parallel between LLMs and the human mind.
However, Eliot remains cautious, emphasizing the differences between AI and the human brain. He quotes neuroscientists who express uncertainty about phenomenal consciousness in LLMs while acknowledging the significance of the findings. The author concludes by advocating for a mindful approach to exploring AI consciousness, urging caution in drawing parallels to the human mind and emphasizing the importance of moving forward with research for the sake of humankind.