a sign that says the truth is out there

For the last couple of years, the AI interpretability community has been chasing a 'holy grail': the truth vector. The idea was simple: if we could find a single linear direction in an LLM's latent space that separates true statements from false ones, we could effectively build a lie detector for AI. It sounded like a win-win for safety and transparency. But a new research perspective is throwing a wrench in the works.

The Tarski Attack

Recent work by Abel Jansma suggests that the search for a universal 'truth direction' is fundamentally flawed. By applying a 'diagonal attack'—inspired by Tarski’s undefinability theorem—Jansma argues that no linear probe can ever truly pin down truth. Tarski famously proved that truth in a language cannot be defined within that same language. When applied to LLMs, this implies that any probe attempting to identify 'truth' is actually just identifying the model's belief about truth, not truth itself.

Beyond Linear Geometry

Previous studies, such as those exploring the 'Geometry of Truth,' claimed that true and false statements separate linearly in high-dimensional space. While these probes work on specific datasets, the Tarski-inspired critique suggests this is an illusion of scale rather than a fundamental law. If truth is not a direction, then 'same-axis self-auditing' is destined to fail. To fix this, researchers are proposing more complex theories, such as Orthogonal Auditing by Rotation (OAR), to move beyond simple linear classifiers.

The Path Forward

This shift marks a transition from optimistic geometry to rigorous logic. We are moving away from the idea that we can simply 'point' to truth in a neural network and toward a more nuanced understanding of how LLMs represent knowledge. The 'truth vector' might have been a useful stepping stone, but the real map of AI cognition is likely far more twisted.

Sources

Media