Causal tests of persistent belief in learned and pretrained agents
View the Project on GitHub Silentpartnercoding/epistemic-amputation
James Siyuan He
Permanent archive: doi:10.5281/zenodo.22233258
Can a learned belief-like representation survive the exact evidence that was supposed to defeat it—and continue controlling a model even after the model says the claim is false?
The preregistered Gemma 4/J-space experiment returned NOT SUPPORTED.
The interesting result is the gap between reading and causing. A model can contain a decodable pattern that resembles a hidden belief without that pattern being the mechanism controlling its decision.
Interpretability tools are often described as revealing what a model “really believes.” This study supplies a stricter rule: a readable representation should not be called a hidden belief unless controlled intervention shows that it causes the relevant prediction or action.
This is a working paper and public research release, not a peer-reviewed claim about consciousness, subjective experience, or human religion.