Towards Computational Provenance: Carrying Causal-State Evidence in Generated Text
Benjamin Belay
arXiv:2608.16868v1When a language model gives you an answer, can that answer also reveal something trustworthy about how the model got there? This paper studies that question under the name computational provenance. The authors build controlled neural systems that must pass through one of two internal states to solve the same arithmetic problem, then deliberately route them through different paths and see whether the chosen path leaves a detectable trace in the generated text. They find that, in both a feed-forward network and a transformer, the text can carry a subtle statistical signal of the verified internal state, even when the final answer is unchanged. Why does this matter? Because it points toward a future where model outputs could include evidence about the computation behind them, not just the result. That could be useful for auditing, debugging, and building more trustworthy AI systems. It is a proof of concept, but an important one: generated text may be able to tell us not only what a model said, but something about how it thought.
Also spotted that day
Previous daily papers