Position: Semantic Uncertainty Measures Disagreement, Not Reliability

Joseph Hoche    Maxime Corlay    David Brellmann    Andrei Bursuc    Pavel Izmailov    Angela Yao    Gianni Franchi

NeurIPS 2026

Paper  

Position: Semantic Uncertainty Measures Disagreement, Not Reliability

Sources of uncertainty in LLMs and LVLMs, from semantic ambiguity to decoding strategies.


Abstract

Large Language Models and their multi-modal variants see increasingly rapid adoption and deployment, including in settings where reliability matters. However their uncertainty remains difficult to assess as classic or early token-level approaches cannot deal with the open-ended outputs of such models. Semantic uncertainty arises as a promising solution to the limits of token-level confidence by measuring disagreement among multiple generated responses. This position paper argues that this framing is incomplete: semantic uncertainty measures semantic disagreement, not actual reliability. A model can be uncertain while producing several valid answers, or certain while repeatedly producing the same wrong answer. We introduce a semantic bias-uncertainty decomposition to show that reliability depends both on variability across meanings and systematic deviation from correct or grounded meanings. This perspective reveals that common hallucination-detection evaluations conflate uncertainty with error. We argue for reliability-centered evaluation that separates semantic disagreement, correctness, and enables a more fine-grained characterization of uncertainty beyond the level of the full answer.



BibTeX

@inproceedings{hoche2026semanticuncertainty,
  title     = {Position: Semantic Uncertainty Measures Disagreement, Not Reliability},
  author    = {Hoche, Joseph and Corlay, Maxime and Brellmann, David and Bursuc, Andrei and Izmailov, Pavel and Yao, Angela and Franchi, Gianni},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
  year      = {2026}
}