Artificial IntelligencearXiv — cs.CLWed, Jun 3, 2026, 4:00 AMNeutral

Evaluating and Calibrating LLM Confidence on Questions with Multiple Correct Answers

A recent study has introduced MACE, a benchmark consisting of 12,000 factual questions across six domains, to evaluate the confidence calibration of large language models (LLMs) when faced with multiple correct answers. The research reveals that existing training-free calibration methods often misestimate confidence levels due to disagreements among equally valid responses.

WPN Brief

  • What Happened

    A recent study has introduced MACE, a benchmark consisting of 12,000 factual questions across six domains, to evaluate the confidence calibration of large language models (LLMs) when faced with multiple correct answers. The research reveals that existing training-free calibration methods often misestimate confidence levels due to disagreements among equally valid responses.

  • Why It Matters

    This development is significant as it highlights the limitations of current LLMs in accurately assessing their confidence, which is crucial for their reliability in applications requiring precise information.

  • The Bigger Picture

    The findings underscore a broader issue in AI regarding the calibration of confidence in responses, as other studies have shown that LLMs can exhibit overconfidence and inconsistencies in their probabilistic beliefs, raising concerns about their trustworthiness in various contexts.

Ask WPN AI