Artificial IntelligencearXiv — cs.CLWed, May 20, 2026, 4:00 AMNeutral

LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening

LLMEval-Logic has been introduced as a new Chinese benchmark for evaluating logical reasoning in large language models (LLMs), utilizing a pipeline that combines expert audits and adversarial hardening to enhance the reliability of assessments. This benchmark includes a Base subset of 246 items and a Hard subset of 190 items, with a focus on realistic situational scenarios.

WPN Brief

  • What Happened

    LLMEval-Logic has been introduced as a new Chinese benchmark for evaluating logical reasoning in large language models (LLMs), utilizing a pipeline that combines expert audits and adversarial hardening to enhance the reliability of assessments. This benchmark includes a Base subset of 246 items and a Hard subset of 190 items, with a focus on realistic situational scenarios.

  • Why It Matters

    The development of LLMEval-Logic is significant as it addresses the limitations of existing logical reasoning benchmarks, which often lack rigorous formal annotations and are quickly outpaced by advanced reasoning models. By implementing a solver verification process with Z3, the benchmark aims to provide a more robust evaluation framework for LLMs.

  • The Bigger Picture

    This initiative reflects a broader trend in AI research towards improving the evaluation of LLMs through more sophisticated methodologies, such as decomposing reasoning efficiency and enhancing interactive capabilities. These advancements highlight the ongoing challenges in aligning LLM outputs with human judgment and the need for reliable benchmarks in the rapidly evolving landscape of AI.

Ask WPN AI