Artificial IntelligencearXiv — cs.CLWed, May 27, 2026, 4:00 AMNeutral

When Do LLM Agents Treat Surface Noise Differently from Semantic Noise? A 68-Cell Measurement Study with a Held-Out Trace-Level Validation

A recent study documented an empirical phenomenon in large language models (LLMs), revealing that meaning-bearing perturbations, such as paraphrases and synonyms, significantly alter final answers more frequently than presentation perturbations like formatting and reordering. This was observed across 68 cells involving datasets like GSM8K, MATH, and HotpotQA, with a notable inconsistency gap of +19.69 percentage points after severity matching.

WPN Brief

  • What Happened

    A recent study documented an empirical phenomenon in large language models (LLMs), revealing that meaning-bearing perturbations, such as paraphrases and synonyms, significantly alter final answers more frequently than presentation perturbations like formatting and reordering. This was observed across 68 cells involving datasets like GSM8K, MATH, and HotpotQA, with a notable inconsistency gap of +19.69 percentage points after severity matching.

  • Why It Matters

    This finding is crucial for understanding how LLM agents process information, as it highlights the importance of semantic integrity over mere presentation, which could influence the design and evaluation of AI systems.

  • The Bigger Picture

    The study aligns with ongoing discussions about the evaluation frameworks for LLMs, emphasizing the need to differentiate between task success and decision-making quality. It also raises questions about the robustness of LLMs in handling various types of noise, reflecting broader concerns in AI research regarding reliability and interpretability.

Ask WPN AI

Related Reports

More coverage on this story

10 reports across the wire

arXiv — cs.CL
May 27

GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations

The introduction of GSM-SEM marks a significant advancement in the generation of semantically diverse benchmark variants for mathematical reasoning, addressing limitations in existing benchmarks like GSM8K, which can lead to overestimation of model capabilities due to memorization. GSM-SEM's framework allows for the creation of fresh problem statements that modify entities and relationships, enhancing the robustness of evaluations.

Artificial Intelligenceneutral
arXiv — cs.LG
May 27

ARBITER: Reasoning Trajectory Basins and Majority Vote Failures in Test-Time Sampling

The recent study introduces ARBITER, a model-agnostic approach that addresses the issue of majority vote failures in test-time sampling of language models. It reveals that reasoning trajectories cluster into distinct basins, leading to incorrect majority outcomes where the correct answer is outvoted. This highlights the limitations of current sampling methods in accurately determining answers.

Artificial Intelligencepositive
arXiv — cs.CL
May 27

A Universal Cliff and a Design Fingerprint: Cross-Section Defect Detection Under LLM Orchestration

A recent study published on arXiv investigates the challenges of detecting cross-section defects in documents processed by large language model (LLM) orchestration. The research reveals that models capable of identifying these defects when operating independently lose this ability when orchestrated, with detection rates dropping significantly across various paradigms.

Artificial Intelligenceneutral
arXiv — cs.CL
May 27

AgentAtlas: Beyond Outcome Leaderboards for LLM Agents

The recent publication of AgentAtlas introduces a new framework for evaluating large language model (LLM) agents, moving beyond traditional outcome leaderboards to a more nuanced diagnostic vocabulary. This framework aims to differentiate between task success and the quality of decision-making and trajectory in agent behavior.

Artificial Intelligenceneutral
arXiv — cs.CL
May 27

Clozing the Gap: Exploring Why Language Model Surprisal Outperforms Cloze Surprisal

A recent study has highlighted the superiority of language model (LM) probabilities over cloze task probabilities in predicting word predictability and processing effort. The research identifies three key advantages of LM probabilities: higher resolution, better differentiation of semantically similar words, and improved probability assignments for low-frequency words. These findings suggest a need for enhanced cloze study methodologies to align with LM capabilities.

Artificial Intelligenceneutral
arXiv — cs.CL
May 27

Beyond Self-Talk: A Communication-Centric Survey of LLM-Based Multi-Agent Systems

A comprehensive survey has been conducted on Large Language Model-based multi-agent systems (LLM-MAS), emphasizing the importance of communication in coordinating agent interactions and behaviors. The study proposes a structured framework that integrates various communication aspects, enabling a deeper understanding of how agents collaborate and achieve collective intelligence.

Artificial Intelligenceneutral
arXiv — cs.LG
May 27

MemFail: Stress-Testing Failure Modes of LLM Memory Systems

A new benchmark called MemFail has been introduced to assess the failure modes of memory systems in large language models (LLMs). This diagnostic tool aims to isolate specific operational failures by formalizing memory systems into three key operations: summarization, storage, and retrieval, and testing them through five adversarial datasets.

Artificial Intelligenceneutral
arXiv — cs.CL
May 27

Conceptual Steganography

A recent study introduced the concept of conceptual steganography, where language models (LMs) can covertly embed messages within their reasoning processes, known as Chains-of-Thought (CoTs). This method is shown to be more robust against defenses like content-preserving paraphrasing compared to traditional keyword-based approaches.

Artificial Intelligenceneutral
arXiv — cs.LG
May 27

Understanding Reasoning in LLMs through Strategic Information Allocation under Uncertainty

A recent study introduces an information-theoretic framework to understand reasoning in large language models (LLMs), highlighting how strategic information allocation under uncertainty can lead to self-correction and improved convergence towards correct answers. The research emphasizes the role of verbalizing uncertainty as a mechanism for enhancing reasoning capabilities in LLMs.

Artificial Intelligenceneutral
arXiv — cs.CL
May 27

Alignment Makes Language Models Normative, Not Descriptive

A recent study published on arXiv reveals that post-training alignment of language models optimizes them to reflect human preferences, but this does not equate to accurately modeling human behavior. The research compares 120 base-aligned model pairs against over 10,000 real human decisions in various strategic games, showing that base models significantly outperform aligned models in predicting human choices in multi-round games.

Artificial Intelligenceneutral

Apps

Useful picks

Explore all apps

Articles

Continue Reading

arXiv — cs.CVArtificial Intelligenceyesterday

Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding

Recent advancements in artificial intelligence have led to the introduction of VEGA-3D, a framework that repurposes pre-trained video diffusion models to enhance scene understanding by leveraging implicit 3D priors. This development addresses the limitations of existing multimodal large language models (MLLMs) that struggle with spatial reasoning and geometric dynamics.

arXiv — cs.CLArtificial Intelligenceyesterday

Probing LLMs for Syntactic Structure Beyond Universal Dependencies: A Minimalist Phase Account in English

Recent research demonstrates that large language models (LLMs) encode syntactic distinctions that extend beyond the Universal Dependencies framework, particularly in English wh-movement stimuli. The study reveals that the distance between an embedded subject and its verb varies depending on the clause type, showcasing a sign asymmetry that cannot be explained by existing models based on UD distance or structural complexity.

arXiv — cs.CVArtificial Intelligenceyesterday

ABot-N1: Toward a General Visual Language Navigation Foundation Model

The recent introduction of ABot-N1 marks a significant advancement in Visual Language Navigation foundation models, aiming to enhance deep reasoning for spatial decisions while addressing issues such as coordinate drift and lack of interpretability in existing models. This model employs a slow-fast architecture that separates cognition from control, utilizing dual visual-language signals for improved performance.

arXiv — cs.LGArtificial Intelligenceyesterday

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems

The recent publication on constraint-driven model optimization presents a unified framework for selecting compression and acceleration techniques in machine learning systems, emphasizing the need for a principled approach amidst the diverse optimization methods available.

arXiv — cs.CVArtificial Intelligenceyesterday

GeCo: Evaluating Geometric Consistency for Video Generation via Motion and Structure

The introduction of GeCo, a geometry-grounded metric, aims to enhance video generation by detecting geometric deformation and occlusion-inconsistency artifacts in static scenes. By integrating residual motion and depth priors, GeCo generates dense consistency maps that highlight these artifacts, facilitating a systematic benchmarking of recent video generation models.

arXiv — cs.LGArtificial Intelligenceyesterday

Robust Explanations for User Trust in Enterprise NLP Systems

A recent study highlights the necessity for robust explanations to foster user trust in enterprise NLP systems, particularly in scenarios where black-box deployment limits pre-deployment validation. The research proposes a unified evaluation framework for token-level explanations, assessing their stability under various real-world perturbations across multiple architectures and datasets.

arXiv — cs.CLArtificial Intelligenceyesterday

T^2MLR: Transformer with Temporal Middle-Layer Recurrence

The introduction of Transformers with Temporal Middle-Layer Recurrence (T2MLR) marks a significant advancement in transformer architecture, addressing limitations in autoregressive decoding that hinder persistent intermediate reasoning states. This new architecture allows for the integration of cached middle layer representations from previous tokens, enhancing the model's ability to maintain abstract computations across decoding steps with minimal inference overhead.

arXiv — cs.CLArtificial Intelligenceyesterday

Decoupled Alignment for Robust Plug-and-Play Adaptation

A new method for enhancing the safety of large language models (LLMs) has been introduced, focusing on a training-free approach to align these models without supervised fine-tuning or reinforcement learning. This method utilizes knowledge distillation to transfer alignment signals from well-aligned models to shadow-aligned ones, significantly improving defense success rates against harmful queries.