Look It Up: Analysing Internal Web Search Capabilities of Modern LLMs

arXiv — cs.CL•Tuesday, November 25, 2025 at 5:00:00 AM

NeutralArtificial Intelligence

Modern large language models (LLMs) like GPT-5-mini and Claude Haiku 4.5 have been evaluated for their internal web search capabilities, revealing that while web access improves accuracy for static queries, it does not effectively enhance performance on dynamic queries due to poor query formulation. This assessment introduces a benchmark to measure the necessity and effectiveness of web searches in real-time responses.
The findings highlight significant implications for the development of LLMs, as they underscore the need for improved calibration in recognizing when to utilize web search. This could lead to advancements in how these models are trained and deployed, ultimately enhancing user experience and trust in AI-generated responses.
The ongoing exploration of LLM capabilities reflects a broader trend in AI research, where benchmarks like Bench360 and metrics such as ConCISE are being developed to evaluate various aspects of model performance. These efforts aim to address common challenges in AI, such as the balance between accuracy and conciseness, and the impact of training data on model behavior.

— via World Pulse Now AI Editorial System

Read Original

Was this article worth reading? Share it

Grasp.info

Extract key insights instantly from any article, video, or document.

AI & DataTry the app

Kwrds

Discover high-performing keywords and popular questions people are searching for.

AI & DataTry the app

Supametas.AI

Extract and structure unstructured data for seamless LLM RAG integration.

AI & DataTry the app

Continue Readings

Analytics India Magazine16 hours ago

Cornell Tech Secures $7 Million From NASA and Schmidt Sciences to Modernise arXiv

PositiveArtificial Intelligence

Cornell Tech has secured a $7 million investment from NASA and Schmidt Sciences aimed at modernizing arXiv, a preprint repository for scientific papers. This funding will facilitate the migration of arXiv to cloud infrastructure, upgrade its outdated codebase, and develop new tools to enhance the discovery of relevant preprints for researchers.

Read full article

via Analytics India Magazine

arXiv — cs.CLa day ago

Speech Recognition Model Improves Text-to-Speech Synthesis using Fine-Grained Reward

PositiveArtificial Intelligence

Recent advancements in text-to-speech (TTS) technology have led to the development of a new model called Word-level TTS Alignment by ASR-driven Attentive Reward (W3AR), which utilizes fine-grained reward signals from automatic speech recognition (ASR) systems to enhance TTS synthesis. This model addresses the limitations of traditional evaluation methods that often overlook specific problematic words in utterances.

Read full article

via arXiv — cs.CL

arXiv — cs.CVa day ago

Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation

PositiveArtificial Intelligence

A new framework called Task-aware Virtual View Exploration (TVVE) has been introduced to enhance robotic manipulation by integrating virtual view exploration with task-specific representation learning. This approach addresses limitations in existing vision-language-action models that rely on static viewpoints, improving 3D perception and reducing task interference.

Read full article

via arXiv — cs.CV

arXiv — cs.CLa day ago

For Those Who May Find Themselves on the Red Team

NeutralArtificial Intelligence

A recent position paper emphasizes the need for literary scholars to engage with research on large language model (LLM) interpretability, suggesting that the red team could serve as a platform for this ideological struggle. The paper argues that current interpretability standards are insufficient for evaluating LLMs.

Read full article

via arXiv — cs.CL

arXiv — cs.CLa day ago

Generating Reading Comprehension Exercises with Large Language Models for Educational Applications

PositiveArtificial Intelligence

A new framework named Reading Comprehension Exercise Generation (RCEG) has been proposed to leverage large language models (LLMs) for automatically generating personalized English reading comprehension exercises. This framework utilizes fine-tuned LLMs to create content candidates, which are then evaluated by a discriminator to select the highest quality output, significantly enhancing the educational content generation process.

Read full article

via arXiv — cs.CL

arXiv — cs.CLa day ago

Representational Stability of Truth in Large Language Models

NeutralArtificial Intelligence

Recent research has introduced the concept of representational stability in large language models (LLMs), focusing on how these models encode distinctions between true, false, and neither-true-nor-false content. The study assesses this stability by training a linear probe on LLM activations to differentiate true from not-true statements and measuring shifts in decision boundaries under label changes.

Read full article

via arXiv — cs.CL

arXiv — cs.CVa day ago

PRISM-Bench: A Benchmark of Puzzle-Based Visual Tasks with CoT Error Detection

PositiveArtificial Intelligence

PRISM-Bench has been introduced as a new benchmark for evaluating multimodal large language models (MLLMs) through puzzle-based visual tasks that assess both problem-solving capabilities and reasoning processes. This benchmark specifically requires models to identify errors in a step-by-step chain of thought, enhancing the evaluation of logical consistency and visual reasoning.

Read full article

via arXiv — cs.CV

arXiv — cs.LGa day ago

PocketLLM: Ultimate Compression of Large Language Models via Meta Networks

PositiveArtificial Intelligence

A novel approach named PocketLLM has been introduced to address the challenges of compressing large language models (LLMs) for efficient storage and transmission on edge devices. This method utilizes meta-networks to project LLM weights into discrete latent vectors, achieving significant compression ratios, such as a 10x reduction for Llama 2-7B, while maintaining accuracy.

Read full article

via arXiv — cs.LG