IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household Tasks

arXiv — cs.CV•Monday, December 8, 2025 at 5:00:00 AM

NeutralArtificial Intelligence

The introduction of IS-Bench marks a significant advancement in evaluating the interactive safety of VLM-driven embodied agents, particularly in household tasks. This benchmark addresses the limitations of existing evaluation paradigms by simulating dynamic risks and assessing an agent's ability to perceive and mitigate these risks effectively.
This development is crucial as it enhances the safety and reliability of VLM-driven agents, which are increasingly being deployed in real-world scenarios. By ensuring that these agents can navigate complex environments safely, IS-Bench could facilitate broader adoption in various applications.
The ongoing discourse surrounding the reliability and safety of visual language models (VLMs) highlights a critical need for robust evaluation frameworks. As advancements in models like GPT-4o and Gemini-2.5 continue, the focus on interactive safety and risk mitigation becomes paramount, reflecting a broader trend towards ensuring AI systems can operate safely in dynamic environments.

— via World Pulse Now AI Editorial System

Read Original

Was this article worth reading? Share it

LucidQuery AI

Combines diffusion reasoning with autoregressive LLM for advanced AI analysis.

AI & DataView app details

Chattermate

Build and deploy AI support agents without writing any code.

AI & DataView app details

AIvilization

Create an AI agent to learn, work, and socialize in a self-running multiplayer town.

Lifestyle & HealthView app details

Continue Readings

arXiv — cs.CL2 days ago

Shrinking the Generation-Verification Gap with Weak Verifiers

PositiveArtificial Intelligence

A new framework named Weaver has been introduced to enhance the performance of language model verifiers by combining multiple weak verifiers into a stronger ensemble. This approach addresses the existing performance gap between general-purpose verifiers and oracle verifiers, which have perfect accuracy. Weaver utilizes weak supervision to estimate the accuracy of each verifier, allowing for a more reliable scoring of generated responses.

Read full article

via arXiv — cs.CL

arXiv — cs.CL2 days ago

SimSUM: Simulated Benchmark with Structured and Unstructured Medical Records

NeutralArtificial Intelligence

SimSUM has been introduced as a benchmark dataset comprising 10,000 simulated patient records that connect unstructured clinical notes with structured background variables, specifically in the context of respiratory diseases. The dataset aims to enhance clinical information extraction by incorporating tabular data generated from a Bayesian network, with clinical notes produced by a large language model, GPT-4o.

Read full article

via arXiv — cs.CL

arXiv — cs.CV2 days ago

Towards Effective and Efficient Long Video Understanding of Multimodal Large Language Models via One-shot Clip Retrieval

PositiveArtificial Intelligence

A new paradigm called One-shot video-Clip based Retrieval AuGmentation (OneClip-RAG) has been proposed to enhance the efficiency of Multimodal Large Language Models (MLLMs) in processing long videos, addressing the limitations of existing models that can only handle a limited number of frames due to memory constraints.

Read full article

via arXiv — cs.CV

arXiv — cs.CV3 days ago

Geo3DVQA: Evaluating Vision-Language Models for 3D Geospatial Reasoning from Aerial Imagery

NeutralArtificial Intelligence

Geo3DVQA has been introduced as a benchmark for evaluating vision-language models in 3D geospatial reasoning using RGB-only aerial imagery, addressing challenges in urban planning and environmental assessment that traditional sensor-based methods face. The benchmark includes 110,000 curated question-answer pairs across 16 task categories, emphasizing realistic scenarios that integrate various 3D cues.

Read full article

via arXiv — cs.CV

arXiv — cs.CV3 days ago

GeoShield: Safeguarding Geolocation Privacy from Vision-Language Models via Adversarial Perturbations

PositiveArtificial Intelligence

GeoShield has been introduced as a novel adversarial framework aimed at protecting geolocation privacy from Vision-Language Models (VLMs) like GPT-4o, which can infer users' locations from publicly shared images. This framework includes three modules designed to enhance the robustness of geoprivacy protection in real-world scenarios.

Read full article

via arXiv — cs.CV

arXiv — cs.CV3 days ago

VRSA: Jailbreaking Multimodal Large Language Models through Visual Reasoning Sequential Attack

NeutralArtificial Intelligence

The introduction of the Visual Reasoning Sequential Attack (VRSA) highlights vulnerabilities in Multimodal Large Language Models (MLLMs), which are increasingly used for their advanced cross-modal capabilities. This method decomposes harmful text into sequential sub-images, allowing MLLMs to externalize harmful intent more effectively.

Read full article

via arXiv — cs.CV

arXiv — cs.CL3 days ago

Policy-based Sentence Simplification: Replacing Parallel Corpora with LLM-as-a-Judge

PositiveArtificial Intelligence

A new approach to sentence simplification has been introduced, utilizing Large Language Models (LLMs) as judges to create policy-aligned training data, eliminating the need for expensive human annotations or parallel corpora. This method allows for tailored simplification systems that can adapt to various policies, enhancing readability while maintaining meaning.

Read full article

via arXiv — cs.CL

arXiv — cs.CL3 days ago

Living the Novel: A System for Generating Self-Training Timeline-Aware Conversational Agents from Novels

PositiveArtificial Intelligence

The Living Novel system has been developed to transform literary works into immersive conversational experiences, addressing challenges such as persona drift and narrative coherence in large language models (LLMs). This innovative approach employs a two-stage training pipeline, including Deep Persona Alignment and Coherence and Robustness Enhancing stages, to ensure characters remain true to their narratives.

Read full article

via arXiv — cs.CL