Artificial IntelligencearXiv — cs.CVThu, Jun 11, 2026, 4:00 AMPositive

Diffusion-based Cumulative Adversarial Purification for Vision Language Models

A recent study has introduced DiffCAP, a diffusion-based purification strategy designed to enhance the reliability of Vision Language Models (VLMs) by neutralizing adversarial perturbations that can significantly distort model outputs. This approach theoretically establishes a recovery region in the forward diffusion process, demonstrating that adversarial effects diminish as diffusion progresses.

WPN Brief

  • What Happened

    A recent study has introduced DiffCAP, a diffusion-based purification strategy designed to enhance the reliability of Vision Language Models (VLMs) by neutralizing adversarial perturbations that can significantly distort model outputs. This approach theoretically establishes a recovery region in the forward diffusion process, demonstrating that adversarial effects diminish as diffusion progresses.

  • Why It Matters

    The development of DiffCAP is crucial for improving the robustness of VLMs, which are increasingly utilized in real-world applications where accuracy is paramount. By effectively addressing vulnerabilities to adversarial attacks, this innovation can enhance trust in AI systems that rely on VLMs for multimodal understanding.

  • The Bigger Picture

    The introduction of DiffCAP aligns with ongoing efforts in the AI community to bolster the resilience of VLMs against adversarial threats. Other recent advancements, such as methods for continual machine unlearning and frameworks for risk awareness injection, reflect a broader trend towards enhancing the safety and utility of AI models, highlighting the importance of developing robust solutions in the face of evolving challenges.

Ask WPN AI

Related Reports

More coverage on this story

8 reports across the wire

arXiv — cs.CV
May 26

Closed-Loop Bidirectional Prompting for Adversarial Robustness of Vision Language Models

arXiv:2605.25922v1 Announce Type: new Abstract: Vision Language Models adapt well to downstream tasks but are highly vulnerable to adversarial perturbations that disrupt cross-modal semantic alignment. Existing defenses are largely unidirectional or structural, failing to exploit bidirectional cross-modal complementarity and instance-wise adaptive protection. To overcome the limitations of unidirectional and static defenses in adversarial settings, we propose Closed-Loop Bidirectional Prompting, casting robust adaptation as cross-modal agreement recovery via a dynamic feedback loop on frozen encoders. A Semantic Anchor is introduced as a stable prior to constrain cyclic updates and mitigate perturbation-induced feature corruption. Through anchor-based bootstrapping, textual semantics denoise visual representations, while the refined visuals enable instance-adaptive prompt updating, yielding a rectified and robust consensus. Extensive evaluations across 11 datasets validate state-of-the-art robustness and strong base-to-new generalization, while maintaining a favorable trade-off between computational cost and accuracy.

Artificial Intelligence
arXiv — cs.LG
May 19

CATA: Continual Machine Unlearning via Conflict-Averse Task Arithmetic

A recent study introduces CATA, a method for continual machine unlearning in vision-language models (VLMs), addressing challenges such as effectively removing target knowledge while preserving model utility and preventing knowledge re-emergence. This marks a significant advancement in the field of artificial intelligence, particularly in the context of VLMs, which are increasingly utilized in various applications.

Artificial Intelligencepositive
arXiv — cs.CV
Jun 1

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation

A recent study has re-evaluated the fine-tuning of Vision Language Models (VLMs) by utilizing a fully controlled data generation and annotation pipeline, which aims to produce bias-free data with balanced distribution and clean annotations. This approach addresses the common issues of biases and errors in traditional data collection methods.

Artificial Intelligencepositive
arXiv — cs.CV
Mar 17

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models

A novel framework called LLMind has been introduced, which focuses on bio-inspired training-free adaptive visual representations for Vision-Language Models (VLMs). This approach contrasts with traditional VLMs that apply uniform spatial fidelity across visual inputs, instead mimicking human vision's adaptive and selective nature through a Bio-inspired Adaptive Sampling Strategy (BASS).

Artificial Intelligencepositive
arXiv — cs.CV
Jun 9

Collaborative Edge-to-Server Inference for Vision-Language Models

A new collaborative edge-to-server inference framework for vision-language models (VLMs) has been proposed, aiming to reduce communication costs while preserving inference accuracy. This framework allows servers to perform initial inference on downsized images and request detailed data only when necessary, addressing the challenges of high communication overhead and potential accuracy loss from excessive image compression.

Artificial Intelligencepositive
arXiv — cs.CV
May 29

Mitigating State Aliasing in Vision-Language-Action Models via Inverse Dynamics Learning

Recent advancements in Vision-Language-Action (VLA) models have introduced inverse dynamics learning as a method to mitigate state aliasing, which occurs when visually similar states require different actions. This approach aims to enhance the performance of VLA models by directly supervising the vision encoder, addressing a critical limitation in current models that often misinterpret visual distinctions.

Artificial Intelligenceneutral
arXiv — cs.CV
May 28

Towards Unified Vision-Language Models with Incomplete Multi-Modal Inputs

A new study has introduced a unified video-language model capable of processing incomplete multi-modal inputs, addressing the limitations of existing Video-Language Models (VLMs) that require complete data. This model aims to enhance performance in real-world applications where sensor deactivation may lead to incomplete data.

Artificial Intelligenceneutral
arXiv — cs.CV
May 14

DistractMIA: Black-Box Membership Inference on Vision-Language Models via Semantic Distraction

A new framework named DistractMIA has been introduced to conduct black-box membership inference on vision-language models (VLMs) by utilizing semantic distraction techniques. This approach addresses the challenges faced by auditors who typically only observe generated textual responses, allowing for a more effective analysis of training data that may contain sensitive information.

Artificial Intelligenceneutral

Apps

Useful picks

Explore all apps

Articles

Continue Reading

arXiv — cs.CLArtificial Intelligenceyesterday

Evaluation of Multilingual Ability to Use Spatial Deictic Expressions in Vision-Language Models

A recent study evaluates the multilingual ability of vision-language models (VLMs) to use spatial deictic expressions, which are context-dependent terms like 'this' and 'that'. The research highlights the necessity for VLMs to integrate language and visual context to accurately interpret these expressions across different languages.

arXiv — cs.CVArtificial Intelligenceyesterday

FMMC: Harnessing the Power of Foundation Models for Accurate Material Classification

A novel framework has been introduced by FMMC to enhance material classification accuracy by leveraging vision-language foundation models (VLMs). This approach addresses the challenges posed by limited annotated data, which has historically hindered the effectiveness of material recognition tasks in computer vision and graphics.

arXiv — cs.CVArtificial Intelligenceyesterday

From My View to Yours: Learning Egocentric Cues from Exocentric Video using Privileged Egocentric Supervision

A new framework called Ego2ExoVLM has been proposed to enhance Vision Language Models (VLMs) by enabling them to infer egocentric properties from exocentric video observations, addressing limitations in understanding human-object interactions crucial for Activities of Daily Living (ADL) monitoring.

arXiv — cs.CVArtificial Intelligenceyesterday

T2T-VICL: Cross-Task Visual In-Context Learning via Implicit Text-Driven VLMs

The paper introduces T2T-VICL, a framework for cross-task visual in-context learning (VICL) using vision-language models (VLMs). This approach addresses the challenge of mismatched visual demonstrations and queries, enabling VLMs to generate implicit textual guidance without explicitly naming tasks.

arXiv — cs.LGArtificial Intelligenceyesterday

When Prompts Ignore Structure: Graph-Based Attribute Reasoning for Calibrated VLMs

A new approach called ARGTCA has been proposed to enhance the calibration of vision-language models (VLMs) by utilizing a Graph Attention Network (GAT) to represent class and attribute pairs as nodes in a Symbolic Attribute Graph. This method aims to improve confidence estimation during test-time adaptation, addressing issues of overconfidence that arise from traditional prompt tuning methods.

arXiv — cs.CVArtificial Intelligence2 days ago

AirflowAttack: Thermal-Airflow Adversarial Perturbations against Infrared Remote-Sensing Vision-Language Models

A new adversarial attack method named AirflowAttack has been introduced, targeting infrared remote-sensing vision-language models (VLMs). This approach utilizes thermal-airflow turbulence as a perturbation prior, achieving a significant attack success rate of 48.5% across various CLIP model backbones, which is notably higher than existing physical baselines.

arXiv — cs.CVArtificial Intelligence2 days ago

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models

A new framework called Vision Contextualized Probing (VisCoP) has been introduced to enhance the adaptability of Vision Language Models (VLMs) in novel domains, addressing performance degradation due to distribution shifts from pretraining data. VisCoP employs a compact set of learnable visual probes to learn domain-specific representations while minimizing updates to existing model components.

arXiv — cs.CVArtificial Intelligence2 days ago

PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails

The introduction of PolicyShiftGuard represents a significant advancement in the field of image guardrails, focusing on policy-adaptive mechanisms that evaluate images based on varying safety policies. This approach is crucial as it allows models to adapt to changing policy definitions rather than relying solely on fixed safety standards.