SGM: Safety Glasses for Multimodal Large Language Models via Neuron-Level Detoxification

arXiv — cs.CL•Thursday, December 18, 2025 at 5:00:00 AM

PositiveArtificial Intelligence

A new framework named SGM has been developed to enhance the safety of multimodal large language models (MLLMs) by implementing neuron-level detoxification. This approach selectively recalibrates toxic neurons, significantly reducing harmful outputs from 48.2% to 2.5% while maintaining fluency in generated content. The framework also introduces MM-TOXIC-QA, a multimodal toxicity evaluation system to assess its effectiveness.
The introduction of SGM is crucial for improving the reliability and safety of MLLMs, which are increasingly used in various applications. By addressing the inherent risks associated with toxic and biased outputs, SGM aims to foster greater trust in AI technologies and their deployment across sensitive domains.
This development reflects a growing trend in AI research focused on mitigating biases and enhancing the safety of AI systems. As MLLMs become more prevalent, the need for effective detoxification methods is paramount, especially in light of recent studies highlighting challenges such as hallucinations and visual neglect. The ongoing exploration of frameworks like SGM, V-ITI, and SafePTR indicates a concerted effort within the AI community to establish robust safety measures.

— via World Pulse Now AI Editorial System

Read Original

Was this article worth reading? Share it

LucidQuery AI

Combines diffusion reasoning with autoregressive LLM for advanced AI analysis.

AI & DataView app details

Humanize AI

Transform AI-generated text into undetectable, human-like content effortlessly.

Business & ProductivityView app details

Magicley AI

Access a suite of AI generators for all your creative and productivity tasks.

AI & DataView app details

GPTHumanizer

Bypass AI detection with guaranteed undetectable content generation.

AI & DataView app details

Grubby.AI

Humanize AI text instantly to pass Turnitin and other detectors with ease.

Lifestyle & HealthView app details

Sellm

Track brand mentions across ChatGPT, Perplexity, and other AI platforms.

Marketing & CommerceView app details

Continue Readings

arXiv — cs.CV2 days ago

Where Does Vision Meet Language? Understanding and Refining Visual Fusion in MLLMs via Contrastive Attention

PositiveArtificial Intelligence

A recent study has explored the integration of visual and textual information in Multimodal Large Language Models (MLLMs), revealing that visual-text fusion occurs at specific layers within these models rather than uniformly across the network. The research highlights a late-stage

Read full article

via arXiv — cs.CV

arXiv — cs.CV2 days ago

Incentivizing Cardiologist-Like Reasoning in MLLMs for Interpretable Echocardiographic Diagnosis

PositiveArtificial Intelligence

A novel approach has been proposed to enhance echocardiographic diagnosis through the integration of a Cardiac Reasoning Template (CRT) and CardiacMind, aimed at improving the reasoning capabilities of multimodal large language models (MLLMs). This method addresses the challenges faced by existing models in capturing the relationship between quantitative measurements and clinical manifestations in cardiac screening.

Read full article

via arXiv — cs.CV

arXiv — cs.CV2 days ago

UR-Bench: A Benchmark for Multi-Hop Reasoning over Ultra-High-Resolution Images

NeutralArtificial Intelligence

The introduction of the Ultra-high-resolution Reasoning Benchmark (UR-Bench) aims to evaluate the reasoning capabilities of multimodal large language models (MLLMs) specifically on ultra-high-resolution images, which have been largely unexplored in existing visual question answering benchmarks. This benchmark features two main categories, Humanistic Scenes and Natural Scenes, with images ranging from hundreds of megapixels to gigapixels, accompanied by structured questions.

Read full article

via arXiv — cs.CV

arXiv — cs.CV2 days ago

M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding

PositiveArtificial Intelligence

The introduction of M3CoTBench marks a significant advancement in the evaluation of Chain-of-Thought (CoT) reasoning within Multimodal Large Language Models (MLLMs) specifically for medical image understanding, addressing the limitations of existing benchmarks that focus solely on final answers without considering the reasoning process.

Read full article

via arXiv — cs.CV

Ready to build your own newsroom?

Subscribe to unlock a personalised feed, podcasts, newsletters, and notifications tailored to the topics you actually care about