Artificial IntelligencearXiv — cs.CVTue, May 19, 2026, 4:00 AMNeutral

Are Multimodal LLMs Ready for Surveillance? A Reality Check on Zero-Shot Anomaly Detection in the Wild

Recent evaluations of Multimodal Large Language Models (MLLMs) have highlighted their potential for Video Anomaly Detection (VAD), yet their effectiveness in real-world applications remains uncertain. The study reformulates VAD as a binary classification task under weak temporal supervision, revealing a conservative bias in zero-shot settings where models favor the 'normal' class, leading to high precision but low recall.

WPN Brief

  • What Happened

    Recent evaluations of Multimodal Large Language Models (MLLMs) have highlighted their potential for Video Anomaly Detection (VAD), yet their effectiveness in real-world applications remains uncertain. The study reformulates VAD as a binary classification task under weak temporal supervision, revealing a conservative bias in zero-shot settings where models favor the 'normal' class, leading to high precision but low recall.

  • Why It Matters

    This development is significant as it underscores the limitations of MLLMs in practical surveillance scenarios, raising questions about their reliability in detecting anomalies in diverse environments. The findings suggest that while MLLMs show promise, their current performance may not meet the demands of real-world applications, particularly in safety-critical contexts.

  • The Bigger Picture

    The challenges faced by MLLMs in anomaly detection reflect broader concerns regarding their capabilities, including issues of visual representation degradation during training and the need for improved frameworks to enhance performance in complex visual tasks. These recurring themes highlight the ongoing debate about the readiness of advanced AI models for practical deployment in sensitive areas such as surveillance and safety.

Ask WPN AI