Artificial IntelligencearXiv — cs.CVThu, May 28, 2026, 4:00 AMPositive

From Pixels to Words -- Towards Native One-Vision Models at Scale

A new foundation model named NEO-ov has been introduced, which aims to enhance vision-language models (VLMs) by learning cross-frame and pixel-word correspondence in an end-to-end manner, eliminating the need for external encoders or auxiliary adapters. This model addresses the fragmentation of pixel-level signals and aims to improve performance in multi-image and video understanding.

WPN Brief

  • What Happened

    A new foundation model named NEO-ov has been introduced, which aims to enhance vision-language models (VLMs) by learning cross-frame and pixel-word correspondence in an end-to-end manner, eliminating the need for external encoders or auxiliary adapters. This model addresses the fragmentation of pixel-level signals and aims to improve performance in multi-image and video understanding.

  • Why It Matters

    The development of NEO-ov is significant as it narrows the performance gap with modular counterparts while excelling in fine-grained visual perception, indicating a shift towards more integrated and efficient AI models.

  • The Bigger Picture

    This advancement reflects a broader trend in AI research towards unified models that can handle complex tasks without relying on traditional modular frameworks, as seen in other recent innovations like VidPrism and Self-Prophetic Decoding, which also focus on enhancing multimodal capabilities and efficiency.

Ask WPN AI