Source archive

arXiv — cs.CV

AIarxiv.org100 reports
Open source

Recent reports

Showing 1-12 of 100 reports

Jul 9, 2026arxiv.orgPositive

NavEYE: Vision-Centered Multi-Sensor Fusion-Based Situational Awareness System for Intelligent Surface Vehicles

The NavEYE system has been developed as a vision-centered multi-sensor fusion system aimed at enhancing situational awareness for intelligent surface vehicles (ISVs). By integrating data from various sensors, including AIS, radar, and RGB cameras, NavEYE seeks to improve navigation in complex environments, addressing the challenges posed by data loss and uncertainty.

Jul 9, 2026arxiv.orgNeutral

Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation

Visual generators are facing a significant challenge as they excel at rendering but often fabricate information beyond their training, which is based on fixed corpora. The introduction of SearchGen-20K and SearchGen-Bench, comprising 20,839 prompts across various domains, highlights the limitations of current models, scoring only 21 to 28 out of 100 on the benchmark.

Jul 9, 2026arxiv.orgNeutral

What if? Emulative Simulation with World Models for Situated Reasoning

A new dataset named WanderDream has been introduced to facilitate emulative simulation for situated reasoning, allowing agents to mentally simulate trajectories toward target situations without active exploration. This dataset includes 15.8K panoramic videos and 158K question-answer pairs, aimed at enhancing the reasoning capabilities of models in scenarios where physical exploration is limited.

Jul 9, 2026arxiv.orgPositive

Video-Based Detection of squint and cataract for accessibility-aware adaptive web interface rendering

A new study presents a real-time video-based detection system for squint and cataract, utilizing computer vision and image processing methods to enhance accessibility in web interfaces. The system employs a media-pipe face-mesh model to classify squint and assess cataract severity through video recordings from standard cameras. Experimental results indicate high accuracy rates of 98.39% for squint detection and 96.90% for cataract classification.

Jul 9, 2026arxiv.orgNeutral

Prior-matched evaluation of operational Earth-observation classifiers: a three-number reporting method demonstrated on Sentinel-1 internal-wave detection

The Internal Waves Service has implemented a prior-matched evaluation method for operational classifiers used in detecting internal solitary waves from Sentinel-1 satellite data. This approach addresses the discrepancies between balanced-test precision and real operational performance, revealing a significant gap in reported metrics.

Jul 9, 2026arxiv.orgNeutral

Comparative Study of Domain-adapted VLMs for General Document Visual Question Answering

A comparative study has been conducted on eight open-source pretrained Vision-Language Models (VLMs) for Document Visual Question Answering (DocVQA), evaluating their performance across three document domains: industrial documents, infographics, and presentation slides. The study assesses model capabilities through zero-shot evaluations, supervised finetuning, and few-shot learning. Findings indicate a decline in performance for complex visual layouts despite strong zero-shot baselines.

Jul 9, 2026arxiv.orgPositive

CompDiff: Hierarchical Compositional Diffusion for Fair and Zero-Shot Intersectional Medical Image Generation

A new framework named CompDiff has been introduced, focusing on hierarchical compositional diffusion to enhance the fairness and quality of medical image generation across diverse demographic groups. This approach addresses the imbalanced generator problem, which often leads to subpar image synthesis for underrepresented subgroups in medical datasets.

Jul 9, 2026arxiv.orgPositive

Rail Track Extraction from Rasterized Classified Point Clouds Using a Full-Resolution, Fully Convolutional Recurrent Neural Network

A novel method for rail track extraction from rasterized classified point clouds has been introduced, utilizing a fully convolutional recurrent neural network that maintains full spatial resolution and is trained on synthetically generated data. This approach enhances the quality of per-pixel data, crucial for effective railway asset management and maintenance.

Jul 9, 2026arxiv.orgPositive

Gen4U: Unifying Video Generation and Understanding via Diffusion

A new framework named Gen4U has been introduced to unify video generation and understanding through advanced video diffusion models, addressing previous limitations in capturing high-level semantics. This framework leverages structured latent spaces and attention mechanisms to enhance the decoding of visual representations across varying noise levels.