CARScenes: Semantic VLM Dataset for Safe Autonomous Driving

A new dataset named ClimateIQA has been introduced to enhance the capabilities of Vision-Language Models (VLMs) in analyzing meteorological anomalies. This dataset, which includes 26,280 high-quality images, aims to address the challenges faced by existing models like GPT-4o and Qwen-VL in interpreting complex meteorological heatmaps characterized by irregular shapes and color variations.

Read full article

via arXiv — cs.CV

arXiv — cs.CV2 days ago

HisTrackMap: Global Vectorized High-Definition Map Construction via History Map Tracking

PositiveArtificial Intelligence

HisTrackMap introduces a novel end-to-end tracking framework for global high-definition map construction, addressing the challenges of maintaining consistent temporal perception outcomes in autonomous driving systems. This framework utilizes historical trajectories of map elements to enhance the stability and reliability of map data collection.

Read full article

via arXiv — cs.CV

arXiv — cs.CV2 days ago

LLaVAction: evaluating and training multi-modal large language models for action understanding

PositiveArtificial Intelligence

The research titled 'LLaVAction' focuses on evaluating and training multi-modal large language models (MLLMs) for action understanding, reformulating the EPIC-KITCHENS-100 dataset into a benchmark for MLLMs. The study reveals that leading MLLMs struggle with recognizing correct actions when faced with difficult distractors, highlighting a gap in their fine-grained action understanding capabilities.

Read full article

via arXiv — cs.CV

arXiv — cs.CV2 days ago

DriveRX: A Vision-Language Reasoning Model for Cross-Task Autonomous Driving

PositiveArtificial Intelligence

DriveRX has been introduced as a vision-language reasoning model aimed at enhancing cross-task autonomous driving by addressing the limitations of traditional end-to-end models, which struggle with complex scenarios due to a lack of structured reasoning. This model is part of a broader framework called AutoDriveRL, which optimizes four core tasks through a unified training approach.

Read full article

via arXiv — cs.CV

arXiv — cs.CV2 days ago

Semantic Misalignment in Vision-Language Models under Perceptual Degradation

NeutralArtificial Intelligence

Recent research has highlighted significant semantic misalignment in Vision-Language Models (VLMs) when subjected to perceptual degradation, particularly through controlled visual perception challenges using the Cityscapes dataset. This study reveals that while traditional segmentation metrics show only moderate declines, VLMs exhibit severe failures in downstream tasks, including hallucinations and inconsistent safety judgments.

Read full article

via arXiv — cs.CV

arXiv — cs.CV2 days ago

Unifying Appearance Codes and Bilateral Grids for Driving Scene Gaussian Splatting

PositiveArtificial Intelligence

A novel approach has been proposed to unify appearance codes and bilateral grids for Gaussian Splatting in driving scene reconstruction, significantly enhancing geometric accuracy in dynamic environments. This method addresses the challenges of photometric consistency in real-world scenarios, which has been a limitation in existing neural rendering techniques like NeRF and Gaussian Splatting.

Read full article

via arXiv — cs.CV

arXiv — stat.ML2 days ago

Bayesian Multiobject Tracking With Neural-Enhanced Motion and Measurement Models

PositiveArtificial Intelligence

A new paper introduces a hybrid method for Bayesian Multiobject Tracking (MOT) that integrates neural networks to enhance traditional statistical models, addressing limitations in existing approaches. This development aims to improve performance in various applications, including autonomous driving and aerospace surveillance.

Read full article

via arXiv — stat.ML

Ready to build your own newsroom?

Subscribe to unlock a personalised feed, podcasts, newsletters, and notifications tailored to the topics you actually care about