Artificial IntelligencearXiv — cs.CVMon, Jun 1, 2026, 4:00 AMPositive

Towards Consistent Video Geometry Estimation

A new foundation model named ViGeo has been introduced, designed to recover spatially dense and temporally consistent geometry from video sequences. This model utilizes a plain transformer architecture and features dynamic chunking attention, allowing it to adapt its attention patterns during inference without the need for retraining. Additionally, a completion-based data refinement framework enhances the quality of supervision by conditioning on sparse annotations.

WPN Brief

  • What Happened

    A new foundation model named ViGeo has been introduced, designed to recover spatially dense and temporally consistent geometry from video sequences. This model utilizes a plain transformer architecture and features dynamic chunking attention, allowing it to adapt its attention patterns during inference without the need for retraining. Additionally, a completion-based data refinement framework enhances the quality of supervision by conditioning on sparse annotations.

  • Why It Matters

    The development of ViGeo is significant as it enables improved video geometry estimation, which is crucial for applications in computer vision, robotics, and augmented reality. By providing a unified model capable of handling various video lengths and contexts, ViGeo positions itself as a versatile tool in the AI landscape.

  • The Bigger Picture

    This advancement reflects a growing trend in AI towards models that can efficiently process and analyze video data, addressing challenges such as temporal consistency and depth estimation. The integration of techniques like dynamic attention and data refinement showcases the ongoing innovation in the field, paralleling efforts in related areas such as video super-resolution and scene reconstruction.

Ask WPN AI

Related Reports

More coverage on this story

10 reports across the wire

arXiv — cs.CV
Jun 1

Linear Scaling Video VLMs for Long Video Understanding

A new method called StateKV has been introduced to enhance video vision-language models (VLMs) for long video understanding. This approach allows for linear-time video prefill by utilizing a fixed-capacity, importance-based recurrent state to carry cross-frame context, while maintaining a full per-frame cache for decoding. This innovation addresses the computational challenges posed by traditional spatiotemporal self-attention mechanisms, which increase latency and resource demands as the number of frames grows.

Artificial Intelligencepositive
arXiv — cs.CV
Jun 1

Learning Global Motion with Compact Gaussians for Feed-Forward 4D Reconstruction

A new framework, C4G, has been introduced for dynamic scene reconstruction from monocular video, addressing challenges in predicting 3D Gaussians pixel-wise for each frame. This method utilizes timestamp-conditioned learnable Gaussian query tokens to aggregate features across the full temporal context, enabling globally coherent motion modeling without the need for per-scene optimization.

Artificial Intelligenceneutral
arXiv — cs.CV
Jun 8

GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation

GEM-4D, a new geometry-enhanced video world model, has been introduced to improve robot manipulation by addressing the limitations of existing video world models that struggle with consistent physical point tracking over time. This model incorporates dense 4D correspondence supervision from a pretrained geometry foundation model, allowing it to generate videos that are both visually plausible and physically grounded for reliable action execution.

Artificial Intelligencepositive
arXiv — cs.CV
Jun 1

SLAP: The Semantic Least Action Principle for Variational Video-Language Modeling

A new paradigm in video-language modeling has emerged with the introduction of the Semantic Least Action Principle (SLAP), which addresses the limitations of existing Large Video-Language Models (LVLMs) by enforcing object persistence and semantic consistency over extended temporal horizons. This approach shifts from probabilistic generation to variational mechanics, modeling video trajectories on a Riemannian manifold governed by a Semantic Lagrangian.

Artificial Intelligenceneutral
arXiv — cs.CV
Jun 1

Feature-Optimized Vision for Adaptive 3D Scene Reconstruction

A new study presents an adaptive feature-optimized vision front end for 3D scene reconstruction, enhancing the process by scoring candidate features based on various criteria such as texture and distinctiveness. This method aims to maximize useful tracks while minimizing computational waste in 3D reconstruction tasks.

Artificial Intelligenceneutral
arXiv — cs.CV
Jun 1

DTG-Restore: Training-Free Diffusion Refinement for Generative Video Super-Resolution

A new framework called DTG-Restore has been introduced, which enhances the quality of distorted and low-resolution videos without the need for retraining. This method utilizes Decoupled Time Guidance (DTG) to separate conditional and unconditional signals in time, improving the restoration process in generative video super-resolution.

Artificial Intelligencepositive
arXiv — cs.CV
Jun 1

CameraNoise: Enabling Faithful Camera Control in Video Diffusion through Geometry-Flow-Guided Noise Warping

The recent introduction of CameraNoise presents a novel approach to video diffusion, focusing on precise camera pose control while addressing the challenges of maintaining geometric consistency. This method utilizes a flow-to-noise warping technique that embeds camera poses directly into the noise space, thereby preserving trajectory dynamics and ensuring consistent noise propagation during camera transformations.

Artificial Intelligencepositive
arXiv — cs.CV
Jun 1

DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory

Recent advancements in video generative models have led to the introduction of DecMem, a novel architecture designed to enhance consistent world generation by utilizing a decoupled memory system. This approach addresses key challenges in maintaining spatio-temporal consistency during long-horizon reasoning, significantly outperforming existing methods.

Artificial Intelligencepositive
arXiv — cs.LG
Jun 1

RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video

The introduction of RayDer, a unified feed-forward transformer, marks a significant advancement in self-supervised novel view synthesis (NVS) from real-world video. This model consolidates camera estimation, scene reconstruction, and rendering into a single framework, addressing the challenges of scaling NVS despite the availability of extensive video data.

Artificial Intelligencepositive
arXiv — cs.CV
Jun 1

DGSG-Mind: Dynamic 3D Gaussian Scene Graphs for Long-Term Scene Understanding and Grounding

DGSG-Mind introduces a novel hybrid instance-aware 3D Gaussian dynamic scene graph system aimed at enhancing long-term scene understanding and grounding for robotics. This system integrates open-vocabulary semantic information and employs a probabilistic voxel grid with explicit 3D Gaussians, addressing challenges in instance association and object-level topological changes.

Artificial Intelligencepositive