arXiv:2511.11106v1 Announce Type: cross 
Abstract: Recent advancements in Audio-Video Large Language Models (AV-LLMs) have enhanced their capabilities in tasks like audio-visual question answering and multimodal dialog systems. Video and audio introduce an extended temporal dimension, resulting in a larger key-value (KV) cache compared to static image embedding. A naive optimization strategy is to selectively focus on and retain KV caches of audio or video based on task. However, in the experiment, we observed that the attention of AV-LLMs to various modalities in the high layers is not strictly dependent on the task. In higher layers, the attention of AV-LLMs shifts more towards the video modality. In addition, we also found that directly integrating temporal KV of audio and spatial-temporal KV of video may lead to information confusion and significant performance degradation of AV-LLMs. If audio and video are processed indiscriminately, it may also lead to excessive compression or reservation of a certain modality, thereby disrupting the alignment between modalities. To address these challenges, we propose AccKV, an Adaptive-Focusing and Cross-Calibration KV cache optimization framework designed specifically for efficient AV-LLMs inference. Our method is based on layer adaptive focusing technology, selectively focusing on key modalities according to the characteristics of different layers, and enhances the recognition of heavy hitter tokens through attention redistribution. In addition, we propose a Cross-Calibration technique that first integrates inefficient KV caches within the audio and video modalities, and then aligns low-priority modalities with high-priority modalities to selectively evict KV cache of low-priority modalities. The experimental results show that AccKV can significantly improve the computational efficiency of AV-LLMs while maintaining accuracy.

أدت التطورات الأخيرة في نماذج اللغة الصوتية-المرئية (AV-LLMs) إلى تحسين أدائها في مهام مثل الإجابة على الأسئلة الصوتية-المرئية وأنظمة الحوار متعددة الوسائط. تسلط الدراسة الضوء على أن ذاكرة التخزين المؤقت للقيم الرئيسية (KV) لنماذج AV-LLMs أكبر بسبب البعد الزمني الممتد الذي تقدمه الفيديو والصوت. وقد لوحظ أن انتباه نماذج AV-LLMs يتحول نحو نمط الفيديو في الطبقات العليا، وأن دمج ذاكرات التخزين المؤقت الصوتية والمرئية بشكل عشوائي قد يؤدي إلى تدهور الأداء.

Los recientes avances en los Modelos de Lenguaje de Audio-Vídeo (AV-LLMs) han mejorado su rendimiento en tareas como la respuesta a preguntas audio-visuales y los sistemas de diálogo multimodal. El estudio destaca que la caché de clave-valor (KV) para los AV-LLMs es más grande debido a la dimensión temporal extendida introducida por el video y el audio. Se observó que la atención de los AV-LLMs se desplaza hacia la modalidad de video en las capas superiores, y la integración indiscriminada de las cachés KV de audio y video puede llevar a una degradación del rendimiento.

Les récentes avancées dans les modèles de langage audio-vidéo (AV-LLMs) ont amélioré leur performance dans des tâches telles que la réponse à des questions audio-visuelles et les systèmes de dialogue multimodal. L'étude souligne que le cache clé-valeur (KV) pour les AV-LLMs est plus grand en raison de la dimension temporelle étendue introduite par la vidéo et l'audio. Il a été constaté que l'attention des AV-LLMs se déplace vers la modalité vidéo dans les couches supérieures, et l'intégration indiscriminée des caches KV audio et vidéo peut entraîner une dégradation des performances.

Recent advancements in Audio-Video Large Language Models (AV-LLMs) have improved their performance in tasks such as audio-visual question answering and multimodal dialog systems. The study highlights that the key-value (KV) cache for AV-LLMs is larger due to the extended temporal dimension introduced by video and audio. It was found that the attention of AV-LLMs shifts towards the video modality in higher layers, and integrating audio and video KV caches indiscriminately can lead to performance degradation.

AccKV: Towards Efficient Audio-Video LLMs Inference via Adaptive-Focusing and Cross-Calibration KV Cache Optimization

Guardio is leveraging its experience building browser extensions and apps that scan for malicious and phishing sites to build a tool that looks for artifacts in code and websites made with vibe coding tools.

نجحت شركة Guardio الناشئة في مجال الأمن السيبراني في تأمين 80 مليون دولار من التمويل من ION Crossover Partners. تشتهر الشركة بخبرتها في تطوير إضافات المتصفح والتطبيقات التي تكشف عن المواقع الضارة والمحتالة. تخطط Guardio لاستخدام هذا التمويل لإنشاء أداة تبحث عن الآثار في الشيفرة والمواقع التي تم إنشاؤها باستخدام أدوات البرمجة vibe.

La startup de seguridad Guardio ha asegurado 80 millones de dólares en financiamiento de ION Crossover Partners. La empresa es conocida por su experiencia en el desarrollo de extensiones de navegador y aplicaciones que detectan sitios maliciosos y de phishing. Guardio planea utilizar estos fondos para crear una herramienta que identifique artefactos en el código y sitios web construidos con herramientas de codificación vibe.

La startup de sécurité Guardio a obtenu 80 millions de dollars de financement de la part d'ION Crossover Partners. Connue pour son expertise dans le développement d'extensions de navigateur et d'applications détectant les sites malveillants et de phishing, Guardio prévoit d'utiliser ce financement pour créer un outil identifiant les artefacts dans le code et les sites web construits avec des outils de codage vibe.

Security startup Guardio has secured $80 million in funding from ION Crossover Partners. The company is known for its expertise in developing browser extensions and applications that detect malicious and phishing websites. Guardio plans to utilize this funding to create a tool that identifies artifacts in code and websites built with vibe coding tools.

Security startup Guardio nabs $80M from ION Crossover Partners

A new artificial intelligence startup founded by the creators of <a href="https://opencv.org/">the world&#x27;s most widely used computer vision library</a> has emerged from stealth with technology that generates realistic human-centric videos up to five minutes long — a dramatic leap beyond the capabilities of rivals including OpenAI&#x27;s <a href="https://openai.com/sora/">Sora</a> and Google&#x27;s <a href="https://deepmind.google/models/veo/">Veo</a>.<a href="https://craftstory.com/">CraftStory</a>, which launched Tuesday with $2 million in funding, is introducing Model 2.0, a video generation system that addresses one of the most significant limitations plaguing the nascent AI video industry: duration. While OpenAI&#x27;s <a href="https://openai.com/index/sora-2/">Sora 2</a> tops out at 25 seconds and most competing models generate clips of 10 seconds or less, CraftStory&#x27;s system can produce continuous, coherent video performances that run as long as a typical YouTube tutorial or product demonstration.The breakthrough could unlock substantial commercial value for enterprises struggling to scale video production for training, marketing, and customer education — markets where brief AI-generated clips have proven inadequate despite their visual polish.&quot;If you really try to create a video with one of these video generation systems, you find that a lot of the times you want to implement a certain creative vision, and regardless of how detailed the instructions are, the systems basically ignore a part of your instructions,&quot; said Victor Erukhimov, CraftStory&#x27;s founder and CEO, in an exclusive interview with VentureBeat. &quot;We developed a system that can generate videos basically as long as you need them.&quot;<h3>How parallel processing solves the long-form video problem</h3>CraftStory&#x27;s advance rests on what the company describes as a parallelized diffusion architecture — a fundamentally different approach to how AI models generate video compared to the sequential methods employed by most competitors.Traditional video generation models work by running diffusion algorithms on increasingly large three-dimensional volumes where time represents the third axis. To generate a longer video, these models require proportionally larger networks, more training data, and significantly more computational resources.<a href="https://craftstory.com/">CraftStory</a> instead runs multiple smaller diffusion algorithms simultaneously across the entire duration of the video, with bidirectional constraints connecting them. &quot;The latter part of the video can influence the former part of the video too,&quot; Erukhimov explained. &quot;And this is pretty important, because if you do it one by one, then an artifact that appears in the first part propagates to the second one, and then it accumulates.&quot;Rather than generating eight seconds and then stitching on additional segments, CraftStory&#x27;s system processes all five minutes concurrently through interconnected diffusion processes.Crucially, CraftStory trained its model on proprietary footage rather than relying solely on internet-scraped videos. The company hired studios to shoot actors using high-frame-rate camera systems that capture crisp detail even in fast-moving elements like fingers — avoiding the motion blur inherent in standard 30-frames-per-second YouTube clips.&quot;What we showed is that you don&#x27;t need a lot of data and you don&#x27;t need a lot of training budget to create high quality videos,&quot; Erukhimov said. &quot;You just need high quality data.&quot;Model 2.0 currently operates as a video-to-video system: users upload a still image to animate and a &quot;driving video&quot; containing a person whose movements the AI will replicate. CraftStory provides preset driving videos shot with professional actors, who receive revenue shares when their motion data is used, or users can upload their own footage.The system generates 30-second clips at low resolution in approximately 15 minutes. An advanced lip-sync system synchronizes mouth movements to scripts or audio tracks, while gesture alignment algorithms ensure body language matches speech rhythm and emotional tone.<h3>Fighting a war chest battle with $2 million against billions</h3>CraftStory&#x27;s funding comes almost entirely from <a href="https://finance.yahoo.com/news/2-25-billion-exit-taught-130300997.html">Andrew Filev</a>, who sold his project management software company Wrike to Citrix for <a href="https://techcrunch.com/2021/01/19/citrix-is-acquiring-wrike-from-vista-for-2-25b/">$2.25 billion</a> in 2021 and now runs <a href="https://zencoder.ai/">Zencoder</a>, an AI coding company. The modest raise stands in stark contrast to the billions flowing into competing efforts — OpenAI has <a href="https://www.reuters.com/technology/artificial-intelligence/openai-closes-66-billion-funding-haul-valuation-157-billion-with-investment-2024-10-02/">raised over $6 billion</a> in its latest funding round alone.Erukhimov pushed back on the notion that massive capital is prerequisite for success. &quot;I don&#x27;t necessarily buy the thesis that compute is the path to success,&quot; he said. &quot;It definitely helps if you have compute. But if you raise a billion dollars on a PowerPoint, in the end, no one is happy, neither the founders nor the investors.&quot;Filev defended the David-versus-Goliath approach. &quot;When you invest in startups, you&#x27;re fundamentally betting on people,&quot; he said in an interview with VentureBeat. &quot;To paraphrase Margaret Mead: never underestimate what a small group of thoughtful, committed engineers and scientists can build.&quot;He argued that CraftStory benefits from a focused strategy. &quot;The big labs are in an arms race to build general-purpose video foundation models,&quot; Filev said. &quot;CraftStory is riding that wave and going very deep into a specific format: long-form, engaging, human-centric video.&quot;<h3>Why computer vision expertise matters in generative AI video</h3>Erukhimov&#x27;s credibility stems from his deep roots in computer vision rather than the transformer architectures that have dominated recent AI advances. He was an early contributor to <a href="https://opencv.org/">OpenCV</a> — the Open Source Computer Vision Library that has become the de facto standard for computer vision applications, with over <a href="https://github.com/opencv/opencv">84,000 stars on GitHub</a>.When Intel reduced its support for OpenCV in the mid-2000s, Erukhimov co-founded Itseez with the explicit goal of maintaining and advancing the library. The company expanded OpenCV significantly and pivoted toward automotive safety systems before Intel acquired it in 2016.Filev said this background is precisely what makes Erukhimov well-positioned for video generation. &quot;What people sometimes miss is that generative AI video isn&#x27;t just about the generative part. It&#x27;s about understanding motion, facial dynamics, temporal coherence, and how humans actually move,&quot; Filev said. &quot;Victor has spent his career mastering exactly those problems.&quot;<h3>Enterprise focus targets training videos and product demos</h3>While much of the public excitement around AI video generation has centered on creative tools for consumers, CraftStory is pursuing a decidedly enterprise-focused strategy.&quot;We are definitely thinking about B2B more than consumer,&quot; Erukhimov said. &quot;We&#x27;re thinking about companies, specifically software companies, being able to make cool training videos and product videos and launch videos.&quot;The logic is straightforward: corporate training, product tutorials, and customer education videos often run several minutes and require consistent quality throughout. A 10-second AI clip cannot effectively demonstrate how to use enterprise software or explain a complex product feature.&quot;If you need a longer-form video, then you should go with us,&quot; Erukhimov said. &quot;We can create up to five minutes, consistent video, high quality.&quot;Filev echoed this assessment. &quot;One huge gap in this market is the lack of models that can generate consistent videos over longer sequences — and that&#x27;s extremely important for real-world use,&quot; he said. &quot;If you&#x27;re creating a commercial for your company, a 10-second video, no matter how good it looks, just isn&#x27;t enough. You need 30 seconds, you need two minutes — you need more.&quot;The company anticipates cost savings for customers. Filev suggested that &quot;a small business owner could create content in minutes that previously would have cost $20,000 and taken two months to produce.&quot;CraftStory is also courting creative agencies that produce video content for corporate clients, with the value proposition centered on cost and speed: agencies can record an actor on camera and transform that footage into a finished AI video, rather than managing expensive multi-day shoots.The next major development on CraftStory&#x27;s roadmap is a text-to-video model that would allow users to generate long-form content directly from scripts. The team is also developing support for moving-camera scenarios, including the popular &quot;walk-and-talk&quot; format common in high-end advertising.<h3>Where CraftStory fits in a fragmented competitive landscape</h3>CraftStory enters a crowded and rapidly evolving market. OpenAI&#x27;s <a href="https://openai.com/index/sora-2/">Sora 2</a>, while not yet publicly available, has generated significant buzz. Google&#x27;s <a href="https://deepmind.google/models/veo/">Veo models</a> are advancing quickly. <a href="https://runwayml.com/">Runway</a>, <a href="https://pika.art/login">Pika</a>, and <a href="https://stability.ai/">Stability AI </a>all offer video generation tools with different capabilities.Erukhimov acknowledged the competitive pressure but emphasized that CraftStory serves a distinct niche focused on human-centric videos. He positioned rapid innovation and market capture as the company&#x27;s primary strategy rather than relying on technical moats.Filev sees the market fragmenting into distinct layers, with large tech companies serving as &quot;API providers of powerful, general-purpose generation models&quot; while specialized players like CraftStory focus on specific use cases. &quot;If the big players are building the engines, CraftStory is building the production studio and assembly line on top,&quot; he said.Model 2.0 is available now at app.craftstory.com/model-2.0, with the company offering early access to users and enterprises interested in testing the technology. Whether a lightly-funded startup can capture meaningful market share against deep-pocketed incumbents remains uncertain, but Erukhimov is characteristically confident about the opportunity ahead.&quot;AI-generated video will soon become the primary way companies communicate their stories,&quot; he said.

أطلقت CraftStory، وهي شركة ناشئة جديدة في مجال الذكاء الاصطناعي أسسها مبتكرو OpenCV، نظامًا لتوليد الفيديو قادرًا على إنتاج مقاطع فيديو واقعية تركز على الإنسان تصل مدتها إلى خمس دقائق. تتجاوز هذه التكنولوجيا بشكل كبير قدرات المنافسين مثل Sora من OpenAI وVeo من Google، الذين لديهم حدود زمنية أقصر. حصلت الشركة الناشئة على تمويل بقيمة مليوني دولار لدعم نهجها المبتكر في صناعة الفيديو بالذكاء الاصطناعي.

CraftStory, una nueva startup de IA fundada por los creadores de OpenCV, ha lanzado un sistema de generación de video capaz de producir videos realistas centrados en humanos de hasta cinco minutos de duración. Esta tecnología supera significativamente a competidores como Sora de OpenAI y Veo de Google, que tienen límites de duración más cortos. La startup ha asegurado 2 millones de dólares en financiamiento para apoyar su enfoque innovador en la industria de video de IA.

CraftStory, une nouvelle startup d'IA fondée par les créateurs d'OpenCV, a lancé un système de génération vidéo capable de produire des vidéos réalistes centrées sur l'humain d'une durée allant jusqu'à cinq minutes. Cette technologie surpasse considérablement les concurrents tels que Sora d'OpenAI et Veo de Google, qui ont des limites de durée plus courtes. La startup a sécurisé 2 millions de dollars de financement pour soutenir son approche innovante de l'industrie vidéo IA.

CraftStory, a new AI startup founded by the creators of OpenCV, has launched a video generation system capable of producing realistic human-centric videos up to five minutes long. This technology significantly outpaces competitors like OpenAI's Sora and Google's Veo, which have shorter duration limits. The startup has secured $2 million in funding to support its innovative approach to the AI video industry.

OpenCV founders launch AI video startup to take on OpenAI and Google

<a href="https://petapixel.com/2025/11/19/remote-cameras-may-have-captured-first-recorded-tool-use-by-a-wild-wolf/"><img width="1600" height="840" src="https://petapixel.com/assets/uploads/2025/11/wolve-tool-use.jpg" class="attachment-card-large size-card-large wp-post-image" alt="A wolf stands at the edge of a rocky shoreline, holding an orange and white fishing bobber in its mouth. A fishing net and additional gear are lying on the ground nearby. Rippling water is in the background." decoding="async" fetchpriority="high" /></a>Remote cameras have captured footage of wild wolves pulling crab traps out of the sea by their lines to eat the bait inside -- in the first evidence of possible tool use by the canines.
[<a href="https://petapixel.com/2025/11/19/remote-cameras-may-have-captured-first-recorded-tool-use-by-a-wild-wolf/">Read More</a>]

التقطت الكاميرات عن بُعد لقطات لذئاب برية تستخدم تقنية تتضمن سحب فخاخ السلطعون من البحر للوصول إلى الطعم داخلها. وهذا يمثل أول دليل موثق على استخدام محتمل للأدوات من قبل هذه الكلاب.

Cámaras remotas han grabado a lobos salvajes utilizando una técnica que consiste en sacar trampas de cangrejos del mar para acceder al cebo en su interior. Esto marca la primera evidencia documentada de un posible uso de herramientas por parte de estos caninos.

Des caméras à distance ont enregistré des loups sauvages utilisant une technique qui consiste à tirer des pièges à crabes de la mer pour accéder à l'appât à l'intérieur. Cela marque la première preuve documentée d'un potentiel usage d'outils par ces canidés.

Remote cameras have recorded wild wolves using a technique that involves pulling crab traps from the sea to access the bait inside. This marks the first documented evidence of potential tool use by these canines.

Remote Cameras May Have Captured First Recorded Tool Use by a Wild Wolf

The AI startup Firebird Inc. has received US government approval to export Nvidia Corp. chips to Armenia for a supercomputer project in the country, part of a global push to expand artificial intelligence infrastructure.

حصلت شركة Firebird Inc. الناشئة في مجال الذكاء الاصطناعي على موافقة الحكومة الأمريكية لتصدير شرائح Nvidia Corp. إلى أرمينيا، مما يسهل إنشاء مشروع حاسوب فائق في البلاد. تأتي هذه المبادرة كجزء من جهد عالمي أوسع لتعزيز بنية الذكاء الاصطناعي التحتية.

La startup de IA Firebird Inc. ha recibido la aprobación del gobierno de EE. UU. para exportar chips de Nvidia Corp. a Armenia, facilitando el establecimiento de un proyecto de supercomputadora en el país. Esta iniciativa forma parte de un esfuerzo global más amplio para mejorar la infraestructura de inteligencia artificial.

La startup d'IA Firebird Inc. a obtenu l'approbation du gouvernement américain pour exporter des puces de Nvidia Corp. en Arménie, facilitant ainsi l'établissement d'un projet de superordinateur dans le pays. Cette initiative s'inscrit dans un effort mondial plus large pour améliorer l'infrastructure de l'intelligence artificielle.

AI startup Firebird Inc. has obtained approval from the US government to export Nvidia Corp. chips to Armenia, facilitating the establishment of a supercomputer project in the country. This initiative is part of a broader global effort to enhance artificial intelligence infrastructure.

AI Startup Firebird Gets US Approval to Use Nvidia Chips in Armenian Data Center

Gesture of goodwill temporarily cools EU–China dispute that rattled global car supply chains.
The post <a href="https://www.techrepublic.com/article/news-netherlands-nexperia-chipmaker-control/">Netherlands Pauses Move to Seize Chinese-Owned Chipmaker Nexperia</a> appeared first on <a href="https://www.techrepublic.com">TechRepublic</a>.

أوقفت هولندا مؤقتًا جهودها للاستيلاء على شركة Nexperia لصناعة الرقائق المملوكة للصين، وهو إجراء يُعتبر بمثابة لفتة حسن نية تهدف إلى تهدئة التوترات في النزاع القائم بين الاتحاد الأوروبي والصين. تأتي هذه الخطوة وسط مخاوف بشأن تأثيرها على سلاسل الإمداد العالمية التي تأثرت بالفعل بنقص أشباه الموصلات.

Los Países Bajos han detenido temporalmente sus esfuerzos por apoderarse del fabricante de chips Nexperia, de propiedad china, un gesto que se considera como una buena voluntad para aliviar las tensiones en la disputa entre la UE y China. Esta decisión se produce en medio de preocupaciones sobre el impacto en las cadenas de suministro globales, que ya se han visto afectadas por la escasez de semiconductores.

Les Pays-Bas ont temporairement suspendu leurs efforts pour saisir le fabricant de puces Nexperia, détenu par des Chinois, un geste perçu comme une volonté d'apaiser les tensions dans le conflit en cours entre l'UE et la Chine. Cette décision intervient dans un contexte d'inquiétudes concernant l'impact sur les chaînes d'approvisionnement mondiales, déjà affectées par des pénuries de semi-conducteurs.

The Netherlands has temporarily halted its efforts to seize control of the Chinese-owned chipmaker Nexperia, a move seen as a gesture of goodwill aimed at easing tensions in the ongoing EU-China dispute. This decision comes amid concerns over the impact on global car supply chains, which have been affected by semiconductor shortages.

Netherlands Pauses Move to Seize Chinese-Owned Chipmaker Nexperia

استقال لاري سامرز من مجلس إدارة OpenAI، كما أفاد نيويورك تايمز. تأتي استقالته بعد التدقيق في اتصالاته السابقة مع المدان بجريمة الاعتداء الجنسي جيفري إبستين. تمثل هذه الخطوة انسحابًا كبيرًا لسامرز من الأدوار العامة وسط انتقادات متزايدة.

Larry Summers ha renunciado a la junta de OpenAI, según informa The New York Times. Su renuncia se produce tras el escrutinio sobre sus comunicaciones pasadas con el delincuente sexual condenado Jeffrey Epstein. Esta decisión marca un paso significativo en el retiro de Summers de los roles públicos en medio de una creciente crítica.

Larry Summers a démissionné du conseil d'administration d'OpenAI, selon le New York Times. Sa démission fait suite à un examen minutieux de ses communications passées avec le délinquant sexuel condamné Jeffrey Epstein. Cette décision marque un retrait significatif de Summers de ses rôles publics face à une critique croissante.

Larry Summers has resigned from the board of OpenAI, as reported by The New York Times. His resignation follows scrutiny over his past communications with convicted sex offender Jeffrey Epstein. This decision marks a significant step in Summers' withdrawal from public roles amid growing criticism.

Larry Summers Resigns From OpenAI’s Board

arXiv:2511.14109v1 Announce Type: new 
Abstract: Visual Place Recognition (VPR) aims to match query images against a database using visual cues. State-of-the-art methods aggregate features from deep backbones to form global descriptors. Optimal transport-based aggregation methods reformulate feature-to-cluster assignment as a transport problem, but the standard Sinkhorn algorithm symmetrically treats source and target marginals, limiting effectiveness when image features and cluster centers exhibit substantially different distributions. We propose an asymmetric aggregation VPR method with geometric constraints for locally aggregated descriptors, called $A^2$GC-VPR. Our method employs row-column normalization averaging with separate marginal calibration, enabling asymmetric matching that adapts to distributional discrepancies in visual place recognition. Geometric constraints are incorporated through learnable coordinate embeddings, computing compatibility scores fused with feature similarities, thereby promoting spatially proximal features to the same cluster and enhancing spatial awareness. Experimental results on MSLS, NordLand, and Pittsburgh datasets demonstrate superior performance, validating the effectiveness of our approach in improving matching accuracy and robustness.

$A^2$GC-VPR هو أسلوب جديد للتعرف على الأماكن البصرية (VPR) يتناول قيود أساليب التجميع التقليدية في مطابقة صور الاستعلام مع قاعدة بيانات. من خلال اعتماد نهج تجميع غير متماثل مع قيود هندسية، يعزز هذا الأسلوب فعالية مطابقة الميزات، خاصة عند التعامل مع توزيعات متباينة لميزات الصورة ومراكز التجمع. تستخدم التقنية متوسطات تطبيع الصفوف والأعمدة مع تضمينات إحداثيات قابلة للتعلم لتحسين درجات التوافق لوصفيات التجميع المحلي.

$A^2$GC-VPR es un nuevo método para el Reconocimiento Visual de Lugares (VPR) que aborda las limitaciones de los métodos de agregación tradicionales al emparejar imágenes de consulta con una base de datos. Al emplear un enfoque de agregación asimétrica con restricciones geométricas, este método mejora la efectividad del emparejamiento de características, especialmente cuando se enfrentan a distribuciones variables de características de imagen y centros de clúster. La técnica utiliza promedios de normalización fila-columna y embeddings de coordenadas aprendibles para mejorar las puntuaciones de…

$A^2$GC-VPR est une nouvelle méthode pour la reconnaissance de lieux visuels (VPR) qui s'attaque aux limites des méthodes d'agrégation traditionnelles dans l'appariement d'images de requête à une base de données. En adoptant une approche d'agrégation asymétrique avec des contraintes géométriques, cette méthode améliore l'efficacité de l'appariement des caractéristiques, en particulier lorsqu'il s'agit de distributions variées des caractéristiques d'image et des centres de clusters. La technique utilise une moyenne de normalisation ligne-colonne et des embeddings de coordonnées apprenables pour…

$A^2$GC-VPR is a new method for Visual Place Recognition (VPR) that addresses the limitations of traditional aggregation methods in matching query images to a database. By employing an asymmetric aggregation approach with geometric constraints, this method enhances the effectiveness of feature matching, particularly when dealing with varying distributions of image features and cluster centers. The technique utilizes row-column normalization averaging and learnable coordinate embeddings to improve compatibility scores for locally aggregated descriptors.

$A^2$GC: $A$symmetric $A$ggregation with Geometric Constraints for Locally Aggregated Descriptors

arXiv:2511.14247v1 Announce Type: new 
Abstract: Multi-agents rely on accurate poses to share and align observations, enabling a collaborative perception of the environment. However, traditional GNSS-based localization often fails in GNSS-denied environments, making consistent feature alignment difficult in collaboration. To tackle this challenge, we propose a robust GNSS-free collaborative perception framework based on LiDAR localization. Specifically, we propose a lightweight Pose Generator with Confidence (PGC) to estimate compact pose and confidence representations. To alleviate the effects of localization errors, we further develop the Pose-Aware Spatio-Temporal Alignment Transformer (PASTAT), which performs confidence-aware spatial alignment while capturing essential temporal context. Additionally, we present a new simulation dataset, V2VLoc, which can be adapted for both LiDAR localization and collaborative detection tasks. V2VLoc comprises three subsets: Town1Loc, Town4Loc, and V2VDet. Town1Loc and Town4Loc offer multi-traversal sequences for training in localization tasks, whereas V2VDet is specifically intended for the collaborative detection task. Extensive experiments conducted on the V2VLoc dataset demonstrate that our approach achieves state-of-the-art performance under GNSS-denied conditions. We further conduct extended experiments on the real-world V2V4Real dataset to validate the effectiveness and generalizability of PASTAT.

يقدم المقال إطارًا جديدًا للإدراك التعاوني بدون GNSS باستخدام تحديد المواقع بواسطة LiDAR، حيث يتناول التحديات التي تواجهها البيئات التي تفتقر إلى GNSS. غالبًا ما تواجه طرق تحديد المواقع التقليدية صعوبات في هذه البيئات، مما يعيق التعاون الفعال بين أنظمة الوكلاء المتعددة. تتضمن الحلول المقترحة مولد وضع خفيف الوزن مع ثقة (PGC) لتقدير الأوضاع وتمثيلات الثقة، بالإضافة إلى محول التوافق الزماني المكاني الواعي بالوضع (PASTAT) الذي يقوم بأداء التوافق المكاني مع مراعاة الثقة. كما تم تقديم مجموعة بيانات محاكاة جديدة، V2VLoc، التي يمكن تكييفها لمهام تحديد المواقع بواسطة LiDAR والاكتشاف التعاوني.

El artículo presenta un nuevo marco para la percepción colaborativa sin GNSS utilizando la localización por LiDAR, abordando los desafíos que se enfrentan en entornos sin GNSS. Los métodos de localización tradicionales a menudo tienen dificultades en estos entornos, lo que dificulta la colaboración efectiva entre sistemas multiagente. La solución propuesta incluye un Generador de Pose con Confianza (PGC) para estimar poses y confianza, junto con el Transformador de Alineación Espacio-Temporal Consciente de la Pose (PASTAT) para el alineamiento espacial. Se introduce un nuevo conjunto de datos …

L'article présente un nouveau cadre pour la perception collaborative sans GNSS utilisant la localisation par LiDAR, abordant les défis rencontrés dans les environnements privés de GNSS. Les méthodes de localisation traditionnelles peinent souvent dans ces contextes, entravant la collaboration efficace entre systèmes multi-agents. La solution proposée comprend un générateur de pose léger avec confiance (PGC) pour estimer les poses et la confiance, ainsi qu'un transformateur d'alignement spatio-temporel conscient de la pose (PASTAT) pour l'alignement spatial. Un nouveau jeu de données de simulat…

The article presents a new framework for GNSS-free collaborative perception using LiDAR localization, addressing the challenges faced in GNSS-denied environments. Traditional localization methods often struggle in these settings, hindering effective collaboration among multi-agent systems. The proposed solution includes a lightweight Pose Generator with Confidence (PGC) for estimating poses and confidence, alongside the Pose-Aware Spatio-Temporal Alignment Transformer (PASTAT) for spatial alignment. A new simulation dataset, V2VLoc, is introduced, which supports LiDAR localization and collabor…

V2VLoc: Robust GNSS-Free Collaborative Perception via LiDAR Localization

arXiv:2511.14210v1 Announce Type: cross 
Abstract: We introduce Orion, a visual agent framework that can take in any modality and generate any modality. Using an agentic framework with multiple tool-calling capabilities, Orion is designed for visual AI tasks and achieves state-of-the-art results. Unlike traditional vision-language models that produce descriptive outputs, Orion orchestrates a suite of specialized computer vision tools, including object detection, keypoint localization, panoptic segmentation, Optical Character Recognition, and geometric analysis, to execute complex multi-step visual workflows. The system achieves competitive performance on MMMU, MMBench, DocVQA, and MMLongBench while extending monolithic vision-language models to production-grade visual intelligence. By combining neural perception with symbolic execution, Orion enables autonomous visual reasoning, marking a transition from passive visual understanding to active, tool-driven visual intelligence.

أورايون هو إطار جديد لوكيل بصري قادر على معالجة وتوليد أنماط متعددة. يستخدم إطارًا وكيلًا مع قدرات متعددة لاستدعاء الأدوات، محققًا نتائج رائدة في مهام الذكاء الاصطناعي البصري. على عكس نماذج الرؤية-اللغة التقليدية، يستخدم أورايون أدوات رؤية حاسوبية متخصصة لتنفيذ سير عمل بصري معقد، محققًا أداءً تنافسيًا في معايير مثل MMMU وMMBench وDocVQA وMMLongBench. يمثل هذا النظام تحولًا نحو الاستدلال البصري المستقل، مما يعزز الذكاء البصري.

Orion es un nuevo marco de agente visual capaz de procesar y generar diversas modalidades. Utiliza un marco agentivo con múltiples capacidades de llamada a herramientas, logrando resultados de vanguardia en tareas de IA visual. A diferencia de los modelos tradicionales de visión-lenguaje, Orion emplea herramientas especializadas de visión por computadora para flujos de trabajo visuales complejos, alcanzando un rendimiento competitivo en benchmarks como MMMU, MMBench, DocVQA y MMLongBench. Este sistema marca una transición hacia el razonamiento visual autónomo, mejorando la inteligencia visual.

Orion est un nouveau cadre d'agent visuel capable de traiter et de générer diverses modalités. Il utilise un cadre agentique avec plusieurs capacités d'appel d'outils, atteignant des résultats de pointe dans les tâches d'IA visuelle. Contrairement aux modèles traditionnels de vision-langage, Orion utilise des outils de vision par ordinateur spécialisés pour des flux de travail visuels complexes, obtenant des performances compétitives sur des benchmarks tels que MMMU, MMBench, DocVQA et MMLongBench. Ce système marque un tournant vers le raisonnement visuel autonome, améliorant l'intelligence vi…

Orion is a newly introduced visual agent framework capable of processing and generating various modalities. It employs an agentic framework with multiple tool-calling capabilities, achieving state-of-the-art results in visual AI tasks. Unlike traditional vision-language models, Orion utilizes specialized computer vision tools for complex visual workflows, achieving competitive performance on benchmarks like MMMU, MMBench, DocVQA, and MMLongBench. This system marks a shift towards autonomous visual reasoning, enhancing visual intelligence.

AccKV: Towards Efficient Audio-Video LLMs Inference via Adaptive-Focusing and Cross-Calibration KV Cache Optimization

Was this article worth reading? Share it