Evaluating Large Language Models on Multimodal Chemistry Olympiad Exams

arXiv — cs.CV•Thursday, December 18, 2025 at 5:00:00 AM

NeutralArtificial Intelligence

A systematic evaluation of 40 multimodal large language models (LLMs), including GPT
This evaluation is significant as it highlights the limitations of current LLMs in integrating visual and textual reasoning, which is crucial for fields like chemistry where complex problem
The findings resonate with ongoing discussions about the capabilities of LLMs across various domains, including medical imaging and document understanding, emphasizing the need for improved models that can effectively handle multimodal data.

— via World Pulse Now AI Editorial System

Read Original

Was this article worth reading? Share it

LucidQuery AI

Combines diffusion reasoning with autoregressive LLM for advanced AI analysis.

AI & DataView app details

Airparser

Extract and parse data from documents using GPT-4 automation.

AI & DataView app details

Magicley AI

Access a suite of AI generators for all your creative and productivity tasks.

AI & DataView app details

ChatOne

Chat with multiple AI models like ChatGPT, Claude, and Gemini in one place.

AI & DataView app details

Langfuse

Debug, monitor, and improve your complex LLM applications with ease.

Tech & Developer ToolsView app details

ConsoleX

Connect to all major LLMs in one unified development playground.

Business & ProductivityView app details

Continue Readings

NYT — Technologya day ago

Can A.I. Generate New Ideas?

NeutralArtificial Intelligence

OpenAI has launched GPT-5.2, its latest AI model, which is designed to enhance productivity and has shown mixed results in tests compared to its predecessor, GPT-5.1. This development comes amid increasing competition from Google's Gemini 3, which has rapidly gained a significant user base.

Read full article

via NYT — Technology

arXiv — cs.CL2 days ago

Measuring Iterative Temporal Reasoning with Time Puzzles

NeutralArtificial Intelligence

The introduction of Time Puzzles marks a significant advancement in evaluating iterative temporal reasoning in large language models (LLMs). This task combines factual temporal anchors with cross-cultural calendar relations, generating puzzles that challenge LLMs' reasoning capabilities. Despite the simplicity of the dataset, models like GPT-5 achieved only 49.3% accuracy, highlighting the difficulty of the task.

Read full article

via arXiv — cs.CL

arXiv — cs.CL2 days ago

From Rows to Reasoning: A Retrieval-Augmented Multimodal Framework for Spreadsheet Understanding

PositiveArtificial Intelligence

A new framework called From Rows to Reasoning (FRTR) has been introduced to enhance the reasoning capabilities of Large Language Models (LLMs) when dealing with complex spreadsheets. This framework includes FRTR-Bench, a benchmark featuring 30 enterprise-grade Excel workbooks, which aims to improve the understanding of multimodal data by breaking down spreadsheets into granular components.

Read full article

via arXiv — cs.CL

arXiv — cs.CV2 days ago

KidVis: Do Multimodal Large Language Models Possess the Visual Perceptual Capabilities of a 6-Year-Old?

NeutralArtificial Intelligence

A new benchmark called KidVis has been introduced to evaluate the visual perceptual capabilities of Multimodal Large Language Models (MLLMs), specifically assessing their performance against that of 6-7 year old children across six atomic visual capabilities. The results reveal a significant performance gap, with human children scoring an average of 95.32 compared to GPT-5's score of 67.33.

Read full article

via arXiv — cs.CV

arXiv — cs.CV2 days ago

VideoHEDGE: Entropy-Based Hallucination Detection for Video-VLMs via Semantic Clustering and Spatiotemporal Perturbations

NeutralArtificial Intelligence

A new framework named VideoHEDGE has been introduced to detect hallucinations in video-capable vision-language models (Video-VLMs), addressing the frequent inaccuracies in video question answering. This system employs entropy-based reliability estimation and semantic clustering to evaluate the correctness of generated answers against video-question pairs.

Read full article

via arXiv — cs.CV

Ready to build your own newsroom?

Subscribe to unlock a personalised feed, podcasts, newsletters, and notifications tailored to the topics you actually care about