Findings of the Fourth Shared Task on Multilingual Coreference Resolution: Can LLMs Dethrone Traditional Approaches?

arXiv — cs.CL•Friday, November 7, 2025 at 5:00:00 AM

The fourth edition of the Shared Task on Multilingual Coreference Resolution has showcased exciting advancements in the field, particularly with the introduction of a dedicated Large Language Model (LLM) track. This year's competition, part of the CODI-CRAC 2025 workshop, challenged participants to refine their systems for identifying and clustering mentions based on identity coreference. The focus on LLMs highlights the growing importance of these models in tackling complex linguistic tasks, potentially reshaping how we approach language processing in multilingual contexts.

— via World Pulse Now AI Editorial System

Read Original

Was this article worth reading? Share it

Recommended Readings

arXiv — cs.LG2 days ago

On the Limitations of Language Targeted Pruning: Investigating the Calibration Language Impact in Multilingual LLM Pruning

NeutralArtificial Intelligence

Recent advancements in large language model (LLM) pruning have demonstrated state-of-the-art compression results without the need for post-training or retraining, while still maintaining high predictive performance. However, prior research predominantly focused on English text for calibration, overlooking the multilingual capabilities of modern LLMs. This paper presents a comprehensive empirical study analyzing the effects of different calibration languages on pruning multilingual models, revealing significant insights into performance and internal representation changes.

Read full article

via arXiv — cs.LG

arXiv — cs.LG3 days ago

Intelligence per Watt: Measuring Intelligence Efficiency of Local AI

PositiveArtificial Intelligence

Read full article

via arXiv — cs.LG

arXiv — cs.LG3 days ago

ExPairT-LLM: Exact Learning for LLM Code Selection by Pairwise Queries

PositiveArtificial Intelligence

ExPairT-LLM is introduced as an exact learning algorithm aimed at improving code selection from multiple outputs generated by large language models (LLMs). Traditional code selection algorithms often struggle to identify the correct program due to misidentification of nonequivalent programs or reliance on LLMs that may not always provide accurate outputs. ExPairT-LLM addresses these issues by utilizing pairwise membership and pairwise equivalence queries, enhancing the accuracy of program selection. Evaluations show a significant improvement in success rates over existing algorithms.

Read full article

via arXiv — cs.LG

arXiv — cs.LG3 days ago

Go-UT-Bench: A Fine-Tuning Dataset for LLM-Based Unit Test Generation in Go

PositiveArtificial Intelligence

The Go-UT-Bench dataset, introduced in a recent study, addresses the training data imbalance faced by code LLMs, particularly in Golang. This dataset comprises 5,264 pairs of code and unit tests sourced from 10 permissively licensed Golang repositories. The study demonstrates that fine-tuning LLMs with this dataset significantly enhances their performance, with models outperforming their base versions on over 75% of benchmark tasks.

Read full article

via arXiv — cs.LG

arXiv — cs.LG3 days ago

Experience-Guided Adaptation of Inference-Time Reasoning Strategies

PositiveArtificial Intelligence

The article discusses the Experience-Guided Reasoner (EGuR), a novel AI system designed to adapt its problem-solving strategies based on experiences accumulated during inference time. Unlike existing systems that only modify textual inputs, EGuR generates tailored strategies dynamically, allowing for a more flexible approach to AI reasoning. This advancement addresses the challenge of enabling agentic AI systems to adapt their methodologies post-training.

Read full article

via arXiv — cs.LG

arXiv — cs.CL3 days ago

Benchmarking Retrieval-Augmented Large Language Models in Biomedical NLP: Application, Robustness, and Self-Awareness

NeutralArtificial Intelligence

The paper titled 'Benchmarking Retrieval-Augmented Large Language Models in Biomedical NLP: Application, Robustness, and Self-Awareness' discusses the capabilities of large language models (LLMs) in biomedical natural language processing (NLP) tasks. It highlights the sensitivity of LLMs to demonstration selection and addresses the hallucination issue through retrieval-augmented LLMs (RAL). However, there is a lack of rigorous evaluation of RAL's impact on various biomedical NLP tasks, which complicates understanding its capabilities in this domain.

Read full article

via arXiv — cs.CL