Artificial IntelligencearXiv — cs.LGFri, May 29, 2026, 4:00 AMNeutral

Data filtering methods for training language models

A comparative analysis of two automatic label error detection methods, Confident Learning and Dataset Cartography, was conducted on three Russian text classification corpora, revealing the impact of data quality on machine learning model effectiveness.

WPN Brief

  • What Happened

    A comparative analysis of two automatic label error detection methods, Confident Learning and Dataset Cartography, was conducted on three Russian text classification corpora, revealing the impact of data quality on machine learning model effectiveness.

  • Why It Matters

    The findings underscore the importance of accurate labeling in training datasets, as label errors can introduce noise that diminishes model generalization, which is crucial for the deployment of reliable AI systems.

  • The Bigger Picture

    This research aligns with ongoing discussions in the AI community regarding data organization and the methodologies employed to enhance the training efficiency of language models, highlighting the need for robust frameworks to ensure high-quality data inputs.

Ask WPN AI

Related Reports

More coverage on this story

10 reports across the wire

arXiv — cs.CL
May 29

Calibration Is Not Enough: Evaluating Confidence Estimation Under Language Variations

A recent study published on arXiv introduces a novel evaluation framework for confidence estimation (CE) in large language models, emphasizing the need for robustness, stability, and sensitivity under language variations. This framework aims to ensure that confidence estimates remain consistent across semantically equivalent prompts while varying with changes in answer meaning.

Artificial Intelligenceneutral
arXiv — cs.CL
May 29

Demystifying Data Organization for Enhanced LLM Training

Recent research has highlighted the importance of data organization in enhancing the training efficiency of Large Language Models (LLMs). The study identifies four key guidelines for optimizing data organization and introduces two novel data ordering methods, STR and SAW, which leverage pre-computed sample-level scores to minimize computational overhead.

Artificial Intelligencepositive
arXiv — cs.LG
May 29

Rare Event Analysis of Large Language Models

A recent study published on arXiv presents a comprehensive framework for analyzing rare events in large language models (LLMs), emphasizing the significance of these atypical behaviors during inference. The framework includes theoretical foundations, generation strategies, probability estimation, and error analysis, illustrated with practical examples.

Artificial Intelligenceneutral
arXiv — cs.CV
May 29

Evaluating Dataset Watermarking for Fine-tuning Traceability of Customized Diffusion Models: A Comprehensive Benchmark and Removal Approach

A recent study evaluated dataset watermarking techniques aimed at enhancing the traceability of customized diffusion models, addressing the copyright and security risks associated with fine-tuning these models. The research established a comprehensive evaluation framework that assesses watermarking methods based on Universality, Transmissibility, and Robustness, revealing vulnerabilities in existing approaches.

Artificial Intelligenceneutral
arXiv — cs.LG
May 29

Statistical Consistency and Generalization of Contrastive Representation Learning

A recent study on contrastive representation learning (CRL) has introduced a unified statistical learning theory, addressing key limitations in the understanding of statistical consistency and generalization bounds, particularly as the number of negative samples increases. The research evaluates retrieval quality using an AUC-type population criterion, demonstrating that contrastive loss aligns with optimal ranking.

Artificial Intelligenceneutral
arXiv — cs.LG
May 29

Open World Autoencoding Drift Detection with Novel Class Recognition in Tabular Non-stationary Data Streams

A new study has introduced an unsupervised concept drift detection method that leverages autoencoders to identify shifts in known class distributions and recognize novel class samples in tabular non-stationary data streams. This approach utilizes reconstruction errors and density estimation to adapt to evolving data distributions effectively.

Artificial Intelligencepositive
arXiv — cs.LG
May 29

Feedback-to-Rubrics: Can We Learn Expert Criteria from Inline Comments?

A recent study proposes a method for learning reusable natural-language rubrics from inline comments on various drafts, including those generated by large language models (LLMs). This approach aims to refine the evaluation process by addressing the often tacit and undocumented criteria that influence expert feedback.

Artificial Intelligencepositive
arXiv — cs.CL
May 29

Comparative Evaluation of Machine Translation Systems on Images with Text

A recent study published on arXiv presents a comparative evaluation of machine translation systems applied to images containing text, focusing on three paradigms: modular pipelines, multi-modal large language models (MLLMs), and the end-to-end model Translatotron-V. The research utilized state-of-the-art OCR and multilingual LLMs, conducting experiments on multilingual datasets to assess performance using BLEU, chrF, and TER metrics.

Artificial Intelligencepositive
arXiv — cs.LG
May 29

Leak@$k$: Unlearning Does Not Make LLMs Forget Under Probabilistic Decoding

A recent study published on arXiv highlights the limitations of unlearning methods in large language models (LLMs), revealing that these models do not effectively forget sensitive information when using probabilistic decoding. The research introduces a new metric, leak@$k$, to measure the likelihood of previously learned knowledge resurfacing during text generation.

Artificial Intelligenceneutral
arXiv — cs.LG
May 29

How's it going? Reinforcement learning in language models recruits a functional welfare axis

Recent research indicates that reinforcement learning (RL) significantly influences the internal representations of language models, particularly in how they assess functional welfare. The study involved training language models in a neutral maze environment, revealing that reward and punishment vectors correspond to positive and negative welfare representations, respectively.

Artificial Intelligenceneutral

Articles

Continue Reading

arXiv — cs.LGArtificial Intelligence2 days ago

Gibbs randomness-compression proposition

A new proposition has been introduced that connects randomness and compression through Gibbs entropy, focusing on measurement vectors linked to compression processes. This approach utilizes the performance of learning tasks as a metric for assessing compression across multiple cycles, suggesting that lossy compression can be viewed as directed randomness that retains information within specific Gibbs entropy limits.

arXiv — cs.LGArtificial Intelligence2 days ago

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning

A new framework called Similarity as Reward Alignment (SARA) has been introduced in preference-based reinforcement learning (PbRL), addressing the challenges of labeler errors and adapting to various feedback formats. SARA computes rewards based on the similarity of learned latent representations of preferred samples, demonstrating improved stability and performance in offline reinforcement learning benchmarks.

arXiv — cs.LGArtificial Intelligence2 days ago

Contrastive Conformal Sets

A recent study introduces Contrastive Conformal Sets, enhancing contrastive learning by constructing geometric sets in the semantic feature space, ensuring user-specified coverage of positive samples while maximizing the exclusion of negative samples. This method extends conformal prediction principles to improve the reliability of machine learning models.

arXiv — cs.LGArtificial Intelligence2 days ago

Data Driven Block Replacement Scheduling

A new study has introduced data-driven algorithms for managing independent identical machines under a block replacement policy, focusing on determining the optimal replacement interval based on operational data. The research formulates this challenge as a stochastic multi-armed bandit problem, proposing algorithms that achieve regret matching the Lai–Robbins lower bound.

arXiv — cs.LGArtificial Intelligence2 days ago

Distributionally Robust Optimization via Iterative Algorithms in Continuous Probability Spaces

A recent study has introduced a framework for distributionally robust optimization (DRO) in continuous probability spaces, addressing the computational challenges associated with infinite-dimensional optimization problems. The research leverages Brenier's theorem to define the least favorable distribution as a pushforward of a transport map, leading to a minimax problem in Wasserstein space and proposing an iterative algorithmic framework with global convergence guarantees.

arXiv — cs.LGArtificial Intelligence2 days ago

To Grok Grokking: Provable Grokking in Ridge Regression

A recent study published on arXiv explores the phenomenon of grokking within the context of ridge regression, demonstrating that models can overfit training data initially, yet later achieve significant generalization. The research provides rigorous quantitative bounds on the delay of generalization, termed 'grokking time', and emphasizes the role of hyperparameter tuning in influencing this process.

arXiv — cs.LGArtificial Intelligence2 days ago

Generalized Neural Distributional Regression

The Generalized Neural Distributional Regression (GNDR) framework has been introduced, integrating deep neural networks with classical probability distributions to enhance statistical modeling. This framework employs a semi-parametric estimation procedure to address the non-identifiability of deep architectures, allowing for the extraction of analytical Fisher Information matrices and facilitating rigorous uncertainty quantification.

arXiv — cs.LGArtificial Intelligence2 days ago

Selecting Hyperparameters for Tree-Boosting

A recent study published on arXiv explores various methods for hyperparameter optimization in tree-boosting, a prevalent machine learning technique for tabular data. The research empirically compares methods such as random grid search, SMAC, and Gaussian-process-based Bayesian optimization across 59 datasets, revealing that SMAC consistently outperforms others under a fixed tuning budget.