Artificial IntelligencearXiv — cs.LGMon, May 25, 2026, 4:00 AMPositive

Long-Context Reasoning Through Proxy-Based Chain-of-Thought Tuning

Recent advancements in long-context reasoning have been made with the introduction of ProxyCoT, a training framework designed to enhance the reasoning capabilities of large language models (LLMs) by utilizing short proxy contexts to improve performance on long-context tasks. This approach leverages high-quality reasoning traces obtained through reinforcement learning or distillation from larger models, followed by supervised fine-tuning on full contexts.

WPN Brief

  • What Happened

    Recent advancements in long-context reasoning have been made with the introduction of ProxyCoT, a training framework designed to enhance the reasoning capabilities of large language models (LLMs) by utilizing short proxy contexts to improve performance on long-context tasks. This approach leverages high-quality reasoning traces obtained through reinforcement learning or distillation from larger models, followed by supervised fine-tuning on full contexts.

  • Why It Matters

    The development of ProxyCoT is significant as it addresses the performance disparity observed in LLMs when handling long-context tasks, which are essential for complex reasoning. By improving the models' ability to utilize proxy contexts effectively, this framework could lead to more accurate and efficient applications in various fields, including natural language processing and artificial intelligence.

  • The Bigger Picture

    This innovation reflects ongoing efforts to enhance LLMs' reasoning abilities, particularly in light of challenges such as positional failures and the divergence in reasoning processes among different models. The introduction of frameworks like ProxyCoT, along with other recent advancements in contextual reasoning and multimodal orchestration, underscores a broader trend in AI research aimed at refining model performance and expanding their applicability across diverse tasks.

Ask WPN AI

Related Reports

More coverage on this story

10 reports across the wire

arXiv — cs.LG
May 25

Decoding the Critique Mechanism in Large Reasoning Models

Recent research has unveiled the critique mechanism in Large Reasoning Models (LRMs), highlighting their ability to backtrack and self-verify during complex reasoning tasks. This study demonstrates that LRMs can recover from errors without explicit corrections, suggesting an internal mechanism for self-correction termed 'hidden critique ability.'

Artificial Intelligenceneutral
arXiv — cs.CL
May 25

Convergence Without Understanding: When Language Models Agree on Representations but Disagree on Reasoning

Recent research has demonstrated that large language models (LLMs) exhibit a convergence in internal representations, as outlined by the Platonic Representation Hypothesis, yet they diverge in reasoning processes across various problem types. This study evaluated 16 models on 800 reasoning tasks, revealing notable discrepancies in performance based on problem difficulty and model architecture.

Artificial Intelligenceneutral
arXiv — cs.LG
May 25

Scaling-Aware Adapter for Structure-Grounded LLM Reasoning

A new framework called Cuttlefish has been introduced to enhance large language models (LLMs) by grounding language reasoning in geometric cues and adapting to structural complexity. This unified multimodal LLM aims to address the limitations of existing methods that often compress structural inputs, leading to hallucinations and inefficiencies in reasoning.

Artificial Intelligencepositive
arXiv — cs.LG
May 25

Positional Failures in Long-Context LLMs: A Blind Spot in Reasoning Benchmarks

A recent audit of long-context benchmarks reveals a significant oversight in mainstream reasoning evaluations, as none effectively control the positional placement of target tasks, filler content, and context length. This gap was highlighted in the study titled 'Positional Failures in Long-Context LLMs,' which evaluated 11 benchmarks and proposed a new framework called Context Rot Evaluation (CRE) to address these issues.

Artificial Intelligenceneutral
arXiv — cs.LG
May 25

Self-Improving In-Context Learning

A new approach to in-context learning (ICL) has been proposed, which optimizes the continuous embeddings of a fixed few-shot prompt at test time. This method utilizes log-probabilities from a single forward pass to create a self-supervised confidence proxy, enhancing model performance without the need for finetuning or external data.

Artificial Intelligencepositive
arXiv — cs.CL
May 25

Training-Free Multimodal Large Language Model Orchestration

A new framework called Training-Free Large Language Model Orchestration (LLM Orchestration) has been introduced, allowing for the integration of various modality experts into a unified multimodal system without the need for additional training. This framework includes an LLM controller for user intent inference, a cross-modal memory for efficient evidence retrieval, and an interaction layer for executing routing and memory management.

Artificial Intelligencepositive
arXiv — cs.CV
May 25

CoReVAD: A Contextual Reasoning Framework for Training-Free Video Anomaly Detection

A new framework named CoReVAD has been introduced for training-free video anomaly detection, leveraging a single frozen Vision-Language Model (VLM) to generate anomaly scores and temporal descriptions without the need for extensive training. This approach aims to reduce domain dependency and training costs associated with traditional Video Anomaly Detection methods.

Artificial Intelligencepositive
arXiv — cs.CL
May 25

OnePred: Next-Query Prediction via Recursive Intent Memory in Multi-Turn Conversations

A new study introduces OnePred, a system designed for next-query prediction in multi-turn conversations, addressing the limitations of current large language models (LLMs) that react only after user input. By maintaining a recursively updated memory of user intent, OnePred aims to enhance proactive interaction without the inefficiencies of traditional dialogue history management.

Artificial Intelligencepositive
arXiv — cs.CL
May 25

How Far Will They Go? Red-Teaming Online Influence with Large Language Models

A recent study published on arXiv explores the implications of large language models (LLMs) in online discourse, particularly focusing on their potential to support political influence campaigns. The research introduces a red-teaming framework to evaluate LLMs' Overton Windows, which define the range of political opinions they can express, revealing systematic biases in content generation.

Artificial Intelligenceneutral
arXiv — cs.CL
May 27

How Human-Like Are Large Language Models? A Register-Aware Linguistic Evaluation Framework

A new framework for evaluating the human-likeness of texts generated by Large Language Models (LLMs) has been proposed, focusing on the linguistic features that reflect context-dependent language production. This framework utilizes a two-sample problem to compare LLM outputs with human reference corpora, aiming to assess the adherence to linguistic patterns that characterize human communication.

Artificial Intelligenceneutral

Articles

Continue Reading

arXiv — cs.LGArtificial Intelligence2 days ago

Gibbs randomness-compression proposition

A new proposition has been introduced that connects randomness and compression through Gibbs entropy, focusing on measurement vectors linked to compression processes. This approach utilizes the performance of learning tasks as a metric for assessing compression across multiple cycles, suggesting that lossy compression can be viewed as directed randomness that retains information within specific Gibbs entropy limits.

arXiv — cs.LGArtificial Intelligence2 days ago

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning

A new framework called Similarity as Reward Alignment (SARA) has been introduced in preference-based reinforcement learning (PbRL), addressing the challenges of labeler errors and adapting to various feedback formats. SARA computes rewards based on the similarity of learned latent representations of preferred samples, demonstrating improved stability and performance in offline reinforcement learning benchmarks.

arXiv — cs.LGArtificial Intelligence2 days ago

Contrastive Conformal Sets

A recent study introduces Contrastive Conformal Sets, enhancing contrastive learning by constructing geometric sets in the semantic feature space, ensuring user-specified coverage of positive samples while maximizing the exclusion of negative samples. This method extends conformal prediction principles to improve the reliability of machine learning models.

arXiv — cs.LGArtificial Intelligence2 days ago

Data Driven Block Replacement Scheduling

A new study has introduced data-driven algorithms for managing independent identical machines under a block replacement policy, focusing on determining the optimal replacement interval based on operational data. The research formulates this challenge as a stochastic multi-armed bandit problem, proposing algorithms that achieve regret matching the Lai–Robbins lower bound.

arXiv — cs.LGArtificial Intelligence2 days ago

Distributionally Robust Optimization via Iterative Algorithms in Continuous Probability Spaces

A recent study has introduced a framework for distributionally robust optimization (DRO) in continuous probability spaces, addressing the computational challenges associated with infinite-dimensional optimization problems. The research leverages Brenier's theorem to define the least favorable distribution as a pushforward of a transport map, leading to a minimax problem in Wasserstein space and proposing an iterative algorithmic framework with global convergence guarantees.

arXiv — cs.LGArtificial Intelligence2 days ago

To Grok Grokking: Provable Grokking in Ridge Regression

A recent study published on arXiv explores the phenomenon of grokking within the context of ridge regression, demonstrating that models can overfit training data initially, yet later achieve significant generalization. The research provides rigorous quantitative bounds on the delay of generalization, termed 'grokking time', and emphasizes the role of hyperparameter tuning in influencing this process.

arXiv — cs.LGArtificial Intelligence2 days ago

Generalized Neural Distributional Regression

The Generalized Neural Distributional Regression (GNDR) framework has been introduced, integrating deep neural networks with classical probability distributions to enhance statistical modeling. This framework employs a semi-parametric estimation procedure to address the non-identifiability of deep architectures, allowing for the extraction of analytical Fisher Information matrices and facilitating rigorous uncertainty quantification.

arXiv — cs.LGArtificial Intelligence2 days ago

Selecting Hyperparameters for Tree-Boosting

A recent study published on arXiv explores various methods for hyperparameter optimization in tree-boosting, a prevalent machine learning technique for tabular data. The research empirically compares methods such as random grid search, SMAC, and Gaussian-process-based Bayesian optimization across 59 datasets, revealing that SMAC consistently outperforms others under a fixed tuning budget.