The Initialization Determines Whether In-Context Learning Is Gradient Descent

arXiv — cs.LG•Friday, December 5, 2025 at 5:00:00 AM

PositiveArtificial Intelligence

Recent research has explored the initialization's role in determining whether in-context learning (ICL) in large language models (LLMs) behaves like gradient descent (GD). The study challenges previous assumptions by demonstrating that multi-head linear self-attention (LSA) can approximate GD under more realistic conditions, particularly with non-zero Gaussian prior means in linear regression formulations of ICL.
This development is significant as it enhances the understanding of ICL mechanisms in LLMs, potentially leading to improved model performance and more effective applications in various domains, including natural language processing and machine learning.
The findings contribute to ongoing discussions about the optimization capabilities of LLMs, particularly in relation to their efficiency and adaptability in processing complex data. This aligns with emerging frameworks aimed at enhancing LLMs' performance, such as adaptive context compression and training-free policy violation detection, highlighting the evolving landscape of AI research.

— via World Pulse Now AI Editorial System

Read Original

Was this article worth reading? Share it

LucidQuery AI

Combines diffusion reasoning with autoregressive LLM for advanced AI analysis.

AI & DataTry the app

LCW

An invisible AI copilot that helps you ace every coding interview.

AI & DataTry the app

CodeSpaced

AI tutors that reinforce learning with personalized spaced repetition.

Lifestyle & HealthTry the app

Continue Readings

arXiv — cs.LGa day ago

Convergence of Stochastic Gradient Langevin Dynamics in the Lazy Training Regime

NeutralArtificial Intelligence

A recent study published on arXiv presents a non-asymptotic convergence analysis of stochastic gradient Langevin dynamics (SGLD) in the lazy training regime, demonstrating that SGLD achieves exponential convergence to the empirical risk minimizer under certain conditions. The findings are supported by numerical examples in regression settings.

Read full article

via arXiv — cs.LG

arXiv — cs.CVa day ago

LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling

PositiveArtificial Intelligence

LongVT has been introduced as an innovative framework designed to enhance video reasoning capabilities in large multimodal models (LMMs) by facilitating a process known as 'Thinking with Long Videos.' This approach utilizes a global-to-local reasoning loop, allowing models to focus on specific video clips and retrieve relevant visual evidence, thereby addressing challenges associated with long-form video processing.

Read full article

via arXiv — cs.CV

arXiv — cs.CLa day ago

LangSAT: A Novel Framework Combining NLP and Reinforcement Learning for SAT Solving

PositiveArtificial Intelligence

A novel framework named LangSAT has been introduced, which integrates reinforcement learning (RL) with natural language processing (NLP) to enhance Boolean satisfiability (SAT) solving. This system allows users to input standard English descriptions, which are then converted into Conjunctive Normal Form (CNF) expressions for solving, thus improving accessibility and efficiency in SAT-solving processes.

Read full article

via arXiv — cs.CL

$Geschlechts\"ubergreifende Maskulina im Sprachgebrauch Eine korpusbasierte Untersuchung zu lexemspezifischen Unterschieden$

arXiv — cs.CLa day ago

Geschlechts\"ubergreifende Maskulina im Sprachgebrauch Eine korpusbasierte Untersuchung zu lexemspezifischen Unterschieden

NeutralArtificial Intelligence

A recent study published on arXiv investigates the use of generic masculines (GM) in contemporary German press texts, analyzing their distribution and linguistic characteristics. The research focuses on lexeme-specific differences among personal nouns, revealing significant variations, particularly between passive role nouns and prestige-related personal nouns, based on a corpus of 6,195 annotated tokens.

Read full article

via arXiv — cs.CL

arXiv — cs.CLa day ago

Limit cycles for speech

PositiveArtificial Intelligence

Recent research has uncovered a limit cycle organization in the articulatory movements that generate human speech, challenging the conventional view of speech as discrete actions. This study reveals that rhythmicity, often associated with acoustic energy and neuronal excitations, is also present in the motor activities involved in speech production.

Read full article

via arXiv — cs.CL

arXiv — cs.CLa day ago

Control Illusion: The Failure of Instruction Hierarchies in Large Language Models

NegativeArtificial Intelligence

Recent research highlights the limitations of hierarchical instruction schemes in large language models (LLMs), revealing that these models struggle with consistent instruction prioritization, even in simple cases. The study introduces a systematic evaluation framework to assess how effectively LLMs enforce these hierarchies, finding that the common separation of system and user prompts fails to create a reliable structure.

Read full article

via arXiv — cs.CL

arXiv — cs.LGa day ago

CARL: Critical Action Focused Reinforcement Learning for Multi-Step Agent

PositiveArtificial Intelligence

CARL, a new reinforcement learning algorithm, has been introduced to enhance the performance of multi-step agents by focusing on critical actions rather than treating all actions equally. This approach addresses the limitations of conventional policy optimization methods, which often overlook the varying importance of different actions in achieving desired outcomes.

Read full article

via arXiv — cs.LG

arXiv — cs.CLa day ago

Which Type of Students can LLMs Act? Investigating Authentic Simulation with Graph-based Human-AI Collaborative System

PositiveArtificial Intelligence

Recent advancements in large language models (LLMs) have prompted research into their ability to authentically simulate student behavior, addressing challenges in educational data collection and intervention design. A new three-stage collaborative pipeline has been developed to generate and filter high-quality student agents, utilizing automated scoring and human expert validation to enhance realism in simulations.

Read full article

via arXiv — cs.CL