Artificial IntelligencearXiv — stat.MLWed, May 27, 2026, 4:00 AMNeutral

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models

A recent study on scale vectors in large language models (LLMs) reveals that while these vectors represent a small fraction of model parameters, their removal significantly hampers the pre-training process. The research highlights the role of scale vectors in enhancing optimization through a self-amplifying preconditioning effect, particularly in Pre-Norm architectures.

WPN Brief

  • What Happened

    A recent study on scale vectors in large language models (LLMs) reveals that while these vectors represent a small fraction of model parameters, their removal significantly hampers the pre-training process. The research highlights the role of scale vectors in enhancing optimization through a self-amplifying preconditioning effect, particularly in Pre-Norm architectures.

  • Why It Matters

    Understanding the function of scale vectors is crucial for improving LLM performance, as their optimization capabilities directly influence the effectiveness of these models in various applications, including natural language processing and machine learning tasks.

  • The Bigger Picture

    This development underscores ongoing discussions about the architecture and training of LLMs, particularly regarding the balance between model complexity and performance. As researchers explore the implications of scale vectors, broader concerns about cultural biases, data exposure, and the ethical use of LLMs continue to shape the landscape of artificial intelligence.

Ask WPN AI

Related Reports

More coverage on this story

10 reports across the wire

arXiv — cs.LG
May 27

Emergent Causal-Geometric Dynamics Across Depth in Large Language Models

Recent research has revealed that large language models (LLMs) exhibit structured variations across their depth, highlighting a transition from context-processing to prediction-forming computations. This study integrates geometric analysis with mechanistic interventions to provide a comprehensive understanding of how LLMs evolve their representational structures to produce predictions.

Artificial Intelligenceneutral
arXiv — cs.CL
May 27

How Human-Like Are Large Language Models? A Register-Aware Linguistic Evaluation Framework

A new framework for evaluating the human-likeness of texts generated by Large Language Models (LLMs) has been proposed, focusing on the linguistic features that reflect context-dependent language production. This framework utilizes a two-sample problem to compare LLM outputs with human reference corpora, aiming to assess the adherence to linguistic patterns that characterize human communication.

Artificial Intelligenceneutral
arXiv — cs.CL
May 27

Why LLMs Hallucinate on Structured Knowledge: A Mechanistic Analysis of Reasoning over Linearized Representations

A recent study investigates the mechanisms behind hallucinations in large language models (LLMs), revealing that these errors stem from systematic internal dynamics rather than random noise. The research highlights that attention in LLMs often focuses on shortcut-like cues instead of the full context, leading to failures in semantic grounding.

Artificial Intelligenceneutral
arXiv — cs.CL
Jun 10

Culturally uneven urban perception in large language models

A recent study highlights the culturally uneven urban perception exhibited by large language models (LLMs), revealing that their evaluations of cities are biased towards European and North American cultural framings. This research introduces a measurement framework to assess the cultural neutrality of LLM-generated urban descriptions using a diverse street-view image dataset.

Artificial Intelligencenegative
arXiv — cs.CL
May 27

Cultural Value Alignment Via Latent Activation Steering in Large Language Models

A new framework for cultural evaluation and intervention in Large Language Models (LLMs) has been proposed, focusing on latent activation steering to better align models with human values as mapped by the World Values Survey. This approach transitions from abstract queries to scenario-based behavioral probing, revealing the cultural depth of LLMs.

Artificial Intelligenceneutral
arXiv — cs.LG
May 27

Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications

A recent survey has highlighted the growing concerns surrounding Pretraining Data Exposure (PDE) in Large Language Models (LLMs), emphasizing the need to understand whether specific data was included in the pretraining datasets. This issue intersects with data contamination and membership inference, which have traditionally been studied separately. The paper aims to unify these areas and address the implications for privacy and evaluation integrity in NLP.

Artificial Intelligenceneutral
arXiv — cs.LG
May 27

Faithfulness Evaluation for Decoder-only LLM Attributions with Controlled Retained Information

A new evaluation framework for Large Language Models (LLMs) has been proposed, focusing on attribution methods that control the number of retained words during perturbation. This framework, named $ ext{π}$-Soft-NC and $ ext{π}$-Soft-NS, aims to provide a more accurate comparison of attribution quality by addressing the inflation of scores due to word retention. Additionally, a gradient-based method called Grad-ELLM has been introduced for decoder-only LLMs, enhancing the attribution process during decoding tasks.

Artificial Intelligenceneutral
arXiv — cs.LG
May 27

Benchmark Leakage Trap: Can We Trust LLM-based Recommendation?

A recent study has highlighted the issue of benchmark data leakage in Large Language Model (LLM)-based recommendation systems, revealing that exposure to benchmark datasets during pre-training can lead to misleadingly inflated performance metrics. This phenomenon was validated through experiments simulating various data leakage scenarios, demonstrating that domain-relevant leaked data can create substantial but spurious performance gains.

Artificial Intelligenceneutral
arXiv — cs.LG
May 27

Unified Neural Scaling Laws

A new study introduces a Unified Neural Scaling Law (UNSL) that models the scaling behaviors of deep neural networks across multiple dimensions, including model parameters, dataset size, and training steps. This functional form demonstrates improved accuracy in extrapolating scaling behavior for various architectures and tasks, such as vision, language, and reinforcement learning.

Artificial Intelligenceneutral
arXiv — cs.CL
May 27

Learning to Diagnose and Correct Errors: Towards Moral Sensitivity Acquisition in Large Language Models

A recent study has proposed a pragmatic inference approach aimed at enhancing moral sensitivity acquisition in large language models (LLMs). This approach focuses on enabling LLMs to diagnose and correct moral errors, addressing a critical gap in aligning these models with human moral values.

Artificial Intelligenceneutral

Articles

Continue Reading

arXiv — cs.CVArtificial Intelligenceyesterday

Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding

Recent advancements in artificial intelligence have led to the introduction of VEGA-3D, a framework that repurposes pre-trained video diffusion models to enhance scene understanding by leveraging implicit 3D priors. This development addresses the limitations of existing multimodal large language models (MLLMs) that struggle with spatial reasoning and geometric dynamics.

arXiv — cs.CLArtificial Intelligenceyesterday

Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability

Recent research has explored the emergence of moral biases, specifically the Knobe effect, in finetuned large language models (LLMs). The study revealed that these biases are not only learned during the finetuning process but can also be localized to specific layers within the model, allowing for targeted interventions to mitigate their effects without the need for retraining.

arXiv — cs.CLArtificial Intelligenceyesterday

Probing LLMs for Syntactic Structure Beyond Universal Dependencies: A Minimalist Phase Account in English

Recent research demonstrates that large language models (LLMs) encode syntactic distinctions that extend beyond the Universal Dependencies framework, particularly in English wh-movement stimuli. The study reveals that the distance between an embedded subject and its verb varies depending on the clause type, showcasing a sign asymmetry that cannot be explained by existing models based on UD distance or structural complexity.

arXiv — cs.CLArtificial Intelligenceyesterday

Code-MUE: Measuring Code LLMs' Uncertainty through Execution-based Semantic Interaction Graphs

A new framework named Code-MUE has been introduced to measure the uncertainty of Code Large Language Models (LLMs) through execution-based Semantic Interaction Graphs. This approach addresses the limitations of existing uncertainty estimation methods, which struggle with closed-source models and the unique fragility of code.

arXiv — cs.CVArtificial Intelligenceyesterday

ABot-N1: Toward a General Visual Language Navigation Foundation Model

The recent introduction of ABot-N1 marks a significant advancement in Visual Language Navigation foundation models, aiming to enhance deep reasoning for spatial decisions while addressing issues such as coordinate drift and lack of interpretability in existing models. This model employs a slow-fast architecture that separates cognition from control, utilizing dual visual-language signals for improved performance.

arXiv — cs.LGArtificial Intelligenceyesterday

Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems

The recent publication on constraint-driven model optimization presents a unified framework for selecting compression and acceleration techniques in machine learning systems, emphasizing the need for a principled approach amidst the diverse optimization methods available.

arXiv — cs.CVArtificial Intelligenceyesterday

GeCo: Evaluating Geometric Consistency for Video Generation via Motion and Structure

The introduction of GeCo, a geometry-grounded metric, aims to enhance video generation by detecting geometric deformation and occlusion-inconsistency artifacts in static scenes. By integrating residual motion and depth priors, GeCo generates dense consistency maps that highlight these artifacts, facilitating a systematic benchmarking of recent video generation models.

arXiv — cs.LGArtificial Intelligenceyesterday

Robust Explanations for User Trust in Enterprise NLP Systems

A recent study highlights the necessity for robust explanations to foster user trust in enterprise NLP systems, particularly in scenarios where black-box deployment limits pre-deployment validation. The research proposes a unified evaluation framework for token-level explanations, assessing their stability under various real-world perturbations across multiple architectures and datasets.