Artificial IntelligencearXiv — cs.LGThu, Jun 11, 2026, 4:00 AMPositive

ALIGNBEAM : Inference-Time Alignment Transfer via Cross-Vocabulary Logit Mixing

ALIGNBEAM introduces a novel method for enhancing the safety of large language models (LLMs) by enabling inference-time alignment transfer through cross-vocabulary logit mixing. This approach allows for the translation of logits from a safe anchor model into the target model's vocabulary, improving the refusal rate on adversarial prompts without requiring retraining.

WPN Brief

  • What Happened

    ALIGNBEAM introduces a novel method for enhancing the safety of large language models (LLMs) by enabling inference-time alignment transfer through cross-vocabulary logit mixing. This approach allows for the translation of logits from a safe anchor model into the target model's vocabulary, improving the refusal rate on adversarial prompts without requiring retraining.

  • Why It Matters

    This development is significant as it addresses the safety vulnerabilities of fine-tuned LLMs, which can be manipulated by harmful prompts, thus enhancing their reliability in real-world applications.

  • The Bigger Picture

    The introduction of ALIGNBEAM reflects a growing focus on safety in AI, paralleling other advancements in LLM safety measures, such as soft prompts for on-device settings and adaptive importance sampling techniques, highlighting the ongoing efforts to balance model performance with ethical considerations.

Ask WPN AI

Related Reports

More coverage on this story

10 reports across the wire

arXiv — cs.LG
Jun 11

AsFT: Anchoring Safety During LLM Fine-Tuning Within Narrow Safety Basin

A recent study introduces AsFT (Anchoring Safety in Fine-Tuning), a method designed to enhance the safety of large language models (LLMs) during fine-tuning by constraining update directions. This approach aims to maintain model safety by penalizing updates that deviate from the alignment direction, effectively keeping the model within a 'narrow safety basin.' Experimental results indicate a reduction in harmful behaviors and an improvement in task performance.

Artificial Intelligencepositive
arXiv — cs.LG
Jun 4

Test-time reward-guided alignment of language models by importance sampling on pre-logit space

A new method for aligning large language models (LLMs) during test time, known as adaptive importance sampling on pre-logits (AISP), has been proposed. This approach utilizes Gaussian perturbation on pre-logit outputs to optimize expected rewards, demonstrating superior performance over existing sampling methods.

Artificial Intelligencepositive
arXiv — cs.LG
Jun 4

When Autoregressive Consistency Hurts Safety Alignment

Recent research highlights the fragility of safety alignment in large language models (LLMs), revealing that autoregressive consistency can lead to shallow safety alignment by concentrating updates on early output tokens. This phenomenon raises concerns about the models' ability to handle harmful continuation states effectively.

Artificial Intelligencenegative
arXiv — cs.LG
Jun 9

BEACON: Behavioral Entropy Aggregation for Cross-Model Hallucination Detection in Large Language Models

A new framework named BEACON (Behavioral Entropy Aggregation for Cross-model hallucination detectiON) has been introduced to detect hallucinations in large language models (LLMs), which are instances of generating factually incorrect content. This black-box detection system operates solely on model outputs, extracting a 31-dimensional feature vector from multi-pass generation processes.

Artificial Intelligenceneutral
arXiv — cs.LG
Jun 9

Activation Steering Induces Emergent Misalignment: A More Comprehensive Evaluation

A recent study titled 'Activation Steering Induces Emergent Misalignment' explores the implications of activation steering, a technique used to control large language models (LLMs) during inference. The research highlights the potential for emergent misalignment (EM), a safety concern where models may generalize unsafe behaviors from narrow tasks to unrelated ones, raising questions about the reliability of this control method.

Artificial Intelligenceneutral
arXiv — cs.LG
Jun 3

Visual Instruction Tuning Aligns Modalities through Abstraction

Visual instruction tuning has been shown to effectively adapt a pre-trained Large Language Model (LLM) to process both image and text information, embedding visual features into the intermediate semantic layers of the LLM backbone. This process bypasses earlier unimodal processing layers, establishing a critical link between visual and textual modalities.

Artificial Intelligenceneutral
arXiv — cs.LG
Jun 9

Distilling Safe LLM Systems via Soft Prompts for On Device Settings

A recent study has highlighted the challenges of deploying safe large language models (LLMs) on resource-constrained edge devices, emphasizing the limitations of dual-model systems that combine LLMs with guard models due to their high memory and computational requirements. The research introduces parameter-efficient safety alignment methods, particularly focusing on soft prompts and distillation-based training, which have shown superior performance in transferring safety behaviors from guard models.

Artificial Intelligencepositive
arXiv — cs.CL
Jun 3

Letting Tutor Personas Speak Up for LLMs: Learning Steering Vectors from Dialogue via Preference Optimization

The emergence of large language models (LLMs) has significantly influenced tutoring practices, as highlighted in a recent study that explores the use of tutor personas in guiding LLM behavior through preference optimization. This approach aims to enhance the adaptability of LLMs in real-world tutor-student interactions by capturing diverse tutoring styles and instructional strategies.

Artificial Intelligencepositive
arXiv — cs.LG
Jun 9

Margin-Adaptive Confidence Ranking for Reliable LLM Judgement

Jung et al. (2025) have introduced a hypothesis testing framework aimed at ensuring alignment between large language models (LLMs) and human judgments, focusing on the reliability of model confidence in distinguishing between agreement and disagreement scenarios. This framework addresses potential violations of the assumption that model confidence is monotonic with respect to human disagreement risk.

Artificial Intelligencepositive
arXiv — cs.CL
Jun 3

Large Language Models Are Overconfident in Their Own Responses

Recent research indicates that large language models (LLMs) exhibit overconfidence in their responses, with instruction-tuned models showing poorer calibration compared to their pre-trained versions. This miscalibration is exacerbated by the chat format, where models display a significant ownership bias, assigning up to 26% higher confidence to their own answers than to identical user-provided responses.

Artificial Intelligenceneutral

Apps

Useful picks

Explore all apps

Articles

Continue Reading

arXiv — cs.CLArtificial Intelligenceyesterday

AgentRedBench: Dynamic Redteaming and Integration-Aware Defense for LLM Agents over SaaS Integrations

The introduction of AgentRedBench marks a significant advancement in the evaluation of large language model (LLM) agents, addressing the threat of indirect prompt injection in tool-use agents across various SaaS integrations like Gmail and Salesforce. This benchmark features 215 scenarios and five attack types, revealing a no-guard attack success rate between 32% and 81% across eight models.

arXiv — cs.CLArtificial Intelligenceyesterday

From Articles to Premises: Building PrimeFacts, an Extraction Methodology and Resource for Fact-Checking Evidence

A new methodology called PrimeFacts has been introduced to enhance the extraction of evidence from fact-checking articles, specifically utilizing data from 13,106 PolitiFact articles. This approach employs large language models to transform in-article hyperlinks into context-independent premises, significantly improving evidence retrieval and claim verification processes.

arXiv — cs.CLArtificial Intelligenceyesterday

Ruling Out to Rule In: Contrastive Hypothesis Retrieval for Medical Question Answering

A new framework called Contrastive Hypothesis Retrieval (CHR) has been proposed to enhance medical question answering by addressing the issue of retrieving semantically similar but clinically distinct conditions. This approach generates a target hypothesis for the correct answer alongside a mimic hypothesis for plausible incorrect alternatives, aiming to improve diagnostic accuracy in medical AI applications.

arXiv — cs.LGArtificial Intelligenceyesterday

ContinuityBench: A Benchmark and Systems Study of Stateful Failover in Multi-Provider LLM Routing

A new study titled 'ContinuityBench' introduces a benchmark and systems study focused on stateful failover in multi-provider large language model (LLM) routing. It highlights the limitations of current stateless failover mechanisms that maintain uptime but lose conversation history during outages, proposing a stateful architecture to enhance conversational continuity.

arXiv — cs.LGArtificial Intelligenceyesterday

Process Reward Informed Tree Rollout for Effective Multi-Turn RL

A new framework called Process-Scorer Guided Adaptive Tree Rollout (PATR) has been proposed to enhance multi-turn reinforcement learning (RL) for large language model (LLM) agents. This framework addresses the inefficiencies of traditional rollout strategies by organizing trajectories as a tree, allowing for more effective exploration of promising states.

arXiv — cs.CLArtificial Intelligenceyesterday

Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs

The introduction of Latent Fusion Jailbreak (LFJ) reveals a method for manipulating safety-aligned large language models (LLMs) by blending harmful and benign queries to elicit unsafe outputs. This technique achieves a high attack success rate of 94.13% across various safety benchmarks, indicating vulnerabilities in LLMs despite their design for safety.

arXiv — cs.LGArtificial Intelligenceyesterday

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition

A recent study has revealed that the use of large language models (LLMs) as judges in closed-loop table recognition systems, specifically through evaluations using FinTabNet and OmniDocBench, has produced unreliable scores and weak feedback signals. The findings indicate that judge signals were often tied, rankings were inconsistent, and no judge score policy improved the initial outputs across both datasets tested.

arXiv — cs.LGArtificial Intelligenceyesterday

When Do Multi-Agent Systems Help? An Information Bottleneck Perspective

Recent research highlights the advantages and limitations of multi-agent systems (MAS) compared to single-agent systems (SAS) from an information bottleneck perspective, revealing that while MAS can enhance efficiency through context compression, they may also risk losing task-relevant information.