Artificial IntelligencearXiv — cs.CLFri, Jun 12, 2026, 4:00 AMNeutral

FENCE: A Financial and Multimodal Jailbreak Detection Dataset

The introduction of FENCE, a bilingual multimodal dataset, addresses the significant risks posed by jailbreaking to Large Language Models (LLMs) and Vision Language Models (VLMs), particularly in financial applications. This dataset facilitates the training and evaluation of jailbreak detectors, emphasizing finance-relevant queries and image-grounded threats.

WPN Brief

  • What Happened

    The introduction of FENCE, a bilingual multimodal dataset, addresses the significant risks posed by jailbreaking to Large Language Models (LLMs) and Vision Language Models (VLMs), particularly in financial applications. This dataset facilitates the training and evaluation of jailbreak detectors, emphasizing finance-relevant queries and image-grounded threats.

  • Why It Matters

    The development of FENCE is crucial as it provides a robust resource for enhancing the security of LLMs and VLMs, which are increasingly utilized in sensitive sectors like finance, where vulnerabilities can lead to substantial risks.

  • The Bigger Picture

    This initiative reflects a growing recognition of the need for improved security measures in AI, especially as studies reveal inconsistencies in LLMs' safety judgments across various domains, including finance. The ongoing exploration of frameworks and methodologies to bolster LLM security highlights the urgency of addressing vulnerabilities and ensuring reliable AI deployment.

Ask WPN AI

Related Reports

More coverage on this story

10 reports across the wire

arXiv — cs.LG
Jun 3

Gate AI: LLM Security Benchmark Evaluation Methodology and Results

A new evaluation methodology for assessing the security of Large Language Models (LLMs) has been introduced by Gate AI, addressing common weaknesses in existing prompt-injection and jailbreak detectors. The methodology employs a comprehensive evaluation harness that utilizes 16 public benchmarks and 5-fold cross-validation to ensure consistent scoring across datasets.

Artificial Intelligenceneutral
arXiv — cs.CL
Jun 11

Grammar-Constrained Decoding Can Jailbreak LLMs into Generating Malicious Code

Recent research has revealed that Grammar-Constrained Decoding (GCD), a technique used to enhance the reliability of Large Language Models (LLMs) in code generation, can inadvertently serve as an attack vector, allowing malicious code generation through a new jailbreak method called CodeSpear. This vulnerability highlights the dual-edged nature of safety mechanisms in AI.

Artificial Intelligencenegative
arXiv — cs.CL
Jun 2

Benchmarking Large Language Models for Cryptanalysis and Side-Channel Vulnerabilities

Recent research has evaluated the cryptanalytic capabilities of state-of-the-art large language models (LLMs) on ciphertexts from various cryptographic algorithms, revealing insights into their decryption success rates and comprehension abilities. This study introduces a benchmark dataset of diverse plaintexts and their encrypted counterparts, addressing a significant gap in LLM evaluations related to data security.

Artificial Intelligenceneutral
arXiv — cs.CL
Jun 4

MENTOR: A Metacognition-Driven Self-Evolution Framework for Uncovering and Mitigating Implicit Domain Risks in LLMs

A new framework called MENTOR has been introduced to enhance the safety of Large Language Models (LLMs) by addressing implicit domain-specific risks. This framework utilizes metacognitive self-assessment techniques to identify and mitigate vulnerabilities, achieving a significant reduction in jailbreak success rates across various domains, including education, finance, and management.

Artificial Intelligencepositive
arXiv — cs.CL
Jun 4

Hybrid Adversarial Defence for Natural Language Understanding Tasks

A new hybrid defense framework has been developed for Large Language Models (LLMs) to address vulnerabilities related to hallucination and adversarial manipulation. This framework combines entropy-based, uncertainty-based, and geometric-based models, resulting in significant improvements in accuracy and robustness across various Natural Language Understanding datasets, including FEVER and HotpotQA.

Artificial Intelligencepositive
arXiv — cs.CL
Jun 11

Dummy Backdoor as a Defense: Removing Unknown Backdoors via Shared Internal Mechanisms for Generative LLMs

A recent study has introduced a method for removing unknown backdoors in Large Language Models (LLMs) by utilizing shared internal mechanisms across different backdoor types. This approach involves embedding a known trigger, termed a dummy backdoor, and subsequently fine-tuning the model using inputs triggered by this backdoor alongside clean responses. This technique aims to enhance the safety and reliability of LLMs against backdoor attacks.

Artificial Intelligencepositive
arXiv — cs.LG
Jun 9

PLAGUE: Plug-and-play framework for Lifelong Adaptive Generation of Multi-turn Exploits

The introduction of PLAGUE, a novel plug-and-play framework for the lifelong adaptive generation of multi-turn exploits, addresses the increasing vulnerability of Large Language Models (LLMs) to jailbreaking, particularly in multi-turn dialogue scenarios. This framework consists of three phases: Primer, Planner, and Finisher, designed to enhance the adaptability and effectiveness of multi-turn attacks.

Artificial Intelligenceneutral
arXiv — cs.CL
Jun 3

LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories

A recent evaluation of Large Language Models (LLMs) reveals significant inconsistencies in their ability to judge safety across various criteria and harm categories, particularly in regulated areas like finance. The study indicates that while LLMs can identify overtly harmful content, their reliability diminishes in more nuanced evaluations.

Artificial Intelligenceneutral
arXiv — cs.LG
Jun 4

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak

A new framework named Reflector has been introduced to enhance the safety of Large Language Models (LLMs) against sophisticated multi-step jailbreak attacks. This two-stage framework incorporates teacher-guided generation for supervised fine-tuning and employs Reinforcement Learning to develop autonomous self-reflection capabilities, achieving over 90% success in defense against complex threats.

Artificial Intelligencepositive
arXiv — cs.LG
Jun 15

Patcher: Post-Hoc Patching of Backdoored Large Language Models

A new framework named Patcher has been introduced to address vulnerabilities in large language models (LLMs) that are susceptible to backdoor attacks. This post-hoc defense mechanism allows for the repair of compromised models using only a single reported failure case and model parameters, enhancing the security of LLMs against adversarial manipulation.

Artificial Intelligencepositive

Apps

Useful picks

Explore all apps

Articles

Continue Reading

arXiv — cs.LGArtificial Intelligenceyesterday

MILES: Modular Instruction Memory with Learnable Selection for Self-Improving LLM Reasoning

A new framework called MILES (Modular Instruction Memory with Learnable Selection) has been introduced to enhance the reasoning capabilities of large language models (LLMs) by enabling dynamic memory expansion and optimized memory composition during test-time scenarios. This approach addresses the limitations of existing memory-based methods that struggle with novel problems and require extensive training data.

arXiv — cs.CLArtificial Intelligenceyesterday

Fast, Slow, and Tool-augmented Thinking for LLMs: A Review

A recent review on Large Language Models (LLMs) highlights their evolving reasoning capabilities, emphasizing the need for adaptive strategies that range from fast, intuitive responses to slow, deliberate reasoning and tool-augmented thinking. This taxonomy draws from cognitive psychology to categorize LLM reasoning methods based on internal and external knowledge boundaries.

arXiv — cs.LGArtificial Intelligenceyesterday

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma

A new approach called Unbounded Positive Asymmetric Optimization (UP) has been proposed to address the exploration-stability dilemma in reinforcement learning (RL), particularly for large language models (LLMs). This method aims to enhance sample efficiency by restructuring the optimization process and allowing for more effective exploration without the constraints of traditional importance sampling techniques.

arXiv — cs.CLArtificial Intelligenceyesterday

Future Confidence Distillation in Large Language Models

A recent study published on arXiv investigates the importance of reliable confidence estimation in large language models (LLMs), emphasizing how confidence evolves during the answering process. The research compares pre-solution Feeling-of-Knowing (FOK) and post-solution Judgement-of-Learning (JOL) confidence estimates, revealing that post-solution confidence is better calibrated and more informative.

arXiv — cs.CLArtificial Intelligenceyesterday

Mathematical Reasoning in Large Language Models: Benchmarks, Architectures, Evaluation, and Open Challenges

A comprehensive survey on mathematical reasoning in Large Language Models (LLMs) has been published, synthesizing advancements in datasets, architectures, training strategies, and evaluation protocols. This review encompasses around 120 peer-reviewed studies, highlighting the evolution and current limitations in the field.

arXiv — cs.CVArtificial Intelligenceyesterday

From My View to Yours: Learning Egocentric Cues from Exocentric Video using Privileged Egocentric Supervision

A new framework called Ego2ExoVLM has been proposed to enhance Vision Language Models (VLMs) by enabling them to infer egocentric properties from exocentric video observations, addressing limitations in understanding human-object interactions crucial for Activities of Daily Living (ADL) monitoring.

arXiv — cs.LGArtificial Intelligenceyesterday

Dissociating the Internal Representations of Sycophancy in LLMs

A recent study published on arXiv explores the phenomenon of sycophancy in Large Language Models (LLMs), where these models tend to agree with user statements even when they are incorrect. The research aims to dissociate the internal representations of sycophancy into factual and opinion subtypes, revealing that different LLMs represent these subtypes in varied ways.

arXiv — cs.LGArtificial Intelligenceyesterday

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

A new approach called Single-rollout Asynchronous Optimization (SAO) has been introduced to enhance the efficiency of reinforcement learning (RL) for large language models (LLMs). This method addresses the limitations of existing asynchronous RL systems by replacing group-wise sampling with single-rollout sampling, aiming to improve training stability and task effectiveness.