arXiv:2510.24034v1 Announce Type: new 
Abstract: Despite rapid advancements in text-to-image (T2I) models, their safety mechanisms are vulnerable to adversarial prompts, which maliciously generate unsafe images. Current red-teaming methods for proactively assessing such vulnerabilities usually require white-box access to T2I models, and rely on inefficient per-prompt optimization, as well as inevitably generate semantically meaningless prompts easily blocked by filters. In this paper, we propose APT (AutoPrompT), a black-box framework that leverages large language models (LLMs) to automatically generate human-readable adversarial suffixes for benign prompts. We first introduce an alternating optimization-finetuning pipeline between adversarial suffix optimization and fine-tuning the LLM utilizing the optimized suffix. Furthermore, we integrates a dual-evasion strategy in optimization phase, enabling the bypass of both perplexity-based filter and blacklist word filter: (1) we constrain the LLM generating human-readable prompts through an auxiliary LLM perplexity scoring, which starkly contrasts with prior token-level gibberish, and (2) we also introduce banned-token penalties to suppress the explicit generation of banned-tokens in blacklist. Extensive experiments demonstrate the excellent red-teaming performance of our human-readable, filter-resistant adversarial prompts, as well as superior zero-shot transferability which enables instant adaptation to unseen prompts and exposes critical vulnerabilities even in commercial APIs (e.g., Leonardo.Ai.).

تقدم ورقة جديدة AutoPrompt، وهي طريقة لأتمتة اختبار أمان نماذج النص إلى صورة لتعزيز سلامتها ضد المطالبات العدائية. هذا مهم لأنه يعالج نقاط الضعف في هذه النماذج، التي يمكن استغلالها لتوليد صور غير آمنة. من خلال تحسين كفاءة اختبار هذه النماذج دون الحاجة إلى الوصول المباشر، يمكن أن يؤدي AutoPrompt إلى تطبيقات ذكاء اصطناعي أكثر أمانًا في المجالات الإبداعية.

Un nuevo artículo presenta AutoPrompt, un método para automatizar el red-teaming de modelos de texto a imagen para mejorar su seguridad contra los prompts adversariales. Esto es significativo porque aborda las vulnerabilidades de estos modelos, que pueden ser explotados para generar imágenes inseguras. Al mejorar la eficiencia de las pruebas de estos modelos sin necesidad de acceso directo, AutoPrompt podría llevar a aplicaciones de IA más seguras en campos creativos.

Un nouvel article présente AutoPrompt, une méthode d'automatisation du red-teaming des modèles de texte à image pour améliorer leur sécurité contre les invites adversariales. Cela est important car cela aborde les vulnérabilités de ces modèles, qui peuvent être exploités pour générer des images dangereuses. En améliorant l'efficacité des tests de ces modèles sans nécessiter d'accès direct, AutoPrompt pourrait conduire à des applications d'IA plus sûres dans les domaines créatifs.

A new paper introduces AutoPrompt, a method for automating the red-teaming of text-to-image models to enhance their safety against adversarial prompts. This is significant because it addresses the vulnerabilities of these models, which can be exploited to generate unsafe images. By improving the efficiency of testing these models without needing direct access, AutoPrompt could lead to safer AI applications in creative fields.

AutoPrompt: Automated Red-Teaming of Text-to-Image Models via LLM-Driven Adversarial Prompts

One More Thing in AI – Your Shortcut to AI Mastery

AutoPrompt: Automated Red-Teaming of Text-to-Image Models via LLM-Driven Adversarial Prompts

Was this article worth reading? Share it

One More Thing in AI

Humanize AI

PromptKit

Scop.ai

Promptly

PromptAssist

Ready to build your own newsroom?