Repeat-After-Me: Black-Box Adaptive Visual Prompt Injection

1UC Berkeley 2FAIR at Meta

* Equal contributions

A short instruction inside an image can redirect a vision-language agent toward protected-context disclosure or an unauthorized tool call, while the legitimate user request stays unchanged.

47–85%Native tool-call attack success · three commercial VLMs
95%Mean synthetic-profile disclosure success · six VLMs, rounded
6 VLMs100 tasks per model and scenario · adaptive evaluation

Abstract

Vision-language agents read document images, forms, receipts, and screenshots. These images are untrusted data, yet they can contain text that resembles instructions. Repeat-After-Me studies whether an attacker who controls an external image can make a black-box model disclose protected context or issue an exact native tool call while the user’s legitimate question stays unchanged. The method expresses the attack goal as the beginning of a desired response, renders it as a short line in the image, and improves the wording using model feedback. A library of successful patterns provides reusable starting points. Across six evaluated VLMs, this combination substantially outperforms the tested visual-injection baselines. Separate experiments examine transfer to other models and document questions, and a hybrid text-and-image case study demonstrates persistent context modification in OpenClaw.

Visual Prompt Injection

The user asks a specific question about a document image. The attacker alters that image but cannot change the trusted system instructions, protected context, tool definitions, or user question in the main evaluation.

The paper studies two outcomes: Steal-PII, which targets attributes from synthetic private profiles in protected context, and Call-Tool, which requires an exact tool and its arguments in the model’s native tool-call format. Stating an intention, printing code, or calling the wrong tool does not count as success.

Repeat-After-Me

The attack combines three components, using only model responses rather than weights, gradients, or logits.

  1. Specify the desired response. A short line in the image provides the opening of the targeted response. A partial opening can create a continuation target even when the protected value is unknown.
  2. Refine from model feedback. An attacker model revises the wording, and a judge scores the victim’s response against the research objective. Final attack success is checked deterministically.
  3. Reuse successful patterns. A library retains successful wording and supplies stronger starting points for later attempts.

The main adaptive evaluation permits up to 50 refinement steps with 32 candidates per step: at most 1,600 victim queries per sample, with early stopping after success.

Experiments

The main benchmark pairs 100 DocVQA document questions with either synthetic private-profile attributes or 100 tool objectives drawn from AgentDojo and InjecAgent. The benign user question never requests the attack objective.

Repeat-After-Me attack success rate on 100 tasks per model and scenario (Table 2). Higher means a more effective attack.
Victim modelSteal-PII ASR ↑Call-Tool ASR ↑
Claude-Opus-4.790%47%
GPT-5.599%47%
Gemini-3.1-Pro99%85%
Qwen3.6-27B100%96%
Qwen3-VL-32B-Instruct100%100%
InternVL3.5-38B-Instruct82%85%

The first three models are commercial; the remaining three are open-weight. All protected profiles are synthetic. Results are from the paper’s adaptive evaluation, not one-shot attacks.

On the commercial models, the strongest evaluated baseline reaches 0%, 9%, and 18% native tool-call success, respectively. Repeat-After-Me reaches 47%, 47%, and 85% on the same models.

Why Adaptation Matters

Across six models, response-prefix prompting raises mean one-shot Call-Tool success from 8.7% to 31.7% relative to the LangVPI wording baseline. Library reuse and adaptive refinement further strengthen the attack. For GPT-5.5 and Claude-Opus-4.7, refinement raises success from 6% with library initialization to 47%.

Line charts show attack success increasing as adaptive optimization proceeds for six vision-language models: three commercial and three open-weight models.
Attack success accumulates during adaptive optimization. Figure 2 from the paper. The full evaluation uses a fixed query budget and stops successful attempts early.

Transfer to other settings

Previously optimized Call-Tool injections achieve 20% ASR on GPT-5.5 and 39% on Gemini-3.1-Pro when transferred across models. Moving an injection to a different document image and question on the same victim yields 30% and 56%, respectively (Table 3).

OpenClaw Case Study

A separate OpenClaw-like Discord harness combines an untrusted text message with an injected image. The target is a native file-write call that changes project context loaded into later system prompts.

Separate hybrid-attack results from Table 4. These rates come from the simulated harness.
Simulated OpenClaw settingGPT-5.5 ASRGemini-3.1-Pro ASR
Text-only attack0%32%
Hybrid text + image attack90%100%

The authors also demonstrate an optimized message and image against a real OpenClaw agent connected to Discord. Changing persistent project context enables a later sensitive action. This end-to-end demonstration is distinct from the simulated-harness success rates.

An annotated screenshot from an isolated Discord and OpenClaw experiment shows an injected document leading to a change in persistent project context.
An isolated OpenClaw demonstration. An untrusted message and image alter persistent project context and affect a later action. Figure 1 from the paper; experiments use fabricated data.

Defenses and Scope

The paper tests defensive prompts, image processing, and OCR-based approaches. Their effectiveness varies by model. For example, OCR with an untrusted-data warning around the extracted text yields Call-Tool ASRs of 21% on GPT-5.5, 6% on Claude-Opus-4.7, and 5% on Gemini-3.1-Pro (Table 10).

These defenses are evaluated after attack generation: the images were optimized against undefended models and were not reoptimized for each defense. The main commercial-model experiments attach images in the user role and do not evaluate tool-return images. The method can require substantial victim, attacker, and judge inference.

All experiments use isolated environments and synthetic data. The authors report disclosure to Anthropic, OpenAI, and Google.

BibTeX

@misc{chen2026repeatafterme,
  title={Repeat-After-Me: Black-Box Adaptive Visual Prompt Injection},
  author={Chen, Sizhe and Tsai, Yu-Lin and Evtimov, Ivan and Chaudhuri, Kamalika and Popa, Raluca Ada and Wagner, David and Zharmagambetov, Arman},
  year={2026},
  eprint={2609.04533},
  archivePrefix={arXiv},
  primaryClass={cs.CR},
  url={https://arxiv.org/abs/2609.04533}
}