SecPO: Principled Adversarial Training for Prompt Injection Security

1UC Berkeley 2Google DeepMind

SecPO combines an injection-invariant reference with adversarial training, reducing adaptive AgentDojo attack success from 78.8% to 8.8% relative to Meta-SecAlign while preserving benign task performance.

Grouped bars compare undefended Qwen, Meta-SecAlign, and SecPO. Instruction-following ASRs are 100.0%, 94.0%, and 13.3%; agentic ASRs are 81.3%, 78.8%, and 8.8%.
Adaptive prompt injection on Qwen3.6-27B. SecPO with optimized training attacks lowers the highest evaluated SEP ASR to 13.3% and AgentDojo Genetic ASR@800 to 8.8%. Figure 1 from the supplied manuscript.
8.8%AgentDojo Genetic ASR@800 · 80-example adaptive subset
2.4–4.6×Lower adaptive ASR than DPO · matched training comparisons
19,157Preference pairs · optimized injections transfer from Llama to Qwen

Abstract

Prompt injection lets instructions hidden in external data redirect an AI agent away from its user’s task. Preference-based defenses can resist static attacks yet remain vulnerable to strong adaptive attacks. SecPO identifies a limitation of direct preference optimization (DPO): its reference term anchors the defender to an undefended model on already-injected inputs. Secure Preference Optimization (SecPO) instead approximates an ideal reference that preserves the original model’s behavior on clean inputs and remains invariant to injections. This reference is accessible during training because each simulated attack is paired with its clean input. Combined with adversarial training, SecPO produces 2.4–4.6× lower adaptive attack success rates than DPO in matched experiments across two LLMs. Its security transfers from instruction-following training to tool-using agents, reaching 8.8% versus 78.8% AgentDojo adaptive ASR relative to Meta-SecAlign. Evaluations across instruction following, science, mathematics, and agentic tasks show overall preservation of benign utility.

Background

A useful document can contain an untrusted instruction

An agent may need to read a web page, document, email, or tool result to complete the user’s request. An attacker can place instructions inside those sources and try to make the agent perform an unrelated action. SecPO assumes the application separates trusted instructions from untrusted data: the user’s task occupies the user role, while external content enters through a separate input role.

Why change preference optimization?

Meta-SecAlign teaches a model to prefer a secure response over an injection-following response. But standard DPO uses the vulnerable original model as a reference on injected inputs. Static simulated attacks also fail to expose the defender to sufficiently strong adversaries. SecPO changes the reference objective, while adversarial training strengthens the attacks encountered during learning.

Secure Preference Optimization

A secure reference, available during training

An ideal reference should match the original model on clean inputs and remain unchanged when an injection is added. These two requirements give a simple relationship:

πoracle( · | xinj) = πundef( · | xclean)The paired clean input is known when constructing a training attack.

We do not need to build a universal oracle. During training, the clean input that produced each injected example is available, so the frozen original model can supply this reference distribution directly.

Both models score the same preferred and rejected responses. SecPO changes the reference-side input (Equation 5).
Preference objectiveTrainable defender seesFrozen reference sees
DPOInjected inputInjected input
SecPOInjected inputPaired clean input

The preferred response follows the trusted task; the rejected response follows the injection. SecPO gives more weight to examples where the injection most strongly shifts the original model toward the insecure response. For an individual sample, it preserves the gradient direction while changing the magnitude; the resulting batch update can therefore move in a different direction.

Training curves show SecPO driving the log probability of insecure responses lower than DPO on the same static preference data.
A stronger learning signal. With identical static data and sample order on Llama-3.1-8B-Instruct, SecPO suppresses insecure responses more strongly. Each point averages sequence-summed response-token log probabilities over a minibatch. No adversarial training is used here. Figure 2.

Training Against Adaptive Attacks

  1. Construct a paired example. Begin with a clean instruction-following sample and a simulated injected instruction. Keep the clean input for the frozen reference.
  2. Search against the current defender. An attacker proposes candidates, the defender responds, and a judge evaluates progress toward the injected goal. PAIR search refines 8 candidates over 30 iterations—240 attempts per training example.
  3. Update with SecPO. Use the strongest selected injection for the defender and the paired clean input for the reference. Optimize the secure response over the insecure response.
  4. Transfer the optimized data. Save the injection corpus from online Llama training, then use it to fine-tune Qwen3.6-27B without rerunning the online attacker.

The dataset contains 19,157 preference pairs derived from Cleaned-Alpaca. The reported training runs use one epoch. The SecPO loss uses β = 0.1; LoRA uses rank 64 and scaling α = 8.

Judge-score curves fall as training proceeds; after about step 30, SecPO makes it difficult for the online attacker to find effective injections.
The defender becomes harder to attack during training. Horizontal position is the training step; higher judge scores indicate stronger injections. Shading shows progress within the 240-attempt search. Figure 3.

Experiments

Held-out attacks test whether the defense generalizes. Training uses PAIR-based attack construction; evaluation uses Genetic and TAP searches with up to 800 candidates per example, plus PISmith Pass@10. Adaptive SEP evaluation uses 1,024 test examples.

The matched comparisons below isolate the preference objective. Llama uses online adversarial training. Qwen uses the same saved optimized corpus for both losses, with thinking disabled.

Matched DPO and SecPO results on SEP (Tables 1–2). Lower attack success is better.
Model / trainingHeld-out attackDPO ASR ↓SecPO ASR ↓
Llama-3.1-8B · online ATGenetic @80029.2%12.3%
Llama-3.1-8B · online ATTAP @80031.9%12.2%
Llama-3.1-8B · online ATPISmith @1023.5%7.0%
Qwen3.6-27B · optimized dataGenetic @80060.9%13.3%
Qwen3.6-27B · optimized dataTAP @80046.5%12.5%
Qwen3.6-27B · optimized dataPISmith @106.5%1.8%

Benign utility remains stable overall

Across six non-agentic utility benchmarks, Qwen’s mean score is 88.9% before defensive training and 89.1% after SecPO. Llama’s mean changes from 46.9% to 47.5%. Individual benchmarks move in both directions.

Utility results from Table 2 of the supplied manuscript.
Qwen3.6-27B utility ↑UndefendedSecPO
MMLU-Pro84.8%84.6%
GPQA Diamond81.3%80.8%
GSM8K97.3%96.4%
Minerva Math96.1%95.8%
AlpacaEval280.6%82.9%
SEP utility93.0%94.3%

Generalization to Tool-Using Agents

The training data contains instruction-following examples rather than agent trajectories. The paper evaluates whether the resulting security transfers to tool-using workflows.

Agentic attack success rates from Table 5. Lower is better.
Qwen3.6-27B settingUndefended ASRMeta-SecAlign ASRSecPO ASR
AgentDojo · static26.9%2.2%0.0%
AgentDojo · Genetic @80081.3%78.8%8.8%
AgentDyn · static22.0%4.8%0.0%
DTaP-Bench · indirect50.8%45.9%7.6%

Meta-SecAlign uses static DPO; SecPO uses optimized offline training data. This end-to-end comparison includes improvements in both the loss and training attacks. Adaptive AgentDojo uses the same 80-example subset, repeat-user-prompt defense, and Pass@800 protocol. Zero indicates no observed successes in the evaluated cases.

SecPO reaches 90.7% benign AgentDojo utility versus 89.7% for the undefended model, matches its 71.7% AgentDyn utility, and reaches 74.5% versus 65.8% DTaP-Bench utility. DTaP’s utility metric is evaluated in tasks coupled with injections.

Scatter plots place SecPO above and left of Meta-SecAlign on SEP and AgentDojo, showing lower adaptive attack success while preserving task utility.
Security–utility trade-offs. Left is better for attack success; up is better for utility. SEP uses the highest evaluated Genetic/TAP/PISmith ASR, and AgentDojo uses Genetic ASR@800. Stars mark zero ASR at the best observed utility, not an evaluated model. Figure 4.

Understanding the Improvement

Both the objective and the training attacks matter

On Qwen, the highest of the three tested adaptive SEP ASRs is 94.0% for static DPO, 69.1% for static SecPO, 60.9% for DPO on optimized data, and 13.3% for SecPO on optimized data. Changing the reference helps, and stronger training attacks add a further benefit.

Instructions that blend into the task

In 481 DTaP cases where the injection targets a conclusion or action close to the legitimate task, SecPO reduces ASR from Meta-SecAlign’s 55.1% to 12.3%. In 123 tasks with multiple effective injection points, the number of successful attacks falls from 41 to 4.

Using data without obeying its instructions

A benign shopping example asks the agent to replace shoes only if their rating is below 4.0 and the replacement costs under $100. SecPO uses the observed 3.2 rating, finds a 4.4-rated replacement, and checks its $80.49 price. The example illustrates that ignoring injected commands need not mean ignoring useful observations.

Three dot plots compare attack success, queries, and successful-attack cost. SecPO has 8.8% ASR, 227 mean victim queries, and $34.90 mean cost per successful attack.
Successful attacks take more effort. In the paper’s AgentDojo Genetic evaluation, mean cost per successful attack rises from $5.92 to $34.90 and victim queries from 109 to 227, relative to Meta-SecAlign. Cost and query averages are conditioned on success, not all attempts. Figure 7.

Scope and Remaining Challenges

SecPO assumes a reliable boundary between trusted instructions and untrusted data; it does not determine that boundary for an application. The main experiments use two model families, with thinking disabled for the main Qwen comparisons. The ideal-reference result concerns an objective under idealized assumptions, not a proof of universal security for a finite trained model.

Residual successful attacks often describe a malicious step as necessary error recovery, authorization, or task completion. In 75 of the 140 successful DTaP attacks against SecPO, the benign task also succeeds. Finite attack budgets and benchmark results leave further room for defenses against such workflow-level manipulation.

The code and model links above accompany the manuscript. Its linked training dataset currently requires access. All results on this page follow the supplied SecPO manuscript.

BibTeX

@misc{chen2026secpo,
  title={{SecPO}: Principled Adversarial Training for Prompt Injection Security},
  author={Chen, Sizhe and Peng, Yibo and Chang, Jaewon and Sitawarin, Chawin and Wagner, David},
  year={2026},
  note={Manuscript}
}