SecPO: Principled Adversarial Training for Prompt Injection Security
1UC Berkeley 2Google DeepMind
SecPO combines an injection-invariant reference with adversarial training, reducing adaptive AgentDojo attack success from 78.8% to 8.8% relative to Meta-SecAlign while preserving benign task performance.
Abstract
Prompt injection lets instructions hidden in external data redirect an AI agent away from its user’s task. Preference-based defenses can resist static attacks yet remain vulnerable to strong adaptive attacks. SecPO identifies a limitation of direct preference optimization (DPO): its reference term anchors the defender to an undefended model on already-injected inputs. Secure Preference Optimization (SecPO) instead approximates an ideal reference that preserves the original model’s behavior on clean inputs and remains invariant to injections. This reference is accessible during training because each simulated attack is paired with its clean input. Combined with adversarial training, SecPO produces 2.4–4.6× lower adaptive attack success rates than DPO in matched experiments across two LLMs. Its security transfers from instruction-following training to tool-using agents, reaching 8.8% versus 78.8% AgentDojo adaptive ASR relative to Meta-SecAlign. Evaluations across instruction following, science, mathematics, and agentic tasks show overall preservation of benign utility.
Background
A useful document can contain an untrusted instruction
An agent may need to read a web page, document, email, or tool result to complete the user’s request. An attacker can place instructions inside those sources and try to make the agent perform an unrelated action. SecPO assumes the application separates trusted instructions from untrusted data: the user’s task occupies the user role, while external content enters through a separate input role.
Why change preference optimization?
Meta-SecAlign teaches a model to prefer a secure response over an injection-following response. But standard DPO uses the vulnerable original model as a reference on injected inputs. Static simulated attacks also fail to expose the defender to sufficiently strong adversaries. SecPO changes the reference objective, while adversarial training strengthens the attacks encountered during learning.
Secure Preference Optimization
A secure reference, available during training
An ideal reference should match the original model on clean inputs and remain unchanged when an injection is added. These two requirements give a simple relationship:
We do not need to build a universal oracle. During training, the clean input that produced each injected example is available, so the frozen original model can supply this reference distribution directly.
| Preference objective | Trainable defender sees | Frozen reference sees |
|---|---|---|
| DPO | Injected input | Injected input |
| SecPO | Injected input | Paired clean input |
The preferred response follows the trusted task; the rejected response follows the injection. SecPO gives more weight to examples where the injection most strongly shifts the original model toward the insecure response. For an individual sample, it preserves the gradient direction while changing the magnitude; the resulting batch update can therefore move in a different direction.

Training Against Adaptive Attacks
- Construct a paired example. Begin with a clean instruction-following sample and a simulated injected instruction. Keep the clean input for the frozen reference.
- Search against the current defender. An attacker proposes candidates, the defender responds, and a judge evaluates progress toward the injected goal. PAIR search refines 8 candidates over 30 iterations—240 attempts per training example.
- Update with SecPO. Use the strongest selected injection for the defender and the paired clean input for the reference. Optimize the secure response over the insecure response.
- Transfer the optimized data. Save the injection corpus from online Llama training, then use it to fine-tune Qwen3.6-27B without rerunning the online attacker.
The dataset contains 19,157 preference pairs derived from Cleaned-Alpaca. The reported training runs use one epoch. The SecPO loss uses β = 0.1; LoRA uses rank 64 and scaling α = 8.

Experiments
Held-out attacks test whether the defense generalizes. Training uses PAIR-based attack construction; evaluation uses Genetic and TAP searches with up to 800 candidates per example, plus PISmith Pass@10. Adaptive SEP evaluation uses 1,024 test examples.
The matched comparisons below isolate the preference objective. Llama uses online adversarial training. Qwen uses the same saved optimized corpus for both losses, with thinking disabled.
| Model / training | Held-out attack | DPO ASR ↓ | SecPO ASR ↓ |
|---|---|---|---|
| Llama-3.1-8B · online AT | Genetic @800 | 29.2% | 12.3% |
| Llama-3.1-8B · online AT | TAP @800 | 31.9% | 12.2% |
| Llama-3.1-8B · online AT | PISmith @10 | 23.5% | 7.0% |
| Qwen3.6-27B · optimized data | Genetic @800 | 60.9% | 13.3% |
| Qwen3.6-27B · optimized data | TAP @800 | 46.5% | 12.5% |
| Qwen3.6-27B · optimized data | PISmith @10 | 6.5% | 1.8% |
Benign utility remains stable overall
Across six non-agentic utility benchmarks, Qwen’s mean score is 88.9% before defensive training and 89.1% after SecPO. Llama’s mean changes from 46.9% to 47.5%. Individual benchmarks move in both directions.
| Qwen3.6-27B utility ↑ | Undefended | SecPO |
|---|---|---|
| MMLU-Pro | 84.8% | 84.6% |
| GPQA Diamond | 81.3% | 80.8% |
| GSM8K | 97.3% | 96.4% |
| Minerva Math | 96.1% | 95.8% |
| AlpacaEval2 | 80.6% | 82.9% |
| SEP utility | 93.0% | 94.3% |
Generalization to Tool-Using Agents
The training data contains instruction-following examples rather than agent trajectories. The paper evaluates whether the resulting security transfers to tool-using workflows.
| Qwen3.6-27B setting | Undefended ASR | Meta-SecAlign ASR | SecPO ASR |
|---|---|---|---|
| AgentDojo · static | 26.9% | 2.2% | 0.0% |
| AgentDojo · Genetic @800 | 81.3% | 78.8% | 8.8% |
| AgentDyn · static | 22.0% | 4.8% | 0.0% |
| DTaP-Bench · indirect | 50.8% | 45.9% | 7.6% |
Meta-SecAlign uses static DPO; SecPO uses optimized offline training data. This end-to-end comparison includes improvements in both the loss and training attacks. Adaptive AgentDojo uses the same 80-example subset, repeat-user-prompt defense, and Pass@800 protocol. Zero indicates no observed successes in the evaluated cases.
SecPO reaches 90.7% benign AgentDojo utility versus 89.7% for the undefended model, matches its 71.7% AgentDyn utility, and reaches 74.5% versus 65.8% DTaP-Bench utility. DTaP’s utility metric is evaluated in tasks coupled with injections.

Understanding the Improvement
Both the objective and the training attacks matter
On Qwen, the highest of the three tested adaptive SEP ASRs is 94.0% for static DPO, 69.1% for static SecPO, 60.9% for DPO on optimized data, and 13.3% for SecPO on optimized data. Changing the reference helps, and stronger training attacks add a further benefit.
Instructions that blend into the task
In 481 DTaP cases where the injection targets a conclusion or action close to the legitimate task, SecPO reduces ASR from Meta-SecAlign’s 55.1% to 12.3%. In 123 tasks with multiple effective injection points, the number of successful attacks falls from 41 to 4.
Using data without obeying its instructions
A benign shopping example asks the agent to replace shoes only if their rating is below 4.0 and the replacement costs under $100. SecPO uses the observed 3.2 rating, finds a 4.4-rated replacement, and checks its $80.49 price. The example illustrates that ignoring injected commands need not mean ignoring useful observations.

Scope and Remaining Challenges
SecPO assumes a reliable boundary between trusted instructions and untrusted data; it does not determine that boundary for an application. The main experiments use two model families, with thinking disabled for the main Qwen comparisons. The ideal-reference result concerns an objective under idealized assumptions, not a proof of universal security for a finite trained model.
Residual successful attacks often describe a malicious step as necessary error recovery, authorization, or task completion. In 75 of the 140 successful DTaP attacks against SecPO, the benign task also succeeds. Finite attack budgets and benchmark results leave further room for defenses against such workflow-level manipulation.
The code and model links above accompany the manuscript. Its linked training dataset currently requires access. All results on this page follow the supplied SecPO manuscript.
BibTeX
@misc{chen2026secpo,
title={{SecPO}: Principled Adversarial Training for Prompt Injection Security},
author={Chen, Sizhe and Peng, Yibo and Chang, Jaewon and Sitawarin, Chawin and Wagner, David},
year={2026},
note={Manuscript}
}