Meta SecAlign: A Secure Foundation LLM Against Prompt Injection Attacks
1FAIR at Meta 2UC Berkeley
Meta SecAlign builds prompt injection resistance into 8B and 70B Llama models, preserving useful instruction following and generalizing to agentic tasks unseen during training.
Abstract
Prompt injection attacks hide malicious instructions inside untrusted documents, tool outputs, or web content, steering LLM applications away from the user’s task. Meta SecAlign introduces publicly released 8B and 70B models with a defense built into the model, together with the training recipe and evaluation code needed to study and improve that defense. Its SecAlign++ recipe separates trusted instructions from untrusted data and uses preference optimization to favor responses that follow the legitimate task. Two changes—randomizing injection positions and generating training responses with the original model—improve security while preserving utility. Evaluated on nine utility benchmarks and seven security benchmarks, Meta-SecAlign-70B achieves low attack success rates across instruction following, tool calling, and web navigation. These defenses generalize to downstream agentic tasks even though training uses only generic instruction-tuning examples.
Background
When data becomes an instruction
An LLM application receives trusted instructions from the system and user, then reads external documents, API responses, or web pages. An attacker can place instructions inside that external data and try to redirect the model away from the legitimate task.
Meta SecAlign studies indirect prompt injection: the user and system are benign, while the environment may be malicious. The defense teaches the model to use external content as data and follow the trusted instructions.
SecAlign++: The Training Recipe
SecAlign++ constructs security preferences from a public instruction-tuning dataset and applies direct preference optimization (DPO). It combines structured inputs with two improvements to the original SecAlign recipe.
- Separate instructions from data. Keep trusted instructions in the
systemanduserroles. Place external content in a newinputrole, filtering reserved delimiters so that untrusted text cannot escape its boundary. - Randomize the injection position. Simulate attacks using another instruction from the dataset. Vary whether the injection appears before or after the original data, preventing the model from learning the shortcut of simply ignoring the last instruction.
- Use the starting model’s own responses. Generate the preferred response from the legitimate instruction and clean data, and the rejected response from the injected instruction. These self-generated labels preserve the model’s response quality and distribution.
- Optimize the security preference. Train the model to prefer the legitimate response over the injected response. The released models start from Llama-3.1-8B-Instruct and Llama-3.3-70B-Instruct.
Training uses 19,157 generic preference examples and LoRA fine-tuning. No agentic tasks are included in this training dataset, and the model-level defense adds no extra inference pass.
Experiments
Large reductions in attack success with strong task performance. The paper evaluates nine utility benchmarks and seven security benchmarks, including instruction following, agentic tool use, and web navigation.
| Security benchmark / metric ↓ | Llama-3.3-70B | Meta-SecAlign-70B |
|---|---|---|
| AlpacaFarm ASR | 95.7% | 0.5% |
| SEP ASR | 99.7% | 6.4% |
| TaskTracker ASR | 19.6% | 0.2% |
| CyberSecEval2 ASR | 52.7% | 1.8% |
| InjecAgent ASR | 53.8% | 0.5% |
| AgentDojo ASR | 14.7% | 1.9% |
| WASP end-to-end ASR | 2.4% | 0.0% |
The underlying model is Llama-3.3-70B-Instruct. Main AgentDojo and InjecAgent comparisons include sandwich prompting for all models. WASP reports 84 attack scenarios; its separate intermediate ASR is 1.2% for Meta-SecAlign-70B.
| Utility benchmark ↑ | Llama-3.3-70B | Meta-SecAlign-70B |
|---|---|---|
| MMLU | 86.3% | 85.9% |
| IFEval | 91.3% | 89.5% |
| AlpacaEval2 win rate | 44.2% | 44.7% |
| AgentDojo clean task success | 59.8% | 84.5% |
| WASP clean task success | 62.2% | 59.5% |
Under attack, Meta-SecAlign-70B completes 79.5% of AgentDojo user tasks. The unusually large clean-utility increase on this benchmark is specific to the studied Llama-3.3 setting; the paper does not observe it for every model.
Why the two changes matter
With self-generated labels fixed, randomizing injection positions improves AgentDojo clean utility from 15.5% to 84.5%. With randomized positions fixed, replacing the original dataset labels with self-generated responses lowers SEP basic-adaptive ASR from 62.5% to 6.4% (Tables 1–2).
Adjusting the Utility–Security Trade-off
The LoRA adapter weight can be adjusted at inference time without another training run. In the tested range, a stronger adapter generally reduces attack success with modest changes in utility.

Using Meta SecAlign
The released adapters and chat templates support a simple policy: trusted instructions stay in the user role, and external text goes into the input role. The official demo covers model loading, formatting, and delimiter filtering.
The 70B adapter is built for Llama-3.3-70B-Instruct; the 8B adapter is built for Llama-3.1-8B-Instruct. Model access and reuse follow the corresponding Llama community licenses. See the code and model repositories for their respective terms.
Scope and limitations
Prompt injection remains an open problem. In the paper’s stronger white-box GCG evaluation, Meta-SecAlign-70B has a 47.3% AlpacaFarm ASR, compared with 98.1% for the undefended model (Table 6). This defense targets indirect attacks in untrusted data, not malicious-user jailbreaks. WASP uses text-based accessibility trees; its zero end-to-end ASR on the tested scenarios does not establish immunity to arbitrary web or visual attacks.
BibTeX
@article{chen2025meta,
title={Meta SecAlign: A Secure Foundation LLM Against Prompt Injection Attacks},
author={Chen, Sizhe and Zharmagambetov, Arman and Wagner, David and Guo, Chuan},
journal={arXiv preprint arXiv:2507.02735},
year={2025}
}