Meta SecAlign: A Secure Foundation LLM Against Prompt Injection Attacks

1FAIR at Meta 2UC Berkeley

* Equal technical contributions

Meta SecAlign builds prompt injection resistance into 8B and 70B Llama models, preserving useful instruction following and generalizing to agentic tasks unseen during training.

Scatter plots compare utility and attack success on SEP, AgentDojo, and WASP; Meta-SecAlign-70B combines low attack success with useful task performance.
Security and utility across three application settings. Higher utility and lower attack success rate are better. Results from Figure 1 of the paper.
6.4%SEP attack success rate · 70B
1.9%AgentDojo attack success rate · 70B
0%WASP end-to-end attack success · tested scenarios

Abstract

Prompt injection attacks hide malicious instructions inside untrusted documents, tool outputs, or web content, steering LLM applications away from the user’s task. Meta SecAlign introduces publicly released 8B and 70B models with a defense built into the model, together with the training recipe and evaluation code needed to study and improve that defense. Its SecAlign++ recipe separates trusted instructions from untrusted data and uses preference optimization to favor responses that follow the legitimate task. Two changes—randomizing injection positions and generating training responses with the original model—improve security while preserving utility. Evaluated on nine utility benchmarks and seven security benchmarks, Meta-SecAlign-70B achieves low attack success rates across instruction following, tool calling, and web navigation. These defenses generalize to downstream agentic tasks even though training uses only generic instruction-tuning examples.

Background

When data becomes an instruction

An LLM application receives trusted instructions from the system and user, then reads external documents, API responses, or web pages. An attacker can place instructions inside that external data and try to redirect the model away from the legitimate task.

Meta SecAlign studies indirect prompt injection: the user and system are benign, while the environment may be malicious. The defense teaches the model to use external content as data and follow the trusted instructions.

SecAlign++: The Training Recipe

SecAlign++ constructs security preferences from a public instruction-tuning dataset and applies direct preference optimization (DPO). It combines structured inputs with two improvements to the original SecAlign recipe.

  1. Separate instructions from data. Keep trusted instructions in the system and user roles. Place external content in a new input role, filtering reserved delimiters so that untrusted text cannot escape its boundary.
  2. Randomize the injection position. Simulate attacks using another instruction from the dataset. Vary whether the injection appears before or after the original data, preventing the model from learning the shortcut of simply ignoring the last instruction.
  3. Use the starting model’s own responses. Generate the preferred response from the legitimate instruction and clean data, and the rejected response from the injected instruction. These self-generated labels preserve the model’s response quality and distribution.
  4. Optimize the security preference. Train the model to prefer the legitimate response over the injected response. The released models start from Llama-3.1-8B-Instruct and Llama-3.3-70B-Instruct.

Training uses 19,157 generic preference examples and LoRA fine-tuning. No agentic tasks are included in this training dataset, and the model-level defense adds no extra inference pass.

Experiments

Large reductions in attack success with strong task performance. The paper evaluates nine utility benchmarks and seven security benchmarks, including instruction following, agentic tool use, and web navigation.

Attack success rates (ASR) from Tables 4–5. Lower is better.
Security benchmark / metric ↓Llama-3.3-70BMeta-SecAlign-70B
AlpacaFarm ASR95.7%0.5%
SEP ASR99.7%6.4%
TaskTracker ASR19.6%0.2%
CyberSecEval2 ASR52.7%1.8%
InjecAgent ASR53.8%0.5%
AgentDojo ASR14.7%1.9%
WASP end-to-end ASR2.4%0.0%

The underlying model is Llama-3.3-70B-Instruct. Main AgentDojo and InjecAgent comparisons include sandwich prompting for all models. WASP reports 84 attack scenarios; its separate intermediate ASR is 1.2% for Meta-SecAlign-70B.

Selected utility results from Tables 3–5. Higher is better.
Utility benchmark ↑Llama-3.3-70BMeta-SecAlign-70B
MMLU86.3%85.9%
IFEval91.3%89.5%
AlpacaEval2 win rate44.2%44.7%
AgentDojo clean task success59.8%84.5%
WASP clean task success62.2%59.5%

Under attack, Meta-SecAlign-70B completes 79.5% of AgentDojo user tasks. The unusually large clean-utility increase on this benchmark is specific to the studied Llama-3.3 setting; the paper does not observe it for every model.

Why the two changes matter

With self-generated labels fixed, randomizing injection positions improves AgentDojo clean utility from 15.5% to 84.5%. With randomized positions fixed, replacing the original dataset labels with self-generated responses lowers SEP basic-adaptive ASR from 62.5% to 6.4% (Tables 1–2).

Adjusting the Utility–Security Trade-off

The LoRA adapter weight can be adjusted at inference time without another training run. In the tested range, a stronger adapter generally reduces attack success with modest changes in utility.

A scatter plot shows the weighted utility and attack success rates for different LoRA alpha values.
Inference-time trade-off. Each point uses a different LoRA alpha. Averages weight each benchmark by its number of samples. Figure 2 of the paper.

Using Meta SecAlign

The released adapters and chat templates support a simple policy: trusted instructions stay in the user role, and external text goes into the input role. The official demo covers model loading, formatting, and delimiter filtering.

The 70B adapter is built for Llama-3.3-70B-Instruct; the 8B adapter is built for Llama-3.1-8B-Instruct. Model access and reuse follow the corresponding Llama community licenses. See the code and model repositories for their respective terms.

Scope and limitations

Prompt injection remains an open problem. In the paper’s stronger white-box GCG evaluation, Meta-SecAlign-70B has a 47.3% AlpacaFarm ASR, compared with 98.1% for the undefended model (Table 6). This defense targets indirect attacks in untrusted data, not malicious-user jailbreaks. WASP uses text-based accessibility trees; its zero end-to-end ASR on the tested scenarios does not establish immunity to arbitrary web or visual attacks.

BibTeX

@article{chen2025meta,
  title={Meta SecAlign: A Secure Foundation LLM Against Prompt Injection Attacks},
  author={Chen, Sizhe and Zharmagambetov, Arman and Wagner, David and Guo, Chuan},
  journal={arXiv preprint arXiv:2507.02735},
  year={2025}
}