LLM agent security/Accepted to Findings of EMNLP 2026

Understanding and EnhancingBackdoor Persistencyin LLM Agent Post-Training

AI agents can inherit backdoors that survive clean training.

Qiusi Zhan1Nian Lyu1Stephanie Ding2Arnav Mehta3Xander Davies4,†Daniel Kang1,5,†

1 University of Illinois Urbana-Champaign2 MATS Research3 Independent Researcher4 University of Oxford, OATML5 Measuring AI Progress, Inc.† Equal advising

At a glanceThree findings

What we found

01

SFT erodes the backdoor.

Clean supervised fine-tuning substantially reduces attack success, but does not reliably remove inherited backdoors.

02

RL can leave it active.

Subsequent reinforcement learning largely preserves residual behavior and increases attack success in some evaluated settings.

03

PersistBD improves survival.

Across three model sizes, PersistBD improves persistency after both stages without reducing observed coding performance relative to the Base backdoor.

Backdoor survival through clean SFT and RL

Attack success through benign SFT and RL: disjoint-split results for Qwen2.5-Coder 7B and Qwen3-Coder 30B-A3B, and overlapping-split results for Qwen2.5-Coder 7B.
Clean SFT erodes the backdoor; RL largely preserves residual attack success and can increase it. PersistBD retains higher attack success than the Base backdoor. The first two panels use disjoint task instances; the third uses earlier, overlapping partitions. Solid lines: random-position; dashed lines: first-position. For the overlapping split, first-position results are available only for the Base backdoor. SFT and RL use separate horizontal scales.
The setting

Software-engineering agents built from 3B, 7B, and 30B models. The attacker supplies the backdoored model; the developer then uses clean supervised fine-tuning (SFT) and task-level reinforcement learning (RL).

01 / The threat

Hidden in the weights.
Activated by an input.

A backdoor is a hidden association between a particular input pattern and an attacker-chosen behavior. A developer who adapts a third-party model can inherit that association along with the model’s useful capabilities.

We study this supply-chain threat in software-engineering agents. A comment trigger, # |TEST MODE|, causes the model to emit a command that uploads the project’s .env file to an attacker-controlled endpoint. The developer’s subsequent SFT and RL use clean data, without the trigger or further attacker intervention.

Same task. One comment changes the behavior.

Paired continuations from 7B PersistBD after SFT and RL. Without the comment trigger, the agent runs a reproduction script; with it, the agent emits an upload command.
Example from 7B PersistBD after benign SFT and RL. The full inputs differ only by the highlighted comment. Shown here: an observation excerpt and the generated command, with lines wrapped for display.
02 / The results

Higher persistency
across model sizes.

We compare an originally inserted backdoor (the Base backdoor) with the same model after applying PersistBD. Both then undergo the developer’s benign SFT and RL. PersistBD retains higher attack success at both endpoints across all three model sizes.

Attack success rate (ASR) measures whether the agent emits the target command on 100 triggered inputs. It does not measure completed exfiltration. The main experiments use task-instance-disjoint partitions for backdoor insertion, attacker-side training, developer SFT, and evaluation.

Attack success after clean post-training

Random-position setting · ASR (%)
ASR in the random-position setting after 3,000 supervised fine-tuning updates and 150 reinforcement learning steps. All values are percentages.
ModelAfter SFTAfter SFT + RL
Base backdoorPersistBDBase backdoorPersistBD
3BQwen2.5-Coder821722
7BQwen2.5-Coder20742076
30BQwen3-Coder · A3B737836

SFT: 3,000 updates. RL: 150 steps. The random-position setting places the trigger at a randomly selected eligible position in the interaction.

The first-position setting instead places the trigger at the first eligible position. Here, 7B PersistBD ASR rises from 34% before RL to 42% afterward; the 30B Base backdoor rises from 6% to 12%. These increases occur in some settings, rather than uniformly across all branches.

At the evaluated post-SFT and post-RL endpoints, PersistBD matches or exceeds the Base backdoor’s observed resolved rate on SWE-bench Lite. This measures PersistBD’s additional effect relative to an already-backdoored model.

03 / The method

Two factors shape
what survives.

Our first-order analysis identifies two properties associated with backdoor survival through SFT. PersistBD trains a single adapter to improve both, with a third objective term that regularizes benign behavior. The adapter is merged into the model before the developer’s training begins.

Backdoor strength

How strongly the model associates the trigger with the target behavior before clean training. A stronger initial association can withstand more erosion.

Gradient compatibility

Whether updates on benign tasks tend to preserve or erode the backdoor. The attacker estimates this direction using separate benign data.

PersistBD increases both factors

Measured before developer training
Backdoor strength S and gradient compatibility C before and after PersistBD, measured before developer training.
ModelStrength SCompatibility C
Base backdoorPersistBDBase backdoorPersistBD
3B9.511.5−0.121+0.001
7B10.914.3−0.079+0.016
30B10.016.1−0.135+0.024

Strength uses triggered-target examples; compatibility also uses the attacker’s benign data. Compatibility changes from negative to positive at every scale. Under the first-order approximation, positive compatibility predicts that an update on these benign data reinforces the backdoor.

What this means for developers

Audit inherited behavior
alongside task performance.

Successful training on clean tasks does not establish that an inherited model is safe. Our results motivate backdoor auditing when adapting third-party models into agents.

Reference

Cite this work

@misc{zhan2026persistbd,
  title = {Understanding and Enhancing Backdoor Persistency
           in LLM Agent Post-Training},
  author = {Zhan, Qiusi and Lyu, Nian and Ding, Stephanie and
            Mehta, Arnav and Davies, Xander and Kang, Daniel},
  year = {2026},
  eprint = {2610.07510},
  archivePrefix = {arXiv},
  primaryClass = {cs.CR},
  url = {https://arxiv.org/abs/2610.07510}
}

Figure

Open original