Understanding and EnhancingBackdoor Persistencyin LLM Agent Post-Training
AI agents can inherit backdoors that survive clean training.
What we found
SFT erodes the backdoor.
Clean supervised fine-tuning substantially reduces attack success, but does not reliably remove inherited backdoors.
RL can leave it active.
Subsequent reinforcement learning largely preserves residual behavior and increases attack success in some evaluated settings.
PersistBD improves survival.
Across three model sizes, PersistBD improves persistency after both stages without reducing observed coding performance relative to the Base backdoor.

Software-engineering agents built from 3B, 7B, and 30B models. The attacker supplies the backdoored model; the developer then uses clean supervised fine-tuning (SFT) and task-level reinforcement learning (RL).
Hidden in the weights.
Activated by an input.
A backdoor is a hidden association between a particular input pattern and an attacker-chosen behavior. A developer who adapts a third-party model can inherit that association along with the model’s useful capabilities.
We study this supply-chain threat in software-engineering agents. A comment trigger, # |TEST MODE|, causes the model to emit a command that uploads the project’s .env file to an attacker-controlled endpoint. The developer’s subsequent SFT and RL use clean data, without the trigger or further attacker intervention.

Higher persistency
across model sizes.
We compare an originally inserted backdoor (the Base backdoor) with the same model after applying PersistBD. Both then undergo the developer’s benign SFT and RL. PersistBD retains higher attack success at both endpoints across all three model sizes.
Attack success rate (ASR) measures whether the agent emits the target command on 100 triggered inputs. It does not measure completed exfiltration. The main experiments use task-instance-disjoint partitions for backdoor insertion, attacker-side training, developer SFT, and evaluation.
Attack success after clean post-training
Random-position setting · ASR (%)| Model | After SFT | After SFT + RL | ||
|---|---|---|---|---|
| Base backdoor | PersistBD | Base backdoor | PersistBD | |
| 3BQwen2.5-Coder | 8 | 21 | 7 | 22 |
| 7BQwen2.5-Coder | 20 | 74 | 20 | 76 |
| 30BQwen3-Coder · A3B | 7 | 37 | 8 | 36 |
SFT: 3,000 updates. RL: 150 steps. The random-position setting places the trigger at a randomly selected eligible position in the interaction.
The first-position setting instead places the trigger at the first eligible position. Here, 7B PersistBD ASR rises from 34% before RL to 42% afterward; the 30B Base backdoor rises from 6% to 12%. These increases occur in some settings, rather than uniformly across all branches.
At the evaluated post-SFT and post-RL endpoints, PersistBD matches or exceeds the Base backdoor’s observed resolved rate on SWE-bench Lite. This measures PersistBD’s additional effect relative to an already-backdoored model.
Two factors shape
what survives.
Our first-order analysis identifies two properties associated with backdoor survival through SFT. PersistBD trains a single adapter to improve both, with a third objective term that regularizes benign behavior. The adapter is merged into the model before the developer’s training begins.
Backdoor strength
How strongly the model associates the trigger with the target behavior before clean training. A stronger initial association can withstand more erosion.
Gradient compatibility
Whether updates on benign tasks tend to preserve or erode the backdoor. The attacker estimates this direction using separate benign data.
PersistBD increases both factors
Measured before developer training| Model | Strength S | Compatibility C | ||
|---|---|---|---|---|
| Base backdoor | PersistBD | Base backdoor | PersistBD | |
| 3B | 9.5 | 11.5 | −0.121 | +0.001 |
| 7B | 10.9 | 14.3 | −0.079 | +0.016 |
| 30B | 10.0 | 16.1 | −0.135 | +0.024 |
Strength uses triggered-target examples; compatibility also uses the attacker’s benign data. Compatibility changes from negative to positive at every scale. Under the first-order approximation, positive compatibility predicts that an update on these benign data reinforces the backdoor.
What this means for developers
Audit inherited behavior
alongside task performance.
Successful training on clean tasks does not establish that an inherited model is safe. Our results motivate backdoor auditing when adapting third-party models into agents.
Reference
Cite this work
@misc{zhan2026persistbd,
title = {Understanding and Enhancing Backdoor Persistency
in LLM Agent Post-Training},
author = {Zhan, Qiusi and Lyu, Nian and Ding, Stephanie and
Mehta, Arnav and Davies, Xander and Kang, Daniel},
year = {2026},
eprint = {2610.07510},
archivePrefix = {arXiv},
primaryClass = {cs.CR},
url = {https://arxiv.org/abs/2610.07510}
}