Trojan Detection in Large Language Models: Insights from TDC 2023
Short version
The NeurIPS 2023 Trojan Detection Challenge (TDC 2023) gave teams an LLM with 1,000 hidden backdoors and the list of their outputs. The job was to recover the hidden inputs (triggers). Finding some input that forces each output was easy: the winning team reached a REASR of 0.987. Finding the actual trigger the attacker planted was not: the best recall was about 0.16. That is no better than a simple baseline that samples random sentences like the training triggers.
What a trojan is here
A trojan (or backdoor) is a hidden (trigger, target) pair trained into the model. When the input contains the trigger, the model outputs the target. The trigger can be gibberish or an innocent-looking sentence. The target can be harmful, for example a shell command such as echo "kernel.panic = 1" >> /etc/sysctl.conf.
The challenge
- Models: Pythia. The large-model subtrack used 6.9B parameters and the base subtrack 1.4B. Our experiments used the 1.4B model.
- Trojans: 1,000 per model, as 100 target strings with 10 triggers each.
- Given: all 100 targets, plus the triggers for 20 of them as a training set.
- Task: predict the triggers for the other 80 targets.
- Rules: no editing the weights, and at most 2 A100 GPU-days of compute.
Two metrics
- Recall: how close the predicted triggers are to the real ones, measured with BLEU (one-sided Chamfer distance).
- REASR (reverse-engineered attack success rate): whether the predicted triggers actually make the model output the target, also measured with BLEU.
The ranking used the average of the two.
Methods compared
All of these are white-box methods that use gradients to search for input tokens:
- UAT (Universal Adversarial Triggers): swaps trigger tokens using a first-order approximation of the loss, with beam search.
- GBDA: optimizes a distribution over tokens with the Gumbel-softmax trick, with soft fluency constraints.
- PEZ (Hard Prompts made EaZy): optimizes continuous embeddings and projects them back to real tokens at each step.
- GCG (Greedy Coordinate Gradient): uses gradients to propose token swaps at all positions, then tests a batch and keeps the best.
- ARCA: an auditing method that jointly optimizes the prompt and the output with coordinate ascent.
Results
| Method | Recall | REASR |
|---|---|---|
| PEZ (baseline) | 0.105 | 0.052 |
| GBDA (baseline) | 0.116 | 0.056 |
| UAT (baseline) | 0.131 | 0.030 |
| GCG | 0.109 | 0.068 |
| ARCA | 0.077 | 0.358 |
| GCG (winning team) | 0.167 | 0.987 |
During the competition, most teams got REASR close to 100%, even with simple black-box evolutionary search. Recall was the hard part. A baseline that samples sentences from a distribution like the training triggers gets 14% to 17% recall, just from accidental n-gram overlap in BLEU. The top recall of about 0.16 is in that range.
Why recall is so hard: unintended triggers
The model does not have one input that produces each target. It has many. Fine-tuning plants the intended trigger, but optimization finds other inputs that force the same output just as well, or better. We tried several objective functions to separate the intended triggers from these unintended ones. None of them worked. On the development-phase models, the intended triggers were not even local optima: the search easily found inputs that scored higher. The organizers improved this for the test phase, and the intended triggers there were more often local optima, but still not consistently.
Tricks from the winning team
- Initialize from other trojans. Starting the search for one target from the trigger of a different, unrelated trojan made convergence much faster. The team kept pools of known triggers and grew them as the search found new ones. This hints at some shared structure between inserted trojans.
- Filter fragile triggers. The search ran in FP16, so some found triggers failed in batch evaluation. A filtering pass removed them.
- Remove near-duplicates. When two triggers for the same target were within a Levenshtein distance threshold, one was dropped before picking the 20 to submit.
The same team proposed two ways to measure how "tight" a trojan insertion is. First, start the search at the intended trigger and see how much local moves improve the objective. Second, search from random starting points and measure how often, or how fast, the search succeeds.
What this means
- Forcing an output is not the same as finding the backdoor. A detector can report a trigger that works without finding the one the attacker planted.
- Real conditions are harder. In the challenge, defenders had the exact list of targets, some known triggers, and white-box access. A real defender facing a careful attacker may have none of these.
- Some trojans may be impossible to find. Prior work shows backdoors that are provably undetectable under cryptographic assumptions, so far only for toy models. If that extends to transformers, trigger search cannot be a complete defense.
- The tools are still useful. The work improved prompt optimization methods. One follow-up uses a small draft model to filter GCG candidates for a 5.6x speedup.
For anyone using third-party model weights, the practical takeaway is that you cannot rely on trigger search to prove a model is clean. Where the weights came from, and how they were trained, matter more.
Limitations
- We ran experiments only on the 1.4B base-subtrack model.
- The organizers may have made the challenge easier than a real attack on purpose, so real trojans may be even harder to recover.
- The competition did not produce a method that recovers intended triggers reliably.
Cite as
@misc{maloyan2024trojan,
title={Trojan Detection in Large Language Models: Insights from the Trojan Detection Challenge},
author={Maloyan, Narek and Verma, Ekansh and Nutfullin, Bulat and Ashinov, Bislan},
year={2024},
eprint={2404.13660},
archivePrefix={arXiv},
primaryClass={cs.CL}
}