Trojan Detection in Large Language Models: Insights from TDC 2023

Paper: Trojan Detection in Large Language Models: Insights from The Trojan Detection Challenge
Authors: N. Maloyan, E. Verma, B. Nutfullin, B. Ashinov
Published: arXiv:2404.13660, 2024
Trojan Detection Backdoor Attacks LLM Safety


Short version

The NeurIPS 2023 Trojan Detection Challenge (TDC 2023) gave teams an LLM with 1,000 hidden backdoors and the list of their outputs. The job was to recover the hidden inputs (triggers). Finding some input that forces each output was easy: the winning team reached a REASR of 0.987. Finding the actual trigger the attacker planted was not: the best recall was about 0.16. That is no better than a simple baseline that samples random sentences like the training triggers.

What a trojan is here

A trojan (or backdoor) is a hidden (trigger, target) pair trained into the model. When the input contains the trigger, the model outputs the target. The trigger can be gibberish or an innocent-looking sentence. The target can be harmful, for example a shell command such as echo "kernel.panic = 1" >> /etc/sysctl.conf.

The challenge

Two metrics

The ranking used the average of the two.

Methods compared

All of these are white-box methods that use gradients to search for input tokens:

Results

MethodRecallREASR
PEZ (baseline)0.1050.052
GBDA (baseline)0.1160.056
UAT (baseline)0.1310.030
GCG0.1090.068
ARCA0.0770.358
GCG (winning team)0.1670.987

During the competition, most teams got REASR close to 100%, even with simple black-box evolutionary search. Recall was the hard part. A baseline that samples sentences from a distribution like the training triggers gets 14% to 17% recall, just from accidental n-gram overlap in BLEU. The top recall of about 0.16 is in that range.

Why recall is so hard: unintended triggers

The model does not have one input that produces each target. It has many. Fine-tuning plants the intended trigger, but optimization finds other inputs that force the same output just as well, or better. We tried several objective functions to separate the intended triggers from these unintended ones. None of them worked. On the development-phase models, the intended triggers were not even local optima: the search easily found inputs that scored higher. The organizers improved this for the test phase, and the intended triggers there were more often local optima, but still not consistently.

Tricks from the winning team

The same team proposed two ways to measure how "tight" a trojan insertion is. First, start the search at the intended trigger and see how much local moves improve the objective. Second, search from random starting points and measure how often, or how fast, the search succeeds.

What this means

For anyone using third-party model weights, the practical takeaway is that you cannot rely on trigger search to prove a model is clean. Where the weights came from, and how they were trained, matter more.

Limitations


Cite as

@misc{maloyan2024trojan,
  title={Trojan Detection in Large Language Models: Insights from the Trojan Detection Challenge},
  author={Maloyan, Narek and Verma, Ekansh and Nutfullin, Bulat and Ashinov, Bislan},
  year={2024},
  eprint={2404.13660},
  archivePrefix={arXiv},
  primaryClass={cs.CL}
}


Narek Maloyan holds a PhD in Computer Science from Lomonosov Moscow State University and works as an AI Research Engineer at Zencoder. His research focuses on AI safety, LLM security, and adversarial machine learning. Learn more