Prompt-Injection Attacks on LLM-as-a-Judge Architectures

Paper: Investigating the Vulnerability of LLM-as-a-Judge Architectures to Prompt-Injection Attacks
Authors: N. Maloyan, B. Ashinov, D. Namiot
Published: International Journal of Open Information Technologies 13 (9), 2025
LLM-as-Judge AI Safety Prompt Injection


Short version

An LLM judge picks the better of two answers. We appended a short optimized suffix to the worse answer and asked: how often does the judge switch its verdict? With the Greedy Coordinate Gradient (GCG) method, the answer was more than 30% of the time on two 3B open models. Random text of the same length flipped the verdict only 1.2% to 1.5% of the time.

The setup

In a pairwise judge, the model gets a question x and two answers a and b, and outputs [[A]] or [[B]]. The attacker controls only answer b. The attacker appends a suffix δ to it and wants the judge to pick [[B]].

This is the same position as anyone whose output gets graded by an LLM: a model in an arena, a submission to an auto-grader, or a response in an RLHF data pipeline.

Two attack objectives

Both use GCG to find the suffix. GCG starts from a random suffix and, at each step, uses gradients to propose token swaps at every position, tests a subset of them, and keeps the best swap. It needs access to the model's logits, so this is a white-box or strong grey-box attack.

Baselines and controls

Experiment

Results

MethodQwen2.5-3BFalcon3-3B
Random suffix (control)1.2%1.5%
Token shuffle (control)2.8%3.1%
Hard prompt5.1%5.4%
JMA (ours)15.2%16.7%
JudgeDeceiver22.8%24.1%
CUA (ours)31.2%32.4%

What the table says:

What this means in practice

If your judge is an open-weight model, an attacker can download the same weights and run this optimization offline. Assume a working suffix exists. At a 30% flip rate, an attacker can move leaderboard rankings or bias a preference dataset.

The paper does not test defenses. It points to known ones as next steps: perplexity filtering and other baseline defenses (GCG suffixes often look like gibberish), SmoothLLM, certification methods, red teaming, and adversarial training. Our companion paper, Adversarial Attacks on LLM-as-a-Judge Systems, does test defenses and finds that a committee of different models works best.

Limitations

📄 Access the Paper: arXiv:2505.13348 · IJOIT · Google Scholar

Cite as

@article{maloyan2025investigating,
  title={Investigating the Vulnerability of LLM-as-a-Judge Architectures to Prompt-Injection Attacks},
  author={Maloyan, Narek and Ashinov, Bislan and Namiot, Dmitry},
  journal={International Journal of Open Information Technologies},
  volume={13},
  number={9},
  year={2025},
  note={arXiv:2505.13348}
}


Narek Maloyan holds a PhD in Computer Science from Lomonosov Moscow State University and works as an AI Research Engineer at Zencoder. His research focuses on AI safety, LLM security, and adversarial machine learning. Learn more