DIALOG-22 RuATD Generated Text Detection

Authors: N. Maloyan, B. Nutfullin, E. Ilyushin
Published: Computational Linguistics and Intellectual Technologies (Dialogue 2022), issue 21, pp. 394-401
Generated Text Detection Russian NLP AI Detection


Short version

RuATD 2022 was a shared task at the Dialogue conference on detecting machine-generated Russian text. We won 1st place in the binary task (human vs. machine) with 0.82995 accuracy on the private test set, using a stacked ensemble of five fine-tuned transformers. We placed 4th in the multiclass task (which of 13 models wrote the text, or a human) with 0.62856 accuracy, using a single model.

The two tasks

The organizers built the machine-generated part with models fine-tuned for machine translation, paraphrasing, summarization, simplification, and free text generation. Human texts came from open sources in several domains. The example texts in the paper are single short sentences.

Our pipeline

We did no text preprocessing. All models came from the Hugging Face transformers library.

Binary task: stacking

  1. We merged the train and validation sets and split them into 5 folds.
  2. We fine-tuned five models with 5-fold cross-validation: sberbank-ai/sbert_large_nlu_ru, sberbank-ai/ruBert-large, IlyaGusev/mbart_ru_sum_gazeta, MoritzLaurer/mDeBERTa-v3-base-mnli-xnli, and DeepPavlov/xlm-roberta-large-en-ru-mnli.
  3. Each model produced out-of-fold predictions for the training data. For the test set, we averaged the predictions of its 5 fold models.
  4. A logistic regression meta-model, trained on the out-of-fold predictions, produced the final label.

Multiclass task: single models

We fine-tuned DeepPavlov/rubert-base-cased, DeepPavlov/xlm-roberta-large-en-ru-mnli, and IlyaGusev/mbart_ru_sum_gazeta on the train set only, without cross-validation, and submitted the best single model.

Results

Binary task

ModelAccuracy
sberbank-ai/sbert_large_nlu_ru0.79986 ± 0.003
sberbank-ai/ruBert-large0.80154 ± 0.002
IlyaGusev/mbart_ru_sum_gazeta0.80566 ± 0.001
MoritzLaurer/mDeBERTa-v3-base-mnli-xnli0.80710 ± 0.001
DeepPavlov/xlm-roberta-large-en-ru-mnli0.81708 ± 0.002
Ensemble0.82995

The single models sit within 2 points of each other. The stacked ensemble scored about 1.3 points above the best single model, and that was enough for first place.

Multiclass task

ModelValidationKaggle public / private
IlyaGusev/mbart_ru_sum_gazeta0.61420.61459 / 0.61092
DeepPavlov/rubert-base-cased0.60450.60433 / 0.60472
DeepPavlov/xlm-roberta-large-en-ru-mnli0.62420.62856 / 0.62644

On both tasks, the best single model was the XLM-RoBERTa model fine-tuned on English-Russian MNLI. Telling 13 generators apart is much harder than spotting machine text at all: about 62% vs. 83%.

Limitations and what still applies

What still holds: fine-tuned multilingual encoders were strong baselines for Russian detection, and a simple stacking layer on top of diverse models gave a real gain.


Cite as

@inproceedings{maloyan2022dialog,
  title={DIALOG-22 RuATD Generated Text Detection},
  author={Maloyan, Narek and Nutfullin, Bulat and Ilyushin, Eugene},
  booktitle={Computational Linguistics and Intellectual Technologies},
  number={21},
  pages={394--401},
  year={2022},
  doi={10.28995/2075-7182-2022-21-394-401}
}


Narek Maloyan holds a PhD in Computer Science from Lomonosov Moscow State University and works as an AI Research Engineer at Zencoder. His research focuses on AI safety, LLM security, and adversarial machine learning. Learn more