DIALOG-22 RuATD Generated Text Detection
Short version
RuATD 2022 was a shared task at the Dialogue conference on detecting machine-generated Russian text. We won 1st place in the binary task (human vs. machine) with 0.82995 accuracy on the private test set, using a stacked ensemble of five fine-tuned transformers. We placed 4th in the multiclass task (which of 13 models wrote the text, or a human) with 0.62856 accuracy, using a single model.
The two tasks
- Binary: label a text as human-written (H) or machine-generated (M).
- Multiclass: label a text with one of 14 classes: Human, or one of 13 generators. The generators were M2M-100, OPUS-MT, M-BART50, M-BART, ruGPT3-Small, ruGPT3-Medium, ruGPT3-Large, ruGPT2-Large, mT5-Small, mT5-Large, ruT5-Base, ruT5-Base-Multitask, and ruT5-Large.
The organizers built the machine-generated part with models fine-tuned for machine translation, paraphrasing, summarization, simplification, and free text generation. Human texts came from open sources in several domains. The example texts in the paper are single short sentences.
Our pipeline
We did no text preprocessing. All models came from the Hugging Face transformers library.
Binary task: stacking
- We merged the train and validation sets and split them into 5 folds.
- We fine-tuned five models with 5-fold cross-validation: sberbank-ai/sbert_large_nlu_ru, sberbank-ai/ruBert-large, IlyaGusev/mbart_ru_sum_gazeta, MoritzLaurer/mDeBERTa-v3-base-mnli-xnli, and DeepPavlov/xlm-roberta-large-en-ru-mnli.
- Each model produced out-of-fold predictions for the training data. For the test set, we averaged the predictions of its 5 fold models.
- A logistic regression meta-model, trained on the out-of-fold predictions, produced the final label.
Multiclass task: single models
We fine-tuned DeepPavlov/rubert-base-cased, DeepPavlov/xlm-roberta-large-en-ru-mnli, and IlyaGusev/mbart_ru_sum_gazeta on the train set only, without cross-validation, and submitted the best single model.
Results
Binary task
| Model | Accuracy |
|---|---|
| sberbank-ai/sbert_large_nlu_ru | 0.79986 ± 0.003 |
| sberbank-ai/ruBert-large | 0.80154 ± 0.002 |
| IlyaGusev/mbart_ru_sum_gazeta | 0.80566 ± 0.001 |
| MoritzLaurer/mDeBERTa-v3-base-mnli-xnli | 0.80710 ± 0.001 |
| DeepPavlov/xlm-roberta-large-en-ru-mnli | 0.81708 ± 0.002 |
| Ensemble | 0.82995 |
The single models sit within 2 points of each other. The stacked ensemble scored about 1.3 points above the best single model, and that was enough for first place.
Multiclass task
| Model | Validation | Kaggle public / private |
|---|---|---|
| IlyaGusev/mbart_ru_sum_gazeta | 0.6142 | 0.61459 / 0.61092 |
| DeepPavlov/rubert-base-cased | 0.6045 | 0.60433 / 0.60472 |
| DeepPavlov/xlm-roberta-large-en-ru-mnli | 0.6242 | 0.62856 / 0.62644 |
On both tasks, the best single model was the XLM-RoBERTa model fine-tuned on English-Russian MNLI. Telling 13 generators apart is much harder than spotting machine text at all: about 62% vs. 83%.
Limitations and what still applies
- Cost. Five large models times five folds is heavy. The paper says plainly that this pipeline is too slow for real-time detection.
- Old generators. The test data came from 2022 models such as ruGPT-3, mT5, and mBART. The paper does not test texts from newer chat models, so do not expect the same accuracy on them.
- Closed set. The detectors learned the specific generators in the training data. A text from an unseen model is a different problem.
What still holds: fine-tuned multilingual encoders were strong baselines for Russian detection, and a simple stacking layer on top of diverse models gave a real gain.
Cite as
@inproceedings{maloyan2022dialog,
title={DIALOG-22 RuATD Generated Text Detection},
author={Maloyan, Narek and Nutfullin, Bulat and Ilyushin, Eugene},
booktitle={Computational Linguistics and Intellectual Technologies},
number={21},
pages={394--401},
year={2022},
doi={10.28995/2075-7182-2022-21-394-401}
}