DN at SemEval-2023 Task 12: Low-Resource Language Text Classification
Summary
This is our system for AfriSenti-SemEval 2023 Task 12: three-class sentiment (positive, neutral, negative) on tweets in 14 African languages. The task had 12 monolingual tracks, one multilingual track and two zero-shot tracks. Our best result was 3rd place on the multilingual track (Subtask B, Track 16). Several monolingual tracks were much weaker.
Method
- Model:
Davlan/afro-xlmr-large, an XLM-R-large model adapted by masked language modeling to 17 African languages plus Arabic, French and English, with a classification head. - Ensemble: 5-fold stratified split per track, one model per fold, majority vote at test time.
- Training: Adam, learning rate 2e-5, weight decay 0.1, 5 epochs, max length 128.
- Zero-shot tracks: we took the monolingual model with the best validation score. That was the Amharic model for Tigrinya and the Hausa model for Oromo.
We chose the model on Hausa. Weighted F1 was 0.70 for XLM-R-large and 0.82 for AfroXLMR-large (base 0.79, small 0.77, mini 0.71). Text cleanup (links, @user tags, repeated quotes and dots, spaces around emoji) gave 0.82 against 0.81 without it, so the final system does not use it.
Results
| Track | Language | Our F1 | Best F1 | Place |
|---|---|---|---|---|
| 1 | Hausa | 81.09 | 82.62 | 6 |
| 2 | Yoruba | 72.07 | 80.16 | 18 |
| 3 | Igbo | 74.51 | 82.96 | 25 |
| 4 | Nigerian Pidgin | 64.89 | 75.96 | 24 |
| 5 | Amharic | 57.34 | 78.42 | 17 |
| 6 | Algerian Arabic | 65.81 | 74.20 | 17 |
| 7 | Moroccan Arabic/Darija | 57.20 | 64.83 | 11 |
| 8 | Swahili | 62.51 | 65.68 | 10 |
| 9 | Kinyarwanda | 71.91 | 72.63 | 4 |
| 10 | Twi | 55.53 | 68.28 | 29 |
| 11 | Mozambican Portuguese | 69.09 | 74.98 | 14 |
| 12 | Xitsonga | 46.62 | 60.67 | 27 |
| 16 | Multilingual | 72.55 | 75.06 | 3 |
| 17 | Zero-shot Tigrinya | 68.93 | 70.86 | 8 |
| 18 | Zero-shot Oromo | 41.45 | 46.23 | 15 |
Cite as
@inproceedings{homskiy-maloyan-2023-dn,
title={{DN} at {S}em{E}val-2023 Task 12: Low-Resource Language Text Classification via Multilingual Pretrained Language Model Fine-tuning},
author={Homskiy, Daniil and Maloyan, Narek},
booktitle={Proceedings of the 17th International Workshop on Semantic Evaluation (SemEval-2023)},
pages={1537--1541},
address={Toronto, Canada},
publisher={Association for Computational Linguistics},
year={2023},
doi={10.18653/v1/2023.semeval-1.212}
}