DN at SemEval-2023 Task 12: Low-Resource Language Text Classification

Authors: D. Homskiy, N. Maloyan
Published: SemEval-2023 (ACL), pp. 1537-1541, Toronto, July 2023. Preprint arXiv:2305.02607, May 2023
Low-Resource NLP Multilingual Sentiment Analysis


Summary

This is our system for AfriSenti-SemEval 2023 Task 12: three-class sentiment (positive, neutral, negative) on tweets in 14 African languages. The task had 12 monolingual tracks, one multilingual track and two zero-shot tracks. Our best result was 3rd place on the multilingual track (Subtask B, Track 16). Several monolingual tracks were much weaker.

Method

We chose the model on Hausa. Weighted F1 was 0.70 for XLM-R-large and 0.82 for AfroXLMR-large (base 0.79, small 0.77, mini 0.71). Text cleanup (links, @user tags, repeated quotes and dots, spaces around emoji) gave 0.82 against 0.81 without it, so the final system does not use it.

Results

Track Language Our F1 Best F1 Place
1 Hausa 81.09 82.62 6
2 Yoruba 72.07 80.16 18
3 Igbo 74.51 82.96 25
4 Nigerian Pidgin 64.89 75.96 24
5 Amharic 57.34 78.42 17
6 Algerian Arabic 65.81 74.20 17
7 Moroccan Arabic/Darija 57.20 64.83 11
8 Swahili 62.51 65.68 10
9 Kinyarwanda 71.91 72.63 4
10 Twi 55.53 68.28 29
11 Mozambican Portuguese 69.09 74.98 14
12 Xitsonga 46.62 60.67 27
16 Multilingual 72.55 75.06 3
17 Zero-shot Tigrinya 68.93 70.86 8
18 Zero-shot Oromo 41.45 46.23 15

Cite as

@inproceedings{homskiy-maloyan-2023-dn,
  title={{DN} at {S}em{E}val-2023 Task 12: Low-Resource Language Text Classification via Multilingual Pretrained Language Model Fine-tuning},
  author={Homskiy, Daniil and Maloyan, Narek},
  booktitle={Proceedings of the 17th International Workshop on Semantic Evaluation (SemEval-2023)},
  pages={1537--1541},
  address={Toronto, Canada},
  publisher={Association for Computational Linguistics},
  year={2023},
  doi={10.18653/v1/2023.semeval-1.212}
}


Narek Maloyan holds a PhD in Computer Science from Lomonosov Moscow State University and works as an AI Research Engineer at Zencoder. His research focuses on AI safety, LLM security, and adversarial machine learning. Learn more