This repository includes code for register labeling (multi-label classification) using huggingface transformers and datasets and pytorch as well as code for making files for translation and all the used data: originals, downsampled and translated.
The goal is to compare cross-lingual transfer using XLM-Roberta (from en, fi, swe, fre to small languages pt, spa, jp, zh) and translation using DeepL to the small languages using XLM-Roberta and monolingual models.
It is possible to use either only upper labels (8) which tells the other upper category to which the text belongs to or full simplified labels (24) which use sublabels to further identify the genre/register of the text.
| Source | Target |
|---|---|
| en, fi, swe, fre | pt |
| en, fi, swe, fre | spa |
| en, fi, swe, fre | jp |
| en, fi, swe, fre | zh |
| en, swe, fre | fi |
| en, fi, fre | swe |
| translated pt | pt |
| translated spa | spa |
| translated jp | jp |
| translated zh | zh |
| translated fi | fi |
| translated swe | swe |
Results can be found here: https://docs.google.com/spreadsheets/d/18PIjD6Zvd-OPBe97tV2bbysrUFTlTBIeseL5uY3Jpbk/edit?usp=sharing
TLDR: cross-lingual transfer and translation work equally well with XLMR, japanese was the exception and got better results when using upper labels.
register-multilabel.py can be used for example as follows:
python3 register-multilabel.py --train_set [FILE/FILES] --test_set [FILE/FILES] [--full] --batch [NUM] --treshold [NUM] --epochs [NUM] --learning [NUM] --checkpoint [PATH] --model [MODEL]
There are a few other arguments that can be used but only --train_set and --test_set are required. To get more information about the possible arguments, try using the -h flag:
python3 register-multilabel.py -h