Efficient Architectures For Low-Resource Machine Translation

Warning

This publication doesn't include Institute of Computer Science. It includes Faculty of Informatics. Official publication website can be found on muni.cz.
Authors

SIGNORONI Edoardo RYCHLÝ Pavel SIGNORONI Ruggero

Year of publication 2025
Type Paper in proceedings
Conference Proceedings of the First Workshop on Advancing NLP for Low-Resource Languages associated with the International Conference RANLP 2025
MU Faculty or unit

Faculty of Informatics

Citation
web https://acl-bg.org/proceedings/2025/LowResNLP%202025/pdf/2025.lowresnlp-1.6.pdf
Keywords Machine Translation; Low-Resource Languages
Attached files
Description Low-resource Neural Machine Translation is highly sensitive to hyperparameters and needs careful tuning to achieve the best results with small amounts of training data. We focus on exploring the impact of changes in the Transformer architecture on downstream translation quality, and propose a metric to score the computational efficiency of such changes. By experimenting on English-Akkadian, German-Lower Sorbian, English-Italian, and English-Manipuri, we confirm previous finding in low-resource machine translation optimization, and show that smaller and more parameter-efficient models can achieve the same translation quality of larger and unwieldy ones at a fraction of the computational cost. Optimized models have around 95% less parameters, while dropping only up to 14.8% ChrF. We compile a list of optimal ranges for each hyperparameter.
Related projects:

You are running an old browser version. We recommend updating your browser to its latest version.

More info