Evaluating Training Data Construction Strategies for Token-Level Language Identification

Warning

This publication doesn't include Institute of Computer Science. It includes Faculty of Informatics. Official publication website can be found on muni.cz.
Authors

BEDNAŘÍKOVÁ Emma RYCHLÝ Pavel

Year of publication 2025
Type Paper in proceedings
Conference Recent Advances in Slavonic Natural Language Processing, RASLAN 2025
MU Faculty or unit

Faculty of Informatics

Citation
web Proceedings of the Nineteenth Workshop on Recent Advances in Slavonic Natural Languages Processing, RASLAN 2025.
Keywords Languageidentification; Code-switching; Data augmentation
Description This paper contributes to research on developing token-level language identification tool. The tool is designed to recognize Czech and languages frequently spoken in Czechia, such as Slovak and Ukrainian, while also covering additional languages. In this study, multiple datasets are created using three distinct strategies. The datasets are further used to fine-tune a pre-trained language model, and the resulting models are evaluated on datasets containing code-switching.
Related projects:

You are running an old browser version. We recommend updating your browser to its latest version.

More info