Evaluating Training Data Construction Strategies for Token-Level Language Identification
| Authors | |
|---|---|
| Year of publication | 2025 |
| Type | Paper in proceedings |
| Conference | Recent Advances in Slavonic Natural Language Processing, RASLAN 2025 |
| MU Faculty or unit | |
| Citation | |
| web | Proceedings of the Nineteenth Workshop on Recent Advances in Slavonic Natural Languages Processing, RASLAN 2025. |
| Keywords | Languageidentification; Code-switching; Data augmentation |
| Description | This paper contributes to research on developing token-level language identification tool. The tool is designed to recognize Czech and languages frequently spoken in Czechia, such as Slovak and Ukrainian, while also covering additional languages. In this study, multiple datasets are created using three distinct strategies. The datasets are further used to fine-tune a pre-trained language model, and the resulting models are evaluated on datasets containing code-switching. |
| Related projects: |