KréyoLID: From Language Identification Towards Language Mining
Type de ressource
Manuscript
Auteurs/contributeurs
- Dent, Rasul (Author)
- Ortiz Suarez, Pedro (Author)
- Clérice, Thibault (Author)
- Sagot, Benoît (Author)
Title
KréyoLID: From Language Identification Towards Language Mining
Abstract
Automatic language identification is frequently framed as a multi-class classification problem. However, when creating digital corpora for less commonly written languages, it may be more appropriate to consider it a data mining problem. For these varieties, one knows ahead of time that the vast majority of documents are of little interest. By minimizing resources spent on classifying such documents, we can create corpora much faster and with better coverage than using established pipelines. To demonstrate the effectiveness of the language mining perspective, we introduce a new pipeline and corpora for several French-based Creoles.
Date
2025-03
Accessed
25/08/2026 09:04
Short Title
KréyoLID
Library Catalog
HAL
Notes
8 main pages
Référence
Dent, R., Ortiz Suarez, P., Clérice, T., & Sagot, B. (2025). KréyoLID: From Language Identification Towards Language Mining. https://inria.hal.science/hal-04986402
Langue
Type de papier
Lien vers cette notice