Identifying Rare Languages in Common Crawl Data is a Needles-in-a-Haystack Problem
Type de ressource
Conference Paper
Auteurs/contributeurs
- Dent, Rasul (Author)
- Ortiz Suarez, Pedro (Author)
- Clérice, Thibault (Author)
- Sagot, Benoît (Author)
Title
Identifying Rare Languages in Common Crawl Data is a Needles-in-a-Haystack Problem
Abstract
Automatic language identification is frequently framed as a multi-class classification problem. However, when creating digital corpora for less commonly written languages, it may be more appropriate to consider it a data mining problem. For these varieties, one knows ahead of time that the vast majority of documents are of little interest. By minimizing resources spent on classifying such documents, we can create corpora covering previously overlooked languages faster than existing pipelines. To demonstrate the effectiveness of the targeted mining perspective, we introduce a new pipeline that can filter a single snapshot in two hours. We also provide web corpora for several French-based Creoles.
Proceedings Title
Findings of the Association for Computational Linguistics: EMNLP 2025
Publisher
Association for Computational Linguistics
Place
Suzhou, China
Date
2025-11
Pages
1460-1473
Accessed
25/08/2026 09:03
Library Catalog
HAL
Référence
Dent, R., Ortiz Suarez, P., Clérice, T., & Sagot, B. (2025). Identifying Rare Languages in Common Crawl Data is a Needles-in-a-Haystack Problem. Findings of the Association for Computational Linguistics: EMNLP 2025, 1460–1473. https://doi.org/10.18653/v1/2025.findings-emnlp.77
Langue
Type de papier
Lien vers cette notice