Votre recherche
Résultats 5 ressources
-
Cet article dresse un état de l'art de la traduction automatique et de son évaluation pour les langues à variation dialectale, et en particulier pour les continuums dialectaux. Pour illustrer cet état de l'art, nous proposons une série d'expériences préliminaires sur le continuum occitan, afin de dresser un état des performances des systèmes existants pour la traduction depuis et vers plusieurs variétés d'occitan. Nos résultats indiquent d'une part des performances globalement satisfaisantes pour la traduction vers le français et l'anglais. D'autre part, des analyses mélangées à des outils d'identification de langues sur les prédictions vers l'occitan mettent en lumière la capacité de la plupart des systèmes évalués à générer des textes dans cette langue (y compris en zero-shot ), mais révèlent aussi des limitations en termes d'évaluation de la diversité dialectale dans les traductions proposées.
-
Occitan is a Romance language spoken mostly in the South of France and characterised by rich dialectal variation, which can pose problems for certain NLP tools. This shortfall is largely attributable to the scarcity of dialect-annotated corpora, in a context where linguistic classification within the Occitan dialect continuum is still debated and major nomenclatures, such as ISO 639, fail to provide granular codes for varieties below the generic “Occitan” label. In this paper, we introduce OcWikiDialects, a new dataset comprising articles from the Occitan Wikipedia. The corpus features rich metadata, including dialect labels, and is segmented at both paragraph and sentence levels. Combined with previously released datasets, we explore approaches for Occitan dialect identification by training three types of model on up to 8 labels: linear SVM classifiers based on word and character n-grams, FastText classifiers based on pretrained vectors, and BERT-based neural classifiers adapted through fine-tuning. Evaluations across in- and out-of-domain test sets demonstrate the substantial impact of our new dataset for the task. However, a peak macro-averaged F1 score of 58.15 underscores persistent challenges for underrepresented Occitan varieties, supported by our per-dialect analysis. Code, dataset and models are available: https://github.com/DEFI-COLaF/OcWikiDialects.
-
Description du processus de développement et d’anonymisation d’un corpus en occitan issu d’un forum en ligne et riche en métadonnées liées à la variation dialectale.
-
We introduce ForumOccitania, a new Occitan corpus of posts from an online forum, covering a range of topics and dialects. While some existing datasets for this low-resource language include labels of varieties within the dialect continuum, we go one step further by providing metadata pertaining to sociolinguistic factors of language variation (dialect, geographical location, age, proficiency), extracted from self-declared user profiles. We carry out statistical and qualitative analyses, as well as preliminary experiments on unsupervised dialect identification. Our results show that (i) most of the contents is written in Occitan, with the classical spelling conventions, and by young new speakers, (ii) posts display a strong presence of dialectal features from four major Occitan varieties (Lemosin, Lengadocian, Gascon, Provençau), and (iii) a simple topic modelling approach introduced by Kuparinen and Scherrer (2024) effectively detects salient features of these four varieties, but also reveals finer-grained diatopical variation tendencies.