Votre recherche

Dans les auteurs ou contributeurs
  • Automatic language identification is frequently framed as a multi-class classification problem. However, when creating digital corpora for less commonly written languages, it may be more appropriate to consider it a data mining problem. For these varieties, one knows ahead of time that the vast majority of documents are of little interest. By minimizing resources spent on classifying such documents, we can create corpora covering previously overlooked languages faster than existing pipelines. To demonstrate the effectiveness of the targeted mining perspective, we introduce a new pipeline that can filter a single snapshot in two hours. We also provide web corpora for several French-based Creoles.

  • For decades, the pidgin-creole life cycle hypothesis was considered the standard account for the development of both pidgins and creoles. Key aspects of this account include the belief that pidgins emerge from communicative necessity in multilingual environments, while creoles are the outcomes of the stabilization of pidgins, especially via nativization. This account has been critiqued from various directions, with French-based Creoles often serving as the crux of the debate. For these languages, a major point of contention has been the similarity between the languages of the Americas and the Indian Ocean despite the apparent absence of a stable precursor pidgin. For example, McWhorter (2000) defends the life cycle model by hypothesizing that French- based Creoles ultimately descend from an unattested 17th century slave castle pidgin. In contrast, Mufwene (2020) maintains that pidgins mostly formed after the development of creoles during the during a colonial transitional period in Africa, Asia, and the Pacific. Despite these canonical debates regarding French-based Creoles, French-based pidgins remain underexplored. While Skirgård (2013) has compiled and analyzed a corpus of Français Tirailleur, a variety closely associated with West African soldiers, there is still debate about whether it represents a “real” contact variety from the late 19th and early 20th centuries Parkvall (2018) or a stereotypical projection of otherness (Vigouroux (2017). In contrast, there is more consensus that Tây Bồi, or Vietnamese Pidgin French, developed in French Indochina during roughly the same time period and fell out of use when the French military presence ended. Although Français Tirailleur and Tây Bồi share many features, they are often assumed to be mostly independent developments, and to our knowledge, no extensive comparison has been conducted. In this work, we first review descriptions of Tây Bồi and Français Tirailleur to highlight their close structural affinity, Then, we introduce textual data showing that they were largely consolidated by 1890. Finally, we revisit some contemporary accounts of language education in French colonies. Taken together, these documents suggest that, far from being spontaneous solutions to communicative barriers, Français Tirailleur and Tây Bồi were heavily shaped by top- down instruction of ostensibly simplified French. Rather than disproving the existence of pidginized varieties, the codevelopment of regionalized Frenches in West Africa and Southeast Asia underscores the contradictions and instability of thecolonial-era two-track approach to language education.

  • Automatic language identification is frequently framed as a multi-class classification problem. However, when creating digital corpora for less commonly written languages, it may be more appropriate to consider it a data mining problem. For these varieties, one knows ahead of time that the vast majority of documents are of little interest. By minimizing resources spent on classifying such documents, we can create corpora much faster and with better coverage than using established pipelines. To demonstrate the effectiveness of the language mining perspective, we introduce a new pipeline and corpora for several French-based Creoles.

Dernière mise à jour depuis la base de données : 25/08/2026 09:59 (UTC)

Explorer

Langue

Type de papier