Tag your Middle French text
You can find more information on this specific model here: https://zenodo.org/records/20522515
You can find more information on this specific model here: https://zenodo.org/records/20522515
Please remember that corpus creation and software engineering is valid research, so please cite these resources when you use this lemmatizer for your research: this includes the wonderful original research by E. Manjavacas, M. Kestemont and Á. Kádár as well as the software wrapping built to handle pre- and post-processing.
For each models, a bibliography and potentially other citable works are given, such as models and datasets are given.
@software{thibault_clerice_2020_3883590,
author = {Clérice, Thibault},
title = {Pie Extended, an extension for Pie with pre-processing and post-processing},
month = jun,
year = 2020,
publisher = {Zenodo},
doi = {10.5281/zenodo.3883589},
url = {https://doi.org/10.5281/zenodo.3883589}
}
@inproceedings{manjavacas-etal-2019-improving,
title = "Improving Lemmatization of Non-Standard Languages with Joint Learning",
author = "Manjavacas, Enrique and
K{\'a}d{\'a}r, {\'A}kos and
Kestemont, Mike",
booktitle = "Proceedings of the 2019 Conference of the North {A}merican Chapter of the Association for Computational
Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)",
month = jun,
year = "2019",
address = "Minneapolis, Minnesota",
publisher = "Association for Computational Linguistics",
url = "https://www.aclweb.org/anthology/N19-1153",
doi = "10.18653/v1/N19-1153",
pages = "1493--1503",}
@dataset{dugaz_2026_20522515,
author = {Dugaz, Lucien and Ing, Lucence and Vidal-Gorène, Chahan and Duval, Frédéric},
title = {Pie Model for Lemmatization, POS Tagging, and Morphological Analysis of Middle French},
year = 2026,
publisher = {Zenodo},
version = {v1},
doi = {10.5281/zenodo.20522515},
url = {https://doi.org/10.5281/zenodo.20522515}
}
This model provides support for the lemmatization, part-of-speech tagging and morphological analysis (case, degree, gender, mood, number, person, tense) of Middle French texts. The training corpus spans 1300–1529 and contains 115,554 wordforms from several sources and genres. The models require pre-tokenized input and were developed through the Centre de ressources computationnelles pour les langues à variation graphique at the École nationale des chartes, with support from the ANR (ANR-21-ESRE-0005) via the Biblissima+ programme.
Model and data are published on Zenodo, under a Creative Commons Attribution 4.0 International (CC-BY-4.0) licence: https://zenodo.org/records/20522515 (DOI: 10.5281/zenodo.20522515).
The training dataset is available at https://github.com/chartes/mf-corpus.
| Task | Accuracy |
|---|---|
| POS tagging | 0.9743 |
| Number, tense, person, mode | 0.9873–0.9957 |
| Lemmatization | 0.9167 |
This lemmatizer is provided to you thanks to the data of the LASLA, the software of Emmanuel Manjavacas and Mike Kestemont and some engineering from the École nationale des chartes. If you want to cite them :