Tag your Middle French text

You can find more information on this specific model here: https://zenodo.org/records/20522515

Cite with the following

Please remember that corpus creation and software engineering is valid research, so please cite these resources when you use this lemmatizer for your research: this includes the wonderful original research by E. Manjavacas, M. Kestemont and Á. Kádár as well as the software wrapping built to handle pre- and post-processing.

For each models, a bibliography and potentially other citable works are given, such as models and datasets are given.

@software{thibault_clerice_2020_3883590,
  author       = {Clérice, Thibault},
  title        = {Pie Extended, an extension for Pie with pre-processing and post-processing},
  month        = jun,
  year         = 2020,
  publisher    = {Zenodo},
  doi          = {10.5281/zenodo.3883589},
  url          = {https://doi.org/10.5281/zenodo.3883589}
}
@inproceedings{manjavacas-etal-2019-improving,
    title = "Improving Lemmatization of Non-Standard Languages with Joint Learning",
    author = "Manjavacas, Enrique  and
      K{\'a}d{\'a}r, {\'A}kos  and
      Kestemont, Mike",
    booktitle = "Proceedings of the 2019 Conference of the North {A}merican Chapter of the Association for Computational
      Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)",
    month = jun,
    year = "2019",
    address = "Minneapolis, Minnesota",
    publisher = "Association for Computational Linguistics",
    url = "https://www.aclweb.org/anthology/N19-1153",
    doi = "10.18653/v1/N19-1153",
    pages = "1493--1503",}
@dataset{dugaz_2026_20522515,
  author       = {Dugaz, Lucien and Ing, Lucence and Vidal-Gorène, Chahan and Duval, Frédéric},
  title        = {Pie Model for Lemmatization, POS Tagging, and Morphological Analysis of Middle French},
  year         = 2026,
  publisher    = {Zenodo},
  version      = {v1},
  doi          = {10.5281/zenodo.20522515},
  url          = {https://doi.org/10.5281/zenodo.20522515}
}

Information about the model

This model provides support for the lemmatization, part-of-speech tagging and morphological analysis (case, degree, gender, mood, number, person, tense) of Middle French texts. The training corpus spans 1300–1529 and contains 115,554 wordforms from several sources and genres. The models require pre-tokenized input and were developed through the Centre de ressources computationnelles pour les langues à variation graphique at the École nationale des chartes, with support from the ANR (ANR-21-ESRE-0005) via the Biblissima+ programme.

Model and data are published on Zenodo, under a Creative Commons Attribution 4.0 International (CC-BY-4.0) licence: https://zenodo.org/records/20522515 (DOI: 10.5281/zenodo.20522515).

The training dataset is available at https://github.com/chartes/mf-corpus.

Task Accuracy
POS tagging0.9743
Number, tense, person, mode0.9873–0.9957
Lemmatization0.9167

Bibliography

This lemmatizer is provided to you thanks to the data of the LASLA, the software of Emmanuel Manjavacas and Mike Kestemont and some engineering from the École nationale des chartes. If you want to cite them :

  • E. Manjavacas & Á. Kádár & M. Kestemont, « Improving Lemmatization of Non-Standard Languages with Joint Learning », Proceedings of the 2019 Conference of the North {A}merican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), Special issue on "Natural Language Processing and Ancient Languages", 2019, pp. 493--1503.
  • Enrique Manjavacas & Mike Kestemont. (2019, January 17). emanjavacas/pie v0.1.3 (Version v0.1.3). Zenodo. http://doi.org/10.5281/zenodo.2542537 Check the latest version here :Zenodo DOI