Votre recherche

Dans les auteurs ou contributeurs
  • Automatic Text Recognition (ATR) has become a key component of digital editorial workflows, enabling the conversion of scanned documents into TEI. Beyond text recognition, ATR pipelines rely on Document Layout Analysis (DLA) to segment pages into labeled zones that support document reconstruction. Controlled vocabularies such as SegmOnto have already been widely used for this purpose, notably in projects like Gallicorpora and SETA. This paper presents LADaS (Layout Analysis Dataset with SegmOnto), an extension of SegmOnto designed for deeper alignment with the TEI Guidelines. LADaS introduces a two-level annotation system combining SegmOnto broad layout zones with visually identifiable subzones mapped on the TEI. The paper describes the process of building this controlled vocabulary and its documentation, as well as the associated annotated dataset and a proposed pipeline to convert historical scanned documents into deeper encoded TEI files.

  • We present a novel, open-access dataset designed for semantic layout analysis, built to support document recreation workflows through mapping with the Text Encoding Initiative (TEI) standard. This dataset includes 7,254 annotated pages spanning a large temporal range (1600-2024) of digitised and born-digital materials across diverse document types (magazines, papers from sciences and humanities, PhD theses, monographs, plays, administrative reports, etc.) sorted into modular subsets. By incorporating content from different periods and genres, it addresses varying layout complexities and historical changes in document structure. The modular design allows domain-specific configurations. We evaluate object detection models on this dataset, examining the impact of input size and subset-based training. Results show that a 1280-pixel input size for YOLO is optimal and that training on subsets generally benefits from incorporating them into a generic model rather than fine-tuning pre-trained weights.

  • We introduce an open-access and openly-licensed diachronic dataset for Document Layout Analysis (DLA) targeting heterogeneous historical and contemporary materials 4 . The dataset comprises 8,287 manually annotated pages spanning 1600-2024, covering both digitised and born-digital documents across diverse genres including magazines, academic papers, monographs, plays, and administrative records. Annotations follow a two-level vocabulary derived from SegmOnto, extended to paragraph-level semantic regions consistent with the Text Encoding Initiative (TEI), and provide 45 fine-grained classes over 115,007 instances. The data is organised in modular subsets, enabling domain-specific configurations. We establish baselines using two families of object detectors (YOLOv11 and D-FINE) and study three regimes: in-domain training, label-set simplification, and large-scale pretraining (Objects365). We additionally conduct a cross-dataset transfer study against DocLayNet and M 6 Doc using a unified label mapping and a per-decade temporal evaluation. Results show that models trained on existing modern-centric DLA datasets degrade sharply on pre-20th century material (mean F1 dropping by up to 0.25), while in-domain training on our dataset preserves performance across centuries. This highlights the need for diachronic benchmarks and demonstrates the value of combining CV-based detection with DH-oriented annotation standards.

Dernière mise à jour depuis la base de données : 25/08/2026 09:59 (UTC)

Explorer

Type de papier