Votre recherche

Dans les auteurs ou contributeurs
Type de papier
  • We introduce an open-access and openly-licensed diachronic dataset for Document Layout Analysis (DLA) targeting heterogeneous historical and contemporary materials 4 . The dataset comprises 8,287 manually annotated pages spanning 1600-2024, covering both digitised and born-digital documents across diverse genres including magazines, academic papers, monographs, plays, and administrative records. Annotations follow a two-level vocabulary derived from SegmOnto, extended to paragraph-level semantic regions consistent with the Text Encoding Initiative (TEI), and provide 45 fine-grained classes over 115,007 instances. The data is organised in modular subsets, enabling domain-specific configurations. We establish baselines using two families of object detectors (YOLOv11 and D-FINE) and study three regimes: in-domain training, label-set simplification, and large-scale pretraining (Objects365). We additionally conduct a cross-dataset transfer study against DocLayNet and M 6 Doc using a unified label mapping and a per-decade temporal evaluation. Results show that models trained on existing modern-centric DLA datasets degrade sharply on pre-20th century material (mean F1 dropping by up to 0.25), while in-domain training on our dataset preserves performance across centuries. This highlights the need for diachronic benchmarks and demonstrates the value of combining CV-based detection with DH-oriented annotation standards.

Dernière mise à jour depuis la base de données : 25/08/2026 09:59 (UTC)

Explorer

Type de papier