Romance-MoLE: Rich Cousins and Their Benefits. An Approach to Low-Resource Language Modelling

Type de ressource
Thesis
Auteur/contributeur
Title
Romance-MoLE: Rich Cousins and Their Benefits. An Approach to Low-Resource Language Modelling
Type
Master thesis
University
University of Gothenburg
Place
Gothenburg, Sweden
Date
2026
# of Pages
69
Citation Key
gogaRomanceMoLERichCousins2026
Short Title
Romance-MoLE
Language
en
Library Catalog
Zotero
Notes

Adapt Llama3.1-Carballo (pretrained on Ibero-Romance languages) to Occitan (Languedocian):

  1. extract and TIES-merge French and Catalan LoRA adapters
  2. train an Occitan expert
  3. train a Hierarchical Mixture of LoRA Experts (“Romance-MoLE”) (dynamic switch between French/Catalan/Occitan)

Largely inspired by (Kunz et al, 2025), who did something similar for Faroese (using Icelandic and Danish as cousins)

Methodology

Tokeniser adaptation: Trans-Tokenisation

Map new Occitan tokens to the average of the most similar French and Catalan counterparts.

Selected words to inject: elisions, diactrics, frequently fragmented

Data

Gold dataset

HPLT-V3 (2025, mix of Common Crawl and Internet Archive)

+ cleaning pipeline

~75k samples (documents or sentences?) after cleaning, i.e. ~48M tokens

Synthetic data

  • Use Gemini
  • Prompt includes rules to enforce Alibert orthography and Languedocien dialect
  • Generate multiple datasets

    • “Morphological Stress” dataset --> tokenisation edge-cases
    • “Spoken Dialogue” dataset --> clitic pronouns, complex stacked combinations
    • “FLORES”
    • “Alpaca” --> 2k instruction-response pairs

      • extract from Catalan
      • cultural adaptation via Gemini (e.g. geography)
  • Validation by an Occitan pedagogical expert (general comments based on a representative subset of all generated samples)

    • generally good, yet occasionally uses calques (grammar e.g. passé composé vs prétérit, or lexicon e.g. téner vs aver)
    • errors are more frequent in out-of-distribution contexts

Occitan fine-tuning

  • System prompt to enforce output in Occitan lengadocian
  • Curriculum : HPLT, then synthetic

    • use NEFtune to avoid overfitting

MoE

  • train with Alpaca (50% French, 25% Catalan, 25% Occitan)

    • tag each instruction with the language code

Evaluation

Test sets

FLORES-200

Référence
Goga, M. (2026). Romance-MoLE: Rich Cousins and Their Benefits. An Approach to Low-Resource Language Modelling [Master thesis, University of Gothenburg]. https://gupea.ub.gu.se/server/api/core/bitstreams/8ba73144-bdb9-4ee5-9f3d-dce0a52c9c6d/content
Langue