Romance-MoLE: Rich Cousins and Their Benefits. An Approach to Low-Resource Language Modelling
- Goga, Mihaela (Author)
Adapt Llama3.1-Carballo (pretrained on Ibero-Romance languages) to Occitan (Languedocian):
- extract and TIES-merge French and Catalan LoRA adapters
- train an Occitan expert
- train a Hierarchical Mixture of LoRA Experts (“Romance-MoLE”) (dynamic switch between French/Catalan/Occitan)
Largely inspired by (Kunz et al, 2025), who did something similar for Faroese (using Icelandic and Danish as cousins)
Methodology
Tokeniser adaptation: Trans-Tokenisation
Map new Occitan tokens to the average of the most similar French and Catalan counterparts.
Selected words to inject: elisions, diactrics, frequently fragmented
Data
Gold dataset
HPLT-V3 (2025, mix of Common Crawl and Internet Archive)
+ cleaning pipeline
~75k samples (documents or sentences?) after cleaning, i.e. ~48M tokens
Synthetic data
- Use Gemini
- Prompt includes rules to enforce Alibert orthography and Languedocien dialect
-
Generate multiple datasets
- “Morphological Stress” dataset --> tokenisation edge-cases
- “Spoken Dialogue” dataset --> clitic pronouns, complex stacked combinations
- “FLORES”
-
“Alpaca” --> 2k instruction-response pairs
- extract from Catalan
- cultural adaptation via Gemini (e.g. geography)
-
Validation by an Occitan pedagogical expert (general comments based on a representative subset of all generated samples)
- generally good, yet occasionally uses calques (grammar e.g. passé composé vs prétérit, or lexicon e.g. téner vs aver)
- errors are more frequent in out-of-distribution contexts
Occitan fine-tuning
- System prompt to enforce output in Occitan lengadocian
-
Curriculum : HPLT, then synthetic
- use NEFtune to avoid overfitting
MoE
-
train with Alpaca (50% French, 25% Catalan, 25% Occitan)
- tag each instruction with the language code
Evaluation
Test sets
FLORES-200