- JOB
- France
Job Information
- Organisation/Company
- Inria, the French national research institute for the digital sciences
- Research Field
- Computer science
- Researcher Profile
- Recognised Researcher (R2)
- Application Deadline
- Country
- France
- Type of Contract
- Temporary
- Job Status
- Full-time
- Hours Per Week
- 38.5
- Offer Starting Date
- Is the job funded through the EU Research Framework Programme?
- Not funded by a EU programme
- Reference Number
- 2026-10555
- Is the Job related to staff position within a Research Infrastructure?
- No
Offer Description
This position is part of the ANR project ROMAM², led by Thibault Clérice at Inria Paris within the ALMAnaCH project-team (Automatic Language Modelling and Analysis & Computational Humanities, led by Benoît Sagot). ROMAM² treats pre-editorial normalisation (PEN) of graphemic automatic text recognition (ATR) output as a dedicated NLP task. PEN is a traceable process in which every editorial inference (abbreviation expansion, post-correction of recognition errors, spelling normalisation) stays anchored to the manuscript it comes from. The project works on two medieval languages: Latin, which is heavily abbreviated, and Old French, whose spelling varies in linguistically meaningful ways.
A central claim of the project is that editorial normalisation is not neutral. Printed critical editions silently expand abbreviations and regularise spelling. In doing so, they erase variation that is evidence for the history of the language. Dees' quantitative geography of Old French and its successors rest largely on such editions, and Morin has shown how this can distort dialectal conclusions. However, no study has yet measured this distortion on a controlled parallel corpus. This postdoctoral position is designed to produce that study and the gold data it requires.
The postdoctoral researcher will be supervised by Thibault Clérice. They will work closely with:
- the project's PhD candidate in NLP, co-supervised by Thibault Clérice, Benoît Sagot and Rachel Bawden. The PhD candidate will use the gold data and evaluation framework produced by the postdoc to train and evaluate normalisation models.
- David Smith (Northeastern University), a specialist in aligning noisy historical data.
- the ANR JCJC project Phil•IA, coordinated by Ariane Pinche at CIHAM (UMR 5648, ENS de Lyon). A regular collaboration is expected on Old French graphemic transcription, digital editing and TEI encoding.
The position is based at Inria Paris, within ALMAnaCH. The team brings together researchers in NLP, language modelling and computational humanities, and offers a rare environment for philologists and linguists who want to work directly with NLP researcher. The postdoc will also benefit from existing community resources developed by the team: the CATMuS dataset (the largest ATR dataset for medieval manuscripts), the CoMMA corpus (3.3 billion tokens of Latin and Old French from over 32,000 manuscripts) and the upcoming work on Biblissima-Textes.
The position is for 18 months, starting March 2027.
Bibliography
- Clérice, T., Bawden, R., Glaise, A., Pinche, A., & Smith, D. (2026). Pre-Editorial Normalization for Automatically Transcribed Medieval Manuscripts in Old French and Latin. In Proceedings of the Fourth Workshop on Language Technologies for Historical and Ancient Languages (LT4HALA) @ LREC 2026. https://arxiv.org/abs/2602.13905
- Clérice, T., Pinche, A., Vlachou-Efstathiou, M., Chagué, A., Camps, J.-B., et al. (2024). CATMuS Medieval: A multilingual large-scale cross-century dataset in Latin script for handwritten text recognition and beyond. In Proceedings of ICDAR 2024 (LNCS 14806, pp. 174–194). Springer. https://doi.org/10.1007/978-3-031-70543-4_11
- Clérice, T., Gabay, S., Vlachou-Efstathiou, M., Pinche, A., & Sagot, B. (2026). CoMMA, a Large-scale Corpus of Multilingual Medieval Archives. In Proceedings of the Fifteenth Language Resources and Evaluation Conference. ELRA. https://inria.hal.science/hal-05299220
- Dees, A. (1985). Dialectes et scriptae à l'époque de l'ancien français. Revue de Linguistique Romane, 49(193–194), 87–117.
- Morin, Y. C. (2006). Histoire du corpus d'Amsterdam : le traitement des données dialectales. In Le Nouveau Corpus d'Amsterdam. Actes de l'atelier de Lauterbad.
- Scheer, T., & Brun-Trigaud, G. (2022). L'atlas Dees électronique. Concordial, Grenoble. https://hal.science/hal-03912660
- Kuparinen, O., & Scherrer, Y. (2024). Corpus-based dialectometry with topic models. Journal of Linguistic Geography, 12(1), 1–12.
- Pinche, A. (2022). Guide de transcription pour les manuscrits du Xe au XVe siècle. https://hal.archives-ouvertes.fr/hal-03697382
- Duval, F. (2012). Transcrire le français médiéval : de l'« Instruction » de Paul Meyer à la description linguistique contemporaine. Bibliothèque de l'École des chartes, 170(2), 321–342.
The postdoctoral researcher, trained in philology or historical linguistics with skills in digital humanities, will build the philological foundations and evaluation resources of ROMAM² and carry out a controlled study of how editorial practices affect the dialectometry of Old French. The work involves:
- Building a multi-layer gold corpus from the Nouveau Corpus d'Amsterdam (NCA, formerly Dees' corpus). Using Tobias Scheer's Atlas Dees Électronique mapping between NCA editions and their source manuscripts, the postdoc will sample each text (~500 words per document, ~100,000 words in total) and transcribe them in eScriptorium, align the edited text with a graphemic transcription of the manuscript, and annotate tokens as abbreviated or not in XML-TEI.
- Abbreviation-aware dialectometry. The postdoc will quantify abbreviation practices across the corpus and compare the resulting feature maps and dialectal distances with those of Dees and his successors (e.g. Scherrer's Dialektkarten). They will assess how far ambiguous abbreviation resolution, of the kind Morin identified in Floovant, changes the conclusions, using several dialectometric methods.
- Auditing and augmenting the training data. In collaboration with David Smith, the postdoc will audit the automatically aligned corpus from the prototype PEN work. The goal is to separate valid alignments (identity, ATR post-correction, abbreviation expansion) from invalid ones (literary variants, spelling variants), and to explore inter- and intra-manuscript alignment between witnesses. This work also yields a corpus for studying abbreviation practices across textual traditions.
- Contributing to the evaluation framework, jointly with the PhD candidate. This includes lossless conversion between ALTO, plain text and TEI that preserves uncertainty markup (<choice>, <abbr>, <expan>), stage-specific metrics, and a fine-grained error taxonomy (overnormalisation, variant insertion, hallucination, morphosyntactic errors, ambiguity collapse, etc.). This builds on the expertise of Ariane Pinche and the Phil•IA project.
- Supervising annotation work carried out by hourly-paid annotators, with the PI, to ensure the linguistic and editorial quality of the gold data (Old French and Latin if the applicant knows Latin).
The research component lies at the intersection of historical linguistics, material philology and NLP evaluation. The postdoc is expected to publish results in both communities: a paper on the impact of normalisation on dialectal attribution in Old French, open datasets (the TEI NCA sample and the new gold PEN dataset), and a proposed panel at the International Medieval Congress (Leeds).
The main activities of the applicant will include:
- carrying out research on the topic outlined in the job descriptions, including developing new ideas, positioning the work with respect to related work in historical linguistics and NLP, and validating the methodology through corpus construction, experiments and analysis
- producing and releasing open, reusable datasets in XML-TEI using the ParamHTRs interface (TEI NCA sample with abbreviation annotation, gold PEN data for Old French and Latin)
- working closely with the project's PhD candidate (starting September 2027), so that the gold data and evaluation framework directly support model development
- collaborating with the ANR Phil•IA project (CIHAM, Lyon) and with the project's external experts (David Smith, Ariane Pinche)
- supervising and checking annotation work carried out by annotators
- presenting work both internally and externally in conference, journal and workshop papers, in NLP and humanities venues (e.g. Revue de linguistique romane, IMC Leeds, CHR, LREC)
- Organizing and taking part in the project's workshops and exchanging with colleagues on NLP and philological topics
Where to apply
Requirements
Required:
- PhD in philology, historical linguistics, medieval studies or a related field
- Strong knowledge of Old French
- Training in palaeography, with experience reading medieval manuscripts
- Working knowledge of XML-TEI
- Ability to do statistics and to process or parse structured data with at least one programming language (Python or R)
Highly appreciated:
- Experience in dialectometry, or a demonstrated interest in the dialects and scriptae of medieval French
- Knowledge of medieval Latin
Soft skills:
- Ability to work in an interdisciplinary team with NLP researchers
- Good written and oral communication in English; French is an asset
- Good organization skills
The ideal candidate is a philologist or historical linguist who enjoys going back to the manuscript. They are curious about what editions hide, and see scribal abbreviations and spelling variation as evidence, not noise. They like to test big claims about the history of French against carefully built data, and they are comfortable being both a careful annotator and a critical analyst of quantitative results.
They should be at ease in an interdisciplinary environment and willing to learn from NLP researchers. They should also be able to explain philological constraints clearly to people who build models. Rigour, patience with detailed corpus work, and a taste for collaboration matter as much as technical skill. They will work on a weekly basis at least with a PhD student, and regularly with colleagues in Lyon and abroad.
A background in Old French and palaeography, together with an interest in medieval French dialects, is the core of the profile. Programming and statistics can be at an intermediate level, as long as the candidate wants to improve. Knowledge of medieval Latin would be a real asset.
- Languages
- FRENCH
- Level
- Basic
- Languages
- ENGLISH
- Level
- Good
Additional Information
- Subsidized meals
- Partial reimbursement of public transport costs
- Leave: 7 weeks of annual leave + 10 extra days off due to RTT (statutory reduction in working hours) + possibility of exceptional leave (sick children, moving home, etc.)
- Possibility of teleworking and flexible organization of working hours
- Professional equipment available (videoconferencing, loan of computer equipment, etc.)
- Social, cultural and sports events and activities
- Access to vocational training
- Social security coverage
- Website for additional job details
Work Location(s)
- Number of offers available
- 1
- Company/Institute
- Inria
- Country
- France
- City
- Paris
Contact
- City
- LE CHESNAY CEDEX
- Website
- Street
- Domaine de Voluceau - Rocquencourt
- Postal Code
- 78153