
Accepted papers
| A finite-state model of Innu verbal inflection Loïc Daignault-Pichette and François Lareau |
| A Kazakh–Russian Corpus of Child-Directed Language for Low-Resource Languages Albina Mukusheva, Achille Fusco and Cristiano Chesi |
| A Massive Open-Source Corpus for Valencian: Over 4.7 Billion Tokens for Low-Resource Language Modelling Yoan Gutiérrez, Juan Pablo Consuegra‐Ayala and Robiert Sepúlveda-Torres |
| A Multilingual Sentence-Transformer Baseline for Serbian Newspaper Topic Classification Sasa Petalinkar, Milica Ikonić Nešić, Ranka Stankovic and Jelena Graovac |
| A Neuro-Symbolic RAG System for Marine Environmental Law: from Domain Ontology to Knowledge Graphs and Context-Enriched Legal Question Answering Hadiza Dite Maa Assoumana Souley, Marie Bonnin, Youssef Al Mouatamid and Jihad Zahir |
| Agentic retrieval for low-resource languages: An exploratory system design Malithi P. Alahapperuma, Andreas Vlachidis and Antonis Bikakis |
| Arabic Dialect-to-MSA Translation: A Comparative Evaluation of Large Language Models and Neural Machine Translation Maram I. Alharbi, Jihad R’Baiti, Ruslan Mitkov, Tharindu Ranasinghe and Hansi Hettiarachchi |
| Assessment of Human-in-the-Loop Multi-LLMs for Low-Resource Educational Data Expansion Christophe Friezas Gonçalves, Hedi Tebourbi, Christoph Schommer and salima lamsiyah |
| Automating Mwotlap morphology using formal grammar Fabio Meroni, Alexandre François and Max Silberztein |
| Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR Khalil Hennara, Muhammad Hreden, Mohamed Motasim Hamed, Ahmad Bastati, Zeina Aldallal, Sara Chrouf and Safwan AlModhayan |
| Benchmarking Large Language Models on Mandarin Proverb Explanation and Contextual Matching Xiaojing Zhao, salima lamsiyah, Emmanuele Chersoni and Han Xu |
| Benchmarking Parameter-Efficient Fine-Tuning of Large Language Models for Low-Resource Tajik Text Generation with the Tajik Web Corpus Mullosharaf Kurbonovich Arabov, Karomatullo Habibullozoda, Nurali Shirinov and Svetlana Khaybullina |
| Benchmarking Speech Foundation Models and LLMs for Sentiment Analysis in Low-Resource Najdi Arabic of Saudi Arabia Nadia Ghezaiel, Maram I. Alharbi and Ruslan Mitkov |
| Beyond RAG: A Multi-Graph, Multi-Agent, Recursive Retrieval Architecture for Traceable and Source-Grounded Legal Answers Assia Bouamir, MARIE Bonnin, Youssef Al Mouatamid and Jihad Zahir |
| Building an Aligned Speech Corpus for Northern Russian Kseniya Protonina and Daniil Ignatev |
| Clinical Entity Recognition from Electronic Health Records and Linking to Biomedical Knowledge Bases in Low-resource Language Settings Ivelina Nikolova-Koleva, Svetla Boytcheva, Kris Collins and Eno-Martin Lotman |
| Creating the RozMuz Corpus: Applying Ethics and Technology in Language Research Alicja Helena Derych, Bartłomiej Alberski, Hubert Jankowski and Paweł Dembowski |
| Cross-Lingual Hate Speech Detection in Low-Resource Persian and Kurdish Shahin Yousefi, Ernesto Luis Estevanell-Valladares and Ruslan Mitkov |
| Cross-Lingual Transfer from Portuguese to Nheengatu: Evidence for Contact-Induced Convergence as a Computational Bridge Rafael Macario Fernandes |
| Cross-Lingual Transfer Learning For Moroccan Dialect Semantic Textual Similarity Said Belbachir, Mohamed Chahhou, Mohammed El Mohajir and Ouafae Nahli |
| Data Curation, Annotation Quality, and Error Patterns in Old English Automatic Lemmatisation Javier Martín Arista |
| Dependency Parsing Across the Resource Spectrum: Evaluating Architectures on High and Low-Resource Languages Kevin Guan, Happy Buzaaba and Christiane Fellbaum |
| Distil and Evolve: Compact Adaptive Document Classification with Continual Learning for Low-Resource Settings Mehdi Soufiane, Mourad Hammale, Kawtar Zerhouni, Adil Ahidar-Coutrix and Rachid Benouini |
| Do Large Language Models Style-Shift? Register-Conditioned Stylistic Fronting in AI-Generated Icelandic Anton Karl Ingason, Johanna Mechler and Lilja Björk Stefánsdóttir |
| Enhancing Urdu ASR with Whisper v3: Fine-Tuning on Latest Datasets and Realistic Multi-Speaker Evaluation with SLM Post-Processing Zehra Ahmed, Farah Inayat, Zuha Aqib and Sajjad Haider |
| Evaluating Model-Task Fit in Arabic Word Sense Disambiguation Yousef Younes, Abdelhalim Hafedh Dahou and Brigitte Mathiak |
| Fact over Fiction: Detection of Pathological Hallucinations in Sinhala-to-English Neural Machine Translation Navam Imanjith Obeyskeara and Nevidu Jayatilleke |
| FLICK: Few-Label Incremental Learning for Low-Resource Dialects Ali Almutairi, Abdullah Alsuhaibani, Shoaib Jameel, Aditya Joshi, Gelareh Mohammadi and Imran Razzak |
| From Amnesia to Hallucination: Mapping the Affective Entropy of Large Language Models Victoria Portnaya, Taras Baraniuk and Roman Kyslyi |
| Generator-Guided Amount Recovery for Voice-Based Financial Record-Keeping in Mooré-French Code-Switched Speech Maimouna Ouattara, El-Hacen Diallo, Fred Philippy, Abdoul Kader Kaboré, Jacques Klein and Tegawendé F. Bissyandé |
| Improving OCR for a Latvian Pronunciation Dictionary Viesturs Jūlijs Lasmanis |
| Is Overconfidence Language-Specific? Cross-Lingual Calibration and Recalibration Transfer in Base and Instruction-Tuned LLMs Divya Divya and Ruslan Mitkov |
| Knowledge Tracing for Early Childhood Learners with Minimal Telemetry from Low-Connectivity Environments: Evidence from India’s Anganwadi Ecosystem Badmavasan Kirouchenassamy, Rahul Singh and Sudeep Gowrishankar |
| Kuwain 1.5B: An Arabic SLM via Language Injection Khalil Hennara, Sara Chrouf, Mohamed Motasim Hamed, Zeina Aldallal, Mohammad Omar Hadid and Safwan AlModhayan |
| Linguistic Proximity Enables Pivot-Based Machine Translation for an Under-Resourced Tribal Language Pooja Singh, M Kaab Bin Shahid, Atai Waris Khan, Aryan Kumar Jha and Sandeep Kumar |
| Linguistic Specialization of Arabic in Text Embedding Models via Multi-Teacher Knowledge Distillation Sara Chrouf, Safwan AlModhayan, Zeina Aldallal and Ahmad Abdelfattah |
| LLM-Based Post-Correction for Arabic Handwritten Text Recognition: A Systematic Evaluation omar Arjafellah, abdellah yousfi and azhar hadmi |
| LUXDIAG-RAG: Diagnostic Evaluation of Retrieval-Augmented Generation for Luxembourgish Reading Comprehension Keerthana Murugaraj, Hedi Tebourbi, Christophe Friezas Goncalves and salima lamsiyah |
| LuxIT: A Luxembourgish Instruction Tuning Dataset from Monolingual Seed Data Julian Valline, Cedric Lothritz, Siwen Guo and Jordi Cabot Sagrera |
| LëtzCross: A Cross-Lingual Page-Level Benchmark for Multimodal Retrieval over Luxembourgish Documents Omar El Bachyr, Fred Philippy, Laura Maria Bernardy, Saad Ezzini, Jacques Klein and TEGAWENDE BISSYANDE |
| Morphological Operations in the Persian Verbal System Marzieh RABIEI and Max Silberztein |
| Multilingual Text-to-Speech Model for Low-Resource Languages Mikhail S. Vdovichenko and Rustam S. Azimov |
| Probing Geographic Performance Gaps in Moroccan ASR with ALAMA: Annotating Local Audio for Moroccan Arabic Avery Cole Kanel, Christian Schuler, Bouazza Laracha, Yassine Chaouri, Imrane Lbouhli, Yusser Al Ghussin and Timo Baumann |
| Retrieval-Augmented Translation for Bohairic Coptic: A Pilot Benchmark and Hallucination Analysis So Miyagawa |
| Safe in English, Critical Failures in Hindi: Cross-lingual LLM Responses to Mental Health Queries Gavin Abercrombie |
| SAMTRAMOD – Saaho Machine Translation Model Jama Musse Jama |
| SimuLe-Lux: A Neuro-Symbolic Teaching System for Diagnosing L2 Reading Comprehension Christophe Friezas Gonçalves, Christoph Schommer and salima lamsiyah |
| Spatial Entity Extraction Methods from Arabic Texts in the Context of Epidemiological Surveillance: A Comparative Study Fatima Ezzahra El Houbri, Najlae IDRISSI, Mathieu Roche and Sarah Valentin |
| The United Nations Sixth Committee English–Arabic Parallel Corpus: A Domain-Specific Open-Access Dataset for Legal Translation Hanem El-Farahaty, Asma Alduhaim and Nouran Khallaf |
| Topic Modeling for Moroccan Darija: A Comparative Study of Classical Machine Learning Approaches SALMA MEKAOUI, Ilham CHAKER, Arsalane Zarghili and Nikola S. Nikolov |
| Towards a Universal Dependencies Treebank for Amazigh: A Tarifit Pilot Study Azzeddine Afrouni, Fadoua Ataa Allah and Jamal Abarnous |
| Towards Readability Assessment for Under-Resourced Arabic Dialects: A Study of Moroccan Darija Houdaifa Atou, salima lamsiyah, Nouran Khallaf and Ruslan Mitkov |
| Wasm: A Pipeline for Constructing Structured Arabic Interleaved Multimodal Corpora Khalil Hennara, Ahmad Bastati, Muhammad Hreden, Mohamed Motasim Hamed, Zeina Aldallal, Sara Chrouf and Safwan AlModhayan |
| Worldview Annotation for Low-Resource Language Tetiana Ilman |
| ‘What bites you is inside your clothes!’ Safety Evaluation of LLMs in Swahili Ima Emanueli Mahenge, Gavin Abercrombie and Silviana Lazaro Swallo |
