A finite-state model of Innu verbal inflection
  Loïc Daignault-Pichette and François Lareau
A Kazakh–Russian Corpus of Child-Directed Language for Low-Resource Languages
  Albina Mukusheva, Achille Fusco and Cristiano Chesi
A Massive Open-Source Corpus for Valencian: Over 4.7 Billion Tokens for Low-Resource Language Modelling
  Yoan Gutiérrez, Juan Pablo Consuegra‐Ayala and Robiert Sepúlveda-Torres
A Multilingual Sentence-Transformer Baseline for Serbian Newspaper Topic Classification
  Sasa Petalinkar, Milica Ikonić Nešić, Ranka Stankovic and Jelena Graovac
A Neuro-Symbolic RAG System for Marine Environmental Law: from Domain Ontology to Knowledge Graphs and Context-Enriched Legal Question Answering
  Hadiza Dite Maa Assoumana Souley, Marie Bonnin, Youssef Al Mouatamid and Jihad Zahir
Agentic retrieval for low-resource languages: An exploratory system design
  Malithi P. Alahapperuma, Andreas Vlachidis and Antonis Bikakis
Arabic Dialect-to-MSA Translation: A Comparative Evaluation of Large Language Models and Neural Machine Translation
  Maram I. Alharbi, Jihad R’Baiti, Ruslan Mitkov, Tharindu Ranasinghe and Hansi Hettiarachchi
Assessment of Human-in-the-Loop Multi-LLMs for Low-Resource Educational Data Expansion
  Christophe Friezas Gonçalves, Hedi Tebourbi, Christoph Schommer and salima lamsiyah
Automating Mwotlap morphology using formal grammar
  Fabio Meroni, Alexandre François and Max Silberztein
Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR
  Khalil Hennara, Muhammad Hreden, Mohamed Motasim Hamed, Ahmad Bastati, Zeina Aldallal, Sara Chrouf and Safwan AlModhayan
Benchmarking Large Language Models on Mandarin Proverb Explanation and Contextual Matching
  Xiaojing Zhao, salima lamsiyah, Emmanuele Chersoni and Han Xu
Benchmarking Parameter-Efficient Fine-Tuning of Large Language Models for Low-Resource Tajik Text Generation with the Tajik Web Corpus
  Mullosharaf Kurbonovich Arabov, Karomatullo Habibullozoda, Nurali Shirinov and Svetlana Khaybullina
Benchmarking Speech Foundation Models and LLMs for Sentiment Analysis in Low-Resource Najdi Arabic of Saudi Arabia
  Nadia Ghezaiel, Maram I. Alharbi and Ruslan Mitkov
Beyond RAG: A Multi-Graph, Multi-Agent, Recursive Retrieval Architecture for Traceable and Source-Grounded Legal Answers
  Assia Bouamir, MARIE Bonnin, Youssef Al Mouatamid and Jihad Zahir
Building an Aligned Speech Corpus for Northern Russian
  Kseniya Protonina and Daniil Ignatev
Clinical Entity Recognition from Electronic Health Records and Linking to Biomedical Knowledge Bases in Low-resource Language Settings
  Ivelina Nikolova-Koleva, Svetla Boytcheva, Kris Collins and Eno-Martin Lotman
Creating the RozMuz Corpus: Applying Ethics and Technology in Language Research
  Alicja Helena Derych, Bartłomiej Alberski, Hubert Jankowski and Paweł Dembowski
Cross-Lingual Hate Speech Detection in Low-Resource Persian and Kurdish
  Shahin Yousefi, Ernesto Luis Estevanell-Valladares and Ruslan Mitkov
Cross-Lingual Transfer from Portuguese to Nheengatu: Evidence for Contact-Induced Convergence as a Computational Bridge
  Rafael Macario Fernandes
Cross-Lingual Transfer Learning For Moroccan Dialect Semantic Textual Similarity
  Said Belbachir, Mohamed Chahhou, Mohammed El Mohajir and Ouafae Nahli
Data Curation, Annotation Quality, and Error Patterns in Old English Automatic Lemmatisation
  Javier Martín Arista
Dependency Parsing Across the Resource Spectrum: Evaluating Architectures on High and Low-Resource Languages
  Kevin Guan, Happy Buzaaba and Christiane Fellbaum
Distil and Evolve: Compact Adaptive Document Classification with Continual Learning for Low-Resource Settings
  Mehdi Soufiane, Mourad Hammale, Kawtar Zerhouni, Adil Ahidar-Coutrix and Rachid Benouini
Do Large Language Models Style-Shift? Register-Conditioned Stylistic Fronting in AI-Generated Icelandic
  Anton Karl Ingason, Johanna Mechler and Lilja Björk Stefánsdóttir
Enhancing Urdu ASR with Whisper v3: Fine-Tuning on Latest Datasets and Realistic Multi-Speaker Evaluation with SLM Post-Processing
  Zehra Ahmed, Farah Inayat, Zuha Aqib and Sajjad Haider
Evaluating Model-Task Fit in Arabic Word Sense Disambiguation
  Yousef Younes, Abdelhalim Hafedh Dahou and Brigitte Mathiak
Fact over Fiction: Detection of Pathological Hallucinations in Sinhala-to-English Neural Machine Translation
  Navam Imanjith Obeyskeara and Nevidu Jayatilleke
FLICK: Few-Label Incremental Learning for Low-Resource Dialects
  Ali Almutairi, Abdullah Alsuhaibani, Shoaib Jameel, Aditya Joshi, Gelareh Mohammadi and Imran Razzak
From Amnesia to Hallucination: Mapping the Affective Entropy of Large Language Models
  Victoria Portnaya, Taras Baraniuk and Roman Kyslyi
Generator-Guided Amount Recovery for Voice-Based Financial Record-Keeping in Mooré-French Code-Switched Speech
  Maimouna Ouattara, El-Hacen Diallo, Fred Philippy, Abdoul Kader Kaboré, Jacques Klein and Tegawendé F. Bissyandé
Improving OCR for a Latvian Pronunciation Dictionary
  Viesturs Jūlijs Lasmanis
Is Overconfidence Language-Specific? Cross-Lingual Calibration and Recalibration Transfer in Base and Instruction-Tuned LLMs
  Divya Divya and Ruslan Mitkov
Knowledge Tracing for Early Childhood Learners with Minimal Telemetry from Low-Connectivity Environments: Evidence from India’s Anganwadi Ecosystem
  Badmavasan Kirouchenassamy, Rahul Singh and Sudeep Gowrishankar
Kuwain 1.5B: An Arabic SLM via Language Injection
  Khalil Hennara, Sara Chrouf, Mohamed Motasim Hamed, Zeina Aldallal, Mohammad Omar Hadid and Safwan AlModhayan
Linguistic Proximity Enables Pivot-Based Machine Translation for an Under-Resourced Tribal Language
  Pooja Singh, M Kaab Bin Shahid, Atai Waris Khan, Aryan Kumar Jha and Sandeep Kumar
Linguistic Specialization of Arabic in Text Embedding Models via Multi-Teacher Knowledge Distillation
  Sara Chrouf, Safwan AlModhayan, Zeina Aldallal and Ahmad Abdelfattah
LLM-Based Post-Correction for Arabic Handwritten Text Recognition: A Systematic Evaluation
  omar Arjafellah, abdellah yousfi and azhar hadmi
LUXDIAG-RAG: Diagnostic Evaluation of Retrieval-Augmented Generation for Luxembourgish Reading Comprehension
  Keerthana Murugaraj, Hedi Tebourbi, Christophe Friezas Goncalves and salima lamsiyah
LuxIT: A Luxembourgish Instruction Tuning Dataset from Monolingual Seed Data
  Julian Valline, Cedric Lothritz, Siwen Guo and Jordi Cabot Sagrera
LëtzCross: A Cross-Lingual Page-Level Benchmark for Multimodal Retrieval over Luxembourgish Documents
  Omar El Bachyr, Fred Philippy, Laura Maria Bernardy, Saad Ezzini, Jacques Klein and TEGAWENDE BISSYANDE
Morphological Operations in the Persian Verbal System
  Marzieh RABIEI and Max Silberztein
Multilingual Text-to-Speech Model for Low-Resource Languages
  Mikhail S. Vdovichenko and Rustam S. Azimov
Probing Geographic Performance Gaps in Moroccan ASR with ALAMA: Annotating Local Audio for Moroccan Arabic
  Avery Cole Kanel, Christian Schuler, Bouazza Laracha, Yassine Chaouri, Imrane Lbouhli, Yusser Al Ghussin and Timo Baumann
Retrieval-Augmented Translation for Bohairic Coptic: A Pilot Benchmark and Hallucination Analysis
  So Miyagawa
Safe in English, Critical Failures in Hindi: Cross-lingual LLM Responses to Mental Health Queries
  Gavin Abercrombie
SAMTRAMOD – Saaho Machine Translation Model
  Jama Musse Jama
SimuLe-Lux: A Neuro-Symbolic Teaching System for Diagnosing L2 Reading Comprehension
  Christophe Friezas Gonçalves, Christoph Schommer and salima lamsiyah
Spatial Entity Extraction Methods from Arabic Texts in the Context of Epidemiological Surveillance: A Comparative Study
  Fatima Ezzahra El Houbri, Najlae IDRISSI, Mathieu Roche and Sarah Valentin
The United Nations Sixth Committee English–Arabic Parallel Corpus: A Domain-Specific Open-Access Dataset for Legal Translation
  Hanem El-Farahaty, Asma Alduhaim and Nouran Khallaf
Topic Modeling for Moroccan Darija: A Comparative Study of Classical Machine Learning Approaches
  SALMA MEKAOUI, Ilham CHAKER, Arsalane Zarghili and Nikola S. Nikolov
Towards a Universal Dependencies Treebank for Amazigh: A Tarifit Pilot Study
  Azzeddine Afrouni, Fadoua Ataa Allah and Jamal Abarnous
Towards Readability Assessment for Under-Resourced Arabic Dialects: A Study of Moroccan Darija
  Houdaifa Atou, salima lamsiyah, Nouran Khallaf and Ruslan Mitkov
Wasm: A Pipeline for Constructing Structured Arabic Interleaved Multimodal Corpora
  Khalil Hennara, Ahmad Bastati, Muhammad Hreden, Mohamed Motasim Hamed, Zeina Aldallal, Sara Chrouf and Safwan AlModhayan
Worldview Annotation for Low-Resource Language
  Tetiana Ilman
‘What bites you is inside your clothes!’ Safety Evaluation of LLMs in Swahili
  Ima Emanueli Mahenge, Gavin Abercrombie and Silviana Lazaro Swallo