Department of Computer Science
PhD Defence by Esther Ploeger

Room 1.001
A. C. Meyers Vænge 15
2450 København
26.11.2025 13:00 - 16:30
English
On location
Room 1.001
A. C. Meyers Vænge 15
2450 København
26.11.2025 13:00 - 16:30
English
On location
Department of Computer Science
PhD Defence by Esther Ploeger

Room 1.001
A. C. Meyers Vænge 15
2450 København
26.11.2025 13:00 - 16:30
English
On location
Room 1.001
A. C. Meyers Vænge 15
2450 København
26.11.2025 13:00 - 16:30
English
On location
Fakta
All interested parties are welcome. After the defence the Department of Computer Science will host a small reception.
Abstract
The languages of the world vary considerably on a structural level. Yet, language technology has mostly been developed with the English language in mind. Assumptions based on a single language do not necessarily generalize. For example, popular tokenization strategies that work well for English and morphologically similar languages such as Danish, may not be equally applicable to more morphologically complex languages such as Kalaallisut. Yet, defining what are structurally ‘similar’ or ‘distant’ languages is non-trivial: a systematic approach to the structural variation between languages has not been systematically integrated into natural language processing (NLP) evaluation.
The aim of this thesis is to increase such systematicity, leveraging insights from the field of linguistic typology. It addresses the selection, processing and downstream application of typological information in NLP research. Firstly, we assess the reliability of the most widely used typological database in NLP, finding that it may not be well-suited for computational applications. We then propose a novel way to measure typological similarity and diversity of language selections, enabling a more systematic treatment of language diversity in NLP. Lastly, we show how this affects the generalizability of downstream task performance, with a focus on machine translation.
Ultimately, we find that so far the role of linguistic typology in NLP has mostly been limited to a database-centric view, where information from typological databases is often integrated with limited attention to the theoretical and methodological foundations of typology itself. I argue that the use of linguistic typology in NLP benefits from a more interdisciplinary perspective. Beyond leveraging databases, there are common methodological challenges in typology and NLP, such as language stratification and sampling. By integrating insights from both disciplines, we propose a more principled framework that promotes robust and inclusive multilingual language technologies.
This Ph.D. project was made possible under the project Multilingual Modelling for Resource-Poor Languages through the Carlsberg Foundation, grant CF21-0454.
Attendees
- Associate Professor A. Seza Doğruöz, Ghent University, Belgium.
- Associate Professor Christian Hardmeier, IT University of Copenhagen, Denmark.
- Associate Professor Kristine Bundgaard (chair), Department of Culture and Learning, Aalborg University, Denmark.
- Professor Johannes Bjerva, Department of Computer Science, Aalborg University, Denmark.
- Docent Robert Östling, Department of Linguistics, Stockholm University, Sweden.
- Associate Professor Andrés Masegosa, Department of Computer Science, Aalborg University, Denmark.