Gülsen Eryigit

dblp:57/1920 · DBLP profile ↗
← Back
29ranked-venue papers
7as first author
10since 2021 · last 2026
0000-0003-4607-7305ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 7 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Typology-Aware Multilingual Morphosyntactic Parsing with Joint Abstract Node Modeling
abstract
The UniDive 2025 Morphosyntactic Parsing (MSP) shared task introduces a representation unifying dependency structure, morphological features, and unrealized arguments.Unlike Universal Dependencies, MSP encodes abstract nodes (e.g., dropped subjects, implicit pronouns) as labels projected onto content words, which standard UD parsers cannot model.To address this challenge, in this paper we present a multilingual, typology-aware joint system integrating word-type prediction, content-only parsing, morphological tagging, and an abstract-node component within a single architecture.The model combines the baseline joint framework with typology-conditioned adapters and progressive weighting for abstract supervision.On the MSP test set, our model outperforms the leading submission by 3.23 percentage points in MSLAS, 3.35 in LAS, and 1.78 in FEATS macro F1, demonstrating the effectiveness of typology-sensitive multi-task learning in MSP.
Kutay Acar, Gülsen Eryigit
ACL (1)2
2026 A Parallel Cross-Lingual Benchmark for Multimodal Idiomaticity Understanding
abstract
Potentially idiomatic expressions (PIEs) carry meanings inherently tied to the everyday experience of a given language community. As such, they constitute an interesting challenge for assessing the linguistic (and to some extent cultural) capabilities of NLP systems. In this paper, we present XMPIE, a parallel multilingual and multimodal dataset of potentially idiomatic expressions. The dataset, containing 34 languages and over ten thousand items, allows comparative analyses of idiomatic patterns among language-specific realisations and preferences in order to gather insights about shared cultural aspects. This parallel dataset allows evaluation of language model performance for a given PIE in different languages and whether idiomatic understanding in one language can be transferred to another. Moreover, the dataset supports the study of PIEs across textual and visual modalities, to measure to what extent PIE understanding in one modality transfers or implies in understanding in another modality (text vs. image). The data was created by language experts, with both textual and visual components crafted under multilingual guidelines, and each PIE is accompanied by five images representing a spectrum from idiomatic to literal meanings, including semantically related and random distractors. The result is a high-quality benchmark for evaluating multilingual and multimodal idiomatic language understanding.
Dilara Torunoglu-Selamet, Dogukan Arslan, Rodrigo Wilkens, Wei He 0017, Doruk Eryigit, Thomas Pickard, Adriana S. Pagano, Aline Villavicencio, Gülsen Eryigit, Ágnes Abuczki, Aida Cardoso, Alesia Lazarenka, Dina Almassova, Amália Mendes, Anna Kanellopoulou, Antoni Brosa-Rodríguez, Baiba Valkovska, Beata Wojtowicz, Bolette Pedersen, Carlos Manuel Hidalgo-Ternero, Chaya Liebeskind, Danka Jokic, Diego Alves, Eleni Triantafyllidi, Erik Velldal, Fred Philippy, Giedre Valunaite Oleskeviciene, Ieva Rizgeliene, Inguna Skadina, Irina Lobzhanidze, Isabell Stinessen Haugen, Jauza Akbar Krito, Jelena M. Markovic, Johanna Monti, Josue Alejandro Sauca, Kaja Dobrovoljc, Kingsley O. Ugwuanyi, Laura Rituma, Lilja Øvrelid, Maha Tufail Agro, Manzura Abjalova, Maria Chatzigrigoriou, María del Mar Sánchez Ramos, Marija Pendevska, Masoumeh Seyyedrezaei, Mehrnoush Shamsfard, Momina Ahsan, Muhammad Ahsan Riaz Khan, Nathalie Carmen Hau Norman, Nilay Erdem Ayyildiz, Nina Hosseini-Kivanani, Noémi Ligeti-Nagy, Numaan Naeem, Olha Kanishcheva, Olha Yatsyshyna, Daniil Orel, Petra Giommarelli, Petya Osenova, Radovan Garabík, Regina E. Semou, Rozane Rebechi, Salsabila Zahirah Pranida, Samia Touileb, Sanni Nimb, Sarvinoz Sharipova, Shahar Golan, Shaoxiong Ji, Sopuruchi Christian Aboh, Srdjan Sucur, Stella Markantonatou, Sussi Olsen, Vahideh Tajalli, Veronika Lipp, Voula Giouli, Yelda Yesildal Eraydin, Zahra Saaberi, Zhuohan Xie
LREC9
2026 CorefInst: Leveraging LLMs for Multilingual Coreference Resolution
abstract
Abstract Coreference Resolution (CR) is a crucial yet challenging task in natural language understanding, often constrained by task-specific architectures and encoder-based language models that demand extensive training and lack adaptability. This study introduces the first multilingual CR methodology which leverages decoder-only LLMs to handle both overt and zero mentions. The article explores how to model the CR task for LLMs via five different instruction sets using a controlled inference method. The approach is evaluated across three LLMs: Llama 3.1, Gemma 2, and Mistral 0.3. The results indicate that LLMs, when instruction-tuned with a suitable instruction set, can surpass state-of-the-art task-specific architectures. Specifically, our best model, a fully fine-tuned Llama 3.1 for multilingual CR, outperforms the leading multilingual CR model (i.e., Corpipe 24 single stage variant) by 2 percentage points on average across all languages in the CorefUD v1.2 dataset collection.
Tugba Pamay Arslan, Emircan Erol, Gülsen Eryigit
Trans. Assoc. Comput. Linguistics3
2025 Enhancing Turkish Coreference Resolution: Insights from deep learning, dropped pronouns, and multilingual transfer learning
Tugba Pamay Arslan, Gülsen Eryigit
Comput. Speech Lang.2
2025 Gamified Crowd-sourcing for Word Sense Disambiguation of Turkish
abstract
Word sense disambiguation (WSD) is the process of determining the correct meaning of a word based on its context in a sentence, a task that remains one of the core challenges in natural language processing (NLP) despite the advancements made by LLMs. Although LLMs have improved WSD performance, challenges remain–especially in nuanced contexts, low-resource languages, and domain-specific terminology. These gaps can hinder the accuracy of LLMs in high-stakes applications and multilingual settings. In response to the need for WSD, our study introduces a novel approach to WSD data collection through gamified crowdsourcing, which, to our knowledge, has not been previously applied in this field. A messaging bot method has been used to engage a wide, diverse audience to generate high-quality WSD data. Using a multiplayer game format, native speakers provide examples for different senses of ambiguous words and rate others’ contributions. Together with the suggested enhancements, this platform attracted a diverse range of participants–spanning ages, backgrounds, and genders–20 times more than a similar crowdsourcing method applied to a different NLP task, and sustained their engagement over an extended period. Unlike conventional academic crowdsourcing pools, this approach focuses on drawing individuals from varied backgrounds beyond academia or AI-focused communities. It not only gathered data but also encouraged participants to evaluate each other’s contributions, creating a rich and reliable dataset. Our findings suggest that gamified crowdsourcing can be a powerful tool for creating WSD corpora. Our approach not only supports future WSD research but also contributes valuable training data for developing more precise and resilient language models.
Dilara Torunoglu-Selamet, Ali Sentas, Gülsen Eryigit
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2024 Abstract meaning representation of Turkish
abstract
Abstract Abstract meaning representation (AMR) is a graph-based sentence-level meaning representation that has become highly popular in recent years. AMR is a knowledge-based meaning representation heavily relying on frame semantics for linking predicate frames and entity knowledge bases such as DBpedia for linking named entity concepts. Although it is originally designed for English, its adaptation to non-English languages is possible by defining language-specific divergences and representations. This article introduces the first AMR representation framework for Turkish, which poses diverse challenges for AMR due to its typological differences compared to English; agglutinative, free constituent order, morphologically highly rich resulting in fewer word surface forms in sentences. The introduced solutions to these peculiarities are expected to guide the studies for other similar languages and speed up the construction of a cross-lingual universal AMR framework. Besides this main contribution, the article also presents the construction of the first AMR corpus of 700 sentences, the first AMR parser (i.e., a tree-to-graph rule-based AMR parser) used for semi-automatic annotation, and the evaluation of the introduced resources for Turkish.
Elif Oral, Ali Acar, Gülsen Eryigit
Nat. Lang. Eng.3
2023 Gamified crowdsourcing for idiom corpora construction
abstract
Learning idiomatic expressions is seen as one of the most challenging stages in second language learning because of their unpredictable meaning. A similar situation holds for their identification within natural language processing applications such as machine translation and parsing. The lack of high-quality usage samples exacerbates this challenge not only for humans but also for artificial intelligence systems. This article introduces a gamified crowdsourcing approach for collecting language learning materials for idiomatic expressions; a messaging bot is designed as an asynchronous multiplayer game for native speakers who compete with each other while providing idiomatic and nonidiomatic usage examples and rating other players' entries. As opposed to classical crowdprocessing annotation efforts in the field, for the first time in the literature, a crowdcreating & crowdrating approach is implemented and tested for idiom corpora construction. The approach is language independent and evaluated on two languages in comparison to traditional data preparation techniques in the field. The reaction of the crowd is monitored under different motivational means (namely, gamification affordances and monetary rewards). The results reveal that the proposed approach is powerful in collecting the targeted materials, and although being an explicit crowdsourcing approach, it is found entertaining and useful by the crowd. The approach has been shown to have the potential to speed up the construction of idiom corpora for different natural languages to be used as second language learning material, training data for supervised idiom identification systems, or samples for lexicographic studies.
Gülsen Eryigit, Ali Sentas, Johanna Monti
Nat. Lang. Eng.1
2022 Fusion of visual representations for multimodal information extraction from unstructured transactional documents
Berke Oral, Gülsen Eryigit
Int. J. Document Anal. Recognit.2
2021 Evaluation of Wizard-of-Oz and Self-Play Data Collection Techniques for Turkish Goal-Oriented Dialogue Agents
abstract
As with all natural language processing tasks, the lack of open-source training data required for the development of dialogue agents is a major obstacle to research studies in the field. Especially languages that are not widely studied, such as Turkish, suffer more from this problem. This article introduces a comparison of Wizard-of-Oz and self-play data collection techniques for Turkish goal-oriented dialogue system generation. Three data sets have been prepared and introduced to the researchers by using these techniques. Being the first publicly available human-to-human Turkish dialogue data sets, although open for development, the created resources from the restaurant domain are very valuable for further research on Turkish dialogue systems. The mentioned methods are quantitatively compared on the produced data sets, in terms of dialog act classification and slot identification scores. Since it is costly to collect data with methods like Wizard-Of-Oz in every domain, an open-source flexible and easy-to-use framework is also provided implementing self-play which may be used to create machine-to-machine dialogue outlines and speed data collection for low-resource languages like Turkish. Besides, designed templates of annotation screens for crowdsourcing are provided for future studies.
Dogukan Arslan, Gülsen Eryigit
INISTA2
2021 Development of Goal-Oriented Dialogue Systems for Customer Services in Automotive Industry
abstract
Being the largest sector in Turkey in the field of export, the automotive sector is one of the most important sectors in Turkey. Companies that manage their customer relations through live support channels are faced with problems such as long waiting times, access through limited channels, the live support team losing time while accessing the inquiry tools, and the inability of the customers to convey a problem through visual channels that cannot be verbally expressed. This paper introduces a Goal-Oriented Dialogue System (GODS) for Customer Services in the Automotive Industry. Although there exist some studies for in-car dialogue systems, this is the first study for GODS in the mentioned field, and we think that the results obtained in this paper on Turkish will shed light on future studies for other languages. Our paper defines 6 general Intents and 16 Entities in the automotive customer services and 8 Dialog act classes. The paper also provides the ways to extend the system for emerging intent classes in live data and for the needs of different service companies, as well as the experimental results on a preliminary data-set.
Merve Tunçer, Mehmet Fahri Bilici, Gülsen Eryigit
INISTA3
2020 Constructing Multimodal Language Learner Texts Using LARA: Experiences with Nine Languages
abstract
LARA (Learning and Reading Assistant) is an open source platform whose purpose is to support easy conversion of plain texts into multimodal online versions suitable for use by language learners. This involves semi-automatically tagging the text, adding other annotations and recording audio. The platform is suitable for creating texts in multiple languages via crowdsourcing techniques that can be used for teaching a language via reading and listening. We present results of initial experiments by various collaborators where we measure the time required to produce substantial LARA resources, up to the length of short novels, in Dutch, English, Farsi, French, German, Icelandic, Irish, Swedish and Turkish. The first results are encouraging. Although there are some startup problems, the conversion task seems manageable for the languages tested so far. The resulting enriched texts are posted online and are freely available in both source and compiled form.
Elham Akhlaghi, Branislav Bédi, Fatih Bektas, Harald Berthelsen, Matt Butterweck, Cathy Chua, Catia Cucchiarini, Gülsen Eryigit, Johanna Gerlach, Hanieh Habibi, Neasa Ní Chiaráin, Manny Rayner, Steinþór Steingrímsson, Helmer Strik
LREC8
2020 Information Extraction from Text Intensive and Visually Rich Banking Documents
Berke Oral, Erdem Emekligil, Seçil Arslan, Gülsen Eryigit
Inf. Process. Manag.4
2018 Turkish Coreference Resolution
abstract
This paper presents the state-of-the-art results in Turkish coreference resolution (CR) which is a task of determining sets of mentions which identify the same real-world entity (e.g. a person, a place, a thing, an event). The proposed system uses support vector machines and solves the CR task with a mention-pair model that basically accepts mention couples and decides on whether they are coreferential with each other or not. The results are evaluated on Marmara Turkish Coreference Corpus by using well-known evaluation metrics (viz. MUC, B3, BLANC and LEA). The introduced approach obtains F1 scores of 90.68% (MUC), 86.89% (B3), 85.13% (BLANC) and 78.34% (LEA) yielding an improvement of 9.12, 16.06, 13.08 and 12.57 percentage points respectively over a recent baseline system on Turkish CR. The paper introduces the system setup (SVM parameters and negative sampling strategy) as well as the selected features and analyzes the impact of these features on the Turkish CR task.
Tugba Pamay, Gülsen Eryigit
INISTA2
2017 Multiword Expression Processing: A Survey
abstract
Multiword expressions (MWEs) are a class of linguistic forms spanning conventional word boundaries that are both idiosyncratic and pervasive across different languages. The structure of linguistic processing that depends on the clear distinction between words and phrases has to be re-thought to accommodate MWEs. The issue of MWE handling is crucial for NLP applications, where it raises a number of challenges. The emergence of solutions in the absence of guiding principles motivates this survey, whose aim is not only to provide a focused review of MWE processing, but also to clarify the nature of interactions between MWE processing and downstream applications. We propose a conceptual framework within which challenges and research contributions can be positioned. It offers a shared understanding of what is meant by “MWE processing,” distinguishing the subtasks of MWE discovery and identification. It also elucidates the interactions between MWE processing and two use cases: Parsing and machine translation. Many of the approaches in the literature can be differentiated according to how MWE processing is timed with respect to underlying use cases. We discuss how such orchestration choices affect the scope of MWE-aware systems. For each of the two MWE processing subtasks and for each of the two use cases, we conclude on open issues and research perspectives.
Matthieu Constant, Gülsen Eryigit, Johanna Monti, Lonneke van der Plas, Carlos Ramisch, Mike Rosner, Amalia Todirascu
Comput. Linguistics2
2017 Social media text normalization for Turkish
abstract
Abstract Text normalization is an indispensable stage in processing noncanonical language from natural sources, such as speech, social media or short text messages. Research in this field is very recent and mostly on English. As is known from different areas of natural language processing, morphologically rich languages (MRLs) pose many different challenges when compared to English. Turkish is a strong representative of MRLs and has particular normalization problems that may not be easily solved by a single-stage pure statistical model. This article introduces the first work on the social media text normalization of an MRL and presents the first complete social media text normalization system for Turkish. The article conducts an in-depth analysis of the error types encountered in Web 2.0 Turkish texts, categorizes them into seven groups and provides solutions for each of them by dividing the candidate generation task into separate modules working in a cascaded architecture. For the first time in the literature, two manually normalized Web 2.0 datasets are introduced for Turkish normalization studies. The exact match scores of the overall system on the provided datasets are 70.40 per cent and 67.37 per cent (77.07 per cent with a case insensitive evaluation).
Gülsen Eryigit, Dilara Torunoglu-Selamet
Nat. Lang. Eng.1
2016 Universal Dependencies for Turkish
abstract
The Universal Dependencies (UD) project was conceived after the substantial recent interest in unifying annotation schemes across languages. With its own annotation principles and abstract inventory for parts of speech, morphosyntactic features and dependency relations, UD aims to facilitate multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. This paper presents the Turkish IMST-UD Treebank, the first Turkish treebank to be in a UD release. The IMST-UD Treebank was automatically converted from the IMST Treebank, which was also recently released. We describe this conversion procedure in detail, complete with mapping tables. We also present our evaluation of the parsing performances of both versions of the IMST Treebank. Our findings suggest that the UD framework is at least as viable for Turkish as the original annotation framework of the IMST Treebank.
Umut Sulubacak, Memduh Gokirmak, Francis M. Tyers, Çagri Çöltekin, Joakim Nivre, Gülsen Eryigit
COLING6
2016 Building machine-readable knowledge representations for Turkish sign language generation
abstract
This article proposes a representation scheme for depicting the Turkish Sign Language (TID) electronically for use in an automated machine translation system whose basic aim is to translate the Turkish primary school educational materials to TID. The main contribution of the article is the introduction of a machine-readable knowledge representation for TID for the first time in the literature. Like many resource-poor languages, TID lacks electronic language resources usable in computerized systems. The utilization of the proposed scheme for resource creation is also provided in this article by two means: an interactive online dictionary platform for TID and an ELAN add-on for corpus creation.
Cihat Eryigit, Hatice Kose-Bagci, Meltem Kelepir, Gülsen Eryigit
Knowl. Based Syst.4
2014 ITU Turkish NLP Web Service
abstract
We present a natural language processing (NLP) platform, namely the "ITU Turkish NLP Web Service" by the natural language processing group of Istanbul Technical University.The platform (available at tools.nlp.itu.edu.tr)operates as a SaaS (Software as a Service) and provides the researchers and the students the state of the art NLP tools in many layers: preprocessing, morphology, syntax and entity recognition.The users may communicate with the platform via three channels: 1. via a user friendly web interface, 2. by file uploads and 3. by using the provided Web APIs within their own codes for constructing higher level applications.
Gülsen Eryigit
EACL1
2012 Initial Explorations on using CRFs for Turkish Named Entity Recognition
Gökhan Akin Seker, Gülsen Eryigit
COLING2
2012 Word Alignment for English-Turkish Language Pair
Mehmet Talha Çakmak, Süleyman Acar, Gülsen Eryigit
LREC3
2012 The Impact of Automatic Morphological Analysis & Disambiguation on Dependency Parsing of Turkish
Gülsen Eryigit
LREC1
2008 Dependency Parsing of Turkish
abstract
The suitability of different parsing methods for different languages is an important topic in syntactic parsing. Especially lesser-studied languages, typologically different from the languages for which methods have originally been developed, pose interesting challenges in this respect. This article presents an investigation of data-driven dependency parsing of Turkish, an agglutinative, free constituent order language that can be seen as the representative of a wider class of languages of similar type. Our investigations show that morphological structure plays an essential role in finding syntactic relations in such a language. In particular, we show that employing sublexical units called inflectional groups, rather than word forms, as the basic parsing units improves parsing accuracy. We test our claim on two different parsing methods, one based on a probabilistic model with beam search and the other based on discriminative classifiers and a deterministic parsing strategy, and show that the usefulness of sublexical units holds regardless of the parsing method. We examine the impact of morphological and lexical information in detail and show that, properly used, this kind of information can improve parsing accuracy substantially. Applying the techniques presented in this article, we achieve the highest reported accuracy for parsing the Turkish Treebank.
Gülsen Eryigit, Joakim Nivre, Kemal Oflazer
Comput. Linguistics1
2008 Dependency Parsing of Turkish
abstract
The suitability of different parsing methods for different languages is an important topic in syntactic parsing. Especially lesser-studied languages, typologically different from the languages for which methods have originally been developed, pose interesting challenges in this respect. This article presents an investigation of data-driven dependency parsing of Turkish, an agglutinative, free constituent order language that can be seen as the representative of a wider class of languages of similar type. Our investigations show that morphological structure plays an essential role in finding syntactic relations in such a language. In particular, we show that employing sublexical units called inflectional groups, rather than word forms, as the basic parsing units improves parsing accuracy. We test our claim on two different parsing methods, one based on a probabilistic model with beam search and the other based on discriminative classifiers and a deterministic parsing strategy, and show that the usefulness of sublexical units holds regardless of the parsing method. We examine the impact of morphological and lexical information in detail and show that, properly used, this kind of information can improve parsing accuracy substantially. Applying the techniques presented in this article, we achieve the highest reported accuracy for parsing the Turkish Treebank.
Gülsen Eryigit, Joakim Nivre, Kemal Oflazer
Comput. Linguistics1
2007 Single Malt or Blended? A Study in Multilingual Parser Optimization
Johan Hall, Jens Nilsson 0001, Joakim Nivre, Gülsen Eryigit, Beáta Megyesi, Mattias Nilsson 0003, Markus Saers
EMNLP-CoNLL4
2007 MaltParser: A language-independent system for data-driven dependency parsing
abstract
Parsing unrestricted text is useful for many language technology applications but requires parsing methods that are both robust and efficient. MaltParser is a language-independent system for data-driven dependency parsing that can be used to induce a parser for a new language from a treebank sample in a simple yet flexible manner. Experimental evaluation confirms that MaltParser can achieve robust, efficient and accurate parsing for a wide range of languages without language-specific enhancements and with rather limited amounts of training data.
Joakim Nivre, Johan Hall, Jens Nilsson 0001, Atanas Chanev, Gülsen Eryigit, Sandra Kübler, Svetoslav Marinov, Erwin Marsi
Nat. Lang. Eng.5
2006 Labeled Pseudo-Projective Dependency Parsing with Support Vector Machines
Joakim Nivre, Johan Hall, Jens Nilsson 0001, Gülsen Eryigit, Svetoslav Marinov
CoNLL4
2006 Statistical Dependency Parsing for Turkish
Gülsen Eryigit, Kemal Oflazer
EACL1
2005 Improvements to penalty-based evolutionary algorithms for the multi-dimensional knapsack problem using a gene-based adaptive mutation approach
abstract
Knapsack problems are among the most common problems in literature tackled with evolutionary algorithms (EA). Their major advantage lies in the fact that they are relatively simple to implement while they allow generalizations for a wide range of real world problems. The multi-dimensional knapsack problem (MKP), which belongs to the class of NP-complete combinatorial optimization problems, is one of the variations of the knapsack problem. The MKP has a wide range of real world applications such as cargo loading, selecting projects to fund, budget management, cutting stock, etc. The MKP has been studied quite extensively in the EA community. Due to the constrained nature of the problem, constraint handling techniques gain great importance in the performance of the proposed EA approaches. In this study, the applicability of a generational EA that uses a penalty-based constraint handling technique and a gene locus based, asymmetric, adaptive mutation scheme is explored for the MKP. The effects of the parameters of the explored approach is determined through tests. Further experiments, using large MKP instances from commonly used benchmarks available through the Internet are performed. Comparison tables are given for the performance of the explored approach and other good performing EAs found in literature for the MKP. Results show that performance improves greatly when compared with other penalty-based techniques, but the explored approach is still not the best performer among all. However, unlike the explored technique, the EAs using the other constraint handling techniques require a great amount of extra computational effort and need heuristic information specific to the optimization problem. Based on these observations, and the fact that the performance difference between the explored scheme and the better performers is not too high, research on improving the explored approach is still in progress.
A. Sima Etaner-Uyar, Gülsen Eryigit
GECCO2
2004 A Gene Based Adaptive Mutation Strategy for Genetic Algorithms
A. Sima Etaner-Uyar, Sanem Sariel, Gülsen Eryigit
GECCO (2)3