VLDB 2026 Research / reviewers in the wild / expert
Gözde Gül Sahin
dblp:216/7164 · also Gözde Gül Isgüder, Gözde Gül Isgüder-Sahin
· DBLP profile ↗
15ranked-venue papers
8as first author
7since 2021 · last 2025
0000-0002-0332-1657ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 8 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorTheory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Deep learning architectures and training · 25% Information extraction and text analysis · 19% Learning theory · 14% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Learning theory
curse of dimensionality |
0.6 | 1 | 2022 | On the Rate of Convergence of a Classifier Based on a Transformer Encoder · IEEE Trans. Inf. Theory 2022 |
Machine learning › Deep learning architectures and training › transformer
transformer encoder |
0.6 | 1 | 2022 | On the Rate of Convergence of a Classifier Based on a Transformer Encoder · IEEE Trans. Inf. Theory 2022 |
Machine learning › Deep learning architectures and training › feedforward neural network
invertible neural network |
0.4 | 1 | 2020 | Two Birds with One Stone: Investigating Invertible Neural Networks for Inverse Problems in Morphology · AAAI 2020 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
metalinguistic reasoning |
0.4 | 1 | 2020 | PuzzLing Machines: A Challenge on Learning From Small Data · ACL 2020 |
Natural language and speech › Language models and text generation
natural language understanding |
0.4 | 1 | 2020 | PuzzLing Machines: A Challenge on Learning From Small Data · ACL 2020 |
Natural language and speech › Information extraction and text analysis
morphological analysis |
0.3 | 1 | 2018 | Character-Level Models versus Morphology in Semantic Role Labeling · ACL (1) 2018 |
Natural language and speech › Information extraction and text analysis
semantic role labeling |
0.3 | 1 | 2018 | Character-Level Models versus Morphology in Semantic Role Labeling · ACL (1) 2018 |
Machine learning › Generative modeling › synthetic data generation
text data augmentation |
0.3 | 1 | 2018 | Data Augmentation via Dependency Tree Morphing for Low-Resource Languages · EMNLP 2018 |
Natural language and speech › Language models and text generation › language modeling
character-level language modeling |
0.1 | 1 | 2018 | Character-Level Models versus Morphology in Semantic Role Labeling · ACL (1) 2018 |
Natural language and speech › Information extraction and text analysis
low-resource NLP |
0.1 | 1 | 2018 | Data Augmentation via Dependency Tree Morphing for Low-Resource Languages · EMNLP 2018 |
Methods — techniques the papers use, named apart from their topics
transformer encoder · 0.6statistical learning theory · 0.6statistical algorithms · 0.4invertible neural network · 0.4deep neural models · 0.4error analysis · 0.3character-level sequence tagging · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Zero-Shot Open-Vocabulary Pipeline for Dialogue UnderstandingabstractAbdulfattah Safa, Gözde Gül Şahin. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Abdulfattah Safa, Gözde Gül Sahin |
NAACL (Long Papers) | 2 |
| 2023 | Lessons Learned from a Citizen Science Project for Natural Language ProcessingabstractJan-Christoph Klie, Ji-Ung Lee, Kevin Stowe, Gözde Şahin, Nafise Sadat Moosavi, Luke Bates, Dominic Petrak, Richard Eckart De Castilho, Iryna Gurevych. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Jan-Christoph Klie, Ji-Ung Lee, Kevin Stowe, Gözde Gül Sahin, Nafise Sadat Moosavi, Luke Bates, Dominic Petrak, Richard Eckart de Castilho, Iryna Gurevych |
EACL | 4 |
| 2023 | MetaQA: Combining Expert Agents for Multi-Skill Question AnsweringabstractThe recent explosion of question-answering (QA) datasets and models has increased the interest in the generalization of models across multiple domains and formats by either training on multiple datasets or combining multiple models.Despite the promising results of multidataset models, some domains or QA formats may require specific architectures, and thus the adaptability of these models might be limited.In addition, current approaches for combining models disregard cues such as questionanswer compatibility.In this work, we propose to combine expert agents with a novel, flexible, and training-efficient architecture that considers questions, answer predictions, and answer-prediction confidence scores to select the best answer among a list of answer predictions.Through quantitative and qualitative experiments, we show that our model i) creates a collaboration between agents that outperforms previous multi-agent and multi-dataset approaches, ii) is highly data-efficient to train, and iii) can be adapted to any QA format.We release our code and a dataset of answer predictions from expert agents for 16 QA datasets to foster future research of multi-agent systems 1 . Haritz Puerto, Gözde Gül Sahin, Iryna Gurevych |
EACL | 2 |
| 2023 | Benchmarking Procedural Language Understanding for Low-Resource Languages: A Case Study on TurkishabstractArda Uzunoglu, Gözde Şahin. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Arda Uzunoglu, Gözde Gül Sahin |
IJCNLP (1) | 2 |
| 2023 | Metric-Based In-context Learning: A Case Study in Text SimplificationabstractIn-context learning (ICL) for large language models has proven to be a powerful approach for many natural language processing tasks.However, determining the best method to select examples for ICL is nontrivial as the results can vary greatly depending on the quality, quantity, and order of examples used.In this paper, we conduct a case study on text simplification (TS) to investigate how to select the best and most robust examples for ICL.We propose Metric-Based in-context Learning (MBL) method that utilizes commonly used TS metrics such as SARI, compression ratio, and BERT-Precision for selection.Through an extensive set of experiments with various-sized GPT models on standard TS benchmarks such as TurkCorpus and ASSET, we show that examples selected by the top SARI scores perform the best on larger models such as GPT-175B, while the compression ratio generally performs better on smaller models such as GPT-13B and GPT-6.7B.Furthermore, we demonstrate that MBL is generally robust to example orderings and out-ofdomain test sets, and outperforms strong baselines and state-of-the-art finetuned language models.Finally, we show that the behaviour of large GPT models can be implicitly controlled by the chosen metric.Our research provides a new framework for selecting examples in ICL, and demonstrates its effectiveness in text simplification tasks, breaking new ground for more accurate and efficient NLG systems. Subhadra Vadlamannati, Gözde Gül Sahin |
INLG | 2 |
| 2022 | To Augment or Not to Augment? A Comparative Study on Text Augmentation Techniques for Low-Resource NLPabstractAbstract Data-hungry deep neural networks have established themselves as the de facto standard for many NLP tasks, including the traditional sequence tagging ones. Despite their state-of-the-art performance on high-resource languages, they still fall behind their statistical counterparts in low-resource scenarios. One methodology to counterattack this problem is text augmentation, that is, generating new synthetic training data points from existing data. Although NLP has recently witnessed several new textual augmentation techniques, the field still lacks a systematic performance analysis on a diverse set of languages and sequence tagging tasks. To fill this gap, we investigate three categories of text augmentation methodologies that perform changes on the syntax (e.g., cropping sub-sentences), token (e.g., random word insertion), and character (e.g., character swapping) levels. We systematically compare the methods on part-of-speech tagging, dependency parsing, and semantic role labeling for a diverse set of language families using various models, including the architectures that rely on pretrained multilingual contextualized language models such as mBERT. Augmentation most significantly improves dependency parsing, followed by part-of-speech tagging and semantic role labeling. We find the experimented techniques to be effective on morphologically rich languages in general rather than analytic languages such as Vietnamese. Our results suggest that the augmentation techniques can further improve over strong baselines based on mBERT, especially for dependency parsing. We identify the character-level methods as the most consistent performers, while synonym replacement and syntactic augmenters provide inconsistent improvements. Finally, we discuss that the results most heavily depend on the task, language pair (e.g., syntactic-level techniques mostly benefit higher-level tasks and morphologically richer languages), and model type (e.g., token-level augmentation provides significant improvements for BPE, while character-level ones give generally higher scores for char and mBERT based models). Gözde Gül Sahin |
Comput. Linguistics | 1 |
| 2022 | On the Rate of Convergence of a Classifier Based on a Transformer EncoderabstractPattern recognition based on a high-dimensional predictor is considered. A classifier is defined which is based on a Transformer encoder. The rate of convergence of the misclassification probability of the classifier towards the optimal misclassification probability is analyzed. It is shown that this classifier is able to circumvent the curse of dimensionality provided the a posteriori probability satisfies a suitable hierarchical composition model. Furthermore, the difference between the Transformer classifiers theoretically analyzed in this paper and the ones used in practice today is illustrated by means of classification problems in natural language processing. Iryna Gurevych, Michael Kohler, Gözde Gül Sahin |
IEEE Trans. Inf. Theory | 3 |
| 2020 | Two Birds with One Stone: Investigating Invertible Neural Networks for Inverse Problems in Morphology
Gözde Gül Sahin, Iryna Gurevych |
AAAI | 1 |
| 2020 | PuzzLing Machines: A Challenge on Learning From Small DataabstractDeep neural models have repeatedly proved excellent at memorizing surface patterns from large datasets for various ML and NLP benchmarks.They struggle to achieve human-like thinking, however, because they lack the skill of iterative reasoning upon knowledge.To expose this problem in a new light, we introduce a challenge on learning from small data, PuzzLing Machines, which consists of Rosetta Stone puzzles from Linguistic Olympiads for high school students.These puzzles are carefully designed to contain only the minimal amount of parallel text necessary to deduce the form of unseen expressions.Solving them does not require external information (e.g., knowledge bases, visual signals) or linguistic expertise, but meta-linguistic awareness and deductive skills.Our challenge contains around 100 puzzles covering a wide range of linguistic phenomena from 81 languages.We show that both simple statistical algorithms and state-of-the-art deep neural models perform inadequately on this challenge, as expected.We hope that this benchmark, available at https://ukplab.github.io/ PuzzLing-Machines/, inspires further efforts towards a new paradigm in NLP-one that is grounded in human-like reasoning and understanding. Gözde Gül Sahin, Yova Kementchedjhieva, Phillip Rust, Iryna Gurevych |
ACL | 1 |
| 2020 | Linguistic Fundamentals for Natural Language Processing II: 100 Essentials from Semantics and Pragmatics
Gözde Gül Sahin |
Comput. Linguistics | 1 |
| 2020 | LINSPECTOR: Multilingual Probing Tasks for Word RepresentationsabstractDespite an ever-growing number of word representation models introduced for a large number of languages, there is a lack of a standardized technique to provide insights into what is captured by these models. Such insights would help the community to get an estimate of the downstream task performance, as well as to design more informed neural architectures, while avoiding extensive experimentation that requires substantial computational resources not all researchers have access to. A recent development in NLP is to use simple classification tasks, also called probing tasks, that test for a single linguistic feature such as part-of-speech. Existing studies mostly focus on exploring the linguistic information encoded by the continuous representations of English text. However, from a typological perspective the morphologically poor English is rather an outlier: The information encoded by the word order and function words in English is often stored on a subword, morphological level in other languages. To address this, we introduce 15 type-level probing tasks such as case marking, possession, word length, morphological tag count, and pseudoword identification for 24 languages. We present a reusable methodology for creation and evaluation of such tests in a multilingual setting, which is challenging because of a lack of resources, lower quality of tools, and differences among languages. We then present experiments on several diverse multilingual word embedding models, in which we relate the probing task performance for a diverse set of languages to a range of five classic NLP tasks: POS-tagging, dependency parsing, semantic role labeling, named entity recognition, and natural language inference. We find that a number of probing tests have significantly high positive correlation to the downstream tasks, especially for morphologically rich languages. We show that our tests can be used to explore word embeddings or black-box neural models for linguistic cues in a multilingual setting. We release the probing data sets and the evaluation suite LINSPECTOR with https://github.com/UKPLab/linspector . Gözde Gül Sahin, Clara Vania, Ilia Kuznetsov, Iryna Gurevych |
Comput. Linguistics | 1 |
| 2018 | Character-Level Models versus Morphology in Semantic Role LabelingabstractCharacter-level models have become a popular approach specially for their accessibility and ability to handle unseen data.However, little is known on their ability to reveal the underlying morphological structure of a word, which is a crucial skill for high-level semantic analysis tasks, such as semantic role labeling (SRL).In this work, we train various types of SRL models that use word, character and morphology level information and analyze how performance of characters compare to words and morphology for several languages.We conduct an in-depth error analysis for each morphological typology and analyze the strengths and limitations of character-level models that relate to out-of-domain data, training data size, long range dependencies and model complexity.Our exhaustive analyses shed light on important characteristics of character-level models and their semantic capability. Gözde Gül Sahin, Mark Steedman |
ACL (1) | 1 |
| 2018 | Data Augmentation via Dependency Tree Morphing for Low-Resource LanguagesabstractNeural NLP systems achieve high scores in the presence of sizable training dataset.Lack of such datasets leads to poor system performances in the case low-resource languages.We present two simple text augmentation techniques using dependency trees, inspired from image processing.We "crop" sentences by removing dependency links, and we "rotate" sentences by moving the tree fragments around the root.We apply these techniques to augment the training sets of low-resource languages in Universal Dependencies project.We implement a character-level sequence tagging model and evaluate the augmented datasets on part-of-speech tagging task.We show that crop and rotate provides improvements over the models trained with non-augmented data for majority of the languages, especially for languages with rich case marking systems. Gözde Gül Sahin, Mark Steedman |
EMNLP | 1 |
| 2016 | Verb Sense Annotation for Turkish PropBank via Crowdsourcing
Gözde Gül Sahin |
CICLing (1) | 1 |
| 2009 | A New 3-D Automated Computational Method to Evaluate In-Stent Neointimal Hyperplasia in In-Vivo Intravascular Optical Coherence Tomography Pullbacks
Serhan Gurmeric, Gözde Gül Sahin, Stephane G. Carlier, Gozde Unal |
MICCAI (1) | 2 |