Gözde Gül Sahin

dblp:216/7164 · also Gözde Gül Isgüder, Gözde Gül Isgüder-Sahin · DBLP profile ↗
← Back
15ranked-venue papers
8as first author
7since 2021 · last 2025
0000-0002-0332-1657ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 8 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorTheory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Deep learning architectures and training · 25% Information extraction and text analysis · 19% Learning theory · 14%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning theory
curse of dimensionality
0.612022
On the Rate of Convergence of a Classifier Based on a Transformer Encoder · IEEE Trans. Inf. Theory 2022
Machine learning › Deep learning architectures and training › transformer
transformer encoder
0.612022
On the Rate of Convergence of a Classifier Based on a Transformer Encoder · IEEE Trans. Inf. Theory 2022
Machine learning › Deep learning architectures and training › feedforward neural network
invertible neural network
0.412020
Two Birds with One Stone: Investigating Invertible Neural Networks for Inverse Problems in Morphology · AAAI 2020
Knowledge, reasoning and agents › Knowledge representation and reasoning
metalinguistic reasoning
0.412020
PuzzLing Machines: A Challenge on Learning From Small Data · ACL 2020
Natural language and speech › Language models and text generation
natural language understanding
0.412020
PuzzLing Machines: A Challenge on Learning From Small Data · ACL 2020
Natural language and speech › Information extraction and text analysis
morphological analysis
0.312018
Character-Level Models versus Morphology in Semantic Role Labeling · ACL (1) 2018
Natural language and speech › Information extraction and text analysis
semantic role labeling
0.312018
Character-Level Models versus Morphology in Semantic Role Labeling · ACL (1) 2018
Machine learning › Generative modeling › synthetic data generation
text data augmentation
0.312018
Data Augmentation via Dependency Tree Morphing for Low-Resource Languages · EMNLP 2018
Natural language and speech › Language models and text generation › language modeling
character-level language modeling
0.112018
Character-Level Models versus Morphology in Semantic Role Labeling · ACL (1) 2018
Natural language and speech › Information extraction and text analysis
low-resource NLP
0.112018
Data Augmentation via Dependency Tree Morphing for Low-Resource Languages · EMNLP 2018

Methods — techniques the papers use, named apart from their topics

transformer encoder · 0.6statistical learning theory · 0.6statistical algorithms · 0.4invertible neural network · 0.4deep neural models · 0.4error analysis · 0.3character-level sequence tagging · 0.3
YearPublicationVenuePosition
2025 A Zero-Shot Open-Vocabulary Pipeline for Dialogue Understanding
abstract
Abdulfattah Safa, Gözde Gül Şahin. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Abdulfattah Safa, Gözde Gül Sahin
NAACL (Long Papers)2
2023 Lessons Learned from a Citizen Science Project for Natural Language Processing
abstract
Jan-Christoph Klie, Ji-Ung Lee, Kevin Stowe, Gözde Şahin, Nafise Sadat Moosavi, Luke Bates, Dominic Petrak, Richard Eckart De Castilho, Iryna Gurevych. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Jan-Christoph Klie, Ji-Ung Lee, Kevin Stowe, Gözde Gül Sahin, Nafise Sadat Moosavi, Luke Bates, Dominic Petrak, Richard Eckart de Castilho, Iryna Gurevych
EACL4
2023 MetaQA: Combining Expert Agents for Multi-Skill Question Answering
abstract
The recent explosion of question-answering (QA) datasets and models has increased the interest in the generalization of models across multiple domains and formats by either training on multiple datasets or combining multiple models.Despite the promising results of multidataset models, some domains or QA formats may require specific architectures, and thus the adaptability of these models might be limited.In addition, current approaches for combining models disregard cues such as questionanswer compatibility.In this work, we propose to combine expert agents with a novel, flexible, and training-efficient architecture that considers questions, answer predictions, and answer-prediction confidence scores to select the best answer among a list of answer predictions.Through quantitative and qualitative experiments, we show that our model i) creates a collaboration between agents that outperforms previous multi-agent and multi-dataset approaches, ii) is highly data-efficient to train, and iii) can be adapted to any QA format.We release our code and a dataset of answer predictions from expert agents for 16 QA datasets to foster future research of multi-agent systems 1 .
Haritz Puerto, Gözde Gül Sahin, Iryna Gurevych
EACL2
2023 Benchmarking Procedural Language Understanding for Low-Resource Languages: A Case Study on Turkish
abstract
Arda Uzunoglu, Gözde Şahin. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Arda Uzunoglu, Gözde Gül Sahin
IJCNLP (1)2
2023 Metric-Based In-context Learning: A Case Study in Text Simplification
abstract
In-context learning (ICL) for large language models has proven to be a powerful approach for many natural language processing tasks.However, determining the best method to select examples for ICL is nontrivial as the results can vary greatly depending on the quality, quantity, and order of examples used.In this paper, we conduct a case study on text simplification (TS) to investigate how to select the best and most robust examples for ICL.We propose Metric-Based in-context Learning (MBL) method that utilizes commonly used TS metrics such as SARI, compression ratio, and BERT-Precision for selection.Through an extensive set of experiments with various-sized GPT models on standard TS benchmarks such as TurkCorpus and ASSET, we show that examples selected by the top SARI scores perform the best on larger models such as GPT-175B, while the compression ratio generally performs better on smaller models such as GPT-13B and GPT-6.7B.Furthermore, we demonstrate that MBL is generally robust to example orderings and out-ofdomain test sets, and outperforms strong baselines and state-of-the-art finetuned language models.Finally, we show that the behaviour of large GPT models can be implicitly controlled by the chosen metric.Our research provides a new framework for selecting examples in ICL, and demonstrates its effectiveness in text simplification tasks, breaking new ground for more accurate and efficient NLG systems.
Subhadra Vadlamannati, Gözde Gül Sahin
INLG2
2022 To Augment or Not to Augment? A Comparative Study on Text Augmentation Techniques for Low-Resource NLP
abstract
Abstract Data-hungry deep neural networks have established themselves as the de facto standard for many NLP tasks, including the traditional sequence tagging ones. Despite their state-of-the-art performance on high-resource languages, they still fall behind their statistical counterparts in low-resource scenarios. One methodology to counterattack this problem is text augmentation, that is, generating new synthetic training data points from existing data. Although NLP has recently witnessed several new textual augmentation techniques, the field still lacks a systematic performance analysis on a diverse set of languages and sequence tagging tasks. To fill this gap, we investigate three categories of text augmentation methodologies that perform changes on the syntax (e.g., cropping sub-sentences), token (e.g., random word insertion), and character (e.g., character swapping) levels. We systematically compare the methods on part-of-speech tagging, dependency parsing, and semantic role labeling for a diverse set of language families using various models, including the architectures that rely on pretrained multilingual contextualized language models such as mBERT. Augmentation most significantly improves dependency parsing, followed by part-of-speech tagging and semantic role labeling. We find the experimented techniques to be effective on morphologically rich languages in general rather than analytic languages such as Vietnamese. Our results suggest that the augmentation techniques can further improve over strong baselines based on mBERT, especially for dependency parsing. We identify the character-level methods as the most consistent performers, while synonym replacement and syntactic augmenters provide inconsistent improvements. Finally, we discuss that the results most heavily depend on the task, language pair (e.g., syntactic-level techniques mostly benefit higher-level tasks and morphologically richer languages), and model type (e.g., token-level augmentation provides significant improvements for BPE, while character-level ones give generally higher scores for char and mBERT based models).
Gözde Gül Sahin
Comput. Linguistics1
2022 On the Rate of Convergence of a Classifier Based on a Transformer Encoder
abstract
Pattern recognition based on a high-dimensional predictor is considered. A classifier is defined which is based on a Transformer encoder. The rate of convergence of the misclassification probability of the classifier towards the optimal misclassification probability is analyzed. It is shown that this classifier is able to circumvent the curse of dimensionality provided the a posteriori probability satisfies a suitable hierarchical composition model. Furthermore, the difference between the Transformer classifiers theoretically analyzed in this paper and the ones used in practice today is illustrated by means of classification problems in natural language processing.
Iryna Gurevych, Michael Kohler, Gözde Gül Sahin
IEEE Trans. Inf. Theory3
2020 Two Birds with One Stone: Investigating Invertible Neural Networks for Inverse Problems in Morphology
Gözde Gül Sahin, Iryna Gurevych
AAAI1
2020 PuzzLing Machines: A Challenge on Learning From Small Data
abstract
Deep neural models have repeatedly proved excellent at memorizing surface patterns from large datasets for various ML and NLP benchmarks.They struggle to achieve human-like thinking, however, because they lack the skill of iterative reasoning upon knowledge.To expose this problem in a new light, we introduce a challenge on learning from small data, PuzzLing Machines, which consists of Rosetta Stone puzzles from Linguistic Olympiads for high school students.These puzzles are carefully designed to contain only the minimal amount of parallel text necessary to deduce the form of unseen expressions.Solving them does not require external information (e.g., knowledge bases, visual signals) or linguistic expertise, but meta-linguistic awareness and deductive skills.Our challenge contains around 100 puzzles covering a wide range of linguistic phenomena from 81 languages.We show that both simple statistical algorithms and state-of-the-art deep neural models perform inadequately on this challenge, as expected.We hope that this benchmark, available at https://ukplab.github.io/ PuzzLing-Machines/, inspires further efforts towards a new paradigm in NLP-one that is grounded in human-like reasoning and understanding.
Gözde Gül Sahin, Yova Kementchedjhieva, Phillip Rust, Iryna Gurevych
ACL1
2020 Linguistic Fundamentals for Natural Language Processing II: 100 Essentials from Semantics and Pragmatics
Gözde Gül Sahin
Comput. Linguistics1
2020 LINSPECTOR: Multilingual Probing Tasks for Word Representations
abstract
Despite an ever-growing number of word representation models introduced for a large number of languages, there is a lack of a standardized technique to provide insights into what is captured by these models. Such insights would help the community to get an estimate of the downstream task performance, as well as to design more informed neural architectures, while avoiding extensive experimentation that requires substantial computational resources not all researchers have access to. A recent development in NLP is to use simple classification tasks, also called probing tasks, that test for a single linguistic feature such as part-of-speech. Existing studies mostly focus on exploring the linguistic information encoded by the continuous representations of English text. However, from a typological perspective the morphologically poor English is rather an outlier: The information encoded by the word order and function words in English is often stored on a subword, morphological level in other languages. To address this, we introduce 15 type-level probing tasks such as case marking, possession, word length, morphological tag count, and pseudoword identification for 24 languages. We present a reusable methodology for creation and evaluation of such tests in a multilingual setting, which is challenging because of a lack of resources, lower quality of tools, and differences among languages. We then present experiments on several diverse multilingual word embedding models, in which we relate the probing task performance for a diverse set of languages to a range of five classic NLP tasks: POS-tagging, dependency parsing, semantic role labeling, named entity recognition, and natural language inference. We find that a number of probing tests have significantly high positive correlation to the downstream tasks, especially for morphologically rich languages. We show that our tests can be used to explore word embeddings or black-box neural models for linguistic cues in a multilingual setting. We release the probing data sets and the evaluation suite LINSPECTOR with https://github.com/UKPLab/linspector .
Gözde Gül Sahin, Clara Vania, Ilia Kuznetsov, Iryna Gurevych
Comput. Linguistics1
2018 Character-Level Models versus Morphology in Semantic Role Labeling
abstract
Character-level models have become a popular approach specially for their accessibility and ability to handle unseen data.However, little is known on their ability to reveal the underlying morphological structure of a word, which is a crucial skill for high-level semantic analysis tasks, such as semantic role labeling (SRL).In this work, we train various types of SRL models that use word, character and morphology level information and analyze how performance of characters compare to words and morphology for several languages.We conduct an in-depth error analysis for each morphological typology and analyze the strengths and limitations of character-level models that relate to out-of-domain data, training data size, long range dependencies and model complexity.Our exhaustive analyses shed light on important characteristics of character-level models and their semantic capability.
Gözde Gül Sahin, Mark Steedman
ACL (1)1
2018 Data Augmentation via Dependency Tree Morphing for Low-Resource Languages
abstract
Neural NLP systems achieve high scores in the presence of sizable training dataset.Lack of such datasets leads to poor system performances in the case low-resource languages.We present two simple text augmentation techniques using dependency trees, inspired from image processing.We "crop" sentences by removing dependency links, and we "rotate" sentences by moving the tree fragments around the root.We apply these techniques to augment the training sets of low-resource languages in Universal Dependencies project.We implement a character-level sequence tagging model and evaluate the augmented datasets on part-of-speech tagging task.We show that crop and rotate provides improvements over the models trained with non-augmented data for majority of the languages, especially for languages with rich case marking systems.
Gözde Gül Sahin, Mark Steedman
EMNLP1
2016 Verb Sense Annotation for Turkish PropBank via Crowdsourcing
Gözde Gül Sahin
CICLing (1)1
2009 A New 3-D Automated Computational Method to Evaluate In-Stent Neointimal Hyperplasia in In-Vivo Intravascular Optical Coherence Tomography Pullbacks
Serhan Gurmeric, Gözde Gül Sahin, Stephane G. Carlier, Gozde Unal
MICCAI (1)2