VLDB 2026 Research / reviewers in the wild / expert
A. Seza Dogruöz
dblp:136/9077
· DBLP profile ↗
19ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0003-2589-5894ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 3 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Are they lovers or friends? Evaluating LLMs' Social Reasoning in English and Korean DialoguesabstractEunsu Kim, Junyeong Park, Juhyun Oh, Kiwoong Park, Seyoung Song, A. Seza Doğruöz, Alice Oh, Najoung Kim. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Eunsu Kim, Junyeong Park, Juhyun Oh, Kiwoong Park, Seyoung Song 0001, A. Seza Dogruöz, Alice Oh, Najoung Kim |
ACL (1) | 6 |
| 2026 | A Typology of Synthetic Datasets for Dialogue Processing in Clinical ContextsabstractSynthetic data sets are used across linguistic domains and NLP tasks, particularly in scenarios where authentic data is limited (or even non-existent). One such domain is that of clinical (healthcare) contexts, where there exist significant and long-standing challenges (e.g., privacy, anonymization, and data governance) which have led to the development of an increasing number of synthetic datasets. One increasingly important category of clinical dataset is that of clinical dialogues which are especially sensitive and difficult to collect, and as such are commonly synthesized. While such synthetic datasets have been shown to be sufficient in some situations, little theory exists to inform how they may be best used and generalized to new applications. In this paper, we provide an overview of how synthetic datasets are created, evaluated and being used for dialogue related tasks in the medical domain. Additionally, we propose a novel typology for use in classifying types and degrees of data synthesis, to facilitate comparison and evaluation. Steven Bedrick, A. Seza Dogruöz, Sergiu Nisioi |
LREC | 2 |
| 2026 | Simple Additions, Substantial Gains: Expanding Scripts, Languages, and Lineage Coverage in URIEL+abstractThe URIEL+ linguistic knowledge base supports multilingual research by encoding languages through geographic, genetic, and typological vectors. However, data sparsity (e.g. missing feature types, incomplete language entries, and limited genealogical coverage) remains prevalent. This limits the usefulness of URIEL+ in cross-lingual transfer, particularly for supporting low-resource languages. To address this sparsity, we extend URIEL+ by introducing script vectors to represent writing system properties for 7,488 languages, integrating Glottolog to add 18,710 additional languages, and expanding lineage imputation for 26,449 languages by propagating typological and script features across genealogies. These improvements reduce feature sparsity by 14% for script vectors, increase language coverage by up to 19,015 languages (1,007%), and boost imputation quality metrics by up to 35%. Our benchmark on cross-lingual transfer tasks (oriented around low-resource languages) shows occasionally divergent performance compared to URIEL+, with performance gains up to 6% in certain setups. Mason Shipton, York Hay Ng, Aditya Khan, Phuong Hanh Hoang, A. Seza Dogruöz, Annie En-Shiun Lee |
LREC | 6 |
| 2025 | URIEL+: Enhancing Linguistic Inclusion and Usability in a Typological and Multilingual Knowledge BaseabstractURIEL is a knowledge base offering geographical, phylogenetic, and typological vector representations for 7970 languages. It includes distance measures between these vectors for 4005 languages, which are accessible via the lang2vec tool. Despite being frequently cited, URIEL is limited in terms of linguistic inclusion and overall usability. To tackle these challenges, we introduce URIEL+, an enhanced version of URIEL and lang2vec that addresses these limitations. In addition to expanding typological feature coverage for 2898 languages, URIEL+ improves the user experience with robust, customizable distance calculations to better suit the needs of users. These upgrades also offer competitive performance on downstream tasks and provide distances that better align with linguistic distance studies. Aditya Armaan Khan, Mason Shipton, David Anugraha, Kaiyao Duan, Phuong Hanh Hoang, Eric Khiu, A. Seza Dogruöz, Annie En-Shiun Lee |
COLING | 7 |
| 2024 | Who Is Bragging More Online? A Large Scale Analysis of Bragging in Social MediaabstractBragging is the act of uttering statements that are likely to be positively viewed by others and it is extensively employed in human communication with the aim to build a positive self-image of oneself. Social media is a natural platform for users to employ bragging in order to gain admiration, respect, attention and followers from their audiences. Yet, little is known about the scale of bragging online and its characteristics. This paper employs computational sociolinguistics methods to conduct the first large scale study of bragging behavior on Twitter (U.S.) by focusing on its overall prevalence, temporal dynamics and impact of demographic factors. Our study shows that the prevalence of bragging decreases over time within the same population of users. In addition, younger, more educated and popular users in the U.S. are more likely to brag. Finally, we conduct an extensive linguistics analysis to unveil specific bragging themes associated with different user traits. Mali Jin, Daniel Preotiuc-Pietro, A. Seza Dogruöz, Nikolaos Aletras |
LREC/COLING | 3 |
| 2024 | Is Spoken Hungarian Low-resource?: A Quantitative Survey of Hungarian Speech Data SetsabstractEven though various speech data sets are available in Hungarian, there is a lack of a general overview about their types and sizes. To fill in this gap, we provide a survey of available data sets in spoken Hungarian in five categories (e.g., monolingual, Hungarian part of multilingual, pathological, child-related and dialectal collections). In total, the estimated size of available data is about 2800 hours (across 7500 speakers) and it represents a rich spoken language diversity. However, the distribution of the data and its alignment to real-life (e.g. speech recognition) tasks is far from optimal indicating the need for additional larger-scale natural language speech data sets. Our survey presents an overview of available data sets for Hungarian explaining their strengths and weaknesses which is useful for researchers working on Hungarian across disciplines. In addition, our survey serves as a starting point towards a unified foundational speech model specific to Hungarian. Péter Mihajlik, Katalin Mády, Anna Kohári, Fruzsina Sára Fruzsina, Gábor Kiss, Tekla Etelka Gráczi, A. Seza Dogruöz |
LREC/COLING | 7 |
| 2023 | Investigating Reproducibility at Interspeech Conferences: A Longitudinal and Comparative PerspectiveabstractReproducibility is a key aspect for scientific advancement across disciplines, and reducing barriers for open science is a focus area for the theme of Interspeech 2023. Availability of source code is one of the indicators that facilitates reproducibility. However, less is known about the rates of reproducibility at Interspeech conferences in comparison to other conferences in the field. In order to fill this gap, we have surveyed 27,717 papers at seven conferences across speech and language processing disciplines. We find that despite having a close number of accepted papers to the other conferences, Interspeech has up to 40% less source code availability. In addition to reporting the difficulties we have encountered during our research, we also provide recommendations and possible directions to increase reproducibility for further studies. Mohammad Arvan, A. Seza Dogruöz, Natalie Parde |
INTERSPEECH | 2 |
| 2023 | The Open-domain Paradox for Chatbots: Common Ground as the Basis for Human-like DialogueabstractThere is a surge in interest in the development of open-domain chatbots, driven by the recent advancements of large language models.The "openness" of the dialogue is expected to be maximized by providing minimal information to the users about the common ground they can expect, including the presumed joint activity.However, evidence suggests that the effect is the opposite.Asking users to "just chat about anything" results in a very narrow form of dialogue, which we refer to as the open-domain paradox.In this position paper, we explain this paradox through the theory of common ground as the basis for human-like communication.Furthermore, we question the assumptions behind open-domain chatbots and identify paths forward for enabling common ground in human-computer dialogue. Gabriel Skantze, A. Seza Dogruöz |
SIGDIAL | 2 |
| 2022 | Automatic Identification and Classification of Bragging in Social MediaabstractBragging is a speech act employed with the goal of constructing a favorable self-image through positive statements about oneself.It is widespread in daily communication and especially popular in social media, where users aim to build a positive image of their persona directly or indirectly.In this paper, we present the first large scale study of bragging in computational linguistics, building on previous research in linguistics and pragmatics.To facilitate this, we introduce a new publicly available data set of tweets annotated for bragging and their types.We empirically evaluate different transformerbased models injected with linguistic information in (a) binary bragging classification, i.e., if tweets contain bragging statements or not; and (b) multi-class bragging type prediction including not bragging.Our results show that our models can predict bragging with macro F1 up to 72.42 and 35.95 in the binary and multi-class classification tasks respectively.Finally, we present an extensive linguistic and error analysis of bragging prediction to guide future research on this topic.1 Mali Jin, Daniel Preotiuc-Pietro, A. Seza Dogruöz, Nikolaos Aletras |
ACL (1) | 3 |
| 2021 | A Survey of Code-switching: Linguistic and Social Perspectives for Language TechnologiesabstractA. Seza Doğruöz, Sunayana Sitaram, Barbara E. Bullock, Almeida Jacqueline Toribio. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. A. Seza Dogruöz, Sunayana Sitaram, Barbara Bullock, Almeida Jacqueline Toribio |
ACL/IJCNLP (1) | 1 |
| 2021 | How "open" are the conversations with open-domain chatbots? A proposal for Speech Event based evaluationabstractOpen-domain chatbots are supposed to converse freely with humans without being restricted to a topic, task or domain.However, the boundaries and/or contents of opendomain conversations are not clear.To clarify the boundaries of "openness", we conduct two studies: First, we classify the types of "speech events" encountered in a chatbot evaluation data set (i.e., Meena by Google) and find that these conversations mainly cover the "small talk" category and exclude the other speech event categories encountered in real life human-human communication.Second, we conduct a small-scale pilot study to generate online conversations covering a wider range of speech event categories between two humans vs. a human and a state-of-the-art chatbot (i.e., Blender by Facebook).A human evaluation of these generated conversations indicates a preference for human-human conversations, since the human-chatbot conversations lack coherence in most speech event categories.Based on these results, we suggest (a) using the term "small talk" instead of "opendomain" for the current chatbots which are not that "open" in terms of conversational abilities yet, and (b) revising the evaluation methods to test the chatbot conversations against other speech events. A. Seza Dogruöz, Gabriel Skantze |
SIGDIAL | 1 |
| 2017 | One Size Does Not Fit All: Profiling Personalized Time-Evolving User BehaviorsabstractGiven the set of social interactions of a user, how can we detect changes in interaction patterns over time? While most previous work has focused on studying network-wide properties and spotting outlier users, the dynamics of individual user interactions remain largely unexplored. This work sets out to explore those dynamics in a way that is minimally invasive to privacy, thus, avoids to rely on the textual content of user posts---except for validation. Our contributions are two-fold. First, in contrast to previous studies, we challenge the use of a fixed interval of observation. We introduce and empirically validate the "Temporal Asymmetry Hypothesis", which states that appropriate observation intervals should vary both among users and over time for the same user. We validate this hypothesis using eight different datasets, including email, messaging, and social networks data. Second, we propose iNET, a comprehensive analytic and visualization framework which provides personalized insights into user behavior and operates in a streaming fashion. iNET learns personalized baseline behaviors of users and uses them to identify events that signify changes in user behavior. We evaluate the effectiveness of iNET by analyzing more than half a million interactions from Facebook users. Labeling of the identified changes in user behavior showed that iNET is able to capture a wide spectrum of exogenous and endogenous events, while the baselines are less diverse in nature and capture only 66% of that spectrum. Furthermore, iNET exhibited the highest precision (95%) compared to all competing approaches. Pravallika Devineni, Evangelos E. Papalexakis, Danai Koutra, A. Seza Dogruöz, Michalis Faloutsos |
ASONAM | 4 |
| 2017 | Integrating Meaning into Quality Evaluation of Machine TranslationabstractOsman Başkaya, Eray Yildiz, Doruk Tunaoğlu, Mustafa Tolga Eren, A. Seza Doğruöz. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017. Osman Baskaya, Eray Yildiz, Doruk Tunaoglu, Mustafa Tolga Eren, A. Seza Dogruöz |
EACL (1) | 5 |
| 2017 | Text based user comments as a signal for automatic language identification of online videosabstractIdentifying the audio language of online videos is crucial for industrial multi-media applications. Automatic speech recognition systems can potentially detect the language of the audio. However, such systems are not available for all languages. Moreover, background noise, music and multi-party conversations make audio language identification hard. Instead, we utilize text based user comments as a new signal to identify audio language of YouTube videos. First, we detect the language of the text based comments. Augmenting this information with video meta-data features, we predict the language of the videos with an accuracy of 97% on a set of publicly available videos. The subject matter discussed in this research is patent pending. A. Seza Dogruöz, Natalia Ponomareva 0001, Sertan Girgin, Reshu Jain, Christoph Oehler |
ICMI | 1 |
| 2016 | Computational Sociolinguistics: A SurveyabstractLanguage is a social phenomenon and variation is inherent to its social nature. Recently, there has been a surge of interest within the computational linguistics (CL) community in the social dimension of language. In this article we present a survey of the emerging field of “computational sociolinguistics” that reflects this increased interest. We aim to provide a comprehensive overview of CL research on sociolinguistic themes, featuring topics such as the relation between language and social identity, language use in social interaction, and multilingual communication. Moreover, we demonstrate the potential for synergy between the research communities involved, by showing how the large-scale data-driven methods that are widely used in CL can complement existing sociolinguistic studies, and how sociolinguistics can inform and challenge the methods and assumptions used in CL studies. We hope to convey the possible benefits of a closer collaboration between the two communities and conclude with a discussion of open challenges. Dong Nguyen 0002, A. Seza Dogruöz, Carolyn P. Rosé, Franciska de Jong |
Comput. Linguistics | 2 |
| 2014 | Why Gender and Age Prediction from Tweets is Hard: Lessons from a Crowdsourcing Experiment
Dong Nguyen 0002, Dolf Trieschnigg, A. Seza Dogruöz, Rilana Gravel, Mariët Theune, Theo Meder, Franciska de Jong |
COLING | 3 |
| 2014 | Modeling the Use of Graffiti Style Features to Signal Social Relations within a Multi-Domain Learning ParadigmabstractIn this paper, we present a series of experiments in which we analyze the usage of graffiti style features for signaling personal gang identification in a large, online street gangs forum, with an accuracy as high as 83% at the gang alliance level and 72% for the specific gang.We then build on that result in predicting how members of different gangs signal the relationship between their gangs within threads where they are interacting with one another, with a predictive accuracy as high as 66% at this thread composition prediction task.Our work demonstrates how graffiti style features signal social identity both in terms of personal group affiliation and between group alliances and oppositions.When we predict thread composition by modeling identity and relationship simultaneously using a multi-domain learning framework paired with a rich feature representation, we achieve significantly higher predictive accuracy than state-of-the-art baselines using one or the other in isolation. Mario Piergallini, A. Seza Dogruöz, Phani Gadde, David Adamson, Carolyn P. Rosé |
EACL | 2 |
| 2014 | Predicting Dialect Variation in Immigrant Contexts Using Light Verb ConstructionsabstractLanguages spoken by immigrants change due to contact with the local languages.Capturing these changes is problematic for current language technologies, which are typically developed for speakers of the standard dialect only.Even when dialectal variants are available for such technologies, we still need to predict which dialect is being used.In this study, we distinguish between the immigrant and the standard dialect of Turkish by focusing on Light Verb Constructions.We experiment with a number of grammatical and contextual features, achieving over 84% accuracy (56% baseline). A. Seza Dogruöz, Preslav Nakov |
EMNLP | 1 |
| 2013 | Word Level Language Identification in Online Multilingual CommunicationabstractMultilingual speakers switch between languages in online and spoken communication.Analyses of large scale multilingual data require automatic language identification at the word level.For our experiments with multilingual online discussions, we first tag the language of individual words using language models and dictionaries.Secondly, we incorporate context to improve the performance.We achieve an accuracy of 98%.Besides word level accuracy, we use two new metrics to evaluate this task. Dong Nguyen 0002, A. Seza Dogruöz |
EMNLP | 2 |