Dong Nguyen 0002

dblp:91/102-2 · DBLP profile ↗
← Back
34ranked-venue papers
14as first author
13since 2021 · last 2025
0000-0002-6062-3117ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 11 first-author · 12 since 2021Databases, data management, data science and information retrieval · 10 · 5 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author
YearPublicationVenuePosition
2025 Disentangling the Roles of Representation and Selection in Data Pruning
abstract
Data pruning, selecting small but impactful subsets, offers a promising way to efficiently scale NLP model training. However, existing methods often involve many different design choices, which have not been systematically studied. This limits future developments. In this work, we decompose data pruning into two key components: the data representation and the selection algorithm, and we systematically analyze their influence on the selection of instances. Our theoretical and empirical results highlight the crucial role of representations: better representations, e.g., training gradients, generally lead to a better selection of instances, regardless of the chosen selection algorithm. Furthermore, different selection algorithms excel in different settings, and none consistently outperforms the others. Moreover, the selection algorithms do not always align with their intended objectives: for example, algorithms designed for the same objective can select drastically different instances, highlighting the need for careful evaluation.
Yupei Du, Yingjin Song, Hugh Mee Wong, Daniil Ignatev, Albert Gatt, Dong Nguyen 0002
ACL (1)6
2025 FTFT: Efficient and Robust Fine-Tuning by Transferring Training Dynamics
abstract
Despite the massive success of fine-tuning Pre-trained Language Models (PLMs), they remain susceptible to out-of-distribution input. Dataset cartography is a simple yet effective dual-model approach that improves the robustness of fine-tuned PLMs. It involves fine-tuning a model on the original training set (i.e. reference model), selecting a subset of important training instances based on the training dynamics, % of the reference model, and fine-tuning again only on these selected examples (i.e. main model). However, this approach requires fine-tuning the same model twice, which is computationally expensive for large PLMs. In this paper, we show that 1) training dynamics are highly transferable across model sizes and pre-training methods, and that 2) fine-tuning main models using these selected training instances achieves higher training efficiency than empirical risk minimization (ERM). Building on these observations, we propose a novel fine-tuning approach: Fine-Tuning by transFerring Training dynamics (FTFT). Compared with dataset cartography, FTFT uses more efficient reference models and aggressive early stopping. FTFT achieves robustness improvements over ERM while lowering the training cost by up to ~50%
Yupei Du, Albert Gatt, Dong Nguyen 0002
COLING3
2025 We Need to Measure Data Diversity in NLP - Better and Broader
abstract
Although diversity in NLP datasets has received growing attention, the question of how to measure it remains largely underexplored.This opinion paper examines the conceptual and methodological challenges of measuring data diversity and argues that interdisciplinary perspectives are essential for developing more fine-grained and valid measures.
Dong Nguyen 0002, Esther Ploeger
EMNLP1
2024 What's Mine becomes Yours: Defining, Annotating and Detecting Context-Dependent Paraphrases in News Interview Dialogs
abstract
Best practices for high conflict conversations like counseling or customer support almost always include recommendations to paraphrase the previous speaker.Although paraphrase classification has received widespread attention in NLP, paraphrases are usually considered independent from context, and common models and datasets are not applicable to dialog settings.In this work, we investigate paraphrases across turns in dialog (e.g., Speaker 1: "That book is mine."becomes Speaker 2: "That book is yours.").We provide an operationalization of context-dependent paraphrases, and develop a training for crowd-workers to classify paraphrases in dialog.We introduce ContextDeP, a dataset with utterance pairs from NPR and CNN news interviews annotated for contextdependent paraphrases.To enable analysis on label variation, the dataset contains 5,581 annotations on 600 utterance pairs.We present promising results with in-context learning and with token classification models for automatic paraphrase detection in dialog.What? Shortened Examples Clear Contextual Equivalence ⊆
Anna Wegmann, Tijs A. van den Broek, Dong Nguyen 0002
EMNLP3
2024 General-Purpose User Modeling with Behavioral Logs: A Snapchat Case Study
abstract
Learning general-purpose user representations based on user behavioral logs is an increasingly popular user modeling approach. It benefits from easily available, privacy-friendly yet expressive data, and does not require extensive re-tuning of the upstream user model for different downstream tasks. While this approach has shown promise in search engines and e-commerce applications, its fit for instant messaging platforms, a cornerstone of modern digital communication, remains largely uncharted. We explore this research gap using Snapchat data as a case study. Specifically, we implement a Transformer-based user model with customized training objectives and show that the model can produce high-quality user representations across a broad range of evaluation tasks, among which we introduce three new downstream tasks that concern pivotal topics in user research: user safety, engagement and churn. We also tackle the challenge of efficient extrapolation of long sequences at inference time, by applying a novel positional encoding method.
Qixiang Fang, Zhihan Zhou 0001, Francesco Barbieri, Yozen Liu, Leonardo Neves, Dong Nguyen 0002, Daniel L. Oberski, Maarten W. Bos, Ron Dotsch
SIGIR6
2023 Measuring the Instability of Fine-Tuning
abstract
Fine-tuning pre-trained language models on downstream tasks with varying random seeds has been shown to be unstable, especially on small datasets.Many previous studies have investigated this instability and proposed methods to mitigate it.However, most studies only used the standard deviation of performance scores (SD) as their measure, which is a narrow characterization of instability.In this paper, we analyze SD and six other measures quantifying instability at different levels of granularity.Moreover, we propose a systematic framework to evaluate the validity of these measures.Finally, we analyze the consistency and difference between different measures by reassessing existing instability mitigation methods.We hope our results will inform the development of better measurements of fine-tuning instability. 1
Yupei Du, Dong Nguyen 0002
ACL (1)2
2023 Perceived Algorithmic Fairness using Organizational Justice Theory: An Empirical Case Study on Algorithmic Hiring
abstract
Growing concerns about the fairness of algorithmic decision-making systems have prompted a proliferation of mathematical formulations aimed at remedying algorithmic bias. Yet, integrating mathematical fairness alone into algorithms is insufficient to ensure their acceptance, trust, and support by humans. It is also essential to understand what humans perceive as fair. In this study, we, therefore, conduct an empirical user study into crowdworkers’ algorithmic fairness perceptions, focusing on algorithmic hiring. We build on perspectives from organizational justice theory, which categorizes fairness into distributive, procedural, and interactional components. By doing so, we find that algorithmic fairness perceptions are higher when crowdworkers are provided not only with information about the algorithmic outcome but also about the decision-making process. Remarkably, we observe this effect even when the decision-making process can be considered unfair, when gender, a sensitive attribute, is used as a main feature. By showing realistic trade-offs between fairness criteria, we moreover find a preference for equalizing false negatives over equalizing selection rates amongst groups. Our findings highlight the importance of considering all components of algorithmic fairness, rather than solely treating it as an outcome distribution problem. Importantly, our study contributes to the literature on the connection between mathematical– and perceived algorithmic fairness, and highlights the potential benefits of leveraging organizational justice theory to enhance the evaluation of perceived algorithmic fairness.
Guusje Juijn, Niya Stoimenova, João Reis 0001, Dong Nguyen 0002
AIES4
2022 Template-based Abstractive Microblog Opinion Summarisation
abstract
Abstract We introduce the task of microblog opinion summarization (MOS) and share a dataset of 3100 gold-standard opinion summaries to facilitate research in this domain. The dataset contains summaries of tweets spanning a 2-year period and covers more topics than any other public Twitter summarization dataset. Summaries are abstractive in nature and have been created by journalists skilled in summarizing news articles following a template separating factual information (main story) from author opinions. Our method differs from previous work on generating gold-standard summaries from social media, which usually involves selecting representative posts and thus favors extractive summarization models. To showcase the dataset’s utility and challenges, we benchmark a range of abstractive and extractive state-of-the-art summarization models and achieve good performance, with the former outperforming the latter. We also show that fine-tuning is necessary to improve performance and investigate the benefits of using different sample sizes.
Iman Munire Bilal, Bo Wang 0034, Adam Tsakalidis, Dong Nguyen 0002, Rob Procter, Maria Liakata
Trans. Assoc. Comput. Linguistics4
2021 HateCheck: Functional Tests for Hate Speech Detection Models
abstract
Detecting online hate is a difficult task that even state-of-the-art models struggle with. Typically, hate speech detection models are evaluated by measuring their performance on held-out test data using metrics such as accuracy and F1 score. However, this approach makes it difficult to identify specific model weak points. It also risks overestimating generalisable model performance due to increasingly well-evidenced systematic gaps and biases in hate speech datasets. To enable more targeted diagnostic insights, we introduce HateCheck, a suite of functional tests for hate speech detection models. We specify 29 model functionalities motivated by a review of previous research and a series of interviews with civil society stakeholders. We craft test cases for each functionality and validate their quality through a structured annotation process. To illustrate HateCheck's utility, we test near-state-of-the-art transformer models as well as two popular commercial models, revealing critical model weaknesses.
Paul Röttger, Bertie Vidgen, Dong Nguyen 0002, Zeerak Talat, Helen Z. Margetts, Janet B. Pierrehumbert
ACL/IJCNLP (1)3
2021 Assessing the Reliability of Word Embedding Gender Bias Measures
abstract
Various measures have been proposed to quantify human-like social biases in word embeddings.However, bias scores based on these measures can suffer from measurement error.One indication of measurement quality is reliability, concerning the extent to which a measure produces consistent results.In this paper, we assess three types of reliability of word embedding gender bias measures, namely testretest reliability, inter-rater consistency and internal consistency.Specifically, we investigate the consistency of bias scores across different choices of random seeds, scoring rules and words.Furthermore, we analyse the effects of various factors on these measures' reliability scores.Our findings inform better design of word embedding gender bias measures.Moreover, we urge researchers to be more critical about the application of such measures.1
Yupei Du, Qixiang Fang, Dong Nguyen 0002
EMNLP (1)3
2021 Does It Capture STEL? A Modular, Similarity-based Linguistic Style Evaluation Framework
abstract
Style is an integral part of natural language.However, evaluation methods for style measures are rare, often task-specific and usually do not control for content.We propose the modular, fine-grained and contentcontrolled similarity-based STyle EvaLuation framework (STEL) to test the performance of any model that can compare two sentences on style.We illustrate STEL with two general dimensions of style (formal/informal and simple/complex) as well as two specific characteristics of style (contrac'tion and numb3r substitution).We find that BERT-based methods outperform simple versions of commonly used style measures like 3-grams, punctuation frequency and LIWC-based approaches.We invite the addition of further tasks and task instances to STEL and hope to facilitate the improvement of style-sensitive measures.
Anna Wegmann, Dong Nguyen 0002
EMNLP (1)2
2021 On learning and representing social meaning in NLP: a sociolinguistic perspective
abstract
The field of NLP has made substantial progress in building meaning representations.However, an important aspect of linguistic meaning, social meaning, has been largely overlooked.We introduce the concept of social meaning to NLP and discuss how insights from sociolinguistics can inform work on representation learning in NLP.We also identify key challenges for this new line of research.
Dong Nguyen 0002, Laura Rosseel, Jack Grieve
NAACL-HLT1
2021 Introducing CAD: the Contextual Abuse Dataset
abstract
Bertie Vidgen, Dong Nguyen, Helen Margetts, Patricia Rossini, Rebekah Tromble. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Bertie Vidgen, Dong Nguyen 0002, Helen Z. Margetts, Patrícia G. C. Rossini, Rebekah Tromble
NAACL-HLT2
2020 tBERT: Topic Models and BERT Joining Forces for Semantic Similarity Detection
abstract
Semantic similarity detection is a fundamental task in natural language understanding.Adding topic information has been useful for previous feature-engineered semantic similarity models as well as neural models for other tasks.There is currently no standard way of combining topics with pretrained contextual representations such as BERT.We propose a novel topic-informed BERT-based architecture for pairwise semantic similarity detection and show that our model improves performance over strong neural baselines across a variety of English language datasets.We find that the addition of topics to BERT helps particularly with resolving domain-specific cases.
Nicole Peinelt, Dong Nguyen 0002, Maria Liakata
ACL2
2020 Do Word Embeddings Capture Spelling Variation?
abstract
Analyses of word embeddings have primarily focused on semantic and syntactic properties.However, word embeddings have the potential to encode other properties as well.In this paper, we propose a new perspective on the analysis of word embeddings by focusing on spelling variation.In social media, spelling variation is abundant and often socially meaningful.Here, we analyze word embeddings trained on Twitter and Reddit data.We present three analyses using pairs of word forms covering seven types of spelling variation in English.Taken together, our results show that word embeddings encode spelling variation patterns of various types to some extent, even embeddings trained using the skipgram model which does not take spelling into account.Our results also suggest a link between the intentionality of the variation and the distance of the non-conventional spellings to their conventional spellings.
Dong Nguyen 0002, Jack Grieve
COLING1
2019 Aiming beyond the Obvious: Identifying Non-Obvious Cases in Semantic Similarity Datasets
abstract
Existing datasets for scoring text pairs in terms of semantic similarity contain instances whose resolution differs according to the degree of difficulty.This paper proposes to distinguish obvious from non-obvious text pairs based on superficial lexical overlap and ground-truth labels.We characterise existing datasets in terms of containing difficult cases and find that recently proposed models struggle to capture the non-obvious cases of semantic similarity.We describe metrics that emphasise cases of similarity which require more complex inference and propose that these are used for evaluating systems for semantic similarity.
Nicole Peinelt, Maria Liakata, Dong Nguyen 0002
ACL (1)3
2019 Room to Glo: A Systematic Comparison of Semantic Change Detection Approaches with Word Embeddings
abstract
Philippa Shoemark, Farhana Ferdousi Liza, Dong Nguyen, Scott Hale, Barbara McGillivray. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Philippa Shoemark, Farhana Ferdousi Liza, Dong Nguyen 0002, Scott A. Hale, Barbara McGillivray
EMNLP/IJCNLP (1)3
2018 Comparing Automatic and Human Evaluation of Local Explanations for Text Classification
abstract
Text classification models are becoming increasingly complex and opaque, however for many applications it is essential that the models are interpretable.Recently, a variety of approaches have been proposed for generating local explanations.While robust evaluations are needed to drive further progress, so far it is unclear which evaluation approaches are suitable.This paper is a first step towards more robust evaluations of local explanations.We evaluate a variety of local explanation approaches using automatic measures based on word deletion.Furthermore, we show that an evaluation using a crowdsourcing experiment correlates moderately with these automatic measures and that a variety of other factors also impact the human judgements.
Dong Nguyen 0002
NAACL-HLT1
2017 A Kernel Independence Test for Geographical Language Variation
abstract
Quantifying the degree of spatial dependence for linguistic variables is a key task for analyzing dialectal variation. However, existing approaches have important drawbacks. First, they are based on parametric models of dependence, which limits their power in cases where the underlying parametric assumptions are violated. Second, they are not applicable to all types of linguistic data: Some approaches apply only to frequencies, others to boolean indicators of whether a linguistic variable is present. We present a new method for measuring geographical language variation, which solves both of these problems. Our approach builds on Reproducing Kernel Hilbert Space (RKHS) representations for nonparametric statistics, and takes the form of a test statistic that is computed from pairs of individual geotagged observations without aggregation into predefined geographical bins. We compare this test with prior work using synthetic data as well as a diverse set of real data sets: a corpus of Dutch tweets, a Dutch syntactic atlas, and a data set of letters to the editor in North American newspapers. Our proposed test is shown to support robust inferences across a broad range of scenarios and types of data.
Dong Nguyen 0002, Jacob Eisenstein
Comput. Linguistics1
2016 Computational Sociolinguistics: A Survey
abstract
Language is a social phenomenon and variation is inherent to its social nature. Recently, there has been a surge of interest within the computational linguistics (CL) community in the social dimension of language. In this article we present a survey of the emerging field of “computational sociolinguistics” that reflects this increased interest. We aim to provide a comprehensive overview of CL research on sociolinguistic themes, featuring topics such as the relation between language and social identity, language use in social interaction, and multilingual communication. Moreover, we demonstrate the potential for synergy between the research communities involved, by showing how the large-scale data-driven methods that are widely used in CL can complement existing sociolinguistic studies, and how sociolinguistics can inform and challenge the methods and assumptions used in CL studies. We hope to convey the possible benefits of a closer collaboration between the two communities and conclude with a discussion of open challenges.
Dong Nguyen 0002, A. Seza Dogruöz, Carolyn P. Rosé, Franciska de Jong
Comput. Linguistics1
2016 Predicting relevance based on assessor disagreement: analysis and practical applications for search evaluation
Thomas Demeester, Robin Aly, Djoerd Hiemstra, Dong Nguyen 0002, Chris Develder
Inf. Retr. J.4
2015 #SupportTheCause: Identifying Motivations to Participate in Online Health Campaigns
abstract
We consider the task of automatically identifying participants' motivations in the public health campaign Movember and investigate the impact of the different motivations on the amount of campaign donations raised.Our classification scheme is based on the Social Identity Model of Collective Action (van Zomeren et al., 2008).We find that automatic classification based on Movember profiles is fairly accurate, while automatic classification based on tweets is challenging.Using our classifier, we find a strong relation between types of motivations and donations.Our study is a first step towards scaling-up collective action research methods.
Dong Nguyen 0002, Tijs A. van den Broek, Claudia Hauff, Djoerd Hiemstra, Michel L. Ehrenhard
EMNLP1
2015 Audience and the Use of Minority Languages on Twitter
Dong Nguyen 0002, Dolf Trieschnigg, Leonie Cornips
ICWSM1
2014 Using Crowdsourcing to Investigate Perception of Narrative Similarity
abstract
For many applications measuring the similarity between documents is essential. However, little is known about how users perceive similarity between documents. This paper presents the first large-scale empirical study that investigates perception of narrative similarity using crowdsourcing. As a dataset we use a large collection of Dutch folk narratives. We study the perception of narrative similarity by both experts and non-experts by analyzing their similarity ratings and motivations for these ratings. While experts focus mostly on the plot, characters and themes of narratives, non-experts also pay attention to dimensions such as genre and style. Our results show that a more nuanced view is needed of narrative similarity than captured by story types, a concept used by scholars to group similar folk narratives. We also evaluate to what extent unsupervised and supervised models correspond with how humans perceive narrative similarity.
Dong Nguyen 0002, Dolf Trieschnigg, Mariët Theune
CIKM1
2014 Aligning Vertical Collection Relevance with User Intent
abstract
Selecting and aggregating different types of content from multiple vertical search engines is becoming popular in web search. The user vertical intent, the verticals the user expects to be relevant for a particular information need, might not correspond to the vertical collection relevance, the verticals containing the most relevant content. In this work we propose different approaches to define the set of relevant verticals based on document judgments. We correlate the collection-based relevant verticals obtained from these approaches to the real user vertical intent, and show that they can be aligned relatively well. The set of relevant verticals defined by those approaches could therefore serve as an approximate but reliable ground-truth for evaluating vertical selection, avoiding the need for collecting explicit user vertical intent, and vice versa.
Ke Zhou 0003, Thomas Demeester, Dong Nguyen 0002, Djoerd Hiemstra, Dolf Trieschnigg
CIKM3
2014 Why Gender and Age Prediction from Tweets is Hard: Lessons from a Crowdsourcing Experiment
Dong Nguyen 0002, Dolf Trieschnigg, A. Seza Dogruöz, Rilana Gravel, Mariët Theune, Theo Meder, Franciska de Jong
COLING1
2014 Exploiting user disagreement for web search evaluation: an experimental approach
abstract
To express a more nuanced notion of relevance as compared to binary judgments, graded relevance levels can be used for the evaluation of search results. Especially in Web search, users strongly prefer top results over less relevant results, and yet they often disagree on which are the top results for a given information need. Whereas previous works have generally considered disagreement as a negative effect, this paper proposes a method to exploit this user disagreement by integrating it into the evaluation procedure.
Thomas Demeester, Robin Aly, Djoerd Hiemstra, Dong Nguyen 0002, Dolf Trieschnigg, Chris Develder
WSDM4
2013 Snippet-Based Relevance Predictions for Federated Web Search
Thomas Demeester, Dong Nguyen 0002, Dolf Trieschnigg, Chris Develder, Djoerd Hiemstra
ECIR2
2013 Folktale Classification Using Learning to Rank
Dong Nguyen 0002, Dolf Trieschnigg, Mariët Theune
ECIR1
2013 Word Level Language Identification in Online Multilingual Communication
abstract
Multilingual speakers switch between languages in online and spoken communication.Analyses of large scale multilingual data require automatic language identification at the word level.For our experiments with multilingual online discussions, we first tag the language of individual words using language models and dictionaries.Secondly, we incorporate context to improve the performance.We achieve an accuracy of 98%.Besides word level accuracy, we use two new metrics to evaluate this task.
Dong Nguyen 0002, A. Seza Dogruöz
EMNLP1
2013 "How Old Do You Think I Am?" A Study of Language and Age in Twitter
Dong Nguyen 0002, Rilana Gravel, Dolf Trieschnigg, Theo Meder
ICWSM1
2012 Federated search in the wild: the combined power of over a hundred search engines
abstract
Federated search has the potential of improving web search: the user becomes less dependent on a single search provider and parts of the deep web become available through a unified interface, leading to a wider variety in the retrieved search results. However, a publicly available dataset for federated search reflecting an actual web environment has been absent. As a result, it has been difficult to assess whether proposed systems are suitable for the web setting. We introduce a new test collection containing the results from more than a hundred actual search engines, ranging from large general web search engines such as Google and Bing to small domain-specific engines. We discuss the design and analyze the effect of several sampling methods. For a set of test queries, we collected relevance judgements for the top 10 results of each search engine. The dataset is publicly available and is useful for researchers interested in resource selection for web search collections, result merging and size estimation of uncooperative resources.
Dong Nguyen 0002, Thomas Demeester, Dolf Trieschnigg, Djoerd Hiemstra
CIKM1
2010 Exploring the Effectiveness of Social Capabilities and Goal Alignment in Computer Supported Collaborative Learning
Hua Ai, Rohit Kumar 0001, Dong Nguyen 0002, Amrut Nagasunder, Carolyn P. Rosé
Intelligent Tutoring Systems (2)3
2010 DesignWebs: A Tool for Automatic Construction of Interactive Conceptual Maps from Document Collections
Sharad V. Oberoi, Dong Nguyen 0002, Gahgene Gweon, Susan Finger, Carolyn P. Rosé
Intelligent Tutoring Systems (2)2