VLDB 2026 Research / reviewers in the wild / expert
João Sedoc
dblp:175/1172
· DBLP profile ↗
24ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0001-6369-3711ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 2 since 2021Computer networks · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Knowing When Not to Answer: Lightweight KB-Aligned OOD Detection for Safe RAGabstractIlias Triantafyllopoulos, Renyi Qu, Salvatore Giorgi, Brenda Curtis, Lyle Ungar, João Sedoc. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Ilias Triantafyllopoulos, Renyi Qu, Salvatore Giorgi, Brenda Curtis, Lyle H. Ungar, João Sedoc |
ACL (1) | 6 |
| 2025 | The Illusion of Empathy: How AI Chatbots Shape Conversation PerceptionabstractAs AI chatbots increasingly incorporate empathy, understanding user-centered perceptions of chatbot empathy and its impact on conversation quality remains essential yet under-explored. This study examines how chatbot identity and perceived empathy influence users' overall conversation experience. Analyzing 155 conversations from two datasets, we found that while GPT-based chatbots were rated significantly higher in conversational quality, they were consistently perceived as less empathetic than human conversational partners. Empathy ratings from GPT-4o annotations aligned with user ratings, reinforcing the perception of lower empathy in chatbots compared to humans. Our findings underscore the critical role of perceived empathy in shaping conversation quality, revealing that achieving high-quality human-AI interactions requires more than simply embedding empathetic language; it necessitates addressing the nuanced ways users interpret and experience empathy in conversations with chatbots. Salvatore Giorgi, Ankit Aich, Allison Lahnala, Brenda Curtis, Lyle H. Ungar, João Sedoc |
AAAI | 7 |
| 2024 | Towards Authoring Open-Ended Behaviors for Narrative Puzzle Games with Large Language Model SupportabstractDesigning games with branching story lines, object annotations, scene details, and dialog can be challenging due to the intensive authoring required. We investigate the potential for authoring open-ended behaviors for point-and-click narrative games using GPT-3.5, a large language model. In our approach, we extend a behavior tree scripting system with nodes that query GPT-3.5 to generate object descriptions, conversations with characters, and responses to player actions. GPT-3.5 is used to generate content when it hasn’t been scripted manually and to update game state by asking questions about whether a player’s input achieves a particular game goal. We demonstrate our approach with puzzles based on scenes from an episode of Star Trek Voyager. Our approach aims to blend a specific plot with open-ended story elements while keeping the authoring work minimal. Based on a pilot study of 16 participants and our own testing, we find that the generated responses have high coherency and show signs of humor and novelty, but that utterances could be improved to be more interesting and better support the designer’s intent. Britney Ngaw, Grishma Jena, João Sedoc, Aline Normoyle |
FDG | 3 |
| 2024 | Lived Experience Matters: Automatic Detection of Stigma toward People Who Use Substances on Social MediaabstractStigma toward people who use substances (PWUS) is a leading barrier to seeking treatment. Further, those in treatment are more likely to drop out if they experience higher levels of stigmatization. While related concepts of hate speech and toxicity, including those targeted toward vulnerable populations, have been the focus of automatic content moderation research, stigma and, in particular, people who use substances have not. This paper explores stigma toward PWUS using a data set of roughly 5,000 public Reddit posts. We performed a crowd-sourced annotation task where workers are asked to annotate each post for the presence of stigma toward PWUS and answer a series of questions related to their experiences with substance use. Results show that workers who use substances or know someone with a substance use disorder are more likely to rate a post as stigmatizing. Building on this, we use a supervised machine learning framework that centers workers with lived substance use experience to label each Reddit post as stigmatizing. Modeling person-level demographics in addition to comment-level language results in a classification accuracy (as measured by AUC) of 0.69 -- a 17% increase over modeling language alone. Finally, we explore the linguist cues which distinguish stigmatizing content: PWUS substances and those who don't agree that language around othering ("people", "they") and terms like "addict" are stigmatizing, while PWUS (as opposed to those who do not) find discussions around specific substances more stigmatizing. Our findings offer insights into the nature of perceived stigma in substance use. Additionally, these results further establish the subjective nature of such machine learning tasks, highlighting the need for understanding their social contexts. Salvatore Giorgi, Douglas Bellew, Daniel Roy Sadek Habib, João Sedoc, Chase Smitterberg, Amanda Devoto, McKenzie Himelein-Wachowiak, Brenda Curtis |
ICWSM | 4 |
| 2024 | Large Human Language Models: A Need and the ChallengesabstractNikita Soni, H. Andrew Schwartz, João Sedoc, Niranjan Balasubramanian. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Nikita Soni 0002, H. Andrew Schwartz, João Sedoc, Niranjan Balasubramanian |
NAACL-HLT | 3 |
| 2024 | Overview of the Tenth Dialog System Technology Challenge: DSTC10abstractThis article introduces the Tenth Dialog System Technology Challenge (DSTC-10). This edition of the DSTC focuses on applying end-to-end dialog technologies for five distinct tasks in dialog systems, namely 1. Incorporation of Meme images into open domain dialogs, 2. Knowledge-grounded Task-oriented Dialogue Modeling on Spoken Conversations, 3. Situated Interactive Multimodal dialogs, 4. Reasoning for Audio Visual Scene-Aware Dialog, and 5. Automatic Evaluation and Moderation of Open-domainDialogue Systems. This article describes the task definition, provided datasets, baselines, and evaluation setup for each track. We also summarize the results of the submitted systems to highlight the general trends of the state-of-the-art technologies for the tasks. Koichiro Yoshino, Yun-Nung Chen, Paul A. Crook, Satwik Kottur, Jinchao Li, Behnam Hedayatnia, Seungwhan Moon, Zhengcong Fei, Zekang Li, Jinchao Zhang 0001, Yang Feng 0004, Jie Zhou 0016, Seokhwan Kim, Yang Liu 0004, Di Jin 0005, Alexandros Papangelis, Karthik Gopalakrishnan 0001, Dilek Hakkani-Tür, Babak Damavandi, Alborz Geramifard, Chiori Hori, Chen Zhang 0020, Haizhou Li 0001, João Sedoc, Luis Fernando D'Haro, Rafael E. Banchs, Alexander I. Rudnicky |
IEEE ACM Trans. Audio Speech Lang. Process. | 25 |
| 2023 | A Needle in a Haystack: An Analysis of High-Agreement Workers on MTurk for SummarizationabstractLining Zhang, Simon Mille, Yufang Hou, Daniel Deutsch, Elizabeth Clark, Yixin Liu, Saad Mahamood, Sebastian Gehrmann, Miruna Clinciu, Khyathi Raghavi Chandu, João Sedoc. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Lining Zhang, Simon Mille, Yufang Hou 0001, Daniel Deutsch, Elizabeth Clark, Yixin Liu 0003, Saad Mahamood, Sebastian Gehrmann, Miruna-Adriana Clinciu, Khyathi Raghavi Chandu, João Sedoc |
ACL (1) | 11 |
| 2023 | An "Integrative Survey on Mental Health Conversational Agents to Bridge Computer Science and Medical Perspectives"abstractMental health conversational agents (a.k.a. chatbots) are widely studied for their potential to offer accessible support to those experiencing mental health challenges. Previous surveys on the topic primarily consider papers published in either computer science or medicine, leading to a divide in understanding and hindering the sharing of beneficial knowledge between both domains. To bridge this gap, we conduct a comprehensive literature review using the PRISMA framework, reviewing 534 papers published in both computer science and medicine. Our systematic review reveals 136 key papers on building mental health-related conversational agents with diverse characteristics of modeling and experimental design techniques. We find that computer science papers focus on LLM techniques and evaluating response quality using automated metrics with little attention to the application while medical papers use rule-based conversational agents and outcome metrics to measure the health outcomes of participants. Based on our findings on transparency, ethics, and cultural heterogeneity in this review, we provide a few recommendations to help bridge the disciplinary divide and enable the cross-disciplinary development of mental health conversational agents. Sunny Rai, Lyle H. Ungar, João Sedoc, Sharath Chandra Guntuku |
EMNLP | 4 |
| 2023 | Conceptor-Aided Debiasing of Large Language ModelsabstractPre-trained large language models (LLMs) reflect the inherent social biases of their training corpus.Many methods have been proposed to mitigate this issue, but they often fail to debias or they sacrifice model accuracy.We use conceptors-a soft projection method-to identify and remove the bias subspace in LLMs such as BERT and GPT.We propose two methods of applying conceptors (1) bias subspace projection by post-processing by the conceptor NOT operation; and (2) a new architecture, conceptor-intervened BERT (CI-BERT), which explicitly incorporates the conceptor projection into all layers during training.We find that conceptor post-processing achieves state-of-theart (SoTA) debiasing results while maintaining LLMs' performance on the GLUE benchmark.Further, it is robust in various scenarios and can mitigate intersectional bias efficiently by its AND operation on the existing bias subspaces.Although CI-BERT's training takes all layers' bias into account and can beat its postprocessing counterpart in bias mitigation, CI-BERT reduces the language model accuracy.We also show the importance of carefully constructing the bias subspace.The best results are obtained by removing outliers from the list of biased words, combining them (via the OR operation), and computing their embeddings using the sentences from a cleaner corpus. 1 Lyle H. Ungar, João Sedoc |
EMNLP | 3 |
| 2023 | Linear Connectivity Reveals Generalization Strategies
Jeevesh Juneja, Rachit Bansal, Kyunghyun Cho, João Sedoc, Naomi Saphra |
ICLR | 4 |
| 2022 | Piloting Family Health History Chatbot with Crowd-Sourced Data Collection
Michelle H. Nguyen, João Sedoc, Casey Overby Taylor |
AMIA | 2 |
| 2022 | Automatic Document Selection for Efficient Encoder PretrainingabstractBuilding pretrained language models is considered expensive and data-intensive, but must we increase dataset size to achieve better performance?We propose an alternative to larger training sets by automatically identifying smaller yet domain-representative subsets.We extend Cynical Data Selection, a statistical sentence scoring method that conditions on a representative target domain corpus.As an example, we treat the OntoNotes corpus as a target domain and pretrain a RoBERTa-like encoder from a cynically selected subset of the Pile.On both perplexity and across several downstream tasks in the target domain, it consistently outperforms random selection with 20x less data, 3x fewer training iterations, and 2x less estimated cloud compute cost, validating the recipe of automatic document selection for LM pretraining. Yukun Feng, Patrick Xia 0002, Benjamin Van Durme, João Sedoc |
EMNLP | 4 |
| 2021 | Measuring the 'I don't know' Problem through the Lens of Gricean QuantityabstractWe consider the intrinsic evaluation of neural generative dialog models through the lens of Grice's Maxims of Conversation (1975).Based on the maxim of Quantity (be informative), we propose Relative Utterance Quantity (RUQ) to diagnose the 'I don't know' problem, in which a dialog system produces generic responses.The linguistically motivated RUQ diagnostic compares the model score of a generic response to that of the reference response.We find that for reasonable baseline models, 'I don't know' is preferred over the reference the majority of the time, but this can be reduced to less than 5% with hyperparameter tuning.RUQ allows for the direct analysis of the 'I don't know' problem, which has been addressed but not analyzed by prior work. Huda Khayrallah, João Sedoc |
NAACL-HLT | 2 |
| 2021 | MimicNet: fast performance estimates for data center networks with machine learningabstractAt-scale evaluation of new data center network innovations is becoming increasingly intractable. This is true for testbeds, where few, if any, can afford a dedicated, full-scale replica of a data center. It is also true for simulations, which while originally designed for precisely this purpose, have struggled to cope with the size of today's networks. This paper presents an approach for quickly obtaining accurate performance estimates for large data center networks. Our system,MimicNet, provides users with the familiar abstraction of a packet-level simulation for a portion of the network while leveraging redundancy and recent advances in machine learning to quickly and accurately approximate portions of the network that are not directly visible. MimicNet can provide over two orders of magnitude speedup compared to regular simulation for a data center with thousands of servers. Even at this scale, MimicNet estimates of the tail FCT, throughput, and RTT are within 5% of the true results. Qizhen Zhang 0001, Kelvin K. W. Ng, Charles W. Kazer, João Sedoc, Vincent Liu 0001 |
SIGCOMM | 5 |
| 2020 | COD3S: Diverse Generation with Discrete Semantic SignaturesabstractWe present COD3S, a novel method for generating semantically diverse sentences using neural sequence-to-sequence (seq2seq) models.Conditioned on an input, seq2seq models typically produce semantically and syntactically homogeneous sets of sentences and thus perform poorly on one-to-many sequence generation tasks.Our two-stage approach improves output diversity by conditioning generation on locality-sensitive hash (LSH)-based semantic sentence codes whose Hamming distances highly correlate with human judgments of semantic textual similarity.Though it is generally applicable, we apply COD3S to causal generation, the task of predicting a proposition's plausible causes or effects.We demonstrate through automatic and human evaluation that responses produced using our method exhibit improved diversity without degrading task performance. Nathaniel Weir, João Sedoc, Benjamin Van Durme |
EMNLP (1) | 2 |
| 2020 | Incremental Neural Coreference Resolution in Constant MemoryabstractWe investigate modeling coreference resolution under a fixed memory constraint by extending an incremental clustering algorithm to utilize contextualized encoders and neural components.Given a new sentence, our endto-end algorithm proposes and scores each mention span against explicit entity representations created from the earlier document context (if any).These spans are then used to update the entity's representations before being forgotten; we only retain a fixed set of salient entities throughout the document.In this work, we successfully convert a highperforming model (Joshi et al., 2020), asymptotically reducing its memory usage to constant space with only a 0.3% relative loss in F1 on OntoNotes 5.0. Patrick Xia 0002, João Sedoc, Benjamin Van Durme |
EMNLP (1) | 2 |
| 2020 | Learning Word Ratings for Empathy and Distress from Document-Level User ResponsesabstractDespite the excellent performance of black box approaches to modeling sentiment and emotion, lexica (sets of informative words and associated weights) that characterize different emotions are indispensable to the NLP community because they allow for interpretable and robust predictions. Emotion analysis of text is increasing in popularity in NLP; however, manually creating lexica for psychological constructs such as empathy has proven difficult. This paper automatically creates empathy word ratings from document-level ratings. The underlying problem of learning word ratings from higher-level supervision has to date only been addressed in an ad hoc fashion and has not used deep learning methods. We systematically compare a number of approaches to learning word ratings from higher-level supervision against a Mixed-Level Feed Forward Network (MLFFN), which we find performs best, and use the MLFFN to create the first-ever empathy lexicon. We then use Signed Spectral Clustering to gain insights into the resulting words. The empathy and distress lexica are publicly available at: http://www.wwbp.org/lexica.html. João Sedoc, Sven Buechel, Yehonathan Nachmany, Anneke Buffone, Lyle H. Ungar |
LREC | 1 |
| 2019 | Unsupervised Post-Processing of Word Vectors via Conceptor NegationabstractWord vectors are at the core of many natural language processing tasks. Recently, there has been interest in post-processing word vectors to enrich their semantic information. In this paper, we introduce a novel word vector post-processing technique based on matrix conceptors (Jaeger 2014), a family of regularized identity maps. More concretely, we propose to use conceptors to suppress those latent features of word vectors having high variances. The proposed method is purely unsupervised: it does not rely on any corpus or external linguistic database. We evaluate the post-processed word vectors on a battery of intrinsic lexical evaluation tasks, showing that the proposed method consistently outperforms existing state-of-the-art alternatives. We also show that post-processed word vectors can be used for the downstream natural language processing task of dialogue state tracking, yielding improved results in different dialogue domains. Tianlin Liu, Lyle H. Ungar, João Sedoc |
AAAI | 3 |
| 2019 | Comparison of Diverse Decoding Methods from Conditional Language ModelsabstractWhile conditional language models have greatly improved in their ability to output high-quality natural language, many NLP applications benefit from being able to generate a diverse set of candidate sequences.Diverse decoding strategies aim to, within a givensized candidate list, cover as much of the space of high-quality outputs as possible, leading to improvements for tasks that re-rank and combine candidate outputs.Standard decoding methods, such as beam search, optimize for generating high likelihood sequences rather than diverse ones, though recent work has focused on increasing diversity in these methods.In this work, we perform an extensive survey of decoding-time strategies for generating diverse outputs from conditional language models.We also show how diversity can be improved without sacrificing quality by oversampling additional candidates, then filtering to the desired number. Daphne Ippolito, Reno Kriz, João Sedoc, Maria Kustikova, Chris Callison-Burch |
ACL (1) | 3 |
| 2019 | Getting in Shape: Word Embedding SubSpacesabstractMany tasks in natural language processing require the alignment of word embeddings. Embedding alignment relies on the geometric properties of the manifold of word vectors. This paper focuses on supervised linear alignment and studies the relationship between the shape of the target embedding. We assess the performance of aligned word vectors on semantic similarity tasks and find that the isotropy of the target embedding is critical to the alignment. Furthermore, aligning with an isotropic noise can deliver satisfactory results. We provide a theoretical framework and guarantees which aid in the understanding of empirical results. Tianyuan Zhou, João Sedoc, Jordan Rodu |
IJCAI | 2 |
| 2018 | Hierarchical Methods for a Unified Approach to Discourse, Domain, and Style in Neural Conversational Models
João Sedoc |
AAAI | 1 |
| 2018 | Modeling Empathy and Distress in Reaction to News StoriesabstractComputational detection and understanding of empathy is an important factor in advancing human-computer interaction.Yet to date, textbased empathy prediction has the following major limitations: It underestimates the psychological complexity of the phenomenon, adheres to a weak notion of ground truth where empathic states are ascribed by third parties, and lacks a shared corpus.In contrast, this contribution presents the first publicly available gold standard for empathy prediction.It is constructed using a novel annotation methodology which reliably captures empathy assessments by the writer of a statement using multiitem scales.This is also the first computational work distinguishing between multiple forms of empathy, empathic concern, and personal distress, as recognized throughout psychology.Finally, we present experimental results for three different predictive models, of which a CNN performs the best. Sven Buechel, Anneke Buffone, Barry Slaff, Lyle H. Ungar, João Sedoc |
EMNLP | 5 |
| 2018 | Fast Network Simulation Through Approximation or: How Blind Men Can Describe ElephantsabstractNetwork researchers today are unable to test their new ideas at scale before deployment due to the prohibitive costs of custom testbeds and the slow speed of large-scale network simulators. Data center simulation is particularly slow because of the massive amount of bandwidth and high degree of redundant computation incurred in simulating the network stacks of thousands of commodity machines. By using approximation to replace redundant portions of the simulation, we improve computation time while retaining high accuracy. Charles W. Kazer, João Sedoc, Kelvin K. W. Ng, Vincent Liu 0001, Lyle H. Ungar |
HotNets | 2 |
| 2017 | Semantic Word Clusters Using Signed Spectral ClusteringabstractVector space representations of words capture many aspects of word similarity, but such methods tend to produce vector spaces in which antonyms (as well as synonyms) are close to each other.For spectral clustering using such word embeddings, words are points in a vector space where synonyms are linked with positive weights, while antonyms are linked with negative weights.We present a new signed spectral normalized graph cut algorithm, signed clustering, that overlays existing thesauri upon distributionally derived vector representations of words, so that antonym relationships between word pairs are represented by negative weights.Our signed clustering algorithm produces clusters of words that simultaneously capture distributional and synonym relations.By using randomized spectral decomposition (Halko et al., 2011) and sparse matrices, our method is both fast and scalable.We validate our clusters using datasets containing human judgments of word pair similarities and show the benefit of using our word clusters for sentiment prediction. João Sedoc, Jean H. Gallier, Dean P. Foster, Lyle H. Ungar |
ACL (1) | 1 |