VLDB 2026 Research / reviewers in the wild / expert
Katherine Atwell
dblp:295/9795
· DBLP profile ↗
10ranked-venue papers
5as first author
10since 2021 · last 2024
0000-0002-5066-5604ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Combining Discourse Coherence with Large Language Models for More Inclusive, Equitable, and Robust Task-Oriented DialogueabstractLarge language models (LLMs) are capable of generating well-formed responses, but using LLMs to generate responses on the fly is not yet feasible for many task-oriented systems. Modular architectures are often still required for safety and privacy guarantees on the output. We hypothesize that an offline generation approach using discourse theories, formal grammar rules, and LLMs can allow us to generate human-like, coherent text in a more efficient, robust, and inclusive manner within a task-oriented setting. To this end, we present the first discourse-aware multimodal task-oriented dialogue system that combines discourse theories with offline LLM generation. We deploy our bot as an app to the general public and keep track of the user ratings for six months. Our user ratings show an improvement from 2.8 to 3.5 out of 5 with the introduction of discourse coherence theories. We also show that our model reduces misunderstandings in the dialect of African-American Vernacular English from 93% to 57%. While terms of use prevent us from releasing our entire codebase, we release our code in a format that can be integrated into most existing dialogue systems. Katherine Atwell, Mert Inan, Anthony Sicilia, Malihe Alikhani |
LREC/COLING | 1 |
| 2024 | Studying and Mitigating Biases in Sign Language Understanding ModelsabstractEnsuring that the benefits of sign language technologies are distributed equitably among all community members is crucial.Thus, it is important to address potential biases and inequities that may arise from the design or use of these resources.Crowd-sourced sign language datasets, such as the ASL Citizen dataset, are great resources for improving accessibility and preserving linguistic diversity, but they must be used thoughtfully to avoid reinforcing existing biases.In this work, we utilize the rich information about participant demographics and lexical features present in the ASL Citizen dataset to study and document the biases that may result from models trained on crowd-sourced sign datasets.Further, we apply several bias mitigation techniques during model training, and find that these techniques reduce performance disparities without decreasing accuracy.With the publication of this work, we release the demographic information about the participants in the ASL Citizen dataset to encourage future bias mitigation work in this space. Katherine Atwell, Danielle Bragg, Malihe Alikhani |
EMNLP | 1 |
| 2023 | How people talk about each other: Modeling Generalized Intergroup Bias and EmotionabstractVenkata Subrahmanyan Govindarajan, Katherine Atwell, Barea Sinno, Malihe Alikhani, David I. Beaver, Junyi Jessy Li. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Venkata Subrahmanyan Govindarajan, Katherine Atwell, Barea Sinno, Malihe Alikhani, David Beaver, Junyi Jessy Li |
EACL | 2 |
| 2023 | Multilingual Content Moderation: A Case Study on RedditabstractMeng Ye, Karan Sikka, Katherine Atwell, Sabit Hassan, Ajay Divakaran, Malihe Alikhani. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Meng Ye 0002, Karan Sikka, Katherine Atwell, Sabit Hassan, Ajay Divakaran, Malihe Alikhani |
EACL | 3 |
| 2022 | Studying the Effect of Moderator Biases on the Diversity of Online Discussions: A Computational Cross-linguistic Study
Sabit Hassan, Katherine Atwell, Malihe Alikhani |
CogSci | 2 |
| 2022 | The Role of Context and Uncertainty in Shallow Discourse ParsingabstractDiscourse parsing has proven to be useful for a number of NLP tasks that require complex reasoning. However, over a decade since the advent of the Penn Discourse Treebank, predicting implicit discourse relations in text remains challenging. There are several possible reasons for this, and we hypothesize that models should be exposed to more context as it plays an important role in accurate human annotation; meanwhile adding uncertainty measures can improve model accuracy and calibration. To thoroughly investigate this phenomenon, we perform a series of experiments to determine 1) the effects of context on human judgments, and 2) the effect of quantifying uncertainty with annotator confidence ratings on model accuracy and calibration (which we measure using the Brier score (Brier et al, 1950)). We find that including annotator accuracy and confidence improves model accuracy, and incorporating confidence in the model’s temperature function can lead to models with significantly better-calibrated confidence measures. We also find some insightful qualitative results regarding human and model behavior on these datasets. Katherine Atwell, Remi Choi, Junyi Jessy Li, Malihe Alikhani |
COLING | 1 |
| 2022 | APPDIA: A Discourse-aware Transformer-based Style Transfer Model for Offensive Social Media ConversationsabstractUsing style-transfer models to reduce offensiveness of social media comments can help foster a more inclusive environment. However, there are no sizable datasets that contain offensive texts and their inoffensive counterparts, and fine-tuning pretrained models with limited labeled data can lead to the loss of original meaning in the style-transferred text. To address this issue, we provide two major contributions. First, we release the first publicly-available, parallel corpus of offensive Reddit comments and their style-transferred counterparts annotated by expert sociolinguists. Then, we introduce the first discourse-aware style-transfer models that can effectively reduce offensiveness in Reddit text while preserving the meaning of the original text. These models are the first to examine inferential links between the comment and the text it is replying to when transferring the style of offensive Reddit text. We propose two different methods of integrating discourse relations with pretrained transformer models and evaluate them on our dataset of offensive comments from Reddit and their inoffensive counterparts. Improvements over the baseline with respect to both automatic metrics and human evaluation indicate that our discourse-aware models are better at preserving meaning in style-transferred text when compared to the state-of-the-art discourse-agnostic models. Katherine Atwell, Sabit Hassan, Malihe Alikhani |
COLING | 1 |
| 2022 | Political Ideology and Polarization: A Multi-dimensional ApproachabstractBarea Sinno, Bernardo Oviedo, Katherine Atwell, Malihe Alikhani, Junyi Jessy Li. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Barea Sinno, Bernardo Oviedo, Katherine Atwell, Malihe Alikhani, Junyi Jessy Li |
NAACL-HLT | 3 |
| 2022 | PAC-Bayesian domain adaptation bounds for multiclass learnersabstractMulticlass neural networks are a common tool in modern unsupervised domain adaptation, yet an appropriate theoretical description for their non-uniform sample complexity is lacking in the adaptation literature. To fill this gap, we propose the first PAC-Bayesian adaptation bounds for multiclass learners. We facilitate practical use of our bounds by also proposing the first approximation techniques for the multiclass distribution divergences we consider. For divergences dependent on a Gibbs predictor, we propose additional PAC-Bayesian adaptation bounds which remove the need for inefficient Monte-Carlo estimation. Empirically, we test the efficacy of our proposed approximation techniques as well as some novel design-concepts which we include in our bounds. Finally, we apply our bounds to analyze a common adaptation algorithm that uses neural networks. Anthony Sicilia, Katherine Atwell, Malihe Alikhani, Seong Jae Hwang |
UAI | 2 |
| 2021 | Where Are We in Discourse Relation Recognition?abstractDiscourse parsers recognize the intentional and inferential relationships that organize extended texts.They have had a great influence on a variety of NLP tasks as well as theoretical studies in linguistics and cognitive science.However it is often difficult to achieve good results from current discourse models, largely due to the difficulty of the task, particularly recognizing implicit discourse relations.Recent developments in transformer-based models have shown great promise on these analyses, but challenges still remain.We present a position paper which provides a systematic analysis of the state of the art discourse parsers.We aim to examine the performance of current discourse parsing models via gradual domain shift: within the same corpus, on in-domain texts, and on out-of-domain texts, and discuss the differences between the transformer-based models and the previous models in predicting different types of implicit relations both interand intra-sentential.We conclude by describing several shortcomings of the existing models and a discussion of how future work should approach this problem. Katherine Atwell, Junyi Jessy Li, Malihe Alikhani |
SIGDIAL | 1 |