Jad Kabbara

dblp:148/9943 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
11since 2021 · last 2026
0009-0009-5611-4078ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 LLM Targeted Underperformance Disproportionately Impacts Vulnerable Users
abstract
While state-of-the-art large language models (LLMs) have shown impressive performance on many tasks, systematically evaluating undesirable behaviors of these models remains critical. In this work, we investigate how the quality of LLM responses changes in terms of information accuracy, truthfulness, and refusals depending on three user traits: English proficiency, education level, and country of origin. We present extensive experimentation on three state-of-the-art LLMs and two different datasets targeting truthfulness and factuality. Our findings suggest that undesirable behaviors in state-of-the-art LLMs occur disproportionately more for users with lower English proficiency, of lower education status, and originating from outside the US, rendering these models unreliable sources of information towards their most vulnerable users.
Elinor Poole-Dayan, Deb Roy, Jad Kabbara
AAAI3
2026 Common to Whom? Regional Cultural Commonsense and LLM Bias in India
abstract
Sangmitra Madhusudan, Trush Shashank More, Steph Buongiorno, Renata Dividino, Jad Kabbara, Ali Emami. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Sangmitra Madhusudan, Trush Shashank More, Steph Buongiorno, Renata Queiroz Dividino, Jad Kabbara, Ali Emami
ACL (1)5
2025 Bridging Context Gaps: Enhancing Comprehension in Long-Form Social Conversations Through Contextualized Excerpts
abstract
We focus on enhancing comprehension in small-group recorded conversations, which serve as a medium to bring people together and provide a space for sharing personal stories and experiences on crucial social matters. One way to parse and convey information from these conversations is by sharing highlighted excerpts in subsequent conversations. This can help promote a collective understanding of relevant issues, by highlighting perspectives and experiences to other groups of people who might otherwise be unfamiliar with and thus unable to relate to these experiences. The primary challenge that arises then is that excerpts taken from one conversation and shared in another setting might be missing crucial context or key elements that were previously introduced in the original conversation. This problem is exacerbated when conversations become lengthier and richer in themes and shared experiences. To address this, we explore how Large Language Models (LLMs) can enrich these excerpts by providing socially relevant context. We present approaches for effective contextualization to improve comprehension, readability, and empathy. We show significant improvements in understanding, as assessed through subjective and objective evaluations. While LLMs can offer valuable context, they struggle with capturing key social aspects. We release the Human-annotated Salient Excerpts (HSE) dataset to support future work. Additionally, we show how context-enriched excerpts can provide more focused and comprehensive conversation summaries.
Shrestha Mohanty, Sarah Xuan, Jacob Jobraeel, Deb Roy, Jad Kabbara
COLING6
2025 Computational Analysis of Conversation Dynamics through Participant Responsivity
abstract
Growing literature explores toxicity and polarization in discourse, with comparatively less work on characterizing what makes dialogue prosocial and constructive.We explore conversational discourse and investigate a method for characterizing its quality built upon the notion of "responsivity"-whether one person's conversational turn is responding to a preceding turn.We develop and evaluate methods for quantifying responsivity-first through semantic similarity of speaker turns, and second by leveraging state-of-the-art large language models (LLMs) to identify the relation between two speaker turns.We evaluate both methods against a ground truth set of human-annotated conversations.Furthermore, selecting the better performing LLM-based approach, we characterize the nature of the response-whether it responded to that preceding turn in a substantive way or not.We view these responsivity links as a fundamental aspect of dialogue but note that conversations can exhibit significantly different responsivity structures.Accordingly, we then develop conversation-level derived metrics to address various aspects of conversational discourse.We use these derived metrics to explore other conversations and show that they support meaningful characterizations and differentiations across a diverse collection of conversations.
Margaret A. Hughes, Brandon Roy, Elinor Poole-Dayan, Deb Roy, Jad Kabbara
EMNLP5
2024 Leveraging Large Language Models for Learning Complex Legal Concepts through Storytelling
abstract
Hang Jiang, Xiajie Zhang, Robert Mahari, Daniel Kessler, Eric Ma, Tal August, Irene Li, Alex Pentland, Yoon Kim, Deb Roy, Jad Kabbara. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Xiajie Zhang, Robert Mahari, Daniel T. Kessler, Eric Ma, Tal August, Irene Li, Alex Pentland, Deb Roy, Jad Kabbara
ACL (1)11
2024 Confidence Under the Hood: An Investigation into the Confidence-Probability Alignment in Large Language Models
abstract
As the use of Large Language Models (LLMs) becomes more widespread, understanding their self-evaluation of confidence in generated responses becomes increasingly important as it is integral to the reliability of the output of these models.We introduce the concept of Confidence-Probability Alignment, that connects an LLM's internal confidence, quantified by token probabilities, to the confidence conveyed in the model's response when explicitly asked about its certainty.Using various datasets and prompting techniques that encourage model introspection, we probe the alignment between models' internal and expressed confidence.These techniques encompass using structured evaluation scales to rate confidence, including answer options when prompting, and eliciting the model's confidence level for outputs it does not recognize as its own.Notably, among the models analyzed, OpenAI's GPT-4 showed the strongest confidence-probability alignment, with an average Spearman's ρ of 0.42, across a wide range of tasks.Our work contributes to the ongoing efforts to facilitate risk assessment in the application of LLMs and to further our understanding of model trustworthiness. 1
Robert Morabito, Sanzhar Umbet, Jad Kabbara, Ali Emami
ACL (1)4
2024 Fora: A corpus and framework for the study of facilitated dialogue
abstract
Facilitated dialogue is increasingly popular as a method of civic engagement and as a method for gathering social insight, but resources for its study are scant.We present Fora, a unique collection of annotated facilitated dialogues.We compile 262 facilitated conversations that were hosted with partner organizations seeking to engage their members and surface insights regarding issues like education, elections, and public health, primarily through the sharing of personal experience.Alongside this corpus of 39,911 speaker turns, we present a framework for the analysis of facilitated dialogue.We taxonomize key personal sharing behaviors and facilitation strategies in the corpus, annotate a 25% sample (10,000+ speaker turns) of the data accordingly, and evaluate and establish baselines on a number of tasks essential to the identification of these phenomena in dialogue.We describe the data, and relate facilitator behavior to turn-taking and participant sharing.We outline how this research can inform future work in understanding and improving facilitated dialogue, parsing spoken conversation, and improving the behavior of dialogue agents.
Hope Schroeder, Deb Roy, Jad Kabbara
ACL (1)3
2024 On the Relationship between Truth and Political Bias in Language Models
abstract
Suyash Fulay, William Brannon, Shrestha Mohanty, Cassandra Overney, Elinor Poole-Dayan, Deb Roy, Jad Kabbara. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Suyash Fulay, William Brannon, Shrestha Mohanty, Cassandra Overney, Elinor Poole-Dayan, Deb Roy, Jad Kabbara
EMNLP7
2024 Position: Data Authenticity, Consent, & Provenance for AI are all broken: what will it take to fix them?
abstract
New capabilities in foundation models are owed in large part to massive, widely-sourced, and under-documented training data collections. Existing practices in data collection have led to challenges in tracing authenticity, verifying consent, preserving privacy, addressing representation and bias, respecting copyright, and overall developing ethical and trustworthy foundation models. In response, regulation is emphasizing the need for training data transparency to understand foundation models’ limitations. Based on a large-scale analysis of the foundation model training data landscape and existing solutions, we identify the missing infrastructure to facilitate responsible foundation model development practices. We examine the current shortcomings of common tools for tracing data authenticity, consent, and documentation, and outline how policymakers, developers, and data creators can facilitate responsible foundation model development by adopting universal data provenance standards.
Shayne Longpre, Robert Mahari, Naana Obeng-Marnu, William Brannon, Tobin South, Katy Ilonka Gero, Alex Pentland, Jad Kabbara
ICML8
2024 Consent in Crisis: The Rapid Decline of the AI Data Commons
abstract
General-purpose artificial intelligence (AI) systems are built on massive swathes of public web data, assembled into corpora such as C4, RefinedWeb, and Dolma. To our knowledge, we conduct the first, large-scale, longitudinal audit of the consent protocols for the web domains underlying AI training corpora. Our audit of 14,000 web domains provides an expansive view of crawlable web data and how codified data use preferences are changing over time. We observe a proliferation of AI-specific clauses to limit use, acute differences in restrictions on AI developers, as well as general inconsistencies between websites' expressed intentions in their Terms of Service and their robots.txt. We diagnose these as symptoms of ineffective web protocols, not designed to cope with the widespread re-purposing of the internet for AI. Our longitudinal analyses show that in a single year (2023-2024) there has been a rapid crescendo of data restrictions from web sources, rendering ~5\%+ of all tokens in C4, or 28%+ of the most actively maintained, critical sources in C4, fully restricted from use. For Terms of Service crawling restrictions, a full 45% of C4 is now restricted. If respected or enforced, these restrictions are rapidly biasing the diversity, freshness, and scaling laws for general-purpose AI systems. We hope to illustrate the emerging crises in data consent, for both developers and creators. The foreclosure of much of the open web will impact not only commercial AI, but also non-commercial AI and academic research.
Shayne Longpre, Robert Mahari, Ariel Lee, Campbell Lund, Hamidah Oderinwale, William Brannon, Nayan Saxena, Naana Obeng-Marnu, Tobin South, Cole Hunter, Kevin Klyman, Christopher Klamm, Hailey Schoelkopf, Nikhil Singh 0003, Manuel Cherep, Ahmad Anis, An Dinh, Caroline Shamiso Chitongo, Da Yin, Damien Sileo, Deividas Mataciunas, Diganta Misra, Emad A. Alghamdi, Enrico Shippole, Jianguo Zhang 0005, Joanna Materzynska, Kun Qian 0016, Kushagra Tiwary, Lester James V. Miranda, Manan Dey, Minnie Liang, Mohammed Hamdy, Niklas Muennighoff, Seonghyeon Ye, Seungone Kim, Shrestha Mohanty, Vivek Sharma 0001, Minh Chien Vu, Caiming Xiong, Stella Biderman, Daphne Ippolito, Sara Hooker, Jad Kabbara, Alex Pentland
NeurIPS48
2022 Investigating the Performance of Transformer-Based NLI Models on Presuppositional Inferences
abstract
Presuppositions are assumptions that are taken for granted by an utterance, and identifying them is key to a pragmatic interpretation of language. In this paper, we investigate the capabilities of transformer models to perform NLI on cases involving presupposition. First, we present simple heuristics to create alternative “contrastive” test cases based on the ImpPres dataset and investigate the model performance on those test cases. Second, to better understand how the model is making its predictions, we analyze samples from sub-datasets of ImpPres and examine model performance on them. Overall, our findings suggest that NLI-trained transformer models seem to be exploiting specific structural and lexical cues as opposed to performing some kind of pragmatic reasoning.
Jad Kabbara, Jackie Chi Kit Cheung
COLING1
2018 Let's do it "again": A First Computational Approach to Detecting Adverbial Presupposition Triggers
abstract
We introduce the task of predicting adverbial presupposition triggers such as also and again.Solving such a task requires detecting recurring or similar events in the discourse context, and has applications in natural language generation tasks such as summarization and dialogue systems.We create two new datasets for the task, derived from the Penn Treebank and the Annotated English Gigaword corpora, as well as a novel attention mechanism tailored to this task.Our attention mechanism augments a baseline recurrent neural network without the need for additional trainable parameters, minimizing the added computational cost of our mechanism.We demonstrate that our model statistically outperforms a number of baselines, including an LSTM-based language model.
Andre Cianflone, Yulan Feng, Jad Kabbara, Jackie Chi Kit Cheung
ACL (1)3
2017 Relevance effect: Exploiting Bayesian networks to improve supervised learning
abstract
Deductive logic and its variants enjoy the common property of monotonicity. For tasks such as inductive reasoning and belief revision, this was eventually deemed a serious flaw, prompting attempts to construct non-monotonic versions of logic. With the introduction of the idea of probabilistic reasoning to AI, particularly with the advent of Bayesian networks (BNs), the aforementioned monotonicity was no longer an issue: Probability is inherently non-monotonic. In this work, we introduce the notion of relevance effect which bears on exploiting BNs to generate realizations of relevant variables to be used for potentially improving the performance of a learning model on a supervised classification task. We explore the potential of using the relevance effect in the context of Deep Belief Networks (DBNs) with a focus on relational domains. We show that although the idea is at odds with the non-monotonicity of probabilistic reasoning, we attain an improvement in learning performance in different simulations on both synthetic and real-world scenarios. The observation that adopting this notion has improved the performance of a powerful model like DBNs hints to its potential to be practiced so as to enhance the performance of supervised learning methods in general. We furthermore highlight the connections as well as the implications of our work to the psychology literature.
Ardavan Salehi Nobandegani, Jad Kabbara, Ioannis N. Psaromiligkos
IJCNN2
2016 Capturing Pragmatic Knowledge in Article Usage Prediction using LSTMs
abstract
We examine the potential of recurrent neural networks for handling pragmatic inferences involving complex contextual cues for the task of article usage prediction. We train and compare several variants of Long Short-Term Memory (LSTM) networks with an attention mechanism. Our model outperforms a previous state-of-the-art system, achieving up to 96.63% accuracy on the WSJ/PTB corpus. In addition, we perform a series of analyses to understand the impact of various model choices. We find that the gain in performance can be attributed to the ability of LSTMs to pick up on contextual cues, both local and further away in distance, and that the model is able to solve cases involving reasoning about coreference and synonymy. We also show how the attention mechanism contributes to the interpretability of the model’s effectiveness.
Jad Kabbara, Yulan Feng, Jackie Chi Kit Cheung
COLING1
2016 Kernel subspace pursuit for sparse regression
Jad Kabbara, Ioannis N. Psaromiligkos
Pattern Recognit. Lett.1
2014 Improving the tracking ability of KRLS using Kernel Subspace Pursuit
abstract
We present a new Kernel Recursive Least Squares (KRLS) algorithm that is able to efficiently track time-varying systems. In order to alleviate the detrimental effect of a large dictionary size on the algorithm's tracking ability, we decouple the equality between dictionary size and weight vector size, an equality that has been encountered in all previous KRLS algorithms. In the proposed method, the maximum size of the weight vector is fixed and is independent from the dictionary size. We introduce the Kernel Subspace Pursuit algorithm which we use to choose a subset of the dictionary that tracks best the most recent received data samples. The selected dictionary elements are then used in the KRLS iterations. We show through simulations that our algorithm outperforms existing KRLS algorithms in tracking time-varying systems.
Jad Kabbara, Ioannis N. Psaromiligkos
ICASSP1