Ashiqur R. KhudaBukhsh

dblp:29/7442 · also Ashique R. KhudaBukhsh · DBLP profile ↗
← Back
51ranked-venue papers
14as first author
34since 2021 · last 2026
0000-0003-2394-7902ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 45 · 12 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 5 first-author · 22 since 2021Databases, data management, data science and information retrieval · 10 · 3 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021
YearPublicationVenuePosition
2026 How Can You Tell if Your Large Language Model Could Be a Closet Antisemite? An Explainability-Based Audit Framework for Implicit Bias
abstract
Auditing large language models (LLMs) for biases is an ongoing and dynamic process, resembling a proverbial cat-and-mouse game. As researchers identify new vulnerabilities in LLMs, guardrails are updated to address them, prompting the need for innovative approaches to audit the increasingly fortified LLMs for biases. This paper makes three contributions. First, it introduces a scalable, explainable framework to measure biases against various identity groups across multiple open large language models. Second, it conducts a bias audit considering five well-known open LLMs and demonstrates their bias inclinations towards several historically disadvantaged groups. Our audit reveals disturbing antisemitic, Islamophobic, and xenophobic biases present in several well-known LLMs. Finally, we release a dataset of 1,000 probes curated under the supervision of an expert social scientist that can facilitate similar audits.
Arka Dutta 0001, Reza Fayyazi, Shanchieh Jay Yang, Ashiqur R. KhudaBukhsh
AAAI4
2026 NewsLensAI: NER-Guided Summarization for Mitigating Hallucination and Bias in LLM-Based News Summaries (Student Abstract)
abstract
Automated news summarization using large language models (LLMs) offers great potential to enhance information accessibility. However, critical challenges, such as hallucinations, bias, and toxicity, threaten their reliability and societal acceptance. In this paper, we present NewsLensAI, a novel summarization framework explicitly designed to address these trustworthiness concerns through Named Entity Recognition (NER)-guided prompting. By anchoring summaries in key factual entities extracted from source articles, our method significantly reduces factual inaccuracies without altering model weights or architectures. We evaluated NewsLensAI on a dataset of 1,500 real-world news articles using open-source (LLaMA 3) and proprietary (Gemini 1.5) LLMs. Our analysis encompasses factual consistency, political bias shifts, sentiment preservation, and moderation of toxicity. Our results indicate substantial improvements in factual alignment, demonstrated by an average increase in the BERTScore from 0.80 (baseline) to 0.88 (NER-enhanced), and an approximately 60% reduction in hallucinated entities. To capture contextual terms that are relevant beyond the core entities, we use TF-IDF salience scoring to supplement standard NER categories, particularly for legislative terms and event identifiers. Furthermore, we identify and characterize a notable “centrist drift,” wherein summaries tend to moderate extreme biases present in source articles, along with a measurable reduction in toxic or emotionally charged language. Complementing our empirical findings, we introduce a real-time NewsLensAI demo that summarizes live news feeds from the Guardian API, providing dynamic bias and sentiment analysis. This practical implementation underscores the real-world applicability and potential societal benefit of our approach. Finally, we discuss critical ethical implications, including potential impacts on media literacy and information diversity. Our interdisciplinary approach, linking NLP, journalism, and ethical analysis, positions NewsLensAI as a meaningful step towards safer, fairer, and more trustworthy AI-generated news consumption.
Gaurank Maheshwari, Ambika Taploo, Ashiqur R. KhudaBukhsh
AAAI3
2026 Investigating Vaccine Buyer's Remorse: Post-Vaccination Decision Regret in COVID-19 Social Media Using Politically Diverse Human Annotation
abstract
A significant gap exists in datasets regarding post-COVID-19 vaccination experiences, particularly “vaccine buyer's remorse”. Understanding the prevalence and nature of vaccine regret, whether based on personal or vicarious experiences, is vital for addressing vaccine hesitancy and refining public health communication. In this paper, we curate a novel dataset from a large YouTube news corpus capturing COVID-19 vaccination experiences, and construct a benchmark subset focused on vaccine regret, annotated by a politically diverse panel to account for the subjective and often politicized nature of the topic. We utilize large language models (LLMs) to identify posts expressing vaccine regret, analyze the reasons behind this regret, and quantify its occurrence in both first and second-person accounts. This paper aims to (1) quantify the prevalence of vaccine regret; (2) identify common reasons for this sentiment; (3) analyze differences between first-person and vicarious experiences; and (4) assess potential biases introduced by different LLMs. We find that while vaccine buyer's remorse appears in only
Miles Stanley, Soumyajit Datta, Ashiqur R. KhudaBukhsh
AAAI4
2026 What About the Scene With the Hitler Reference? HAUNT: A Framework to Probe LLMs' Self-consistency in Closed Domains Via Adversarial Nudge
abstract
Arka Dutta, Sujan Dutta, Rijul Magu, Soumyajit Datta, Munmun De Choudhury, Ashiqur R. KhudaBukhsh. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Arka Dutta 0001, Sujan Dutta, Rijul Magu, Soumyajit Datta, Munmun De Choudhury, Ashiqur R. KhudaBukhsh
ACL (1)6
2025 All You Need Is S P A C E: When Jailbreaking Meets Bias Audit and Reveals What Lies Beneath the Guardrails (Student Abstract)
abstract
This paper makes a novel combination of a recently proposed bias audit framework and a recently proposed jailbreaking technique for Llama3. On an audit comprising several disadvantaged groups, our experiments reveal that a jailbroken Llama3 exhibits worrisome antisemitism, racism, misogyny, and homophobia (to list a few) much akin to a broad suite of LLMs that were susceptible to similar biases.
Arka Dutta 0001, Aman Priyanshu, Ashiqur R. KhudaBukhsh
AAAI3
2025 ARTICLE: Annotator Reliability Through In-Context Learning
abstract
Ensuring annotator quality in training and evaluation data is a key piece of machine learning in NLP. Tasks such as sentiment analysis and offensive speech detection are intrinsically subjective, creating a challenging scenario for traditional quality assessment approaches because it is hard to distinguish disagreement due to poor work from that due to differences of opinions between sincere annotators. With the goal of increasing diverse perspectives in annotation while ensuring consistency, we propose ARTICLE, an in-context learning (ICL) framework to estimate annotation quality through self-consistency. We evaluate this framework on two offensive speech datasets using multiple LLMs and compare its performance with traditional methods. Our findings indicate that ARTICLE can be used as a robust method for identifying reliable annotators, hence improving data quality.
Sujan Dutta, Deepak Pandita, Tharindu Cyril Weerasooriya, Marcos Zampieri, Christopher Homan, Ashiqur R. KhudaBukhsh
AAAI6
2025 ARTICLE: Annotator Reliability Through In-Context Learning (Student Abstract)
abstract
Ensuring annotator quality in training and evaluation data is a key piece of machine learning in NLP. Tasks such as sentiment analysis and offensive speech detection are intrinsically subjective, creating a challenging scenario for traditional quality assessment approaches because it is hard to distinguish disagreement due to poor work from that due to differences of opinions between sincere annotators. With the goal of increasing diverse perspectives in annotation while ensuring consistency, we propose ARTICLE, an in-context learning (ICL) framework to estimate annotation quality through self-consistency. We evaluate this framework on two offensive speech datasets using multiple LLMs and compare its performance with traditional methods. Our findings indicate that ARTICLE can be used as a robust method for identifying reliable annotators, hence improving data quality.
Sujan Dutta, Deepak Pandita, Tharindu Cyril Weerasooriya, Marcos Zampieri, Christopher Homan, Ashiqur R. KhudaBukhsh
AAAI6
2025 Audience Engagement with Political Messaging on YouTube Shorts (Student Abstract)
abstract
This study investigates user engagement and political polarization on YouTube Shorts, a special category of YouTube videos with a duration of 15-60 seconds. Via a substantial corpus of 38,838 videos gleaned from 100 YouTube channels focusing on political content, we contrast YouTube Shorts with long-form content in terms of user engagement, content toxicity, and polarization. Our analyses reveal that (1) YouTube Shorts receive more likes and views and fewer comments as compared to their long-form video counterparts; (2) YouTube Shorts are more toxic; and (3) considerably more polarized than long-form YouTube videos.
Omkar Narkar, Aman Vohra, Ashiqur R. KhudaBukhsh
AAAI3
2025 When Neutral Summaries Are Not That Neutral: Quantifying Political Neutrality in LLM-Generated News Summaries (Student Abstract)
abstract
In an era where societal narratives are increasingly shaped by algorithmic curation, investigating the political neutrality of LLMs is an important research question. This study presents a fresh perspective on quantifying the political neutrality of LLMs through the lens of abstractive text summarization of polarizing news articles. We consider five pressing issues in current US politics: abortion, gun control/rights, healthcare, immigration, and LGBTQ+ rights. Via a substantial corpus of 20,344 news articles, our study reveals a consistent trend towards pro-Democratic biases in several well-known LLMs, with gun control and healthcare exhibiting the most pronounced biases (max polarization differences of -9.49% and -6.14%, respectively). Further analysis uncovers a strong convergence in the vocabulary of the LLM outputs for these divisive topics (55% overlap for Democrat-leaning representations, 52% for Republican). Being months away from a US election of consequence, we consider our findings important.
Supriti Vijay, Aman Priyanshu, Ashiqur R. KhudaBukhsh
AAAI3
2025 Empathy Between Neighboring Nations: Distance Matters
Avik Chakrabarti, Clay H. Yoo, Ashiqur R. KhudaBukhsh
ASONAM (1)3
2025 Hope vs. Hate: Understanding User Interactions with LGBTQ+ News Content in Mainstream US News Media through the Lens of Hope Speech
abstract
This paper makes three contributions.First, via a substantial corpus of 1,419,047 comments posted on 3,161 YouTube news videos of major US cable news outlets, we analyze how users engage with LGBTQ+ news content.Our analyses focus both on positive and negative content.In particular, we construct a hope speech classifier that detects positive (hope speech), negative, neutral, and irrelevant content.Second, in consultation with a public health expert specializing on LGBTQ+ health, we conduct an annotation study with a balanced and diverse political representation and release a dataset of 3,750 instances with crowd-sourced labels and detailed annotator demographic information.Finally, beyond providing a vital resource for the LGBTQ+ community, our annotation study and subsequent in-the-wild assessments reveal (1) strong association between rater political beliefs and how they rate content relevant to a marginalized community, (2) models trained on individual political beliefs exhibit considerable in-the-wild disagreement, and (3) zero-shot large language models (LLMs) align more with liberal raters.Trigger Warning: this paper contains offensive material that some may find upsetting.From crowdsourced storymapping projects sharing stories of love, loss, and a sense of belonging (Kirby et al., 2021) to safe, anonymous spaces to seek resources (McInroy et al., 2019) and dating platforms (Blackwell et al., 2015) -the internet and modern technologies play a positive role in the health and well-being of the LGBTQ+ community in various ways.However, cyberbullying (Abreu and Kenny, 2018), exposure to dehumanization through news media (Mendelsohn et al., 2020), and more recently, homophobic biases in large language models (LLMs) (Dutta et al., 2024a, 2025a) -* Ashiqur R. KhudaBukhsh is the corresponding author.are some of the modern technology perils the community grapples with.
Jonathan Pofcher, Christopher Homan, Randall Sell, Ashiqur R. KhudaBukhsh
EMNLP4
2025 Towards a Bipartisan Understanding of Peace and Vicarious Interactions
abstract
Human input plays a critical role in modern AI systems. As machines take on increasingly nuanced tasks, it becomes essential for the community to embrace subjectivity and diverse perspectives. However, research on sensitive topics often fails to incorporate diverse and balanced perspectives. This paper makes a key contribution to participatory AI design in the context of conflicts between nuclear adversaries (India and Pakistan); where disagreement between stakeholders is anticipated. The paper explores the notion of hope speech detection -- detecting de-escalating content in the context of nuclear adversaries on the brink of war -- through the lens of participatory AI design and vicarious interactions. We release a dataset of 10,081 social web posts annotated by raters from India and Pakistan and examine the bipartisan nature of the language of de-escalation. Our study reveals that vicarious perspectives can be useful for modeling out-group preferences.
Arka Dutta 0001, Syed Mohammad Sualeh Ali, Usman Naseem, Ashiqur R. KhudaBukhsh
IJCAI4
2025 On the State of NLP Approaches to Modeling Depression in Social Media: A Post-COVID-19 Outlook
abstract
Computational approaches to predicting mental health conditions in social media have been substantially explored in the past years. Multiple reviews have been published on this topic, providing the community with comprehensive accounts of the research in this area. Among all mental health conditions, depression is the most widely studied due to its worldwide prevalence. The COVID-19 global pandemic, starting in early 2020, has had a great impact on mental health worldwide. Harsh measures employed by governments to slow the spread of the virus (e.g., lockdowns) and the subsequent economic downturn experienced in many countries have significantly impacted people's lives and mental health. Studies have shown a substantial increase of above 50% in the rate of depression in the population. In this context, we present a review on natural language processing (NLP) approaches to modeling depression in social media, providing the reader with a post-COVID-19 outlook. This review contributes to the understanding of the impacts of the pandemic on modeling depression in social media. We outline how state-of-the-art approaches and new datasets have been used in the context of the COVID-19 pandemic. Finally, we also discuss ethical issues in collecting and processing mental health data, considering fairness, accountability, and ethics.
Ana-Maria Bucur, Andreea-Codrina Moldovan, Krutika Parvatikar, Marcos Zampieri, Ashiqur R. KhudaBukhsh, Liviu P. Dinu
IEEE J. Biomed. Health Informatics5
2024 Novax or Novak? Estimating Social Media Stance towards Celebrity Vaccine Hesitancy (Student Abstract)
abstract
On 15 January 2022, noted tennis player Novak Djokovic was deported from Australia due to his unvaccinated status for the COVID-19 vaccine. This paper presents a stance classifier and evaluates public reaction to this episode and the impact of this behavior on social media discourse on YouTube. We observed a significant spike of individuals who supported and opposed his behavior at the time of the episode. Supporters outnumbered those who opposed this behavior by over 4x. Our study reports a disturbing trend that following every major Djokovic win, even now, vaccine skeptics often conflate his tennis success as a fitting reply to vaccine mandates.
Madhav Hota, Adel Khorramrouz, Ashiqur R. KhudaBukhsh
AAAI3
2024 Quantifying Political Polarization through the Lens of Machine Translation and Vicarious Offense
abstract
This talk surveys three related research contributions that shed light on the current US political divide: 1. a novel machine-translation-based framework to quantify political polarization; 2. an analysis of disparate media portrayal of US policing in major cable news outlets; and 3. a novel perspective of vicarious offense that examines a timely and important question -- how well do Democratic-leaning users perceive what content would be deemed as offensive by their Republican-leaning counterparts or vice-versa?
Ashiqur R. KhudaBukhsh
AAAI1
2024 Anonymous Dissent in the Digital Age: A YouTube Dislikes Dataset
Sujan Dutta, Mallikarjuna T., Ashiqur R. KhudaBukhsh
ASONAM (3)4
2024 You Must Be a Trump Supporter: Political Identity Projections on the Social Web
Shubh Mittal, Tisha Chawla, Ashiqur R. KhudaBukhsh
ASONAM (1)3
2024 Community Needs and Assets: A Computational Analysis of Community Conversations
abstract
A community needs assessment is a tool used by non-profits and government agencies to quantify the strengths and issues of a community, allowing them to allocate their resources better. Such approaches are transitioning towards leveraging social media conversations to analyze the needs of communities and the assets already present within them. However, manual analysis of exponentially increasing social media conversations is challenging. There is a gap in the present literature in computationally analyzing how community members discuss the strengths and needs of the community. To address this gap, we introduce the task of identifying, extracting, and categorizing community needs and assets from conversational data using sophisticated natural language processing methods. To facilitate this task, we introduce the first dataset about community needs and assets consisting of 3,511 conversations from Reddit, annotated using crowdsourced workers. Using this dataset, we evaluate an utterance-level classification model compared to sentiment classification and a popular large language model (in a zero-shot setting), where we find that our model outperforms both baselines at an F1 score of 94% compared to 49% and 61% respectively. Furthermore, we observe through our study that conversations about needs have negative sentiments and emotions, while conversations about assets focus on location and entities.
Md Towhidul Absar Chowdhury, Naveen Sharma, Ashiqur R. KhudaBukhsh
ICWSM3
2024 Down the Toxicity Rabbit Hole: A Framework to Bias Audit Large Language Models with Key Emphasis on Racism, Antisemitism, and Misogyny
Arka Dutta 0001, Adel Khorramrouz, Sujan Dutta, Ashiqur R. KhudaBukhsh
IJCAI4
2024 A Survival Guide for Iranian Women Prescribed by Iranian Women: Participatory AI to Investigate Intimate Partner Physical Violence in Iran
Adel Khorramrouz, Mahbeigom Fayyazi, Ashiqur R. KhudaBukhsh
IJCAI3
2024 Infrastructure Ombudsman: Mining Future Failure Concerns from Structural Disaster Response
abstract
Current research concentrates on studying discussions on social media related to structural failures to improve disaster response strategies. However, detecting social web posts discussing concerns about anticipatory failures is under-explored. If such concerns are channeled to the appropriate authorities, it can aid in the prevention and mitigation of potential infrastructural failures. In this paper, we develop an infrastructure ombudsman -- that automatically detects specific infrastructure concerns. Our work considers several recent structural failures in the US. We present a first-of-its-kind dataset of 2,662 social web instances for this novel task mined from Reddit and YouTube.
Md Towhidul Absar Chowdhury, Soumyajit Datta, Naveen Sharma, Ashiqur R. KhudaBukhsh
WWW4
2023 Auditing and Robustifying COVID-19 Misinformation Datasets via Anticontent Sampling
abstract
This paper makes two key contributions. First, it argues that highly specialized rare content classifiers trained on small data typically have limited exposure to the richness and topical diversity of the negative class (dubbed anticontent) as observed in the wild. As a result, these classifiers' strong performance observed on the test set may not translate into real-world settings. In the context of COVID-19 misinformation detection, we conduct an in-the-wild audit of multiple datasets and demonstrate that models trained with several prominently cited recent datasets are vulnerable to anticontent when evaluated in the wild. Second, we present a novel active learning pipeline that requires zero manual annotation and iteratively augments the training data with challenging anticontent, robustifying these classifiers.
Clay H. Yoo, Ashiqur R. KhudaBukhsh
AAAI2
2023 Subjective Crowd Disagreements for Subjective Data: Uncovering Meaningful CrowdOpinion with Population-level Learning
abstract
Tharindu Cyril Weerasooriya, Sarah Luger, Saloni Poddar, Ashiqur KhudaBukhsh, Christopher Homan. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Tharindu Cyril Weerasooriya, Sarah K. K. Luger, Saloni Poddar, Ashiqur R. KhudaBukhsh, Christopher Homan
ACL (1)4
2023 Quantifying the Transience of Social Web Datasets
abstract
The social web presents a modern-day instrument to analyze a wide range of behavioral research questions. Of these platforms, Twitter has played a key role in social science research for more than a decade. This paper looks into an underexplored aspect - transience of Twitter datasets and makes the following three contributions. First, via a comprehensive investigation of more than 40 Twitter datasets, we identify that many of these datasets suffer from severe retrieval loss. Second, we demonstrate that the retrieval loss across labels is often imbalanced with inappropriate labels (e.g., misinformation, hate speech) suffering from more retrieval loss. Finally, we demonstrate that imbalanced retrieval loss may impact machine learning models differently than balanced retrieval loss.
Mohammed Afaan Ansari, Jiten Sidhpura, Vivek Kumar Mandal, Ashiqur R. KhudaBukhsh
ASONAM4
2023 Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is Offensive
abstract
Tharindu Weerasooriya, Sujan Dutta, Tharindu Ranasinghe, Marcos Zampieri, Christopher Homan, Ashiqur KhudaBukhsh. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Tharindu Cyril Weerasooriya, Sujan Dutta, Tharindu Ranasinghe, Marcos Zampieri, Christopher Homan, Ashiqur R. KhudaBukhsh
EMNLP6
2023 Partisan US News Media Representations of Syrian Refugees
abstract
We investigate how representations of Syrian refugees (2011-2021) differ across US partisan news outlets. We analyze 47,388 articles from the online US media about Syrian refugees to detail differences in reporting between left- and right-leaning media. We use various NLP techniques to understand these differences. Our polarization and question answering results indicated that left-leaning media tended to represent refugees as child victims, welcome in the US, and right-leaning media cast refugees as Islamic terrorists. We noted similar results with our sentiment and offensive speech scores over time, which detail possibly unfavorable representations of refugees in right-leaning media. A strength of our work is how the different techniques we have applied validate each other. Based on our results, we provide several recommendations. Stakeholders may utilize our findings to intervene around refugee representations, and design communications campaigns that improve the way society sees refugees and possibly aid refugee outcomes.
Marzieh Babaeianjelodar, Yiwen Shi, Kamila Janmohamed, Rupak Sarkar, Ingmar Weber, Thomas Davidson, Munmun De Choudhury, Jonathan Huang, Shweta Yadav 0001, Ashiqur R. KhudaBukhsh, Chris T. Bauch, Preslav Nakov, Orestis Papakyriakopoulos, Koustuv Saha, Kaveh Khoshnood, Navin Kumar 0004
ICWSM11
2023 Disentangling Societal Inequality from Model Biases: Gender Inequality in Divorce Court Proceedings
abstract
Divorce is the legal dissolution of a marriage by a court. Since this is usually an unpleasant outcome of a marital union, each party may have reasons to call the decision to quit which is generally documented in detail in the court proceedings. Via a substantial corpus of 17,306 court proceedings, this paper investigates gender inequality through the lens of divorce court proceedings. To our knowledge, this is the first-ever large-scale computational analysis of gender inequality in Indian divorce, a taboo-topic for ages. While emerging data sources (e.g., public court records made available on the web) on sensitive societal issues hold promise in aiding social science research, biases present in cutting-edge natural language processing (NLP) methods may interfere with or affect such studies. A thorough analysis of potential gaps and limitations present in extant NLP resources is thus of paramount importance. In this paper, on the methodological side, we demonstrate that existing NLP resources required several non-trivial modifications to quantify societal inequalities. On the substantive side, we find that while a large number of court cases perhaps suggest changing norms in India where women are increasingly challenging patriarchy, AI-powered analyses of these court proceedings indicate striking gender inequality with women often subjected to domestic violence.
Sujan Dutta, Parth Srivastava, Vaishnavi Solunke, Swaprava Nath, Ashiqur R. KhudaBukhsh
IJCAI5
2023 For Women, Life, Freedom: A Participatory AI-Based Social Web Analysis of a Watershed Moment in Iran's Gender Struggles
abstract
In this paper, we present a computational analysis of the Persian language Twitter discourse with the aim to estimate the shift in stance toward gender equality following the death of Mahsa Amini in police custody. We present an ensemble active learning pipeline to train a stance classifier. Our novelty lies in the involvement of Iranian women in an active role as annotators in building this AI system. Our annotators not only provide labels, but they also suggest valuable keywords for more meaningful corpus creation as well as provide short example documents for a guided sampling step. Our analyses indicate that Mahsa Amini's death triggered polarized Persian language discourse where both fractions of negative and positive tweets toward gender equality increased. The increase in positive tweets was slightly greater than the increase in negative tweets. We also observe that with respect to account creation time, between the state-aligned Twitter accounts and pro-protest Twitter accounts, pro-protest accounts are more similar to baseline Persian Twitter activity.
Adel Khorramrouz, Sujan Dutta, Ashiqur R. KhudaBukhsh
IJCAI3
2022 'Beach' to 'Bitch': Inadvertent Unsafe Transcription of Kids' Content on YouTube
abstract
Over the last few years, YouTube Kids has emerged as one of the highly competitive alternatives to television for children's entertainment. Consequently, YouTube Kids' content should receive an additional level of scrutiny to ensure children's safety. While research on detecting offensive or inappropriate content for kids is gaining momentum, little or no current work exists that investigates to what extent AI applications can (accidentally) introduce content that is inappropriate for kids. In this paper, we present a novel (and troubling) finding that well-known automatic speech recognition (ASR) systems may produce text content highly inappropriate for kids while transcribing YouTube Kids' videos. We dub this phenomenon as inappropriate content hallucination. Our analyses suggest that such hallucinations are far from occasional, and the ASR systems often produce them with high confidence. We release a first-of-its-kind data set of audios for which the existing state-of-the-art ASR systems hallucinate inappropriate content for kids. In addition, we demonstrate that some of these errors can be fixed using language models.
Krithika Ramesh, Ashiqur R. KhudaBukhsh
AAAI2
2022 A Murder and Protests, the Capitol Riot, and the Chauvin Trial: Estimating Disparate News Media Stance
abstract
In this paper, we analyze the responses of three major US cable news networks to three seminal policing events in the US spanning a thirteen month period--the murder of George Floyd by police officer Derek Chauvin, the Capitol riot, Chauvin's conviction, and his sentencing. We cast the problem of aggregate stance mining as a natural language inference task and construct an active learning pipeline for robust textual entailment prediction. Via a substantial corpus of 34,710 news transcripts, our analyses reveal that the partisan divide in viewership of these three outlets reflects on the network's news coverage of these momentous events. In addition, we release a sentence-level, domain-specific text entailment data set on policing consisting of 2,276 annotated instances.
Sujan Dutta, Daniel S. Nagin, Ashiqur R. KhudaBukhsh
IJCAI4
2022 Conversational Inequality Through the Lens of Political Interruption
abstract
We present a novel dataset of dialogues containing interruption with an aim to conduct a large-scale analysis of interruption patterns of people from diverse backgrounds in terms of gender, race/ethnicity, occupation, and political orientation. Our dataset includes 625,409 dialogues containing interruptions found in 275,420 transcripts from CNN, Fox News, and MSNBC spanning between January 2000 and July 2021. From this large, unlabeled pool of interruptions, we release an annotated dataset consisting of 2,000 dialogues with fine-grained interruption labels. We use this dataset to train an interruption classifier and predict the interruption type of a given dialogue. Our results reveal that male speakers (in our collected samples) tend to talk more than female speakers, while female speakers interrupt more. Moreover, people tend to use less intrusive interruptions when talking to others sharing the same political belief. This pattern becomes more pronounced among news media with stronger political bias.
Clay H. Yoo, Yuxi Luo, Kunal Khadilkar, Ashiqur R. KhudaBukhsh
IJCAI5
2021 An Unfair Affinity Toward Fairness: Characterizing 70 Years of Social Biases in BHollywood (Student Abstract)
abstract
Bollywood, aka the Mumbai film industry, is one of the biggest movie industries in the world with a current movie market share of worth 2.1 billion dollars and a target audience base of 1.2 billion people. While the entertainment impact in terms of lives that Bollywood can potentially touch is mammoth, no NLP study on social biases in Bollywood content exists. We thus seek to understand social biases in a developing country through the lens of popular movies. Our argument is simple -- popular movie content reflects social norms and beliefs in some form or shape. We present our preliminary findings on a longitudinal corpus of English subtitles of popular Bollywood movies focusing on (1) social bias toward a fair skin color (2) gender biases, and (3) gender representation. We contrast our findings with a similar corpus of Hollywood movies. Surprisingly, we observe that much of the biases we report in our preliminary experiments on the Bollywood corpus, also gets reflected in the Hollywood corpus.
Kunal Khadilkar, Ashiqur R. KhudaBukhsh
AAAI2
2021 We Don't Speak the Same Language: Interpreting Polarization through Machine Translation
abstract
Polarization among US political parties, media and elites is a widely studied topic. Prominent lines of prior research across multiple disciplines have observed and analyzed growing polarization in social media. In this paper, we present a new methodology that offers a fresh perspective on interpreting polarization through the lens of machine translation. With a novel proposition that two sub-communities are speaking in two different "languages", we demonstrate that modern machine translation methods can provide a simple yet powerful and interpretable framework to understand the differences between two (or more) large-scale social media discussion data sets at the granularity of words. Via a substantial corpus of 86.6 million comments by 6.5 million users on over 200,000 news videos hosted by YouTube channels of four prominent US news networks, we demonstrate that simple word-level and phrase-level translation pairs can reveal deep insights into the current political divide -- what is "black lives matter" to one can be "all lives matter" to the other.
Ashiqur R. KhudaBukhsh, Rupak Sarkar, Mark S. Kamlet, Tom M. Mitchell
AAAI1
2021 Are Chess Discussions Racist? An Adversarial Hate Speech Data Set (Student Abstract)
abstract
On June 28, 2020, while presenting a chess podcast on Grandmaster Hikaru Nakamura, Antonio Radic's YouTube handle got blocked because it contained ``harmful and dangerous'' content. YouTube did not give further specific reason, and the channel got reinstated within 24 hours. However, Radic speculated that given the current political situation, a referral to ``black against white'', albeit in the context of chess, earned him this temporary ban. In this paper, via a substantial corpus of 681,995 comments, on 8,818 YouTube videos hosted by five highly popular chess-focused YouTube channels, we ask the following research question: \emph{how robust are off-the-shelf hate-speech classifiers to out-of-domain adversarial examples?} We release a data set of 1,000 annotated comments where existing hate speech classifiers misclassified benign chess discussions as hate speech. We conclude with an intriguing analogy result on racial bias with our findings pointing out to the broader challenge of color polysemy.
Rupak Sarkar, Ashiqur R. KhudaBukhsh
AAAI2
2020 Voice for the Voiceless: Active Sampling to Detect Comments Supporting the Rohingyas
abstract
The Rohingya refugee crisis is one of the biggest humanitarian crises of modern times with more than 700,000 Rohingyas rendered homeless according to the United Nations High Commissioner for Refugees. While it has received sustained press attention globally, no comprehensive research has been performed on social media pertaining to this large evolving crisis. In this work, we construct a substantial corpus of YouTube video comments (263,482 comments from 113,250 users in 5,153 relevant videos) with an aim to analyze the possible role of AI in helping a marginalized community. Using a novel combination of multiple Active Learning strategies and a novel active sampling strategy based on nearest-neighbors in the comment-embedding space, we construct a classifier that can detect comments defending the Rohingyas among larger numbers of disparaging and neutral ones. We advocate that beyond the burgeoning field of hate speech detection, automatic detection of help speech can lend voice to the voiceless people and make the internet safer for marginalized communities.
Shriphani Palakodety, Ashiqur R. KhudaBukhsh, Jaime G. Carbonell
AAAI2
2020 Hope Speech Detection: A Computational Analysis of the Voice of Peace
abstract
The recent Pulwama terror attack (February 14, 2019, Pulwama, Kashmir) triggered a chain of escalating events between India and Pakistan adding another episode to their 70-year-old dispute over Kashmir. The present era of ubiquitious social media has never seen nuclear powers closer to war. In this paper, we analyze this evolving international crisis via a substantial corpus constructed using comments on YouTube videos (921,235 English comments posted by 392,460 users out of 2.04 million overall comments by 791,289 users on 2,890 videos). Our main contributions in the paper are three-fold. First, we present an observation that polyglot word-embeddings reveal precise and accurate language clusters, and subsequently construct a document language-identification technique with negligible annotation requirements. We demonstrate the viability and utility across a variety of data sets involving several low-resource languages. Second, we present an analysis on temporal trends of pro-peace and pro-war intent observing that when tensions between the two nations were at their peak, pro-peace intent in the corpus was at its highest point. Finally, in the context of heated discussions in a politically tense situation where two nations are at the brink of a full-fledged war, we argue the importance of automatic identification of user-generated web content that can diffuse hostility and address this prediction task, dubbed \emph{hope-speech detection}.
Shriphani Palakodety, Ashiqur R. KhudaBukhsh, Jaime G. Carbonell
ECAI2
2020 Mining Insights from Large-Scale Corpora Using Fine-Tuned Language Models
abstract
Mining insights from large volume of social media texts with minimal supervision is a highly challenging Natural Language Processing (NLP) task. While Language Models' (LMs) efficacy in several downstream tasks is well-studied, assessing their applicability in answering relational questions, tracking perception or mining deeper insights is under-explored. Few recent lines of work have scratched the surface by studying pre-trained LMs' (e.g., BERT) capability in answering relational questions through "fill-in-the-blank" cloze statements (e.g., [Dante was born in MASK]). BERT predicts the MASK-ed word with a list of words ranked by probability (in this case, BERT successfully predicts Florence with the highest probability). In this paper, we conduct a feasibility study of fine-tuned LMs with a different focus on tracking polls, tracking community perception and mining deeper insights typically obtained through costly surveys. Our main focus is on a substantial corpus of video comments extracted from YouTube videos (6,182,868 comments on 130,067 videos by 1,518,077 users) posted within 100 days prior to the 2019 Indian General Election. Using fill-in-the-blank cloze statements against a recent high-performance language modeling algorithm, BERT, we present a novel application of this family of tools that is able to (1) aggregate political sentiment (2) reveal community perception and (3) track evolving national priorities and issues of interest.
Shriphani Palakodety, Ashiqur R. KhudaBukhsh, Jaime G. Carbonell
ECAI2
2020 The Refugee Experience Online: Surfacing Positivity Amidst Hate
abstract
How can Artificial Intelligence help a stateless minority from online abuse? Research efforts in hate speech detection thus far have largely focused on identifying and subsequently filtering out negative content that specifically targets them. In this paper, we highlight a recent work [8] which tackles a different aspect of web-vulnerability of marginalized communities: sparsity of prominority voices championing their cause. The highlighted paper advocates that blocking hate alone may not be sufficient in these cases as the internet shapes community perception to a great extent in modern times and supportive comments to a vulnerable community serve a different purpose. Using an Active Sampling approach, the paper constructs a nuanced voice-for-the-voiceless classifier that automatically discovers comments supporting a (allegedly) persecuted minority. In the context of the Rohingya refugee crisis, one of the biggest humanitarian crises of modern times, the paper presents promising results that can substantially aid content moderation efforts in finding positive content supporting the Rohingyas.
Shriphani Palakodety, Ashiqur R. KhudaBukhsh, Jaime G. Carbonell
ECAI2
2020 Social Media Attributions in the Context of Water Crisis
abstract
Attribution of natural disasters/collective misfortune is a widely-studied political science problem.However, such studies typically rely on surveys, expert opinions, or external signals such as voting outcomes.In this paper, we explore the viability of using unstructured, noisy social media data to complement traditional surveys through automatically extracting attribution factors.We present a novel prediction task of attribution tie detection of identifying the factors (e.g., poor city planning, exploding population etc.) held responsible for the crisis in a social media document.We focus on the 2019 Chennai water crisis that rapidly escalated into a discussion topic with global importance following alarming water-crisis statistics.On a challenging data set constructed from YouTube comments (72,098 comments posted by 43,859 users on 623 videos relevant to the crisis), we present a neural baseline to identify attribution ties that achieves a reasonable performance (accuracy: 87.34% on attribution detection and 81.37% on attribution resolution).We release the first annotated data set of 2,500 comments in this important domain 1 .
Rupak Sarkar, Sayantan Mahinder, Hirak Sarkar, Ashiqur R. KhudaBukhsh
EMNLP (1)4
2020 Harnessing Code Switching to Transcend the Linguistic Barrier
abstract
Code mixing (or code switching) is a common phenomenon observed in social-media content generated by a linguistically diverse user-base. Studies show that in the Indian sub-continent, a substantial fraction of social media posts exhibit code switching. While the difficulties posed by code mixed documents to further downstream analyses are well-understood, lending visibility to code mixed documents under certain scenarios may have utility that has been previously overlooked. For instance, a document written in a mixture of multiple languages can be partially accessible to a wider audience; this could be particularly useful if a considerable fraction of the audience lacks fluency in one of the component languages. In this paper, we provide a systematic approach to sample code mixed documents leveraging a polyglot embedding based method that requires minimal supervision. In the context of the 2019 India-Pakistan conflict triggered by the Pulwama terror attack, we demonstrate an untapped potential of harnessing code mixing for human well-being: starting from an existing hostility diffusing hope speech classifier solely trained on English documents, code mixed documents are utilized to perform cross-lingual sampling and retrieve hope speech content written in a low-resource but widely used language - Romanized Hindi. Our proposed pipeline requires minimal supervision and holds promise in substantially reducing web moderation efforts. A further exploratory study on a new COVID-19 data set introduced in this paper demonstrates the generalizability of our cross-lingual sampling technique.
Ashiqur R. KhudaBukhsh, Shriphani Palakodety, Jaime G. Carbonell
IJCAI1
2019 Toward Reciprocity-Aware Distributed Learning in Referral Networks
Ashiqur R. KhudaBukhsh, Jaime G. Carbonell
PRICAI (2)1
2019 Expertise drift in referral networks
Ashiqur R. KhudaBukhsh, Jaime G. Carbonell
Auton. Agents Multi Agent Syst.1
2018 Endorsement in Referral Networks
Ashiqur R. KhudaBukhsh, Jaime G. Carbonell
EUMAS1
2018 Market-Aware Proactive Skill Posting
Ashiqur R. KhudaBukhsh, Jong Woo Hong, Jaime G. Carbonell
ISMIS1
2018 Robust learning in expert networks: a comparative analysis
Ashiqur R. KhudaBukhsh, Jaime G. Carbonell, Peter J. Jansen
J. Intell. Inf. Syst.1
2017 Robust Learning in Expert Networks: A Comparative Analysis
Ashiqur R. KhudaBukhsh, Jaime G. Carbonell, Peter J. Jansen
ISMIS1
2016 Distributed Learning in Expert Referral Networks
abstract
Human experts or autonomous agents in a referral network must decide whether to accept a task or refer to a more appropriate expert, and if so to whom. In order for the referral network to improve over time, the experts must learn to estimate the topical expertise of other experts. This paper extends concepts from Reinforcement Learning and Active Learning to referral networks, to learn how to refer at the network level, based on the proposed distributed interval estimation learning (DIEL) algorithm. Diverse Monte Carlo simulations reveal that DIEL improves network performance significantly over both greedy and Q-learning baselines [3], approaching optimal given enough data.
Ashiqur R. KhudaBukhsh, Peter J. Jansen, Jaime G. Carbonell
ECAI1
2016 SATenstein: Automatically building local search SAT solvers from components
Ashiqur R. KhudaBukhsh, Holger H. Hoos, Kevin Leyton-Brown
Artif. Intell.1
2015 Building Effective Query Classifiers: A Case Study in Self-harm Intent Detection
abstract
Query-based triggers play a crucial role in modern search systems, e.g., in deciding when to display direct answers on result pages. We address a common scenario in designing such triggers for real-world settings where positives are rare and search providers possess only a small seed set of positive examples to learn query classification models. We choose the critical domain of self-harm intent detection to demonstrate how such small seed sets can be expanded to create meaningful training data with a sizable fraction of positive examples. Our results show that with our method, substantially more positive queries can be found compared to plain random sampling. Additionally, we explored the effectiveness of traditional active learning approaches on classification performance and found that maximum uncertainty performs the best among several other techniques that we considered.
Ashiqur R. KhudaBukhsh, Paul N. Bennett, Ryen W. White
CIKM1
2014 Detecting Non-Adversarial Collusion in Crowdsourcing
abstract
A group of agents are said to collude if they share information or make joint decisions in a manner contrary to explicit or implicit social rules that results in an unfair advantage over non-colluding agents or other interested parties. For instance, collusion manifests as sharing answers in exams, as colluding bidders in auctions, or as colluding participants (e.g., Turkers) in crowd sourcing. This paper studies the latter, where the goal of the colluding participants is to "earn" money without doing the actual work, for instance by copying product ratings of another colluding participant, adding limited noise as attempted obfuscation. Such collusion not only yields fewer independent ratings, but may also introduce strong biases in aggregate results if undetected. Our proposed unsupervised collusion detection algorithm identifies colluding groups in crowd sourcing with fairly high accuracy both in synthetic and real data, and results in significant bias reduction, such as minimizing shifts from the true mean in rating tasks and recovering the true variance among raters.
Ashiqur R. KhudaBukhsh, Jaime G. Carbonell, Peter J. Jansen
HCOMP1
2009 SATenstein: Automatically Building Local Search SAT Solvers from Components
Ashiqur R. KhudaBukhsh, Holger H. Hoos, Kevin Leyton-Brown
IJCAI1