Arka Dutta 0001

dblp:149/5213-1 · DBLP profile ↗
← Back
5ranked-venue papers
5as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Trustworthy machine learning · 64% Language models and text generation · 29% Information extraction and text analysis · 7%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Computational social science and digital humanities · 100%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › fairness
bias evaluation
2.632026
How Can You Tell if Your Large Language Model Could Be a Closet Antisemite? An Explainability-Based Audit Framework for Implicit Bias · AAAI 2026
All You Need Is S P A C E: When Jailbreaking Meets Bias Audit and Reveals What Lies Beneath the Guardrails (Student Abstract) · AAAI 2025
Down the Toxicity Rabbit Hole: A Framework to Bias Audit Large Language Models with Key Emphasis on Racism, Antisemitism, and Misogyny · IJCAI 2024
Machine learning › Trustworthy machine learning
fairness
2.632026
How Can You Tell if Your Large Language Model Could Be a Closet Antisemite? An Explainability-Based Audit Framework for Implicit Bias · AAAI 2026
All You Need Is S P A C E: When Jailbreaking Meets Bias Audit and Reveals What Lies Beneath the Guardrails (Student Abstract) · AAAI 2025
Down the Toxicity Rabbit Hole: A Framework to Bias Audit Large Language Models with Key Emphasis on Racism, Antisemitism, and Misogyny · IJCAI 2024
Natural language and speech › Language models and text generation › large language model safety
large language model bias
1.922026
How Can You Tell if Your Large Language Model Could Be a Closet Antisemite? An Explainability-Based Audit Framework for Implicit Bias · AAAI 2026
All You Need Is S P A C E: When Jailbreaking Meets Bias Audit and Reveals What Lies Beneath the Guardrails (Student Abstract) · AAAI 2025
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
1.012026
What About the Scene With the Hitler Reference? HAUNT: A Framework to Probe LLMs' Self-consistency in Closed Domains Via Adversarial Nudge · ACL (1) 2026
Natural language and speech › Language models and text generation
large language model evaluation
1.012026
What About the Scene With the Hitler Reference? HAUNT: A Framework to Probe LLMs' Self-consistency in Closed Domains Via Adversarial Nudge · ACL (1) 2026
Natural language and speech › Information extraction and text analysis › abusive language detection
hate speech detection
0.912025
Towards a Bipartisan Understanding of Peace and Vicarious Interactions · IJCAI 2025
Machine learning › Trustworthy machine learning › adversarial machine learning
jailbreak attack
0.912025
All You Need Is S P A C E: When Jailbreaking Meets Bias Audit and Reveals What Lies Beneath the Guardrails (Student Abstract) · AAAI 2025
Natural language and speech › Language models and text generation
large language model
0.812024
Down the Toxicity Rabbit Hole: A Framework to Bias Audit Large Language Models with Key Emphasis on Racism, Antisemitism, and Misogyny · IJCAI 2024
Machine learning › Trustworthy machine learning › AI safety › content safety
toxicity analysis
0.812024
Down the Toxicity Rabbit Hole: A Framework to Bias Audit Large Language Models with Key Emphasis on Racism, Antisemitism, and Misogyny · IJCAI 2024
Design research and methods
participatory design
0.312025
Towards a Bipartisan Understanding of Peace and Vicarious Interactions · IJCAI 2025
Computational social science and digital humanities › platform governance
content moderation
0.212024
Down the Toxicity Rabbit Hole: A Framework to Bias Audit Large Language Models with Key Emphasis on Racism, Antisemitism, and Misogyny · IJCAI 2024

Methods — techniques the papers use, named apart from their topics

vicarious interaction · 2.6annotation study · 2.6bias auditing · 1.5explainability · 1.0bias probing · 1.0adversarial prompting · 1.0jailbreaking · 0.9bias audit · 0.9
YearPublicationVenuePosition
2026 How Can You Tell if Your Large Language Model Could Be a Closet Antisemite? An Explainability-Based Audit Framework for Implicit Bias
abstract
Auditing large language models (LLMs) for biases is an ongoing and dynamic process, resembling a proverbial cat-and-mouse game. As researchers identify new vulnerabilities in LLMs, guardrails are updated to address them, prompting the need for innovative approaches to audit the increasingly fortified LLMs for biases. This paper makes three contributions. First, it introduces a scalable, explainable framework to measure biases against various identity groups across multiple open large language models. Second, it conducts a bias audit considering five well-known open LLMs and demonstrates their bias inclinations towards several historically disadvantaged groups. Our audit reveals disturbing antisemitic, Islamophobic, and xenophobic biases present in several well-known LLMs. Finally, we release a dataset of 1,000 probes curated under the supervision of an expert social scientist that can facilitate similar audits.
Arka Dutta 0001, Reza Fayyazi, Shanchieh Jay Yang, Ashiqur R. KhudaBukhsh
AAAI1
2026 What About the Scene With the Hitler Reference? HAUNT: A Framework to Probe LLMs' Self-consistency in Closed Domains Via Adversarial Nudge
abstract
Arka Dutta, Sujan Dutta, Rijul Magu, Soumyajit Datta, Munmun De Choudhury, Ashiqur R. KhudaBukhsh. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Arka Dutta 0001, Sujan Dutta, Rijul Magu, Soumyajit Datta, Munmun De Choudhury, Ashiqur R. KhudaBukhsh
ACL (1)1
2025 All You Need Is S P A C E: When Jailbreaking Meets Bias Audit and Reveals What Lies Beneath the Guardrails (Student Abstract)
abstract
This paper makes a novel combination of a recently proposed bias audit framework and a recently proposed jailbreaking technique for Llama3. On an audit comprising several disadvantaged groups, our experiments reveal that a jailbroken Llama3 exhibits worrisome antisemitism, racism, misogyny, and homophobia (to list a few) much akin to a broad suite of LLMs that were susceptible to similar biases.
Arka Dutta 0001, Aman Priyanshu, Ashiqur R. KhudaBukhsh
AAAI1
2025 Towards a Bipartisan Understanding of Peace and Vicarious Interactions
abstract
Human input plays a critical role in modern AI systems. As machines take on increasingly nuanced tasks, it becomes essential for the community to embrace subjectivity and diverse perspectives. However, research on sensitive topics often fails to incorporate diverse and balanced perspectives. This paper makes a key contribution to participatory AI design in the context of conflicts between nuclear adversaries (India and Pakistan); where disagreement between stakeholders is anticipated. The paper explores the notion of hope speech detection -- detecting de-escalating content in the context of nuclear adversaries on the brink of war -- through the lens of participatory AI design and vicarious interactions. We release a dataset of 10,081 social web posts annotated by raters from India and Pakistan and examine the bipartisan nature of the language of de-escalation. Our study reveals that vicarious perspectives can be useful for modeling out-group preferences.
Arka Dutta 0001, Syed Mohammad Sualeh Ali, Usman Naseem, Ashiqur R. KhudaBukhsh
IJCAI1
2024 Down the Toxicity Rabbit Hole: A Framework to Bias Audit Large Language Models with Key Emphasis on Racism, Antisemitism, and Misogyny
Arka Dutta 0001, Adel Khorramrouz, Sujan Dutta, Ashiqur R. KhudaBukhsh
IJCAI1