VLDB 2026 Research / reviewers in the wild / expert
Pranav Venkit
dblp:287/9127 · also Pranav Narayanan Venkit
· DBLP profile ↗
13ranked-venue papers
8as first author
13since 2021 · last 2026
0000-0002-5671-0461ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 7 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Tale of Two Identities: An Ethical Audit of AI-Crafted Synthetic PersonasabstractAs LLMs (large language models) are increasingly used to generate synthetic personas, particularly in data-limited domains such as health, privacy, and HCI, it becomes necessary to understand how these narratives represent identity, especially that of minority communities. In this paper, we audit synthetic personas generated by 3 LLMs (GPT4o, Gemini 1.5 Pro, Deepseek v2.5) through the lens of representational harm, focusing specifically on racial identity. Using a mixed-methods approach combining close reading, lexical analysis, and a parameterized creativity framework, we compare 1,512 LLM-generated persona to human-authored responses. Our findings reveal that LLMs disproportionately foreground racial markers, overproduce culturally coded language, and construct personas that are syntactically elaborate yet narratively reductive. These patterns result in a range of sociotechnical harms, including stereotyping, exoticism, erasure, and benevolent bias, that are often obfuscated by superficially positive narrations. We formalize this phenomenon as algorithmic othering, where minoritized identities are rendered hypervisible but less authentic. Pranav Venkit, Yingfan Zhou, Sarah Michele Rajtmajer, Shomir Wilson |
AAAI | 1 |
| 2025 | Can Third Parties Read Our Emotions?abstractNatural Language Processing tasks that aim to infer an author’s private states, e.g., emotions and opinions, from their written text, typically rely on datasets annotated by third-party annotators. However, the assumption that third-party annotators can accurately capture authors’ private states remains largely unexamined. In this study, we present human subjects experiments on emotion recognition tasks that directly compare third-party annotations with first-party (author-provided) emotion labels. Our findings reveal significant limitations in third-party annotations—whether provided by human annotators or large language models (LLMs)—in faithfully representing authors’ private states. However, LLMs outperform human annotators nearly across the board. We further explore methods to improve third-party annotation quality. We find that demographic similarity between first-party authors and third-party human annotators enhances annotation performance. While incorporating first-party demographic information into prompts leads to a marginal but statistically significant improvement in LLMs’ performance. We introduce a framework for evaluating the limitations of third-party annotations and call for refined annotation practices to accurately represent and model authors’ private states. Yingfan Zhou, Pranav Venkit, Halima Binte Islam, Sneha Arya, Shomir Wilson, Sarah Michele Rajtmajer |
ACL (1) | 3 |
| 2024 | Do Generative AI Models Output Harm while Representing Non-Western Cultures: Evidence from A Community-Centered ApproachabstractOur research investigates the impact of Generative Artificial Intelligence (GAI) models, specifically text-to-image generators (T2Is), on the representation of non-Western cultures, with a focus on Indian contexts. Despite the transformative potential of T2Is in content creation, concerns have arisen regarding biases that may lead to misrepresentations and marginalizations. Through a Non-Western community-centered approach and grounded theory analysis of 5 focus groups from diverse Indian subcultures, we explore how T2I outputs to English input prompts depict Indian culture and its subcultures, uncovering novel representational harms such as exoticism and cultural misappropriation. These findings highlight the urgent need for inclusive and culturally sensitive T2I systems. We propose design guidelines informed by a sociotechnical perspective, contributing to the development of more equitable and representative GAI technologies globally. Our work underscores the necessity of adopting a community-centered approach to comprehend the sociotechnical dynamics of these models, complementing existing work in this space while identifying and addressing the potential negative repercussions and harms that may arise as these models are deployed on a global scale. Sourojit Ghosh, Pranav Venkit, Sanjana Gautam, Shomir Wilson, Aylin Caliskan |
AIES (1) | 2 |
| 2024 | LLMs Assist NLP Researchers: Critique Paper (Meta-)ReviewingabstractJiangshu Du, Yibo Wang, Wenting Zhao, Zhongfen Deng, Shuaiqi Liu, Renze Lou, Henry Peng Zou, Pranav Narayanan Venkit, Nan Zhang, Mukund Srinath, Haoran Ranran Zhang, Vipul Gupta, Yinghui Li, Tao Li, Fei Wang, Qin Liu, Tianlin Liu, Pengzhi Gao, Congying Xia, Chen Xing, Cheng Jiayang, Zhaowei Wang, Ying Su, Raj Sanjay Shah, Ruohao Guo, Jing Gu, Haoran Li, Kangda Wei, Zihao Wang, Lu Cheng, Surangika Ranathunga, Meng Fang, Jie Fu, Fei Liu, Ruihong Huang, Eduardo Blanco, Yixin Cao, Rui Zhang, Philip S. Yu, Wenpeng Yin. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Jiangshu Du, Yibo Wang 0001, Wenting Zhao 0006, Zhongfen Deng, Shuaiqi Liu 0002, Renze Lou, Henry Peng Zou, Pranav Venkit, Mukund Srinath, Ranran Haoran Zhang, Tao Li 0039, Fei Wang 0060, Qin Liu 0010, Tianlin Liu, Pengzhi Gao, Congying Xia, Chen Xing, Cheng Jiayang, Zhaowei Wang 0003, Raj Sanjay Shah, Ruohao Guo, Haoran Li 0003, Kangda Wei, Zihao Wang 0001, Lu Cheng 0001, Surangika Ranathunga, Fei Liu 0004, Ruihong Huang, Eduardo Blanco 0002, Yixin Cao 0002, Rui Zhang 0037, Philip S. Yu, Wenpeng Yin 0001 |
EMNLP | 8 |
| 2024 | An Audit on the Perspectives and Challenges of Hallucinations in NLPabstractPranav Narayanan Venkit, Tatiana Chakravorti, Vipul Gupta, Heidi Biggs, Mukund Srinath, Koustava Goswami, Sarah Rajtmajer, Shomir Wilson. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Pranav Venkit, Tatiana Chakravorti, Heidi R. Biggs, Mukund Srinath, Koustava Goswami, Sarah Michele Rajtmajer, Shomir Wilson |
EMNLP | 1 |
| 2024 | Race and Privacy in Broadcast Police CommunicationsabstractRadios are essential for the operations of modern police departments, and they function as both a collaborative communication technology and a sociotechnical system. However, little prior research has examined their usage or their connections to individual privacy and the role of race in policing, two growing topics of concern in the US. As a case study, we examine the Chicago Police Department's (CPD's) use of broadcast police communications (BPC) to coordinate the activity of law enforcement officers (LEOs) in the city. From a recently assembled archive of 80, 775 hours of BPC associated with CPD operations, we analyze human-generated text transcripts of radio transmissions broadcast 9:00 AM to 5:00 PM on August 10th, 2018 in one majority Black, one majority White, and one majority Hispanic area of the city (24 hours of audio) to explore four research questions: (1) Do BPC reflect reported racial disparities in policing? (2) How and when is gender, race/ethnicity, and age mentioned in BPC? (3) To what extent do BPC include sensitive information, and who is put at most risk by this practice? (4) To what extent can large language models (LLMs) heighten this risk? We explore the vocabulary and speech acts used by police in BPC, comparing mentions of personal characteristics to local demographics, the personal information shared over BPC, and the privacy concerns that it poses. Analysis indicates (a) policing professionals in the city of Chicago exhibit disproportionate attention to Black members of the public regardless of context, (b) sociodemographic characteristics like gender, race/ethnicity, and age are primarily mentioned in BPC about event information, and (c) disproportionate attention introduces disproportionate privacy risks for Black members of the public. This study shows BPC can provide a novel window into disproportionate attention (i.e., via radio communications) by law enforcement officers to specific racial groups, leading to increased privacy vulnerability for those groups, particularly Black males. Pranav Venkit, Christopher Graziul, Miranda Ardith Goodman, Samantha Nicole Kenny, Shomir Wilson |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2023 | Towards a Holistic Approach: Understanding Sociodemographic Biases in NLP Models using an Interdisciplinary LensabstractThe rapid growth in the usage and applications of Natural Language Processing (NLP) in various sociotechnical solutions has highlighted the need for a comprehensive understanding of bias and its impact on society. While research on bias in NLP has expanded, several challenges persist that require attention. These include the limited focus on sociodemographic biases beyond race and gender, the narrow scope of analysis predominantly centered on models, and the technocentric implementation approaches. Pranav Venkit |
AIES | 1 |
| 2023 | Unmasking Nationality Bias: A Study of Human Perception of Nationalities in AI-Generated ArticlesabstractWe investigate the potential for nationality biases in natural language processing (NLP) models using human evaluation methods. Biased NLP models can perpetuate stereotypes and lead to algorithmic discrimination, posing a significant challenge to the fairness and justice of AI systems. Our study employs a two-step mixed-methods approach that includes both quantitative and qualitative analysis to identify and understand the impact of nationality bias in a text generation model. Through our human-centered quantitative analysis, we measure the extent of nationality bias in articles generated by AI sources. We then conduct open-ended interviews with participants, performing qualitative coding and thematic analysis to understand the implications of these biases on human readers. Our findings reveal that biased NLP models tend to replicate and amplify existing societal biases, which can translate to harm if used in a sociotechnical setting. The qualitative analysis from our interviews offers insights into the experience readers have when encountering such articles, highlighting the potential to shift a reader’s perception of a country. These findings emphasize the critical role of public perception in shaping AI’s impact on society and the need to correct biases in AI systems. Pranav Venkit, Sanjana Gautam, Ruchi Panchanadikar, Ting-Hao 'Kenneth' Huang, Shomir Wilson |
AIES | 1 |
| 2023 | Privacy Now or Never: Large-Scale Extraction and Analysis of Dates in Privacy Policy TextabstractThe General Data Protection Regulation (GDPR) and other recent privacy laws require organizations to post their privacy policies, and place specific expectations on organisations' privacy practices. Privacy policies take the form of documents written in natural language, and one of the expectations placed upon them is that they remain up to date. To investigate legal compliance with this recency requirement at a large scale, we create a novel pipeline that includes crawling, regex-based extraction, candidate date classification and date object creation to extract updated and effective dates from privacy policies written in English. We then analyze patterns in policy dates using four web crawls and find that only about 40% of privacy policies online contain a date, thereby making it difficult to assess their regulatory compliance. We also find that updates in privacy policies are temporally concentrated around passage of laws regulating digital privacy (such as the GDPR), and that more popular domains are more likely to have policy dates as well as more likely to update their policies regularly. Mukund Srinath, Lee Matheson, Pranav Venkit, Gabriela Zanfir-Fortuna, Florian Schaub, C. Lee Giles, Shomir Wilson |
DocEng | 3 |
| 2023 | Privacy Lost and Found: An Investigation at Scale of Web Privacy Policy AvailabilityabstractLegal jurisdictions around the world require organisations to post privacy policies on their websites. However, in spite of laws such as GDPR and CCPA reinforcing this requirement, organisations sometimes do not comply, and a variety of semi-compliant failure modes exist. To investigate the landscape of web privacy policies, we crawl the privacy policies from 7 million organisation websites with the goal of identifying when policies are unavailable. We conduct a large-scale investigation of the availability of privacy policies and identify potential reasons for unavailability such as dead links, documents with empty content, documents that consist solely of placeholder text, and documents unavailable in the specific languages offered by their respective websites. We estimate the frequencies of these failure modes and the overall unavailability of privacy policies on the web and find that privacy policies URLs are only available in 34% of websites. Further, 1.37% of these URLs are broken links and 1.23% of the valid links lead to pages without a policy. Further, to enable investigation of privacy policies at scale, we use the capture-recapture technique to estimate the total number of English language privacy policies on the web and the distribution of these documents across top level domains and sectors of commerce. We estimate the lower bound on the number of English language privacy policies to be around 3 million. Finally, we release the CoLIPPs Corpus containing around 600k policies and their metadata consisting of policy URL, length, readability, sector of commerce, and policy crawl date. Mukund Srinath, Soundarya Nurani Sundareswara, Pranav Venkit, C. Lee Giles, Shomir Wilson |
DocEng | 3 |
| 2023 | Nationality Bias in Text GenerationabstractPranav Narayanan Venkit, Sanjana Gautam, Ruchi Panchanadikar, Ting-Hao Huang, Shomir Wilson. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Pranav Venkit, Sanjana Gautam, Ruchi Panchanadikar, Ting-Hao 'Kenneth' Huang, Shomir Wilson |
EACL | 1 |
| 2023 | The Sentiment Problem: A Critical Survey towards Deconstructing Sentiment AnalysisabstractPranav Venkit, Mukund Srinath, Sanjana Gautam, Saranya Venkatraman, Vipul Gupta, Rebecca Passonneau, Shomir Wilson. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Pranav Venkit, Mukund Srinath, Sanjana Gautam, Saranya Venkatraman, Rebecca J. Passonneau, Shomir Wilson |
EMNLP | 1 |
| 2022 | A Study of Implicit Bias in Pretrained Language Models against People with DisabilitiesabstractPretrained language models (PLMs) have been shown to exhibit sociodemographic biases, such as against gender and race, raising concerns of downstream biases in language technologies. However, PLMs’ biases against people with disabilities (PWDs) have received little attention, in spite of their potential to cause similar harms. Using perturbation sensitivity analysis, we test an assortment of popular word embedding-based and transformer-based PLMs and show significant biases against PWDs in all of them. The results demonstrate how models trained on large corpora widely favor ableist language. Pranav Venkit, Mukund Srinath, Shomir Wilson |
COLING | 1 |