VLDB 2026 Research / reviewers in the wild / expert
Aparup Khatua
dblp:122/2151
· DBLP profile ↗
8ranked-venue papers in the field
6as first author
4since 2021 · last 2025
0000-0001-8235-1637ORCID · corroborated
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 5 (4 first)Data Mining & Knowledge Discovery · 3 (2 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Evaluating LLMs' (In)ability to Follow Prompts in QA TasksabstractWhile LLMs have achieved impressive performance across various tasks, one under-explored area is evaluating their ability to follow instructions provided in the prompt when generating responses. In the context of question-answering (QA) tasks, a crucial research gap is whether LLMs prioritize their own parametric knowledge or the context provided in the prompt when generating an answer. Ignoring prompts, even when explicitly instructed to follow them, may adversely affect performance and potentially lead to unintended consequences. Additionally, LLMs should be self-reflective (i.e., LLMs should recognize when their knowledge is inadequate) and avoid hallucinations in such scenarios. To address our research question, we propose Oedipus, an evaluation framework to evaluate LLMs' ability to follow prompts. We further note that such abilities could also be influenced by contamination (i.e., exposure to datasets during training) and parametric knowledge. Consequently, we develop a novel QA dataset with four types of contexts- correct, masked, noisy, and absurd contexts with recent questions that LLMs are unlikely to have encountered in pre-training data or corpus and cannot be answered from parametric knowledge. We evaluate eight LLMs through our proposed evaluation framework and observe that LLMs often fail to follow instructions correctly and are not self-reflective. Aparup Khatua, Tobias Kalmbach, Prasenjit Mitra 0001, Sandipan Sikdar |
SIGIR | 1 |
| 2024 | Beyond Accuracy: Investigating Error Types in GPT-4 Responses to USMLE QuestionsabstractGPT-4 demonstrates high accuracy in medical QA tasks, leading with an accuracy of 86.70%, followed by Med-PaLM 2 at 86.50%. However, around 14% of errors remain. Additionally, current works use GPT-4 to only predict the correct option without providing any explanation and thus do not provide any insight into the thinking process and reasoning used by GPT-4 or other LLMs. Therefore, we introduce a new domain-specific error taxonomy derived from collaboration with medical students. Our GPT-4 USMLE Error (G4UE) dataset comprises 4153 GPT-4 correct responses and 919 incorrect responses to the United States Medical Licensing Examination (USMLE) respectively. These responses are quite long (258 words on average), containing detailed explanations from GPT-4 justifying the selected option. We then launch a large-scale annotation study using the Potato annotation platform and recruit 44 medical experts through Prolific, a well-known crowdsourcing platform. We annotated 300 out of these 919 incorrect data points at a granular level for different classes and created a multi-label span to identify the reasons behind the error. In our annotated dataset, a substantial portion of GPT-4's incorrect responses is categorized as a "Reasonable response by GPT-4," by annotators. This sheds light on the challenge of discerning explanations that may lead to incorrect options, even among trained medical professionals. We also provide medical concepts and medical semantic predications extracted using the SemRep tool for every data point. We believe that it will aid in evaluating the ability of LLMs to answer complex medical questions. We make the resources available at https://github.com/roysoumya/usmle-gpt4-error-taxonomy. Soumyadeep Roy, Aparup Khatua, Fatemeh Ghoochani, Uwe Hadler, Wolfgang Nejdl, Niloy Ganguly |
SIGIR | 2 |
| 2023 | Host-Centric Social Connectedness of Migrants in Europe on FacebookabstractExtant literature has explored the social integration process of migrants settling in host communities. However, this literature typically takes a migrant-centric view, implicitly putting the burden of a successful integration on the migrant, and trying to identify the factors that lead to integration along various dimensions. In this paper, we flip this point of view by studying the attributes of natives that govern their propensity to form social ties with migrants.We do so by using anonymous and aggregate social network data provided by Facebook’s advertising platform. More specifically, we look at factors that influence the propensity for a likely-to-be non-Muslim Facebook user to have at least one social connection to a Facebook user who celebrates Ramadan. Given that, in the European context, following Islam is predominantly tied to a migration background, this gives us a lens into cross-cultural native-migrant connectivity. Our study considers demographic attributes of the host population, such as age, gender, and education level, as well as spatial variation across 30 European cities. Our findings suggest that young, educated, and male Facebook users are relatively more likely to build cross-cultural ties, compared to older, less educated, and female Facebook users. We also observe heterogeneity across the analyzed cities. Aparup Khatua, Emilio Zagheni, Ingmar Weber |
ICWSM | 1 |
| 2022 | Unraveling Social Perceptions & Behaviors towards Migrants on Twitter
Aparup Khatua, Wolfgang Nejdl |
ICWSM | 1 |
| 2020 | Matching Recruiters and Jobseekers on TwitterabstractAn efficient job recommendation framework needs to recommend an appropriate jobseeker to a recruiter and vice-versa. Prior studies have mostly considered datasets from commercial job portals such as LinkedIn or CareerBuilder. However, these datasets are proprietary and not publicly available. Moreover, these portals charge their clients for offering customized services. Hence, we explore whether publicly available Twitter data can be a viable alternative to commercial job portals. We have extracted 0.76 million job-related tweets. We have manually annotated tweet-pairs from recruiters and jobseekers in the domain of computer science jobs. Next, we have employed Siamese architecture and considered multiple artificial neural network models with different word embeddings. We have achieved around 97% accuracy for some of our models. Our study demonstrates the potential of the Twitter platform for job recommendations. Aparup Khatua, Wolfgang Nejdl |
ASONAM | 1 |
| 2019 | A tale of two epidemics: Contextual Word2Vec for classifying twitter streams during outbreaks
Aparup Khatua, Apalak Khatua, Erik Cambria |
Inf. Process. Manag. | 1 |
| 2018 | Sounds of Silence Breakers: Exploring Sexual Violence on TwitterabstractGender-based-violence is a serious concern in recent times. Due to the social stigma attached to these assaults, victims rarely come forward. Implementing policy measures to prevent sexual violence get constrained due to lack of crime statistics. However, the recent outcry on the Twitter platform allows us to address this concern. Sexual assaults occur at workplaces, public places, educational institutes and also at home. Policy level approaches and awareness campaign for these assaults would not be similar. So, we want to identify the risk factor associated with these sexual assaults. We extracted 0.7 million tweets during the #MeToo social media movement. Next, we employ deep learning techniques to classify these sexual violences. We observe that sexual assaults by a family member at own home is a more serious concern than harassment by a stranger at public places. This study reveals assaults by a known person are more prevalent than assaults by unknown strangers. Aparup Khatua, Erik Cambria, Apalak Khatua |
ASONAM | 1 |
| 2017 | Cricket World Cup 2015: Predicting User's Orientation through Mix Tweets on Twitter PlatformabstractThe existing literature has explored various latent attributes of Twitter users. It is worth noting that user classification research in the context of gender or the political domain is mostly binary in nature such as male-female or Republican-Democrat. Conversely, in multi-team contexts user classification is a not a binary task. Also, prior studies have mostly ignored tweets which mention more than one orientation (i.e. both Republican and Democrat-related keywords) within a tweet. We consider these tweets as mix tweets. We investigate the relationship between user classification (in a multi-team context) and user-level mix tweeting pattern. To test our proposed model, we have extracted 3.5 million tweets during the Cricket World Cup 2015 (CWC'15), in which 14 cricket-playing nations participated. We employed a logistic regression model, and our empirical evidence strongly confirms our hypothesis. Apalak Khatua, Aparup Khatua |
ASONAM | 2 |