VLDB 2026 Research / reviewers in the wild / expert
Shweta Yadav 0001
dblp:206/5459 · also Shweta 0001
· DBLP profile ↗
23ranked-venue papers
16as first author
10since 2021 · last 2026
0000-0003-0001-4464ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 11 first-author · 5 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 1 since 2021Computer networks · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MAMA-Memeia! Multi-Aspect Multi-Agent Collaboration for Depressive Symptoms Identification in MemesabstractOver the past years, memes have evolved from being exclusively a medium of humorous exchanges to one that allows users to express a range of emotions freely and easily. With the ever-growing utilization of memes in expressing depressive sentiments, we conduct a study on identifying depressive symptoms exhibited by memes shared by users of online social media platforms. We introduce RESTOREx as a vital resource for detecting depressive symptoms in memes on social media through the Large Language Model (LLM) generated and human-annotated explanations. We introduce MAMA-Memeia, a collaborative multi-agent multi-aspect discussion framework grounded in the clinical psychology method of Cognitive Analytic Therapy (CAT) Competencies. MAMA-Memeia improves upon the current state-of-the-art by 7.55% in macro-F1 and is established as the new benchmark compared to over 30 methods. Siddhant Agarwal, Adya Dhuler, Polly Ruhnke, Melvin Speisman, Md. Shad Akhtar, Shweta Yadav 0001 |
AAAI | 6 |
| 2025 | Figurative-cum-Commonsense Knowledge Infusion for Multimodal Mental Health Meme ClassificationabstractThe expression of mental health symptoms through non-traditional means, such as memes, has gained remarkable attention over the past few years, with users often highlighting their mental health struggles through figurative intricacies within memes. While humans rely on commonsense knowledge to interpret these complex expressions, current Multimodal Language Models (MLMs) struggle to capture these figurative aspects inherent in memes. To address this gap, we introduce a novel dataset, AxiOM, derived from the GAD anxiety questionnaire, which categorizes memes into six fine-grained anxiety symptoms. Next, we propose a commonsense and domain-enriched framework, M3H, to enhance MLMs' ability to interpret figurative language and commonsense knowledge. The overarching goal remains to first understand and then classify the mental health symptoms expressed in memes. We benchmark M3H against 6 competitive baselines (with 20 variations), demonstrating improvements in both quantitative and qualitative metrics, including a detailed human evaluation. We observe a clear improvement of 4.20% and 4.66% on weighted-F1 metric. To assess the generalizability, we perform extensive experiments on a public dataset, RESTORE, for depressive symptom identification, presenting an ablation study that highlights the contribution of each module. Our findings reveal limitations in existing models and the advantage of employing commonsense to enhance figurative understanding. Abdullah Mazhar, Zuhair Hasan Shaik, Aseem Srivastava, Polly Ruhnke, Lavanya Vaddavalli, Sri Keshav Katragadda, Shweta Yadav 0001, Md. Shad Akhtar |
WWW | 7 |
| 2023 | Towards Identifying Fine-Grained Depression Symptoms from MemesabstractShweta Yadav, Cornelia Caragea, Chenye Zhao, Naincy Kumari, Marvin Solberg, Tanmay Sharma. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Shweta Yadav 0001, Cornelia Caragea, Chenye Zhao, Naincy Kumari, Marvin Solberg, Tanmay Sharma |
ACL (1) | 1 |
| 2023 | Partisan US News Media Representations of Syrian RefugeesabstractWe investigate how representations of Syrian refugees (2011-2021) differ across US partisan news outlets. We analyze 47,388 articles from the online US media about Syrian refugees to detail differences in reporting between left- and right-leaning media. We use various NLP techniques to understand these differences. Our polarization and question answering results indicated that left-leaning media tended to represent refugees as child victims, welcome in the US, and right-leaning media cast refugees as Islamic terrorists. We noted similar results with our sentiment and offensive speech scores over time, which detail possibly unfavorable representations of refugees in right-leaning media. A strength of our work is how the different techniques we have applied validate each other. Based on our results, we provide several recommendations. Stakeholders may utilize our findings to intervene around refugee representations, and design communications campaigns that improve the way society sees refugees and possibly aid refugee outcomes. Marzieh Babaeianjelodar, Yiwen Shi, Kamila Janmohamed, Rupak Sarkar, Ingmar Weber, Thomas Davidson, Munmun De Choudhury, Jonathan Huang, Shweta Yadav 0001, Ashiqur R. KhudaBukhsh, Chris T. Bauch, Preslav Nakov, Orestis Papakyriakopoulos, Koustuv Saha, Kaveh Khoshnood, Navin Kumar 0004 |
ICWSM | 10 |
| 2023 | Towards Understanding Consumer Healthcare Questions on the Web with Semantically Enhanced Contrastive LearningabstractIn recent years, seeking health information on the web has become a preferred way for healthcare consumers to support their information needs. Generally, healthcare consumers use long and detailed questions with several peripheral details to express their healthcare concerns, contributing to natural language understanding challenges. One way to address this challenge is by summarizing the questions. However, most of the existing abstractive summarization systems generate impeccably fluent yet factually incorrect summaries. In this paper, we present a semantically-enhanced contrastive learning-based framework for generating abstractive question summaries that are faithful and factually correct. We devised multiple strategies based on question semantics to generate the erroneous (negative) summaries, such that the model has the understanding of plausible and incorrect perturbations of the original summary. Our extensive experimental results on two benchmark consumer health question summarization datasets confirm the effectiveness of our proposed method by achieving state-of-the-art performance and generating factually correct and fluent summaries, as measured by human evaluation. Shweta Yadav 0001, Stefan Cobeli, Cornelia Caragea |
WWW | 1 |
| 2022 | Towards Summarizing Healthcare Questions in Low-Resource SettingabstractThe current advancement in abstractive document summarization depends to a large extent on a considerable amount of human-annotated datasets. However, the creation of large-scale datasets is often not feasible in closed domains, such as medical and healthcare domains, where human annotation requires domain expertise. This paper presents a novel data selection strategy to generate diverse and semantic questions in a low-resource setting with the aim to summarize healthcare questions. Our method exploits the concept of guided semantic-overlap and diversity-based objective functions to optimally select the informative and diverse set of synthetic samples for data augmentation. Our extensive experiments on benchmark healthcare question summarization datasets demonstrate the effectiveness of our proposed data selection strategy by achieving new state-of-the-art results. Our human evaluation shows that our method generates diverse, fluent, and informative summarized questions. Shweta Yadav 0001, Cornelia Caragea |
COLING | 1 |
| 2022 | Towards Enhancing Health Coaching Dialogue in Low-Resource SettingsabstractHealth coaching helps patients identify and accomplish lifestyle-related goals, effectively improving the control of chronic diseases and mitigating mental health conditions. However, health coaching is cost-prohibitive due to its highly personalized and labor-intensive nature. In this paper, we propose to build a dialogue system that converses with the patients, helps them create and accomplish specific goals, and can address their emotions with empathy. However, building such a system is challenging since real-world health coaching datasets are limited and empathy is subtle. Thus, we propose a modularized health coaching dialogue with simplified NLU and NLG frameworks combined with mechanism-conditioned empathetic response generation. Through automatic and human evaluation, we show that our system generates more empathetic, fluent, and coherent responses and outperforms the state-of-the-art in NLU tasks while requiring less annotation. We view our approach as a key step towards building automated and more accessible health coaching systems. Barbara Di Eugenio, Brian D. Ziebart, Lisa K. Sharp, Bing Liu 0001, Ben S. Gerber, Nikolaos Agadakos, Shweta Yadav 0001 |
COLING | 8 |
| 2022 | Detecting Optimism in Tweets using Knowledge Distillation and Linguistic Analysis of OptimismabstractFinding the polarity of feelings in texts is a far-reaching task. Whilst the field of natural language processing has established sentiment analysis as an alluring problem, many feelings are left uncharted. In this study, we analyze the optimism and pessimism concepts from Twitter posts to effectively understand the broader dimension of psychological phenomenon. Towards this, we carried a systematic study by first exploring the linguistic peculiarities of optimism and pessimism in user-generated content. Later, we devised a multi-task knowledge distillation framework to simultaneously learn the target task of optimism detection with the help of the auxiliary task of sentiment analysis and hate speech detection. We evaluated the performance of our proposed approach on the benchmark Optimism/Pessimism Twitter dataset. Our extensive experiments show the superior- ity of our approach in correctly differentiating between optimistic and pessimistic users. Our human and automatic evaluation shows that sentiment analysis and hate speech detection are beneficial for optimism/pessimism detection. Stefan Cobeli, Ioan-Bogdan Iordache, Shweta Yadav 0001, Cornelia Caragea, Liviu P. Dinu, Dragos Iliescu |
LREC | 3 |
| 2022 | Question-aware transformer models for consumer health question summarization
Shweta Yadav 0001, Asma Ben Abacha, Dina Demner-Fushman |
J. Biomed. Informatics | 1 |
| 2022 | Relation Extraction From Biomedical and Clinical Text: Unified Multitask Learning FrameworkabstractMOTIVATION: To minimize the accelerating amount of time invested on the biomedical literature search, numerous approaches for automated knowledge extraction have been proposed. Relation extraction is one such task where semantic relations between the entities are identified from the free text. In the biomedical domain, extraction of regulatory pathways, metabolic processes, adverse drug reaction or disease models necessitates knowledge from the individual relations, for example, physical or regulatory interactions between genes, proteins, drugs, chemical, disease or phenotype. RESULTS: In this paper, we study the relation extraction task from three major biomedical and clinical tasks, namely drug-drug interaction, protein-protein interaction, and medical concept relation extraction. Towards this, we model the relation extraction problem in a multi-task learning (MTL)framework, and introduce for the first time the concept of structured self-attentive network complemented with the adversarial learning approach for the prediction of relationships from the biomedical and clinical text. The fundamental notion of MTL is to simultaneously learn multiple problems together by utilizing the concepts of the shared representation. Additionally, we also generate the highly efficient single task model which exploits the shortest dependency path embedding learned over the attentive gated recurrent unit to compare our proposed MTL models. The framework we propose significantly improves over all the baselines (deep learning techniques)and single-task models for predicting the relationships, without compromising on the performance of all the tasks. Shweta Yadav 0001, Srivatsa Ramesh Jayashree, Sriparna Saha 0001, Asif Ekbal |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2020 | Identifying Depressive Symptoms from Tweets: Figurative Language Enabled Multitask Learning FrameworkabstractExisting studies on using social media for deriving mental health status of users focus on the depression detection task.However, for case management and referral to psychiatrists, healthcare workers require practical and scalable depressive disorder screening and triage system.This study aims to design and evaluate a decision support system (DSS) to reliably determine the depressive triage level by capturing fine-grained depressive symptoms expressed in user tweets through the emulation of Patient Health Questionnaire-9 (PHQ-9) that is routinely used in clinical practice.The reliable detection of depressive symptoms from tweets is challenging because the 280-character limit on tweets incentivizes the use of creative artifacts in the utterances and figurative usage contributes to effective expression.We propose a novel BERT based robust multi-task learning framework to accurately identify the depressive symptoms using the auxiliary task of figurative usage detection.Specifically, our proposed novel task sharing mechanism, co-task aware attention, enables automatic selection of optimal information across the BERT layers and tasks by soft-sharing of parameters.Our results show that modeling figurative usage can demonstrably improve the model's robustness and reliability for distinguishing the depression symptoms. Shweta Yadav 0001, Jainish Chauhan, Joy Prakash Sain, Krishnaprasad Thirunarayan, Amit P. Sheth, Jeremiah Schumm |
COLING | 1 |
| 2020 | Medical Knowledge-enriched Textual Entailment FrameworkabstractOne of the cardinal tasks in achieving robust medical question answering systems is textual entailment.The existing approaches make use of an ensemble of pre-trained language models or data augmentation, often to clock higher numbers on the validation metrics.However, two major shortcomings impede higher success in identifying entailment: (1) understanding the focus/intent of the question and (2) ability to utilize the real-world background knowledge to capture the context beyond the sentence.In this paper, we present a novel Medical Knowledge-Enriched Textual Entailment framework that allows the model to acquire a semantic and global representation of the input medical text with the help of a relevant domain-specific knowledge graph.We evaluate our framework on the benchmark MEDIQA-RQE dataset and manifest that the use of knowledgeenriched dual-encoding mechanism help in achieving an absolute improvement of 8.27% over SOTA language models.We have made the source code available here. 1 Shweta Yadav 0001, Vishal Pallagani, Amit P. Sheth |
COLING | 1 |
| 2020 | Assessing the Severity of Health States based on Social Media PostsabstractThe unprecedented growth of Internet users has resulted in an abundance of unstructured information on social media including health forums, where patients request health-related information or opinions from other users. Previous studies have shown that online peer support has limited effectiveness without expert intervention. Therefore, a system capable of assessing the severity of health state from the patients' social media posts can help health professionals (HP) in prioritizing the user's post. In this study, we inspect the efficacy of different aspects of Natural Language Understanding (NLU) to identify the severity of the user's health state in relation to two perspectives(tasks) (a) Medical Condition (i.e., Recover, Exist, Deteriorate, Other) and (b) Medication (i.e., Effective, Ineffective, Serious Adverse Effect, Other) in online health communities. We propose a multiview learning framework that models both the textual content as well as contextual-information to assess the severity of the user's health state. Specifically, our model utilizes the NLU views such as sentiment, emotions, personality, and use of figurative language to extract the contextual information. The diverse NLU views demonstrate its effectiveness on both the tasks and as well as on the individual disease to assess a user's health. Shweta Yadav 0001, Joy Prakash Sain, Amit P. Sheth, Asif Ekbal, Sriparna Saha 0001, Pushpak Bhattacharyya |
ICPR | 1 |
| 2020 | eDarkFind: Unsupervised Multi-view Learning for Sybil Account DetectionabstractDarknet crypto markets are online marketplaces using crypto currencies (e.g., Bitcoin, Monero) and advanced encryption techniques to offer anonymity to vendors and consumers trading for illegal goods or services. The exact volume of substances advertised and sold through these crypto markets is difficult to assess, at least partially, because vendors tend to maintain multiple accounts (or Sybil accounts) within and across different crypto markets. Linking these different accounts will allow us to accurately evaluate the volume of substances advertised across the different crypto markets by each vendor. In this paper, we present a multi-view unsupervised framework (eDarkFind) that helps modeling vendor characteristics and facilitates Sybil account detection. We employ a multi-view learning paradigm to generalize and improve the performance by exploiting the diverse views from multiple rich sources such as BERT, stylometric, and location representation. Our model is further tailored to take advantage of domain-specific knowledge such as the Drug Abuse Ontology to take into consideration the substance information. We performed extensive experiments and demonstrated that the multiple views obtained from diverse sources can be effective in linking Sybil accounts. Our proposed eDarkFind model achieves an accuracy of 98% on three real-world datasets which shows the generality of the approach. Ramnath Kumar, Shweta Yadav 0001, Raminta Daniulaityte, Francois R. Lamy, Krishnaprasad Thirunarayan, Usha Lokala, Amit P. Sheth |
WWW | 2 |
| 2020 | Exploring Disorder-Aware Attention for Clinical Event ExtractionabstractEvent extraction is one of the crucial tasks in biomedical text mining that aims to extract specific information concerning incidents embedded in the texts. In this article, we propose a deep learning framework that aims to identify the attributes (severity, course, temporal expression, and document creation time) associated with the medical concepts extracted from electronic medical records. The bi-directional long short-term memory network assisted by the attention mechanism is utilized to uncover the important aspects of the patient’s medical conditions. The attention mechanism specific to the medical disorder mention can focus on various parts of the sentence when different disorders are considered as input. The proposed methodology is evaluated on benchmark ShARe/CLEF eHealth Evaluation Lab 2014 shared task 2 datasets. In addition to the CLEF dataset, we also used the social media text, especially the medical blog posts. Experimental results of the proposed approach illustrate that our proposed approach achieves significant performance improvements over the state-of-the-art techniques and the highly competitive deep learning--based baseline methods. Shweta Yadav 0001, Pralay Ramteke, Asif Ekbal, Sriparna Saha 0001, Pushpak Bhattacharyya |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2019 | A Unified Multi-task Adversarial Learning Framework for Pharmacovigilance MiningabstractThe mining of adverse drug reaction (ADR) has a crucial role in the pharmacovigilance.The traditional ways of identifying ADR are reliable but time-consuming, non-scalable and offer a very limited amount of ADR relevant information.With the unprecedented growth of information sources in the forms of social media texts (Twitter, Blogs, Reviews etc.), biomedical literature, and Electronic Medical Records (EMR), it has become crucial to extract the most pertinent ADR related information from these free-form texts.In this paper, we propose a neural network inspired multitask learning framework that can simultaneously extract ADRs from various sources.We adopt a novel adversarial learning-based approach to learn features across multiple ADR information sources.Unlike the other existing techniques, our approach is capable to extracting fine-grained information (such as 'Indications', 'Symptoms', 'Finding', 'Disease', 'Drug') which provide important cues in pharmacovigilance.We evaluate our proposed approach on three publicly available realworld benchmark pharmacovigilance datasets, a Twitter dataset from PSB 2016 Social Media Shared Task, CADEC corpus and Medline ADR corpus.Experiments show that our unified framework achieves state-of-the-art performance on individual tasks associated with the different benchmark datasets.This establishes the fact that our proposed approach is generic, which enables it to achieve high performance on the diverse datasets.The source code is available here 1 . Shweta Yadav 0001, Asif Ekbal, Sriparna Saha 0001, Pushpak Bhattacharyya |
ACL (1) | 1 |
| 2019 | Information theoretic-PSO-based feature selection: an application in biomedical entity extraction
Shweta Yadav 0001, Asif Ekbal, Sriparna Saha 0001 |
Knowl. Inf. Syst. | 1 |
| 2019 | Feature assisted stacked attentive shortest dependency path based Bi-LSTM model for protein-protein interaction
Shweta Yadav 0001, Asif Ekbal, Sriparna Saha 0001, Pushpak Bhattacharyya |
Knowl. Based Syst. | 1 |
| 2018 | Medical Sentiment Analysis using Social Media: Towards building a Patient Assisted System
Shweta Yadav 0001, Asif Ekbal, Sriparna Saha 0001, Pushpak Bhattacharyya |
LREC | 1 |
| 2018 | Feature selection for entity extraction from multiple biomedical corpora: A PSO-based approach
Shweta Yadav 0001, Asif Ekbal, Sriparna Saha 0001 |
Soft Comput. | 1 |
| 2017 | Entity Extraction in Biomedical Corpora: An Approach to Evaluate Word Embedding Features with PSO based Feature SelectionabstractShweta Yadav, Asif Ekbal, Sriparna Saha, Pushpak Bhattacharyya. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017. Shweta Yadav 0001, Asif Ekbal, Sriparna Saha 0001, Pushpak Bhattacharyya |
EACL (1) | 1 |
| 2016 | A deep learning architecture for protein-protein Interaction Article identificationabstractIn recent past there has been phenomenal growth in biomedical literature and health care records. Robust text mining techniques are essential in order to properly organize the documents as well as to extract relevant information. Traditional techniques for document classification focus on machine learning algorithms where learning of classifier is decided on the basis of labeled data and the features that are prominent. In this paper we focus on developing an automated technique for classifying biomedical articles containing protein-protein interaction related information against the others. Our proposed approach is based on deep neural network framework. We investigate the role of convolution neural network (CNN) and propose two model variants. We evaluate the proposed approach on the benchmark datasets of BioCreative-II Interaction Article Subtask (IAS) data sets. Effectiveness of our proposed model is evident with the significant performance gains, 2.8 % in terms of F-measure and 5 % in terms of accuracy over the traditional models. Shweta Yadav 0001, Asif Ekbal, Sriparna Saha 0001, Pushpak Bhattacharyya |
ICPR | 1 |
| 2015 | PSO-ASent: Feature Selection Using Particle Swarm Optimization for Aspect Based Sentiment Analysis
Kandula Srikanth Reddy, Shweta Yadav 0001, Asif Ekbal |
NLDB | 3 |