EDBT 2026 Demo / reviewers in the wild / expert
Shaina Raza
dblp:236/9245
· DBLP profile ↗
10ranked-venue papers in the field
6as first author
8since 2021 · last 2026
0000-0003-1061-5845ORCID · corroborated
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 4 (3 first)Data Mining & Knowledge Discovery · 3 (2 first)Information Retrieval & Web Search · 2 (1 first)Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BEADS: Bias Evaluation Across DomainsabstractRecent advances in large language models (LLMs) have substantially improved natural language processing (NLP) applications. However, these models often inherit and amplify biases present in their training data. Although several datasets exist for bias detection, most are limited to one or two NLP tasks, typically classification or evaluation and do not provide broad coverage across diverse task settings. To address this gap, we introduce the Bias Evaluations Across Domains (BEADs) dataset, designed to support a wide range of NLP tasks, including text classification, token classification, bias quantification, and benign language generation. A key contribution of this work is a gold-standard annotation scheme that supports both evaluation and supervised training of language models. Experiments on state-of-the-art models reveal some gaps: some models exhibit systematic bias toward specific demographics, while others apply safety guardrails more strictly or inconsistently across groups. Overall, these results highlight persistent shortcomings in current models and underscore the need for comprehensive bias evaluation. The benchmark will be made publicly available for research ( https://huggingface.co/datasets/shainar/BEAD ). Shaina Raza, Michael R. Zhang |
NLDB | 1 |
| 2025 | Development and Evaluation of an Agentic RAG with Embedded LLM-as-a Judge for Heart Failure Management: A Pilot Study
Broderick Bellard, Xinyi Celine Liu, Raima Lohani, Shaina Raza, Pedro Elkind Velmovitsky, Quynh Pham, Shumit Saha |
IEEE Big Data | 4 |
| 2025 | Fake news detection: comparative evaluation of BERT-like models and large language models with generative AI-annotated data
Shaina Raza, Drai Paulen-Patterson, Chen Ding 0004 |
Knowl. Inf. Syst. | 1 |
| 2024 | Reliability Analysis of Psychological Concept Extraction and Classification in User-Penned TextabstractThe social NLP research community witness a recent surge in the computational advancements of mental health analysis to build responsible AI models for a complex interplay between language use and self-perception. Such responsible AI models aid in quantifying the psychological concepts from user-penned texts on social media. On thinking beyond the low-level (classification) task, we advance the existing binary classification dataset, towards a higher-level task of reliability analysis through the lens of explanations, posing it as one of the safety measures. We annotate the LoST dataset to capture nuanced textual cues that suggest the presence of low self-esteem in the posts of Reddit users. We further state that the NLP models developed for determining the presence of low self-esteem, focus more on three types of textual cues: (i) Trigger: words that triggers mental disturbance, (ii) LoST indicators: text indicators emphasizing low self-esteem, and (iii) Consequences: words describing the consequences of mental disturbance. We implement existing classifiers to examine the attention mechanism in pre-trained language models (PLMs) for a domain-specific psychology-grounded task. Our findings suggest the need of shifting the focus of PLMs from Trigger and Consequences to a more comprehensive explanation, emphasizing LoST indicators while determining low self-esteem in Reddit posts. Muskan Garg, MSVPJ Sathvik, Shaina Raza, Amrit Chadha, Sunghwan Sohn |
ICWSM | 3 |
| 2024 | An Intelligent COVID-19-Related Arabic Text Detection Framework Based on Transfer Learning Using Context RepresentationabstractThe misleading information during the coronavirus disease 2019 (COVID-19) pandemic’s peak time is very sensitive and harmful in our community. Analyzing and detecting COVID-19 information on social media are a crucial task. Early detection of COVID-19 information is very helpful and minimizes the risk of psychological security which leads to inconvenience in daily life. In this paper, a deep ensemble transfer learning framework with an understanding of the context of Arabic text COVID-19 information is proposed. This framework is inspired to spontaneously analyze and recognize the text about COVID-19. The ArCOVID-19Vac dataset has been used to train and test our proposed model. A comprehensive experimental study for each scenario is performed. For the binary classification scenario, the proposed framework records better evaluation results with 83.0%, 84.0%, 83.0%, and 84.0% in terms of accuracy, precision, recall, and F1-score, respectively. For the second scenario (three classes), the overall performance is recorded with an accuracy of 82.0%, precision of 80.0%, recall of 82.0%, and F1-score of 80.0%, respectively. In the last scenario with ten classes, the best evaluation performance results are recorded with an accuracy of 67.0%, a precision of 58.0%, a recall of 67.0%, and F1-score of 59.0%, respectively. In addition, we have applied an ensemble transfer learning model for this scenario to get 64.0%, 66.0%, 66.0%, and 65.0% in terms of accuracy, precision, recall, and F1-score, respectively. The results show that the proposed model through transfer learning provides better results for Arabic text than all state-of-the-art methods. Abdullah Yahya Mohammed Muaad, Shaina Raza, Md Belal Bin Heyat, Amerah A. Alabrah, Hanumanthappa J |
Int. J. Intell. Syst. | 2 |
| 2023 | PLNCC: Leveraging New Data Features for Enhanced Accuracy of Fake News DetectionabstractThe prominence of social media poses a significant threat to information integrity as the spread of fake news increases. It becomes imperative for online media outlets to develop effective strategies to mitigate the spread of fake news. In this research, the PLNCC dataset is introduced as an expansion of two state-of-the-art fake news datasets. The objective is to improve the classification of fake news by extracting additional linguistic and psychological features, as well as user comment data. In this work, a quantitative analysis of the linguistic and psychological features of fake news articles and related user comments is performed. The efficacy of the PLNCC dataset is demonstrated through rigorous evaluation, showcasing its performance against state-of-the-art benchmark datasets. The classification models running on this dataset achieved a significant performance improvement, up to 10%, when compared to the original two datasets. Keshopan Arunthavachelvan, Shaina Raza, Chen Ding 0004 |
ASONAM | 2 |
| 2022 | Incorporating Accuracy and Diversity in a News Recommender SystemabstractThere are certain challenges in news recommender systems that arise due to changing users’ preferences over dynamically generated news articles. It is important to expose users to a variety of information. Diversity is required in a news recommender system not only so that users do not get bored of reading similar news but because so that they do not get trapped in information bubbles. We propose a deep neural network based on a two-tower architecture that learns news representation through a news item tower and users’ representations through a query tower. To learn diversity, we introduce a category loss function that aligns items’ representation of uneven news categories. Experimental results on two news datasets reveal that our proposed architecture is more effective compared to the state-of-the-art methods and achieves a balance between accuracy and diversity. Shaina Raza, Syed Raza Bashir, Usman Naseem, Dora D. Liu, Deepak John Reji |
DSAA | 1 |
| 2021 | Deep Neural Network to Tradeoff between Accuracy and Diversity in a News Recommender SystemabstractThe news recommender systems are marked by a few unique challenges specific to the news domain. These challenges emerge from rapidly evolving readers’ interests over dynamically generated news items that continuously change over time. News reading is driven by a blend of a reader’s long-term and short-term interests. In addition, diversity is required in a news recommender system to keep the reader engaged in the reading process and get them exposed to different views and opinions. This paper proposes a deep neural network that jointly learns news and user representation in a unified framework. It learns the news representation (features) from the headlines, snippets (body) and taxonomy (category, subcategory) of news. The attention mechanism learns a reader’s long-term interests from the complete click history, short-term interests from recent clicks via LSTMs and diverse interests. We also apply different levels of attention to our model. We conduct extensive experiments on two news datasets to demonstrate the effectiveness of our approach. Shaina Raza, Chen Ding 0004 |
IEEE BigData | 1 |
| 2020 | A Regularized Model to Trade-off between Accuracy and Diversity in a News Recommender SystemabstractNews recommender systems are usually designed to provide accurate and personalized recommendations to the readers. The diversity of the recommended results has received much less attention in this field. When it is considered, the current state-of-the-art models often apply the re-ranking mechanisms to promote the diversified results to the individual users. In this work, we propose a latent factor model to achieve the requisite level of accuracy while maintaining a reasonable level of diversity in a news recommender system. The existing latent factor methods mostly rely on Tikhonov regularization to improve the generality of the learnt models. These methods tend to focus mainly on accuracy measures, i.e., generating recommendations highly aligned with a user's past preference, which may cause a decrease in the diversity of information to which news readers are exposed. In our work, we make effective use of elastic-net regression to regularize the model for both the accuracy and the diversity in a single optimization framework. We demonstrate the effectiveness of our model over the state-of-the-art methods by conducting extensive experiments on a real-world news dataset. Shaina Raza, Chen Ding 0004 |
IEEE BigData | 1 |
| 2019 | News Recommender System Considering Temporal Dynamics and News TaxonomyabstractIn the past, news recommender systems have been built to recommend list of news items similar to those that a user has accessed before (content-based); or similar to those that have been read by similar users (collaborative filtering). However, the highly volatile nature of the news content and the dynamic and evolving user preferences are either ignored or not taken into full consideration in these systems. In a news recommender system, it is very likely that a user's short-term interest or preference may have a sudden change due to an emerging social or personal event or breaking news while their long-term interests may change gradually or remain. For these long-term interests of the readers, it is often more appropriate to associate them with news categories than with individual news items. In this paper, we propose a biased matrix factorization model with consideration of both temporal dynamics of user preferences and news taxonomy to build a news recommender system. By conducting an extensive experiment on a collection of news data, we demonstrate the effectiveness of our proposed model against traditional matrix factorization models as well as other neural recommender baselines. The findings from our experiments show that news category is an important factor when readers choose news articles to read, and temporal factors with consideration of different temporal resolution also play a role in this process. Shaina Raza, Chen Ding 0004 |
IEEE BigData | 1 |