EDBT 2026 Demo / reviewers in the wild / expert
Jialin Yu 0001
dblp:167/1075-1
· DBLP profile ↗
5ranked-venue papers in the field
0as first author
5since 2021 · last 2026
0000-0003-1381-2203ORCID · conflict
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 3Information Retrieval & Web Search · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Multi-Label Text Interpretation with Chain-of-Thought Prompting and Contextualized KnowledgeabstractExisting multi-label topic models face several challenges when interpreting texts annotated with multiple labels: (1) they often associate irrelevant text segments with incorrect labels, which negatively impacts both segment and label interpretation; (2) they fail to effectively capture the semantic relationships between tokens and labels within the segment; (3) they do not integrate contextualized knowledge that could improve interpretability. To overcome these issues, we introduce the Contextualized Prompting Topic Model (CPTM). CPTM utilizes Chain-of-Thought (CoT) prompting to better align text segments with their semantically relevant labels. Furthermore, it integrates label-specific token visualization and topic mining procedure to facilitate the interpretation of tokens and labels. Experimental evaluations conducted on three multi-label text datasets show that CPTM significantly outperforms existing models in both segment and label interpretation. Human assessments also verify CPTM's effectiveness in accurately identifying label-relevant tokens within segments and providing insightful token-level interpretation. Rui Wang 0043, Haiping Huang, Jialin Yu 0001, Guozi Sun |
WWW | 4 |
| 2023 | Incorporating Emotions into Health Mention Classification Task on Social MediaabstractThe Health Mention Classification (HMC) task plays a pivotal role in leveraging social media discourse for public health mention monitoring, particularly in identifying and tracking disease proliferation. Despite its potential, the task poses significant challenges, due to the nuanced nature of health-related discussion. Building upon recent insights that emotional context can enhance HMC performance, in this paper, we explore how affective information can be integrated into the HMC process. Our study pioneers two distinct methodological pipelines that are designed to embed emotional nuances into the HMC task: (1) a two-stage fine-tuning process, starting with an implicit affective knowledge injection to initialise the model, followed by intermediate HMC task fine-tuning; and (2) an explicit multi-feature fusion strategy to leverage affective knowledge for HMC prediction. We conducted comprehensive evaluations across five diverse HMC benchmark datasets, encompassing content from Twitter, Reddit, and a blend of other social media platforms. Our empirical findings reveal that our affective-enriched models achieve statistically significant improvements across HMC benchmarks. Notably, the explicit multi-feature fusion method yielded a minimum of 3% improvement in F1 score over established BERT baselines across all tested corpora. Intriguingly, our analysis also suggests that the exclusive consideration of negative emotional indicators does not detrimentally impact HMC efficacy compared to leveraging both positive and negative emotions. Furthermore, our affectiveness-aware models demonstrate promising utility as a viable alternative in scenarios where domain-specific HMC datasets are scarce or non-existent for fine-tuning purposes. The consistent performance uplift across datasets sourced from varied social media channels underscores the generalisability and resilience of our proposed framework, marking a significant step forward in the computational understanding of health-related conversations in the digital sphere. Olanrewaju Tahir Aduragba, Jialin Yu 0001, Alexandra I. Cristea |
IEEE Big Data | 2 |
| 2023 | Religion and Spirituality on Social Media in the Aftermath of the Global PandemicabstractDuring the COVID-19 pandemic, the Church closed its physical doors for the first time in about 800 years, which is, arguably, a cataclysmic event. Other religions have found themselves in a similar situation, and they were practically forced to move online, which is an unprecedented occasion. In this paper, we analyse this sudden change in religious activities twofold: we create and deliver a questionnaire, as well as analyse Twitter data, to understand people’s perceptions and activities related to religious activities online. Importantly, we also analyse the temporal variations in this process, by analysing a period of 3 months: July-September 2020. Additionally to the separate analysis of the two data sources, we also discuss the implications from triangulating the results. Olanrewaju Tahir Aduragba, Jialin Yu 0001, Alexandra I. Cristea |
IEEE Big Data | 2 |
| 2023 | Improving Health Mention Classification Through Emphasising Literal Meanings: A Study Towards Diversity and Generalisation for Public Health SurveillanceabstractPeople often use disease or symptom terms on social media and online forums in ways other than to describe their health. Thus the NLP health mention classification (HMC) task aims to identify posts where users are discussing health conditions literally, not figuratively. Existing computational research typically only studies health mentions within well-represented groups in developed nations. Developing countries with limited health surveillance abilities fail to benefit from such data to manage public health crises. To advance the HMC research and benefit more diverse populations, we present the Nairaland health mention dataset (NHMD), a new dataset collected from a dedicated web forum for Nigerians. NHMD consists of 7,763 manually labelled posts extracted based on four prevalent diseases (HIV/AIDS, Malaria, Stroke and Tuberculosis) in Nigeria. With NHMD, we conduct extensive experiments using current state-of-the-art models for HMC and identify that, compared to existing public datasets, NHMD contains out-of-distribution examples. Hence, it is well suited for domain adaptation studies. The introduction of the NHMD dataset imposes better diversity coverage of vulnerable populations and generalisation for HMC tasks in a global public health surveillance setting. Additionally, we present a novel multi-task learning approach for HMC tasks by combining literal word meaning prediction as an auxiliary task. Experimental results demonstrate that the proposed approach outperforms state-of-the-art methods statistically significantly (p < 0.01, Wilcoxon test) in terms of F1 score over the state-of-the-art and shows that our new dataset poses a strong challenge to the existing HMC methods. Olanrewaju Tahir Aduragba, Jialin Yu 0001, Alexandra I. Cristea, Yang Long 0001 |
WWW | 2 |
| 2022 | Is Unimodal Bias Always Bad for Visual Question Answering? A Medical Domain Study with Dynamic AttentionabstractMedical visual question answering (Med-VQA) is to answer medical questions based on clinical images provided. This field is still in its infancy due to the complexity of the trio formed of questions, multimodal features and expert knowledge. In this paper, we tackle, a ’myth’ in the Natural Language Processing area - that unimodal bias is always considered undesirable in learning models. Additionally, we study the effect of integrating a novel dynamic attention mechanism into such models, inspired by a recent graph deep learning study.Unlike traditional attention, dynamic attention scores are conditioned on different query words in a question and thus enhance the representation learning ability of texts. We propose that some questions are answered more accurately with a reinforcement of question embedding after fusing multimodal features. Extensive experiments have been implemented on the VQA-RAD datasets and demonstrate that our proposed model, reinforCe unimOdal dynamiC Attention (COCA), outperforms the state-of-the-art methods overall and performs competitively at open-ended question answering. Zhongtian Sun, Anoushka Harit, Alexandra I. Cristea, Jialin Yu 0001, Noura Al Moubayed, Lei Shi 0003 |
IEEE Big Data | 4 |