VLDB 2026 Research / reviewers in the wild / expert
Ruofan Hu 0001
dblp:321/0629-1
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2025
0009-0005-6331-8838ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CLID-MU: Cross-Layer Information Divergence Based Meta Update Strategy for Learning with Noisy LabelsabstractLearning with noisy labels (LNL) is essential for training deep neural networks with imperfect data. Meta-learning approaches have achieved success by using a clean unbiased labeled set to train a robust model. However, this approach heavily depends on the availability of a clean labeled meta-dataset, which is difficult to obtain in practice. In this work, we thus tackle the challenge of meta-learning for noisy label scenarios without relying on a clean labeled dataset. Our approach leverages the data itself while bypassing the need for labels. Building on the insight that clean samples effectively preserve the consistency of related data structures across the last hidden and the final layer, whereas noisy samples disrupt this consistency, we design the Cross-layer Information Divergence-based Meta Update Strategy (CLID-MU). CLID-MU leverages the alignment of data structures across these diverse feature spaces to evaluate model performance and use this alignment to guide training. Experiments on benchmark datasets with varying amounts of labels under both synthetic and real-world noise demonstrate that CLID-MU outperforms state-of-the-art methods. The code is released at https://github.com/ruofanhu/CLID-MU. Ruofan Hu 0001, Dongyu Zhang 0005, Huayi Zhang, Elke A. Rundensteiner |
KDD (2) | 1 |
| 2024 | LLM-based Hierarchical Label Annotation for Foodborne Illness Detection on Social MediaabstractFoodborne illnesses pose a threat to public health, leading to morbidity, mortality, and economic burden annually. Social media, while providing a rich timely source for training AI models for surveillance, requires effective tools for annotation. While Large Language Models (LLMs) have shown promise for generating simple labels, here hierarchical labels composed of entity types like food type and symptom (at individual word level) and the foodborne illness event (at complete post level) are required. For this, we introduce ICL2FID, the first LLM-based hierarchical labeling framework designed to annotate social media posts for foodborne illness detection at two levels using only a few demonstration examples. To utilize the interconnection between post and word levels, ICL2FID instructs the LLM to leverage information from one level when predicting the other level. To combat model hallucination and cyclic dependencies, a verification step improves evidence propagation between interconnected word and post-level labeling tasks. Strategies for custom selection of demonstration examples are designed reducing biases and increasing representation. We compare ICL2FID against traditional supervised learning and other LLM methods, demonstrating that it not only achieves superior accuracy but does so at a fraction of the cost and time. These findings highlight ICL2FID’s potential as a viable alternative for hierarchical label generation in scenarios with limited resources and huge data sets. Code is available at https://github.com/zdy93/ICL2FID. Dongyu Zhang 0005, Ruofan Hu 0001, Dandan Tao, Elke A. Rundensteiner |
IEEE Big Data | 2 |
| 2024 | Leveraging LLMs for Integrated Sentiment and Topic Analysis on African Social MediaabstractLeveraging AI to analyze key topics on African social media can enhance public governance. Our study analyzes social media discourse within African society on development concerns by (1) evaluating AI techniques for sentiment, topic, and theme extraction, comparing the accuracy of these methods with human annotations, and (2) extracting key insights from the data to provide policymakers with actionable recommendations for sustainable development. For this study, we utilized a data corpus of 22,036 posts from Twitter and YouTube, all focused on development issues in Africa. We applied topic modeling to extract relevant topics from the corpus and used similarity analysis, powered by Large Language Models, to link these topics to prevalent development themes. Additionally, we leveraged unsupervised models such as VADER and Large Language Models to extract sentiment related to the identified topics. To validate these model-generated sentiments, we conducted a small crowdsourced study to gather human-annotated labels as ground truth. Our sentiment analysis findings show improvements with models like TextBlob, VADER, and Llama. Fine-tuning, partic-ularly with BERT, achieved an impressive Fl score of 0.988. Meanwhile, Llama demonstrated strong precision (0.72) and balanced accuracy (0.55) in capturing contextual sentiment. We identified 304 topics using BERTopic and Llama, with robust coherence (0.81 C-v) and divergence (0.58 IRBO). In theme analysis, the One-vs-Rest classification with ensemble voting performed exceptionally well, with ‘Poverty’ achieving the highest F1 score of 0.89. Our results suggest that African policymakers prioritize addressing corruption, unemployment, drought, and instability, while closely monitoring the positive impacts of policy interventions. Harriet Sibitenda, Ruofan Hu 0001, Elke A. Rundensteiner, Awa Diattara, Assitan Traore, Cheikh Ba |
ICMLA | 2 |
| 2024 | CoLafier: Collaborative Noisy Label Purifier With Local Intrinsic Dimensionality GuidanceabstractDeep neural networks (DNNs) have advanced many machine learning tasks, but their performance is often harmed by noisy labels in real-world data. Addressing this, we introduce CoLafier, a novel approach that uses Local Intrinsic Dimensionality (LID) for learning with noisy labels. CoLafier consists of two subnets: LID-dis and LID-gen. LID-dis is a specialized classifier. Trained with our uniquely crafted scheme, LID-dis consumes both a sample's features and its label to predict the label - which allows it to produce an enhanced internal representation. We observe that LID scores computed from this representation that effectively distinguish between correct and incorrect labels across various noise scenarios. In contrast to LID-dis, LID-gen, functioning as a regular classifier, operates solely on the sample's features. During training, CoLafier utilizes two augmented views per instance to feed both subnets. CoLafier considers the LID scores from the two views as produced by LID-dis to assign weights in an adapted loss function for both subnets. Concurrently, LID-gen, serving as classifier, suggests pseudo-labels. LID-dis then processes these pseudo-labels along with two views to derive LID scores. Finally, these LID scores along with the differences in predictions from the two subnets guide the label update decisions. This dual-view and dual-subnet approach enhances the overall reliability of the framework. Upon completion of the training, we deploy the LID-gen subnet of CoLafier as the final classification model. CoLafier demonstrates improved prediction accuracy, surpassing existing methods, particularly under severe label noise. For more details, see the code at https://github.com/zdy93/CoLafier. Dongyu Zhang 0005, Ruofan Hu 0001, Elke A. Rundensteiner |
SDM | 2 |
| 2023 | UCE-FID: Using Large Unlabeled, Medium Crowdsourced-Labeled, and Small Expert-Labeled Tweets for Foodborne Illness DetectionabstractFoodborne illnesses significantly impact public health. Deep learning surveillance applications using social media data aim to detect early warning signals. However, labeling foodborne illness-related tweets for model training requires extensive human resources, making it challenging to collect a sufficient number of high-quality labels for tweets within a limited budget. The severe class imbalance resulting from the scarcity of foodborne illness-related tweets among the vast volume of social media further exacerbates the problem. Classifiers trained on a classimbalanced dataset are biased towards the majority class, making accurate detection difficult. To overcome these challenges, we propose EGAL, a deep learning framework for foodborne illness detection that uses small expert-labeled tweets augmented by crowdsourced-labeled and massive unlabeled data. Specifically, by leveraging tweets labeled by experts as a reward set, EGAL learns to assign a weight of zero to incorrectly labeled tweets to mitigate their negative influence. Other tweets receive proportionate weights to counter-balance the unbalanced class distribution. Extensive experiments on real-world TWEET-FID data show that EGAL outperforms strong baseline models across different settings, including varying expert-labeled set sizes and class imbalance ratios. A case study on a multistate outbreak of Salmonella Typhimurium infection linked to packaged salad greens demonstrates how the trained model captures relevant tweets offering valuable outbreak insights. EGAL, funded by the U.S. Department of Agriculture (USDA), has the potential to be deployed for real-time analysis of tweet streaming, contributing to foodborne illness outbreak surveillance efforts. Ruofan Hu 0001, Dongyu Zhang 0005, Dandan Tao, Huayi Zhang, Elke A. Rundensteiner |
IEEE Big Data | 1 |
| 2022 | TWEET-FID: An Annotated Dataset for Multiple Foodborne Illness Detection TasksabstractFoodborne illness is a serious but preventable public health problem – with delays in detecting the associated outbreaks resulting in productivity loss, expensive recalls, public safety hazards, and even loss of life. While social media is a promising source for identifying unreported foodborne illnesses, there is a dearth of labeled datasets for developing effective outbreak detection models. To accelerate the development of machine learning-based models for foodborne outbreak detection, we thus present TWEET-FID (TWEET-Foodborne Illness Detection), the first publicly available annotated dataset for multiple foodborne illness incident detection tasks. TWEET-FID collected from Twitter is annotated with three facets: tweet class, entity type, and slot type, with labels produced by experts as well as by crowdsource workers. We introduce several domain tasks leveraging these three facets: text relevance classification (TRC), entity mention detection (EMD), and slot filling (SF). We describe the end-to-end methodology for dataset design, creation, and labeling for supporting model development for these tasks. A comprehensive set of results for these tasks leveraging state-of-the-art single-and multi-task deep learning methods on the TWEET-FID dataset are provided. This dataset opens opportunities for future research in foodborne outbreak detection. Ruofan Hu 0001, Dongyu Zhang 0005, Dandan Tao, Thomas Hartvigsen, Elke A. Rundensteiner |
LREC | 1 |