VLDB 2026 Research / reviewers in the wild / expert
Yibo Hu 0002
dblp:23/3288-2
· DBLP profile ↗
11ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0002-1578-7892ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Survey on the Role of Crowds in Combating Online Misinformation: Annotators, Evaluators, and CreatorsabstractOnline misinformation poses a global risk with significant real-world consequences. To combat misinformation, current research relies on professionals like journalists and fact-checkers for annotating and debunking false information while also developing automated machine learning methods for detecting misinformation. Complementary to these approaches, recent research has increasingly concentrated on utilizing the power of ordinary social media users, a.k.a. “the crowd,” who act as eyes-on-the-ground proactively questioning and countering misinformation. Notably, recent studies show that 96% of counter-misinformation responses originate from them. Acknowledging their prominent role, we present the first systematic and comprehensive survey of research papers that actively leverage the crowds to combat misinformation. In this survey, we first identify 88 papers related to crowd-based efforts, 1 following a meticulous annotation process adhering to the PRISMA framework (preferred reporting items for systematic reviews and meta-analyses). We then present key statistics related to misinformation, counter-misinformation, and crowd input in different formats and topics. Upon holistic analysis of the papers, we introduce a novel taxonomy of the roles played by the crowds in combating misinformation: (i) crowds as annotators who actively identify misinformation; (ii) crowds as evaluators who assess counter-misinformation effectiveness; (iii) crowds as creators who create counter-misinformation. This taxonomy explores the crowd’s capabilities in misinformation detection, identifies the prerequisites for effective counter-misinformation, and analyzes crowd-generated counter-misinformation. In each assigned role, we conduct a detailed analysis to categorize the specific utilization of the crowd. Particularly, we delve into (i) distinguishing individual, collaborative, and machine-assisted labeling for annotators; (ii) analyzing the effectiveness of counter-misinformation through surveys, interviews, and in-lab experiments for evaluators; and (iii) characterizing creation patterns and creator profiles for creators. Finally, we conclude this survey by outlining potential avenues for future research in this field. Bing He 0002, Yibo Hu 0002, Yeon-Chang Lee, Soyoung Oh, Gaurav Verma 0005, Srijan Kumar |
ACM Trans. Knowl. Discov. Data | 2 |
| 2024 | Leveraging Codebook Knowledge with NLI and ChatGPT for Zero-Shot Political Relation ClassificationabstractYibo Hu, Erick Skorupa Parolin, Latifur Khan, Patrick Brandt, Javier Osorio, Vito D’Orazio. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yibo Hu 0002, Erick Skorupa Parolin, Latifur Khan, Patrick T. Brandt, Javier Osorio, Vito D'Orazio |
ACL (1) | 1 |
| 2024 | Better to Ask in English: Cross-Lingual Evaluation of Large Language Models for Healthcare Queries
Yiqiao Jin, Mohit Chandra, Gaurav Verma 0005, Yibo Hu 0002, Munmun De Choudhury, Srijan Kumar |
WWW | 4 |
| 2023 | WokeGPT: Improving Counterspeech Generation Against Online Hate Speech by Intelligently Augmenting Datasets Using a Novel MetricabstractWith hate speech spreading rapidly online, it is increasingly important to respond automatically. However, there are some critical limitations in developing systems which produce these responses, which are known as counterspeeches. First, datasets containing paired instances of a hate speech and its appropriate response are very small. There is an abundance of hate speech on the web and in structured datasets, but quality counterspeeches are rare. Thus, since data is scarce, there is a need for automated methods to intelligently increase the size of existing paired datasets. Another critical challenge is that existing Natural Language Generation (NLG) metrics are not suitable for evaluating such systems, because these metrics do not accurately reflect how a human interprets the relationship between a hate speech and its counterspeech. Lastly, language models trained on internet text often exhibit a large amount of bias, which is unsuitable for sensitive tasks such as counterspeech generation. To address these challenges, we first introduce a technique to intelligently augment a small paired dataset of hate speech and counterspeech to make it substantially larger and varied, through a pairing technique that appropriately matches unpaired instances of hate speech with synthetic and existing counterspeeches. Next, we identified a need for a metric that evaluates counterspeech in the same way humans do, and propose a novel metric called PD-Score that leverages an advanced debating system. We empirically show through a large survey, that existing NLG metrics correlate poorly to human assessment and that our alternative is much more tightly bound to human assessment. Lastly, we curated a large domain-specific text corpus called WokeCorpus which we use to pretrain the language model before finetuning it for producing counterspeeches. We show that this both debiases the language model and aids performance. Sadaf Md. Halim, Saquib Irtiza, Yibo Hu 0002, Latifur Khan, Bhavani Thuraisingham |
IJCNN | 3 |
| 2022 | Multi-CoPED: A Multilingual Multi-Task Approach for Coding Political Event Data on Conflict and Mediation DomainabstractPolitical and social scientists monitor, analyze and predict political unrest and violence, preventing (or mitigating) harm, and promoting the management of global conflict. They do so using event coder systems, which extract structured representations from news articles to design forecast models and event-driven continuous monitoring systems. Existing methods rely on expensive manual annotated dictionaries and do not support multilingual settings. To advance the global conflict management, we propose a novel model, Multi-CoPED (Multilingual Multi-Task Learning BERT for Coding Political Event Data), by exploiting multi-task learning and state-of-the-art language models for coding multilingual political events. This eliminates the need for expensive dictionaries by leveraging BERT models' contextual knowledge through transfer learning. The multilingual experiments demonstrate the superiority of Multi-CoPED over existing event coders, improving the absolute macro-averaged F1-scores by 23.3% and 30.7% for coding events in English and Spanish corpus, respectively. We believe that such expressive performance improvements can help to reduce harms to people at risk of violence. Erick Skorupa Parolin, Seyyed MohammadSaleh Hosseini, Yibo Hu 0002, Latifur Khan, Patrick T. Brandt, Javier Osorio, Vito D'Orazio |
AIES | 3 |
| 2022 | Confli-T5: An AutoPrompt Pipeline for Conflict Related Text AugmentationabstractRecent advances in natural language processing (NLP) and Big Data technologies have been crucial for scientists to analyze political unrest and violence, prevent harm, and promote global conflict management. Government agencies and public security organizations have invested heavily in deep learning-based applications to study global conflicts and political violence. However, such applications involving text classification, information extraction, and other NLP-related tasks require extensive human efforts in annotating/labeling texts. While limited labeled data may drastically hurt the models’ performance (over-fitting), large demands on annotation tasks may turn real-world applications impracticable. To address this problem, we propose Confli-T5, a prompt-based method that leverages the domain knowledge from existing political science ontology to generate synthetic but realistic labeled text samples in the conflict and mediation domain. Our model allows generating textual data from the ground up and employs our novel Double Random Sampling mechanism to improve the quality (coherency and consistency) of the generated samples. We conduct experiments over six standard datasets relevant to political science studies to show the superiority of Confli-T5. Our codes are publicly available1. Erick Skorupa Parolin, Yibo Hu 0002, Latifur Khan, Patrick T. Brandt, Javier Osorio, Vito D'Orazio |
IEEE Big Data | 2 |
| 2022 | Knowledge Mining in Cybersecurity: From Attack to Defense
Khandakar Ashrafi Akbar, Sadaf Md. Halim, Yibo Hu 0002, Anoop Singhal, Latifur Khan, Bhavani Thuraisingham |
DBSec | 3 |
| 2022 | ConfliBERT: A Pre-trained Language Model for Political Conflict and ViolenceabstractYibo Hu, MohammadSaleh Hosseini, Erick Skorupa Parolin, Javier Osorio, Latifur Khan, Patrick Brandt, Vito D’Orazio. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Yibo Hu 0002, Seyyed MohammadSaleh Hosseini, Erick Skorupa Parolin, Javier Osorio, Latifur Khan, Patrick T. Brandt, Vito D'Orazio |
NAACL-HLT | 1 |
| 2021 | Multidimensional Uncertainty-Aware Evidential Neural NetworksabstractTraditional deep neural networks (NNs) have significantly contributed to the state-of-the-art performance in the task of classification under various application domains. However, NNs have not considered inherent uncertainty in data associated with the class probabilities where misclassification under uncertainty may easily introduce high risk in decision making in real-world contexts (e.g., misclassification of objects in roads leads to serious accidents). Unlike Bayesian NN that indirectly infer uncertainty through weight uncertainties, evidential NNs (ENNs) have been recently proposed to explicitly model the uncertainty of class probabilities and use them for classification tasks. An ENN offers the formulation of the predictions of NNs as subjective opinions and learns the function by collecting an amount of evidence that can form the subjective opinions by a deterministic NN from data. However, the ENN is trained as a black box without explicitly considering inherent uncertainty in data with their different root causes, such as vacuity (i.e., uncertainty due to a lack of evidence) or dissonance (i.e., uncertainty due to conflicting evidence). By considering the multidimensional uncertainty, we proposed a novel uncertainty-aware evidential NN called WGAN-ENN (WENN) for solving an out-of-distribution (OOD) detection problem. We took a hybrid approach that combines Wasserstein Generative Adversarial Network (WGAN) with ENNs to jointly train a model with prior knowledge of a certain class, which has high vacuity for OOD samples. Via extensive empirical experiments based on both synthetic and real-world datasets, we demonstrated that the estimation of uncertainty by WENN can significantly help distinguish OOD samples from boundary samples. WENN outperformed in OOD detection when compared with other competitive counterparts. Yibo Hu 0002, Yuzhe Ou, Xujiang Zhao, Jin-Hee Cho, Feng Chen 0001 |
AAAI | 1 |
| 2021 | CoMe-KE: A New Transformers Based Approach for Knowledge Extraction in Conflict and Mediation DomainabstractKnowledge discovery and extraction approaches attract special attention across industries and areas moving toward the 5V Era. In the political and social sciences, scholars and governments dedicate considerable resources to develop intelligent systems for monitoring, analyzing and predicting conflicts and affairs involving political entities across the globe. Such systems rely on background knowledge from external knowledge bases, that conflict experts commonly maintain manually. The high costs and extensive human efforts associated with updating and extending these repositories often compromise their correctness of. Here we introduce CoMe-KE (Conflict and Mediation Knowledge Extractor) to extend automatically knowledge bases about conflict and mediation events. We explore state-of-the-art natural language models to discover new political entities, their roles and status from news. We propose a distant supervised method and propose an innovative zero-shot approach based on a dynamic hypothesis procedure. Our methods leverage pre-trained models through transfer learning techniques to obtain excellent results with no need for a labeled data. Finally, we demonstrate the superiority of our method through a comprehensive set of experiments involving two study cases in the social sciences domain. CoMe-KE significantly outperforms the existing baseline, with (on average) double of the performance retrieving new political entities. Erick Skorupa Parolin, Yibo Hu 0002, Latifur Khan, Javier Osorio, Patrick T. Brandt, Vito D'Orazio |
IEEE BigData | 2 |
| 2021 | Uncertainty-Aware Reliable Text ClassificationabstractDeep neural networks have significantly contributed to the success in predictive accuracy for classification tasks. However, they tend to make over-confident predictions in real-world settings, where domain shifting and out-of-distribution (OOD) examples exist. Most research on uncertainty estimation focuses on computer vision because it provides visual validation on uncertainty quality. However, few have been presented in the natural language process domain. Unlike Bayesian methods that indirectly infer uncertainty through weight uncertainties, current evidential uncertainty-based methods explicitly model the uncertainty of class probabilities through subjective opinions. They further consider inherent uncertainty in data with different root causes, vacuity (i.e., uncertainty due to a lack of evidence) and dissonance (i.e., uncertainty due to conflicting evidence). In our paper, we firstly apply evidential uncertainty in OOD detection for text classification tasks. We propose an inexpensive framework that adopts both auxiliary outliers and pseudo off-manifold samples to train the model with prior knowledge of a certain class, which has high vacuity for OOD samples. Extensive empirical experiments demonstrate that our model based on evidential uncertainty outperforms other counterparts for detecting OOD examples. Our approach can be easily deployed to traditional recurrent neural networks and fine-tuned pre-trained transformers. Yibo Hu 0002, Latifur Khan |
KDD | 1 |