EDBT 2026 Demo / reviewers in the wild / expert
Ruohan Zong
dblp:259/6548
· DBLP profile ↗
11ranked-venue papers in the field
4as first author
8since 2021 · last 2026
0000-0002-6499-3406ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 5 (2 first)Data Mining & Knowledge Discovery · 4 (1 first)Big Data, Cloud & Distributed Data Systems · 2 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mixture of Adaptive Retrieval Experts for Veracity Assessment in the Human-LLM Mixed Generation ParadigmabstractThe growing prevalence of false information originating from human–LLM mixed generation sources presents new challenges for veracity assessment, as different generation sources (e.g., human or LLM) exhibit distinct semantic patterns and retrieval needs. In particular, LLM-generated false information is often more fluent, persuasive, and difficult to detect than traditional human-written content. Existing retrieval-augmented generation (RAG) methods for veracity assessment often use a uniform retrieval strategy for all inputs, limiting their flexibility in adapting to varying content characteristics. Current adaptive RAG methods primarily focus on general input difficulty or model confidence, overlooking the distinct semantic characteristics and retrieval needs of different generation sources. We propose MoARE, a Mixture of Adaptive Retrieval Experts framework that dynamically determines whether to retrieve and how much evidence to retrieve for each post. MoARE leverages a Mixture-of-Experts (MoE) network trained via reinforcement learning to balance veracity assessment accuracy and time cost—without requiring knowledge of the input's generation source. Experiments on recent human–LLM mixed datasets demonstrate that MoARE outperforms state-of-the-art RAG baselines, achieving higher accuracy with lower time cost. Ruohan Zong, Yang Zhang 0031, Zhenrui Yue, Dong Wang 0002 |
WSDM | 1 |
| 2025 | Empowering LLMs to Synthesize AI and Human Intelligence for Explainable Public Health Misinformation Detection on Social MediaabstractThis paper studies a critical problem of explainable public health misinformation detection on social media, where clear explanations are essential for enhancing user understanding and trust, surpassing the limitations of black-box misinformation detection results. To tackle this problem, there is a growing trend of leveraging collective intelligence from diverse intelligence sources, such as deep neural networks (DNNs), human intelligence, and large language models (LLMs). However, integrating hybrid intelligence from different sources remains a challenge: DNNs excel in accurate and efficient classification, crowd workers provide contextual understanding and readable explanations, and LLMs offer extensive domain knowledge and advanced language generation. Moreover, current crowdsourcing and human-AI collaboration methods mainly focus on aggregating misinformation detection labels using traditional measures like consistency, often overlooking more complex and challenging inputs like textual explanations. We propose SynthX, a collective intelligence framework that incorporates a holistic prompting design to harness the language and reasoning capabilities of LLMs for synthesizing diverse detection and explanation results. It also integrates a novel estimation theory-LLM hybrid approach to assess the varying reliability of detection results from different intelligence sources. Our evaluation on a real-world social media misinformation dataset demonstrates that SynthX consistently outperforms a rich set of state-of-the-art baselines in both detection accuracy and explanation quality. Ruohan Zong, Yang Zhang 0031, Dong Wang 0002 |
ICWSM | 1 |
| 2024 | MMAdapt: A Knowledge-guided Multi-source Multi-class Domain Adaptive Framework for Early Health Misinformation DetectionabstractThis paper studies a critical problem of emergent health misinformation detection, aiming to mitigate the spread of misinformation in emergent health domains to support well-informed healthcare decisions towards a Web for good health. Our work is motivated by the lack of timely resources (e.g., medical knowledge, annotated data) during the initial phases of an emergent health event or topic. In this paper, we develop a multi-source domain adaptive framework that jointly exploits medical knowledge and annotated data from different high-resource source domains (e.g., cancer, COVID-19) to detect misleading posts in an emergent target domain (e.g., mpox, polio). Two important challenges exist in developing our solution: 1) how to accurately detect the partially misleading and unverifiable content in an emergent target domain? 2) How to identify the conflicting knowledge facts from different source domains to accurately detect emergent misinformation in the target domain? To address these challenges, we develop MMAdapt, a multi-source multi-class domain adaptive misinformation detection framework that effectively explores diverse knowledge facts from different source domains to accurately detect not only the outright misleading but also the partially misleading or unverifiable posts on the Web. Extensive experimental results on four real-world misinformation datasets demonstrate that MMAdapt substantially outperforms state-of-the-art baselines in accurately detecting misinformation in an emergent health domain. Lanyu Shang, Yang Zhang 0031, Bozhang Chen, Ruohan Zong, Zhenrui Yue, Huimin Zeng 0001, Na Wei 0001, Dong Wang 0002 |
WWW | 4 |
| 2024 | SymLearn: A Symbiotic Crowd-AI Collective Learning Framework to Web-based Healthcare Policy Adherence AssessmentabstractThis paper develops a symbiotic human-AI collective learning framework that explores the complementary strengths of both AI and crowdsourced human intelligence to address a novel Web-based healthcare-policy-adherence assessment (WebHA) problem. In particular, the objective of the WebHA problem is to automatically assess people's public health policy adherence during emergent global health crisis events (e.g., COVID-19, MonkeyPox) by exploring massive social media imagery data. Recent advances in human-AI systems exhibit a significant potential in addressing the intricate imagery-based classification problems like WebHA by leveraging the collective intelligence of both humans and AI. This paper aims to address the limitation of existing human-AI systems that often rely heavily on human intelligence to improve AI model performance while overlooking the fact that humans themselves can be fallible and prone to errors. To address the above limitation, this paper develops SymLearn, a symbiotic human-AI co-learning framework that leverages human intelligence to troubleshoot and fine-tune the AI model while using AI models to guide human crowd workers to reduce the inherent human errors in their labels. Extensive experiments on two real-world WebHA applications show that SymLearn clearly outperforms the state-of-the-art baselines by improving WebHA performance and reducing crowd response delay. Yang Zhang 0031, Ruohan Zong, Lanyu Shang, Huimin Zeng 0001, Zhenrui Yue, Dong Wang 0002 |
WWW | 2 |
| 2023 | CollabEquality: A Crowd-AI Collaborative Learning Framework to Address Class-wise Inequality in Web-based Disaster ResponseabstractWeb-based disaster response (WebDR) is emerging as a pervasive approach to acquire real-time situation awareness of disaster events by collecting timely observations from the Web (e.g., social media). This paper studies a class-wise inequality problem in WebDR applications where the objective is to address the limitation of current WebDR solutions that often have imbalanced classification performance across different classes. To address such a limitation, this paper explores the collaborative strengths of the diversified yet complementary biases of AI and crowdsourced human intelligence to ensure a more balanced and accurate performance for WebDR applications. However, two critical challenges exist: 1) it is difficult to identify the imbalanced AI results without knowing the ground-truth WebDR labels a priori; ii) it is non-trivial to address the class-wise inequality problem using potentially imperfect crowd labels. To address the above challenges, we develop CollabEquality, an inequality-aware crowd-AI collaborative learning framework that carefully models the inequality bias of both AI and human intelligence from crowdsourcing systems into a principled learning framework. Extensive experiments on two real-world WebDR applications demonstrate that CollabEquality consistently outperforms the state-of-the-art baselines by significantly reducing class-wise inequality while improving the WebDR classification accuracy. Yang Zhang 0031, Lanyu Shang, Ruohan Zong, Huimin Zeng 0001, Zhenrui Yue, Dong Wang 0002 |
WWW | 3 |
| 2023 | ContrastFaux: Sparse Semi-supervised Fauxtography Detection on the Web using Multi-view Contrastive LearningabstractThe widespread misinformation on the Web has raised many concerns with serious societal consequences. In this paper, we study a critical type of online misinformation, namely fauxtography, where the image and associated text of a social media post jointly convey a questionable or false sense. In particular, we focus on a sparse semi-supervised fauxtography detection problem, which aims to accurately identify fauxtography by only using the sparsely annotated ground truth labels of social media posts. Our problem is motivated by the key limitation of current fauxtography detection approaches that often require a large amount of expensive and inefficient manual annotations to train an effective fauxtography detection model. We identify two key technical challenges in solving the problem: 1) it is non-trivial to train an accurate detection model given the sparse fauxtography annotations, and 2) it is difficult to extract the heterogeneous and complicated fauxtography features from the multi-modal social media posts for accurate fauxtography detection. To address the above challenges, we propose ContrastFaux, a multi-view contrastive learning framework that jointly explores the sparse fauxtography annotations and the cross-modal fauxtography feature similarity between the image and text in multi-modal posts to accurately detect fauxtography on social media. Evaluation results on two social media datasets demonstrate that ContrastFaux consistently outperforms state-of-the-art deep learning and semi-supervised learning fauxtography detection baselines by achieving the highest fauxtography detection accuracy. Ruohan Zong, Yang Zhang 0031, Lanyu Shang, Dong Wang 0002 |
WWW | 1 |
| 2021 | A deep contrastive learning approach to extremely-sparse disaster damage assessment in social sensingabstractSocial sensing has emerged as a pervasive and scalable sensing paradigm to obtain timely information of the physical world from "human sensors". In this paper, we study a new extremely-sparse disaster damage assessment (DBA) problem in social sensing. The objective is to automatically assess the damage severity of affected areas in a disaster event by leveraging the imagery data reported on online social media with extremely sparse training data (e.g., only 1% of the data samples have labels). Our problem is motivated by the limitation of current DDA solutions that often require a significant amount of high-quality training data to learn an effective DDA model. We identify two critical challenges in solving our problem: i) it remains to be a fundamental challenge on how to effectively train a reliable DDA model given the lack of sufficient damage severity labels; ii) it is a difficult task to capture the excessive and fine-grained damage-related features in each image for accurate damage assessment. In this paper, we propose ContrastDDA, a deep contrastive learning approach to address the extremely-sparse DDA problem by designing an integrated contrastive and augmentative neural network architecture for accurate disaster damage assessment using the extremely sparse training samples. The evaluation results on two real-world DDA applications demonstrate that ContrastDDA clearly outperforms state-of-the-art deep learning and semi-supervised learning baselines with the highest DDA accuracy under different application scenarios. Yang Zhang 0031, Ruohan Zong, Lanyu Shang, Ziyi Kou, Dong Wang 0002 |
ASONAM | 2 |
| 2021 | StreamCollab: A Streaming Crowd-AI Collaborative System to Smart Urban Infrastructure Monitoring in Social SensingabstractSocial sensing has emerged as a pervasive and scalable sensing paradigm to collect observations of the physical world from human sensors. A key advantage of social sensing is its infrastructure-free nature. In this paper, we focus on a streaming urban infrastructure monitoring (Streaming UIM) problem in social sensing. The goal is to automatically detect the urban infrastructure damages from the streaming imagery data posted on social media by exploring the collective power of both AI and human intelligence from crowdsourcing systems. Our work is motivated by the limitation of current AI and crowdsourcing solutions that either fail in many critical time-sensitive UIM application scenarios or are not easily generalizable to monitor the damage of different types of urban infrastructures. We identify two critical challenges in solving our problem: i) it is difficult to dynamically integrate AI and crowd intelligence to effectively identify and fix the failure cases of AI solutions; ii) it is non-trivial to obtain accurate human intelligence from unreliable crowd workers in streaming UIM applications. In this paper, we propose StreamCollab, a streaming crowd-AI collaborative system that explores the collaborative intelligence from AI and crowd to solve the streaming UIM problem. The evaluation results on a real-world urban infrastructure imagery dataset collected from social media demonstrate that StreamCollab consistently outperforms both state-of-the-art AI and crowd-AI baselines in UIM accuracy while maintaining the lowest computational cost. Yang Zhang 0031, Lanyu Shang, Ruohan Zong, Ziyi Kou, Dong Wang 0002 |
HCOMP | 3 |
| 2020 | A Hybrid Transfer Learning Approach to Migratable Disaster Assessment in Social Media SensingabstractSocial media sensing has emerged as a powerful sensing paradigm to collect the observations of the physical world by exploring the “wisdom of crowd”. In this paper, we focus on a migratable disaster damage assessment problem in social media sensing applications. Our goal is to accurately identify the damage severity of affected areas in an unfolding disaster event using unlabeled social media data feeds (e.g., image posts on social media). Two fundamental challenges exist in solving our problem: i) different disaster events often have distinct characteristics (e.g., damage types, affected areas) that cannot be easily migrated; ii) it is non-trivial to modify a damage assessment model from a previous event to adapt to a new event without using the labeled data from the new event. To address the above challenges, we develop SocialTrans, a hybrid deep transfer learning framework, to enable effective model migration for accurate damage assessment without using any training data from the studied disaster event. The evaluation results on four real-world disaster events show that SocialTrans consistently outperforms the state-of-the-art baselines in accurately assessing the damage level of disasters. Yang Zhang 0031, Ruohan Zong, Dong Wang 0002 |
ASONAM | 2 |
| 2020 | On Privileged Information Driven Robust Face Verification: A Siamese Convolutional Neural Network ApproachabstractFace verification is an important task to verify people's identities, which has sample pairs and labels with only side information about whether the two images in a pair are from the same subject or not instead of a specific label for each sample. This paper is motivated by the limitations of previous efforts that the deep neural networks with only the RGB feature are not robust to intra-class noise (e.g., illumination), and the deep networks with both the RGB and depth feature are not robust to data availability of the depth images in applications. To overcome these limitations, we focus on developing a "double robust" face verification architecture using the Siamese convolutional neural network (SCNN). In particular, our goal is to propose an SCNN with privileged information (SCNN+), which is inspired by Learning Using Privileged Information (LUPI), through incorporating additional depth feature along with the RGB feature in the training stage to provide privileged information to restrain the RGB feature's prediction error. We conduct experiments to evaluate the performance of the proposed SCNN+ architecture and compare it with different categories of state-of-the-art baselines on two real-world RGB-D face datasets. The evaluation results demonstrate that SCNN+ significantly outperforms all types of baselines. Ruohan Zong |
IEEE BigData | 1 |
| 2019 | TransLand: An Adversarial Transfer Learning Approach for Migratable Urban Land Usage Classification using Remote SensingabstractUrban land usage classification is a critical task in big data based smart city applications that aim to understand the social-economic land functions and physical land attributes in urban environments. This paper focuses on a migratable urban land usage classification problem using remote sensing data (i.e., satellite images). Our goal is to accurately classify the land usage of locations in a target city where the ground truth land usage data is not available by leveraging a classification model from a source city where such data is available. This problem is motivated by the limitation of current solutions that primarily rely on a rich set of ground-truth data for accurate model training, which encounters high annotation costs. Two important challenges exist in solving our problem: i) the target and source cities often have different urban characteristics that prevent the direct application of a model learned from the source city to the target city; ii) the complex visual features in satellite images make it non-trivial to “translate” the images from the target city to the source city for an accurate classification. To address the above challenges, we develop TransLand, an adversarial transfer learning framework to translate the satellite images from the target city to the source city for accurate land usage classification. We evaluate our scheme on the real-world satellite imagery and land usage datasets collected from live different cities in Europe. The results show that TransLand significantly outperforms the state-of-the-art land usage classification baselines in classifying the land usage of locations in a city. Yang Zhang 0031, Ruohan Zong, Jun Han 0010, Hao Zheng 0006, Qiuwen Lou, Daniel Yue Zhang, Dong Wang 0002 |
IEEE BigData | 2 |