VLDB 2026 Research / reviewers in the wild / expert
Peng Qi 0005
dblp:59/9474-5
· DBLP profile ↗
16ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0002-6458-8320ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mitigating GenAI-Powered Evidence Pollution for Out-Of-Context Misinformation Detection
Zehong Yan, Peng Qi 0005, Wynne Hsu, Mong-Li Lee |
ICDE | 2 |
| 2026 | R3Check: Reinforcement Learning for Iterative Retrieval and Structured Reasoning in Complex Fact CheckingabstractAutomated fact-checking aims to verify the veracity of claims based on related evidence, and has become increasingly important as large language models (LLMs) make it easier to generate and disseminate misinformation at scale. In open settings, effective fact-checking requires models to iteratively retrieve relevant evidence and reason over noisy and incomplete information. While recent LLM-based approaches have shown promising reasoning capabilities, prompt-based methods remain limited by the inherent behaviors of base LLMs, and supervised fine-tuning methods typically require costly annotated reasoning trajectories. In this paper, we propose R3Check, a rule-guided reinforcement learning framework that enables LLMs to perform iterative retrieval–reasoning for multi-hop fact-checking. R3Check formulates the retriever as an external environment and optimizes the model using Group Relative Policy Optimization, relying only on final veracity labels and format-based rewards rather than explicit reasoning annotations. To mitigate the mutual interference between retrieval and reasoning that arises under joint training, we introduce a two-stage curriculum that first trains structured reasoning under closed fact-checking with gold evidence, and then jointly optimizes retrieval and reasoning with real-time retrieval. An importance-based sampling strategy further strengthens effective supervision signals during training. Despite using only a 7B backbone, R3Check outperforms existing baselines and even powerful reasoning LLMs, under both given-evidence and real-time retrieval settings, while producing interpretable reasoning chains. This work demonstrates the potential of pure reinforcement learning to induce effective retrieval–reasoning behaviors for fact-checking under weak supervision. Peng Qi 0005, Wynne Hsu, Mong-Li Lee |
SIGIR | 1 |
| 2025 | Enhancing Fake News Video Detection via LLM-Driven Creative Process SimulationabstractThe emergence of fake news on short video platforms has become a new significant societal concern, necessitating automatic video-news-specific detection. Current detectors primarily rely on pattern-based features to separate fake news videos from real ones. However, limited and less diversified training data lead to biased patterns and hinder their performance. This weakness stems from the complex many-to-many relationships between video material segments and fabricated news events in real-world scenarios: a single video clip can be utilized in multiple ways to create different fake narratives, while a single fabricated event often combines multiple distinct video segments. However, existing datasets do not adequately reflect such relationships due to the difficulty of collecting and annotating large-scale real-world data, resulting in sparse coverage and non-comprehensive learning of the characteristics of potential fake news video creation. To address this issue, we propose a data augmentation framework AgentAug that generates diverse fake news videos by simulating typical creative processes. AgentAug implements multiple LLM-driven pipelines of four fabrication categories for news video creation, combined with an active learning strategy based on uncertainty sampling to select the potentially useful augmented samples during training. Experimental results on two benchmark datasets demonstrate that AgentAug consistently improves the performance of short video fake news detectors. Yuyan Bu, Qiang Sheng 0001, Juan Cao 0001, Shaofei Wang 0004, Peng Qi 0005, Yuhui Shi 0002, Beizhe Hu |
CIKM | 5 |
| 2025 | TRUST-VL: An Explainable News Assistant for General Multimodal Misinformation DetectionabstractMultimodal misinformation, encompassing textual, visual, and cross-modal distortions, poses an increasing societal threat that is amplified by generative AI. Existing methods typically focus on a single type of distortion and struggle to generalize to unseen scenarios. In this work, we observe that different distortion types share common reasoning capabilities while also requiring task-specific skills. We hypothesize that joint training across distortion types facilitates knowledge sharing and enhances the model’s ability to generalize. To this end, we introduce TRUST-VL, a unified and explainable vision-language model for general multimodal misinformation detection. TRUST-VL incorporates a novel Question-Aware Visual Amplifier module, designed to extract task-specific visual features. To support training, we also construct TRUST-Instruct, a large-scale instruction dataset containing 198K samples featuring structured reasoning chains aligned with human fact-checking workflows. Extensive experiments on both in-domain and zero-shot benchmarks demonstrate that TRUST-VL achieves state-of-the-art performance, while also offering strong generalization and interpretability. Zehong Yan, Peng Qi 0005, Wynne Hsu, Mong-Li Lee |
EMNLP | 2 |
| 2025 | Combating Online Misinformation Videos: Characterization, Detection, and PreventionabstractRecent progress of generative AI and the popularity of short-form video-sharing platforms have raised new risks of misinformation video issues, posing a potential threat to online multimedia ecosystems. With the aid of generative AI tools, producing and spreading vivid, persuasive misinformation videos has been easier, while detecting and preventing them has become harder. This tutorial introduces how to characterize, detect, and prevent misinformation videos, which consists of three technical parts: 1) Characterization of AI-generated and human-edited misinformation videos; 2) Detection approaches, covering those tailored for fully generated, manipulated, and human-edited videos; and 3) Prevention strategies, including those effective for the creation and spread phases. This tutorial concludes by discussing the status quo and ongoing challenges and highlighting the promising directions for future research. We expect to bring broader attention to misinformation video issues, gather and communicate with researchers of interest, and facilitate the engagement of those who are new to this field. Qiang Sheng 0001, Peng Qi 0005, Tianyun Yang, Yuyan Bu, Wynne Hsu, Mong-Li Lee, Juan Cao 0001 |
ACM Multimedia | 2 |
| 2024 | Bad Actor, Good Advisor: Exploring the Role of Large Language Models in Fake News DetectionabstractDetecting fake news requires both a delicate sense of diverse clues and a profound understanding of the real-world background, which remains challenging for detectors based on small language models (SLMs) due to their knowledge and capability limitations. Recent advances in large language models (LLMs) have shown remarkable performance in various tasks, but whether and how LLMs could help with fake news detection remains underexplored. In this paper, we investigate the potential of LLMs in fake news detection. First, we conduct an empirical study and find that a sophisticated LLM such as GPT 3.5 could generally expose fake news and provide desirable multi-perspective rationales but still underperforms the basic SLM, fine-tuned BERT. Our subsequent analysis attributes such a gap to the LLM's inability to select and integrate rationales properly to conclude. Based on these findings, we propose that current LLMs may not substitute fine-tuned SLMs in fake news detection but can be a good advisor for SLMs by providing multi-perspective instructive rationales. To instantiate this proposal, we design an adaptive rationale guidance network for fake news detection (ARG), in which SLMs selectively acquire insights on news analysis from the LLMs' rationales. We further derive a rationale-free version of ARG by distillation, namely ARG-D, which services cost-sensitive scenarios without inquiring LLMs. Experiments on two real-world datasets demonstrate that ARG and ARG-D outperform three types of baseline methods, including SLM-based, LLM-based, and combinations of small and large language models. Beizhe Hu, Qiang Sheng 0001, Juan Cao 0001, Yuhui Shi 0002, Yang Li 0196, Danding Wang, Peng Qi 0005 |
AAAI | 7 |
| 2024 | Sniffer: Multimodal Large Language Model for Explainable Out-of-Context Misinformation DetectionabstractMisinformation is a prevalent societal issue due to its potential high risks. Out-Of-Context (OOC) misinformation, where authentic images are repurposed with false text, is one of the easiest and most effective ways to mislead audiences. Current methods focus on assessing image- text consistency but lack convincing explanations for their judgments, which are essential for debunking misinformation. While Multimodal Large Language Models (MLLMs) have rich knowledge and innate capability for visual rea- soning and explanation generation, they still lack sophisti- cation in understanding and discovering the subtle cross- modal differences. In this paper, we introduce Sniffer,a novel multimodal large language model specifically engi- neered for OOC misinformation detection and explanation. Snifferemploys two-stage instruction tuning on Instruct- BLIP. The first stage refines the model's concept alignment of generic objects with news-domain entities and the sec- ond stage leverages OOC-specific instruction data gener- ated by language-only GPT-4 to fine-tune the model's dis- criminatory powers. Enhanced by external tools and re- trieval, Sniffernot only detects inconsistencies between text and image but also utilizes external knowledge for con- textual verification. Our experiments show that Sniffersurpasses the original MLLM by over 40% and outperforms state-of-the-art methods in detection accuracy. Snifferalso provides accurate and persuasive explanations as val- idated by quantitative and human evaluations. Peng Qi 0005, Zehong Yan, Wynne Hsu, Mong-Li Lee |
CVPR | 1 |
| 2024 | FakingRecipe: Detecting Fake News on Short Video Platforms from the Perspective of Creative ProcessabstractAs short-form video-sharing platforms become a significant channel for news consumption, fake news in short videos has emerged as a serious threat in the online information ecosystem, making developing detection methods for this new scenario an urgent need. Compared with that in text and image formats, fake news on short video platforms contains rich but heterogeneous information in various modalities, posing a challenge to effective feature utilization. Unlike existing works mostly focusing on analyzing what is presented, we introduce a novel perspective that considers how it might be created. Through the lens of the creative process behind news video production, our empirical analysis uncovers the unique characteristics of fake news videos in material selection and editing. Based on the obtained insights, we design FakingRecipe, a creative process-aware model for detecting fake news short videos. It captures the fake news preferences in material selection from sentimental and semantic aspects and considers the traits of material editing from spatial and temporal aspects. To improve evaluation comprehensiveness, we first construct FakeTT, an English dataset for this task, and conduct experiments on both FakeTT and the existing Chinese FakeSV dataset. The results show FakingRecipe's superiority in detecting fake news on short video platforms. Yuyan Bu, Qiang Sheng 0001, Juan Cao 0001, Peng Qi 0005, Danding Wang, Jintao Li 0001 |
ACM Multimedia | 4 |
| 2023 | FakeSV: A Multimodal Benchmark with Rich Social Context for Fake News Detection on Short Video PlatformsabstractShort video platforms have become an important channel for news sharing, but also a new breeding ground for fake news. To mitigate this problem, research of fake news video detection has recently received a lot of attention. Existing works face two roadblocks: the scarcity of comprehensive and largescale datasets and insufficient utilization of multimodal information. Therefore, in this paper, we construct the largest Chinese short video dataset about fake news named FakeSV, which includes news content, user comments, and publisher profiles simultaneously. To understand the characteristics of fake news videos, we conduct exploratory analysis of FakeSV from different perspectives. Moreover, we provide a new multimodal detection model named SV-FEND, which exploits the cross-modal correlations to select the most informative features and utilizes the social context information for detection. Extensive experiments evaluate the superiority of the proposed method and provide detailed comparisons of different methods and modalities for future works. Our dataset and codes are available in https://github.com/ICTMCG/FakeSV. Peng Qi 0005, Yuyan Bu, Juan Cao 0001, Wei Ji 0008, Ruihao Shui, Junbin Xiao, Danding Wang, Tat-Seng Chua |
AAAI | 1 |
| 2023 | ERASER: AdvERsArial Sensitive Element Remover for Image Privacy PreservationabstractThe daily practice of online image sharing enriches our lives, but also raises a severe issue of privacy leakage. To mitigate the privacy risks during image sharing, some researchers modify the sensitive elements in images with visual obfuscation methods including traditional ones like blurring and pixelating, as well as generative ones based on deep learning. However, images processed by such methods may be recovered or recognized by models, which cannot guarantee privacy. Further, traditional methods make the images very unnatural with low image quality. Although generative methods produce better images, most of them suffer from insufficiency in the frequency domain, which influences image quality. Therefore, we propose the AdvERsArial Sensitive Element Remover (ERASER) to guarantee both image privacy and image quality. 1) To preserve image privacy, for the regions containing sensitive elements, ERASER guarantees enough difference after being modified in an adversarial way. Specifically, we take both the region and global content into consideration with a Prior Transformer and obtain the corresponding region prior and global prior. Based on the priors, ERASER is trained with an adversarial Difference Loss to make the content in the regions different. As a result, ERASER can reserve the main structure and change the texture of the target regions for image privacy preservation. 2) To guarantee the image quality, ERASER improves the frequency insufficiency of current generative methods. Specifically, the region prior and global prior are processed with Fast Fourier Convolution to capture characteristics and achieve consistency in both pixel and frequency domains. Quantitative analyses demonstrate that the proposed ERASER achieves a balance between image quality and image privacy preservation, while qualitative analyses demonstrate that ERASER indeed reduces the privacy risk from the visual perception aspect. Guang Yang 0031, Juan Cao 0001, Danding Wang, Peng Qi 0005, Jintao Li 0001 |
AAAI | 4 |
| 2023 | Combating Online Misinformation Videos: Characterization, Detection, and Future DirectionsabstractWith information consumption via online video streaming becoming increasingly popular, misinformation video poses a new threat to the health of the online information ecosystem. Though previous studies have made much progress in detecting misinformation in text and image formats, video-based misinformation brings new and unique challenges to automatic detection systems: 1) high information heterogeneity brought by various modalities, 2) blurred distinction between misleading video manipulation and nonmalicious artistic video editing, and 3) new patterns of misinformation propagation due to the dominant role of recommendation systems on online video platforms. To facilitate research on this challenging task, we conduct this survey to present advances in misinformation video detection. We first analyze and characterize the misinformation video from three levels including signal, semantic, and intent. Based on the characterization, we systematically review existing works for detection from features of various modalities to techniques for clue integration. We also introduce existing resources including representative datasets and useful tools. Besides summarizing existing studies, we discuss related areas and outline open issues and future directions to encourage and guide more research on misinformation video detection. The corresponding repository is at https://github.com/ICTMCG/Awesome-Misinfo-Video-Detection. Yuyan Bu, Qiang Sheng 0001, Juan Cao 0001, Peng Qi 0005, Danding Wang, Jintao Li 0001 |
ACM Multimedia | 4 |
| 2022 | DRAG: Dynamic Region-Aware GCN for Privacy-Leaking Image DetectionabstractThe daily practice of sharing images on social media raises a severe issue about privacy leakage. To address the issue, privacy-leaking image detection is studied recently, with the goal to automatically identify images that may leak privacy. Recent advance on this task benefits from focusing on crucial objects via pretrained object detectors and modeling their correlation. However, these methods have two limitations: 1) they neglect other important elements like scenes, textures, and objects beyond the capacity of pretrained object detectors. 2) the correlation among objects is fixed, but a fixed correlation is not appropriate for all the images. To overcome the limitations, we propose the Dynamic Region-Aware Graph Convolutional Network (DRAG) that dynamically finds out crucial regions including objects and other important elements, and model their correlation adaptively for each input image. To find out crucial regions, we cluster spatially-correlated feature channels into several region-aware feature maps. Furthermore, we dynamically model the correlation with the self-attention mechanism and explore the interaction among the regions with a graph convolutional network. The DRAG achieved an accuracy of 87% on the largest dataset for privacy-leaking image detection, which is 10 percentage points higher than the state of the art. The further case study demonstrates that it found out crucial regions containing not only objects but other important elements like textures. The code and more details are in https://github.com/guang-yanng/DRAG. Guang Yang 0031, Juan Cao 0001, Qiang Sheng 0001, Peng Qi 0005, Xirong Li 0001, Jintao Li 0001 |
AAAI | 4 |
| 2022 | A Noise-Aware Framework for Blind Image Super-ResolutionabstractThe real-world image degradation in the super-resolution task is recently considered as a combination of Gaussian blur, down-sampling, and additional white Gaussian noise. To han-dle this degradation, previous methods estimate the Gaussian blur kernel or model the degradation based on a randomly selected image patch. However, these methods cannot han-dle degradations with high-level noise well as they ignore the spatial variability or even the existence of noise. Moreover, using image denoising networks to preprocess low-resolution images also fails due to the loss of important high-frequency information. In this paper, we propose a framework called EASE to flexibly handle real-world degradations. Specifi-cally, we develop a lightweight module to erase noise and blur simultaneously by learning from an image denoising and an image restoration network, which adapts to existing net-works that focus on handling bicubic down-sampling. Exten-sive experiments prove the superiority of our method, espe-cially when handling degradations with high-level noise. Guanqun Liu 0002, Xin Wang 0086, Lei Wang 0135, Daren Zha, Lin Zhao 0006, Zhe Kong, Peng Qi 0005 |
ICME | 7 |
| 2022 | Searching Models with Nested Attention for Blind Super-ResolutionabstractBlind super-resolution task aims to restore low-resolution im“ages with unknown degradations to high-resolution counter-parts. Existing methods rely on degradations estimation to re-construct high-resolution images. However, they need human involvement to obtain the best results as they treat unknown types of degradations as known conditions and manually select corresponding trained models. Moreover, they cannot fully use estimated degradations and generate blurry artifacts as they ignore that the impact of degradations on images is re-lated to images contents. In this paper, we propose HIS-NEST which contains an automatic search strategy HIS and a net-work structure NEST. Specifically, to bypass manual partici-pation, HIS automatically selects the clearest image by esti-mating the qualities of generated images. Furthermore, NEST protects the connection between degradations and images by using no loss functions to limit the degradations estimation and analyzing degradations from the perspective of channel and space. Extensive experiments show that our method out-performs state-of-the-art methods. Guanqun Liu 0002, Xin Wang 0086, Lei Wang 0135, Daren Zha, Lin Zhao 0006, Zhe Kong, Peng Qi 0005 |
ICME | 7 |
| 2021 | Improving Fake News Detection by Using an Entity-enhanced Framework to Fuse Diverse Multimodal CluesabstractRecently, fake news with text and images have achieved more effective diffusion than text-only fake news, raising a severe issue of multimodal fake news detection. Current studies on this issue have made significant contributions to developing multimodal models, but they are defective in modeling the multimodal content sufficiently. Most of them only preliminarily model the basic semantics of the images as a supplement to the text, which limits their performance on detection. In this paper, we find three valuable text-image correlations in multimodal fake news: entity inconsistency, mutual enhancement, and text complementation. To effectively capture these multimodal clues, we innovatively extract visual entities (such as celebrities and landmarks) to understand the news-related high-level semantics of images, and then model the multimodal entity inconsistency and mutual enhancement with the help of visual entities. Moreover, we extract the embedded text in images as the complementation of the original text. All things considered, we propose a novel entity-enhanced multimodal fusion framework, which simultaneously models three cross-modal correlations to detect diverse multimodal fake news. Extensive experiments demonstrate the superiority of our model compared to the state of the art. Peng Qi 0005, Juan Cao 0001, Xirong Li 0001, Huan Liu 0031, Qiang Sheng 0001, Xiaoyue Mi, Yongbiao Lv, Chenyang Guo, Yingchao Yu |
ACM Multimedia | 1 |
| 2019 | Exploiting Multi-domain Visual Information for Fake News DetectionabstractThe increasing popularity of social media promotes the proliferation of fake news. With the development of multimedia technology, fake news attempts to utilize multimedia content with images or videos to attract and mislead readers for rapid dissemination, which makes visual content an important part of fake news. Fake-news images, images attached to fake news posts, include not only fake images that are maliciously tampered but also real images that are wrongly used to represent irrelevant events. Hence, how to fully exploit the inherent characteristics of fake-news images is an important but challenging problem for fake news detection. In the real world, fake-news images may have significantly different characteristics from real-news images at both physical and semantic levels, which can be clearly reflected in the frequency and pixel domain, respectively. Therefore, we propose a novel framework Multi-domain Visual Neural Network (MVNN) to fuse the visual information of frequency and pixel domains for detecting fake news. Specifically, we design a CNN-based network to automatically capture the complex patterns of fake-news images in the frequency domain; and utilize a multi-branch CNN-RNN model to extract visual features from different semantic levels in the pixel domain. An attention mechanism is utilized to fuse the feature representations of frequency and pixel domains dynamically. Extensive experiments conducted on a real world dataset demonstrate that MVNN outperforms existing methods with at least 9.2% in accuracy, and can help improve the performance of multi-modal fake news detection by over 5.2%. Peng Qi 0005, Juan Cao 0001, Tianyun Yang, Junbo Guo, Jintao Li 0001 |
ICDM | 1 |