VLDB 2026 Research / reviewers in the wild / expert
Fuyuan Wei
dblp:299/4285
· DBLP profile ↗
14ranked-venue papers
1as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TR-GLP: A Two-Stage Framework with Temporal Reprogramming and Residual Graph Label Propagation for Short Video Fake News DetectionabstractThe rapid growth of short video platforms has made short video fake news detection a critical security task. This task predicts news authenticity by leveraging multimodal information, including video frames, text, audio, and others. Previous works typically capture temporal cues only implicitly and suffer from passive coarse aggregation that favors global semantics, thereby obscuring subtle tampering traces. In addition, treating videos as isolated instances fails to exploit global correlations, which hinders the refinement of predictions for ambiguous cases. To address these limitations, we propose a two-stage framework with Temporal Reprogramming and Residual Graph Label Propagation (TR-GLP) for short video fake news detection. Specifically, we first employ a Temporal Reprogramming module in Stage 1 to learn temporal pattern prototypes as a reference basis for video dynamics and re-express keyframe features via cross-attention over these prototypes, where prototype attention activations reveal prototype-inconsistent temporal cues and support fine-grained forgery detection. Subsequently, we utilize a Residual Graph Label Propagation Network in Stage 2 to propagate supervision from labeled to unlabeled samples over the graph, and the residual design aggregates global context while preserving instance-level features. Experiments on FakeSV and FakeTT indicate that TR-GLP delivers superior performance compared with representative baseline methods. Junjiang Chen, Wenzhong Yang, Yabo Yin, Hongzhen Lv, Fuyuan Wei, Jingfeng He, Zongxu Luo, Junhang Wu |
ICMR | 5 |
| 2026 | HCG-MPB: Hierarchical Complementary Gating Mechanism with Multimodal Pattern Bank for Hateful Video DetectionabstractThe exponential rise of social media videos necessitates accurate detection of hate speech targeting protected attributes. However, existing multimodal approaches are hindered by two critical limitations: symmetric fusion, which induces modal competition and sensitivity to visual noise, and instance-based retrieval, which suffers from semantic ambiguity and high computational overhead due to reliance on raw data. To address these challenges, we propose the Multimodal Pattern Bank-guided Hierarchical Complementary Gating (HCG-MPB) framework. Specifically, we introduce a Hierarchical Complementary Gating (HCG) mechanism. It initially utilizes text as a semantic anchor to establish a stable context, and subsequently generates dynamic gating weights to selectively integrate audio-visual features, thereby effectively suppressing noise while preserving complementary information. Furthermore, to address semantic and efficiency challenges, we propose the Multimodal Pattern Bank (MPB). Instead of retrieving from vast raw instances, MPB leverages Large Language Models (LLMs) to distill extensive training samples into a compact set of interpretable prototypes. This approach significantly minimizes the storage footprint and retrieval latency, providing robust semantic guidance without the computational burden of traditional methods. Experiments on two public datasets demonstrate that HCG-MPB achieves excellent performance in both detection accuracy and efficiency. Wenzhong Yang, Yabo Yin, Fuyuan Wei, Junhang Wu, Junjiang Chen |
ICMR | 4 |
| 2025 | A Script Event Prediction Method Based on Multi-level Joint Pretraining and Prompt Fine-Tuning
Wenzhong Yang, Fuyuan Wei |
PAKDD (7) | 3 |
| 2025 | Bias-Unlearning in MABSA: A Causal Framework with Cross-Modal Counterfactual Inference
Wenzhong Yang, Yabo Yin, Fuyuan Wei |
PRCV (6) | 4 |
| 2025 | DCCR: Debiasing Cross-Document Event Coreference Resolution with Counterfactual Reasoning
Long Yao, Wenzhong Yang, Yabo Yin, Fuyuan Wei, Hongzhen Lv |
PRCV (1) | 4 |
| 2025 | A document-level relation extraction method based on dual-angle attention transfer fusion
Fuyuan Wei, Wenzhong Yang, Shengquan Liu, Chenghao Fu, Qicai Dai, Danni Chen, Xiaodan Tian, Bo Kong 0002, Liruizhi Jia |
Expert Syst. Appl. | 1 |
| 2025 | MDF-FND: A dynamic fusion model for multimodal fake news detectionabstractFake news detection has received increasing attention from researchers in recent years, especially in the area of multimodal fake news detection involving both text and images. However, many previous studies have simply fed the semantic features of both text and image modalities into a binary classifier after applying basic concatenation or attention mechanisms, where these features often contain a significant amount of inherent noise. This, in turn, leads to both intra- and inter-modal uncertainty. In addition, while methods based on simple concatenation of the two modalities have achieved notable results, they often ignore the drawback of applying fixed weights across modalities, which causes some high-impact features to be ignored. To address these issues, we propose a novel semantic-level m ultimodal d ynamic f usion framework for f ake n ews d etection ( MDF-FND ). To the best of our knowledge, this is the first attempt to develop a dynamic fusion framework for semantic-level multimodal fake news detection. Specifically, our model consists of two main components: (1) the U ncertainty E stimation M odule ( UEM ), which is an uncertainty modeling module that uses a multi-head attention mechanism to model intra-modal uncertainty, and (2) the D ynamic F usion N etwork, which is based on Dempster–Shafer evidence theory ( DFN ) and is designed to dynamically integrate the weights of both text and image modalities. To further enhance the dynamic fusion framework, a graph attention network is employed for inter-modal uncertainty modeling before DFN. Extensive experiments have demonstrated the effectiveness of our model across three datasets, with a performance improvement of up to 4% on the Twitter dataset, achieving state-of-the-art performance. We also conducted a systematic ablation study to gain insights into our motivation and architectural design. Our model is publicly available at https://github.com/CoisiniStar/MDF-FND . Hongzhen Lv, Wenzhong Yang, Yabo Yin, Fuyuan Wei, Jiaren Peng, Haokun Geng |
Knowl. Based Syst. | 4 |
| 2024 | SOIRP: Subject-Object Interaction and Reasoning Path based joint relational triple extraction by table filling
Qicai Dai, Wenzhong Yang, Fuyuan Wei, Meimei Tuo |
Neurocomputing | 4 |
| 2024 | Event causality identification via structure optimization and reinforcement learningabstractEvent causality identification (ECI) aims to identify possible causal relationships between event-mention pairs in a text. In the past, ECI models mainly used classification frameworks and rarely used generative models to solve this task. Although some progress has been made, the existing approaches suffer from the following two problems: (1) In the generative approach of inter-event mention dependency paths, noise and unnecessary sentence components cannot be effectively reduced, thus limiting the ability of the model to capture the critical correlation knowledge between event mentions; and (2) Existing multi-task generative model training which uses the REINFORCE algorithm suffers from a high-variance problem that imposes limitations on capturing critical causal knowledge. Therefore, we propose a novel Structural Optimization strategy Reinforcement Learning algorithm Generation model, GenSORL. The model aims to generate causal relationships from input sentences and includes dependency path generation as a complementary task to improve the causal label prediction performance. Specifically, this approach utilizes a new dependency syntax strategy to optimize dependency-path generation and extract important ECI contextual words between event mentions. Regarding the high-variance problem, a policy gradient with baseline is proposed for training the generative model, further adopting an innovative reward function to measure the accuracy of causal prediction and generation quality. In experiments using two frequently used benchmark datasets, the proposed method outperformed state-of-the-art models. Mingliang Chen 0002, Wenzhong Yang, Fuyuan Wei, Qicai Dai, Mingjie Qiu, Chenghao Fu, Mo Sha 0004 |
Knowl. Based Syst. | 3 |
| 2024 | Prompt for extraction: Multiple templates choice model for event extractionabstractEvent Extraction (EE) is an essential task in natural language processing that aims to mine events occurring in event mentions represent events using event records, which usually consist of event types, trigger words, argument elements corresponding to roles in the event types. Recently, prompt-based generative models have been developed to extract events. However, these prompt-based generative studies have ignored the fact that the strong language comprehension capability of the pre-trained language model (PLM) can analyze extract the potential role relationships in multiple templates for more information that can help extract argument elements. To determine the extended templates that can help the model for event extraction, we propose the multiple template choice model (MTCM), which designs an extended event type mining module to automatically mine the extended event types in the event mention uses the templates corresponding to the extended event types to interact with the template, corresponding to the currently to-be-extracted event type of event mention, in a multi-template information interaction, which gives the model more information guides the PLM for event extraction. To validate our model, we used two widely used datasets in the event extraction domain, ACE2005-EN ERE-EN. The experimental results show that our model achieves state-of-the-art performance on the ACE2005-EN dataset significantly improves the ERE dataset. In addition, according to the results, our model can be effectively adapted to low-resource environments. Jiaren Peng, Wenzhong Yang, Fuyuan Wei, Liang He 0003 |
Knowl. Based Syst. | 3 |
| 2023 | FE-YOLOv5: Feature enhancement network based on YOLOv5 for small object detectionabstractDue to their inherent characteristics, small objects have weaker feature representation after multiple down-sampling and are even annihilated in the background. FPN’s simple feature concatenation does not fully utilize multi-scale information and introduces irrelevant context into the information transfer, further reducing the detection performance of the small object. To address the above issues, we propose the simple but effective FE-YOLOv5. (1) We designed the feature enhancement module (FEM) to capture more discriminative features of the small object. Global attention and high-level global contextual information are used to guide shallow, high-resolution features. Global attention interacts with cross-dimensional feature interaction and reduces information loss. High-level context complements more detailed semantic information by modeling global relationships through non-local networks. (2) We design the spatially aware module (SAM) to filter spatial information and enhance the robustness of features. Deformable convolution performs sparse sampling and adaptive spatial learning to better focus on foreground objects. According to the experimental results, our proposed FE-YOLOv5 outperforms the other architectures in the VisDrone2019 dataset and Tsinghua-Tencent100K dataset. Compared to YOLOv5, the APS was improved by 2.8% and 2.9%, respectively. Wenzhong Yang, Danny Chen 0002, Fuyuan Wei, HaiLaTi KeZiErBieKe, Yuanyuan Liao |
J. Vis. Commun. Image Represent. | 5 |
| 2023 | Multiscale Global-Aware Channel Attention for Person Re-identificationabstractMost person re-identification methods are researched under various assumptions. However, viewpoint variations or occlusions are often encountered in practical scenarios. These are prone to intra-class variance. In this paper, we propose a multiscale global-aware channel attention (MGCA) model to solve this problem. It imitates the process of human visual perception, which tends to observe things from coarse to fine. The core of our approach is a multiscale structure containing two key elements: the global-aware channel attention (GCA) module for capturing the global structural information and the adaptive selection feature fusion (ASFF) module for highlighting discriminative features. Moreover, we introduce a bidirectional guided pairwise metric triplet (BPM) loss to reduce the effect of outliers. Extensive experiments on Market-1501, DukeMTMC-reID, and MSMT17, and achieve the state-of-the-art results on mAP. Especially, our approach exceeds the current best method by 2.0% on the most challenging MSMT17 dataset. Yingjie Zhu, Wenzhong Yang, Danny Chen 0002, Fuyuan Wei, HaiLaTi KeZiErBieKe, Yuanyuan Liao |
J. Vis. Commun. Image Represent. | 6 |
| 2022 | Weighted graph convolution over dependency trees for nontaxonomic relation extraction on public opinion information
Guangyao Wang, Shengquan Liu, Fuyuan Wei |
Appl. Intell. | 3 |
| 2022 | Spatial and long-short temporal attention correlation filters for visual trackingabstractAbstract Discriminative correlation filter is one of the quick and effective ways for studying visual tracking. However, discriminative correlation filter‐based methods still suffer from many challenging questions caused by environmental interferences, such as spatial boundary effect, temporal filter degradation, and tracking drift. A novel appearance optimisation model, named spatial and long–short temporal attention model, has been proposed based on a new spatial regularisation term and a long–short temporal regularisation term for learning the correlation filter to localise the target. On the one hand, our proposed method can improve the classical spatial regularisation term with a new weight matrix to alleviate the spatial boundary effect. On the other hand, two new temporal regularisation terms are designed: a short temporal regularisation term and a long temporal regularisation term. The short temporal regularisation term can enlarge the inner connections of the current frame and all foregoing frames to improve the tracking performances, and the long temporal regularisation term can address the influence of occlusion by using the similarity between the initial filter and the current one. Extensive experiments on various benchmarks illustrate that our proposed tracker performs favourably against several related popular trackers. Jianwei Zhao 0004, Fuyuan Wei, Ningning Chen, Zhenghua Zhou |
IET Image Process. | 2 |