VLDB 2026 Research / reviewers in the wild / expert
Zhiyi Tan 0002
dblp:71/4988-2
· DBLP profile ↗
19ranked-venue papers
2as first author
19since 2021 · last 2026
0000-0002-1209-2817ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 13 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Social Event Prediction via Fourier Graph LearningabstractSocial event prediction has garnered increasing attention in web-centered society. Most existing studies represent web-based event stream as chronological graph sequences, then leverage RNNs and GNNs to model temporal and relational patterns. However, this paradigm is inherently flawed: (1) RNNs struggle to capture long-term temporal dependencies, ignoring those temporally distant but influential events. (2) Spatio-temporal GNNs exhibit high computational complexity on large-scale real-time event streams, which hinders their web applications. To this end, we explore a novel paradigm called Fourier Graph Learning from the perspective of frequency domain. Specifically, we first define a novel data structure called Fourier Graph (FG). In FG, both nodes and edges are complex vectors, with real part encoding semantics and imaginary part representing semantic-specific temporal patterns. These temporal patterns are obtained by semantic-aware frequency filter, which utilizes semantics as guidance to adaptively incorporates both long-term dependency and short-term dynamic. Based on FG, we further propose Fourier Graph Neural Network (FGNN). It replaces time-domain convolution with frequency-domain multiplication for efficient aggregation. FGNN also includes a complex-valued event decoder, which fully leverages semantics and temporal patterns from complex space to predict future event probabilities. Extensive experiments show our superior performance with higher accuracy, less complexity and better interpretability compared with baselines. Mingjie Qiu, Zhiyi Tan 0002, Bing-Kun Bao |
WWW | 2 |
| 2026 | Towards Rare Social Event Prediction via Mediator LearningabstractRare social events are infrequent yet influential incidents. Predicting such events is practically significant yet inherently challenging due to their extreme scarcity in web-based event stream. Existing studies view this task as an imbalanced classification problem and adopt static rebalancing methods to mitigate scarcity. However, they (1) ignore inter-event dependency that represents the interactions between different event streams, failing to capture precursors that lead to rare events and fundamentally limits performance. They (2) overlook intra-event dependency between different time points within single event stream, which prevents the model from adapting to shifting event patterns and degrades its generalization ability. To this end, we propose a novel Mediator Learning (ML) framework, which introduces mediators to explicitly model complex dependencies within web-based event streams. Specifically, we propose (1) Precursor Event Router (PER) that utilizes an information-theoretic routing approach to extract precursor events as mediators from massive event streams. Based on extracted mediators, (2) Conditional Hierarchical Graph Network (CHG) is introduced to model observed events, mediators and rare events into bottom-up graph levels, where its upper-level propagation is conditioned on bottom-level probability distribution. Jointly, these two modules decompose the imbalanced task into two more balanced stages, which not only mitigates the scarcity of rare events but also explicitly model inter-event dependencies, so as to capture precursor events leading to target rare events. Finally, we design (3) Adaptive Information Regularizer (AIR) to optimize the two stages. It dynamically adjusts the information flow between two stages, which models intra-event dependencies and facilitate adaption to drifting event patterns. We theoretically reveal the effectiveness of ML by framing it within Information Bottleneck (IB) principle. Extensive experiments show our superior accuracy and interpretability compared with SOTA. Mingjie Qiu, Zhiyi Tan 0002, Bing-Kun Bao |
WWW | 2 |
| 2026 | AdaBi-Deno: a learnable denoising plug-in for multimedia sequential recommendation
Zhiyi Tan 0002 |
Multim. Syst. | 1 |
| 2026 | MyGO: Modality-incomplete Fake News Video Detection via Prompt-assisted Modality Disentangling ModelabstractFake news video detection has become a pressing concern with the growth of short video platforms. However, previous studies have primarily focused on videos with modality-complete data, failing to handle the uncertain missing modality issue in real-world applications. They fall short in two key aspects: (1) Highly coupled feature fusion hinders the model to learn intra- and inter-modality dependencies, making it difficult to form robust multimodal representations when facing uncertain modality missing. (2) Excessive reliance on discriminative modality combinations toward fake news, which leads to inferior performance on other modality combinations. To this end, we propose a novel model for modality-incomplete fake news video detection called MyGO. It contains three modules: (1) Caption-guided Keyframe Attention (CKA) leverages embedded captions to guide feature extraction, which adaptively excludes irrelevant frames to enhance the learning of intra-modality dependencies, resulting in refined modality features. (2) Based on refined modality features from CKA, Modality Disentangling Network (MDN) is designed to decompose them into shared and specific parts, which captures fine-grained inter-modality dependencies effectively. These two kinds of dependencies help avoid coupled multimodal fusion and bridge information gaps caused by missing modalities. (3) Furthermore, missing prompts are newly introduced to explicitly mark modality combinations within each news video. By integrating missing prompts with aforementioned inter-modality dependencies within Prompt-assisted Modality Aligning (PMA) Module, we alleviate over-reliance on discriminative modality combinations and enhancing the representation of less discriminative ones. Extensive experiments showcase that MyGO achieves 3.79–4.85% improvements in accuracy, demonstrating its performance over state-of-the-art approaches under different missing conditions. Mingjie Qiu, Zhiyi Tan 0002, Bing-Kun Bao |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | BeFA: A General Behavior-driven Feature Adapter for Multimedia RecommendationabstractMultimedia recommender systems focus on utilizing behavioral information and content information to model user preferences. Typically, it employs pre-trained feature encoders to extract content features, then fuses them with behavioral features. However, pre-trained feature encoders often extract features from the entire content simultaneously, including excessive preference-irrelevant details.We speculate that it may result in the extracted features not containing sufficient features to accurately reflect user preferences. To verify our hypothesis, we introduce an attribution analysis method for visually and intuitively analyzing the content features. The results indicate that certain items’ content features exhibit the issues of information drift and information omission, reducing the expressive ability of features. Building upon this finding, we propose an effective and efficient general Behaviordriven Feature Adapter (BeFA) to tackle these issues. This adapter reconstructs the content feature with the guidance of behavioral information, enabling content features accurately reflecting user preferences. Extensive experiments demonstrate the effectiveness of the adapter across all multimedia recommendation methods. Qile Fan, Penghang Yu, Zhiyi Tan 0002, Bing-Kun Bao, Guanming Lu |
AAAI | 3 |
| 2025 | Mind Individual Information! Principal Graph Learning for Multimedia RecommendationabstractGraph Neural Network (GNN)-based methods have recently emerged as effective approaches for multimedia recommendation. Typically, these methods employ message passing on the user-item interaction graph, and model user preferences by exploiting co-occurrence patterns. Despite their effectiveness, we argue that they insufficiently exploit the individual information, potentially limiting recommendation performance. To validate our argument, we first analyze existing methods from spectral graph theory. We identify that existing methods focus on capturing global structural features, but underutilize local structural features that convey individual information. Further detailed experiments reveal that such an underutilization leads to overly similar user preferences modeling. Furthermore, we propose a novel Principal Graph Learning (PGL) framework to address this issue. The idea is to enhance user preference modeling by effectively mining and utilizing principal local structural features. PGL first extracts the principal subgraph from the user-item interaction graph using two novel extraction operators: global-aware and local-aware subgraph extraction. It then employs message passing on the principal subgraph to comprehensively model user perference, with the aim of simultaneously capturing co-occurrence patterns and individual information. Compared to existing methods, PGL achieves an average performance improvement of 9%. Penghang Yu, Zhiyi Tan 0002, Guanming Lu, Bing-Kun Bao |
AAAI | 2 |
| 2025 | Hypergraph denoising neural network for session-based recommendation
Zhiyi Tan 0002, Guanming Lu, Jinsheng Wei |
Appl. Intell. | 2 |
| 2025 | Adaptive discriminant feature learning for GNN-based session recommendation
Zhiyi Tan 0002, Guanming Lu, Jinsheng Wei |
Multim. Syst. | 2 |
| 2025 | LD4MRec: simplifying and powering diffusion model for multimedia recommendation
Jiarui Zhu, Penghang Yu, Zhiyi Tan 0002, Bing-Kun Bao |
Multim. Syst. | 4 |
| 2024 | Inferring the effectiveness of epidemic prevention measures based on spatial heterogeneity modelingabstractFaced with recent outbreaks of various epidemics, governments worldwide have adopted various countermeasures to slow down the virus spread. Accurately evaluating the effectiveness of these interventions is crucial for their efficiency. However, most existing models assume uniform mixing of infected and susceptible populations throughout the country, which becomes inadequate when strict restrictions significantly reduce interregional population mobility. This bias may lead to unreliable estimates of intervention effects. To this end, we propose a multi-scale regeneration model that considers spatial inhomogeneity, for more accurate evaluation of intervention effects. Specially, by leveraging the idea of renormalization groups to model the spread process between regions at different scales, our model is able to simulate the spatial heterogeneity in epidemic spreading, and achieve a reliable estimation of the interventions’ effect. Experiments on real data show that compared with traditional epidemic models, our model can more accurately describe the epidemic, thereby achieving more reliable evaluation of epidemic intervention effects. Mingyu Wu 0008, Zhiyi Tan 0002, Bing-Kun Bao |
ICME | 2 |
| 2024 | MSGNN: Multi-scale Spatio-temporal Graph Neural Network for epidemic forecasting
Mingjie Qiu, Zhiyi Tan 0002, Bing-Kun Bao |
Data Min. Knowl. Discov. | 2 |
| 2024 | Game and reference: efficient policy making for epidemic prevention and control
Zhiyi Tan 0002, Bing-Kun Bao |
Multim. Syst. | 1 |
| 2024 | Improving Graph Collaborative Filtering with Directional Behavior Enhanced Contrastive LearningabstractGraph Collaborative Filtering is a widely adopted approach for recommendation, which captures similar behavior features through Graph Neural Network (GNN). Recently, Contrastive Learning (CL) has been demonstrated as an effective method to enhance the performance of graph collaborative filtering. Typically, CL-based methods first perturb users’ history behavior data (e.g., drop clicked items), then construct a self-discriminating task for behavior representations under different random perturbations. However, for widely existing inactive users, random perturbation makes their sparse behavior information more incomplete, thereby harming the behavior feature extraction. To tackle the above issue, we design a novel directional perturbation-based CL method to improve the graph collaborative filtering performance. The idea is to perturb node representations through directionally enhancing behavior features. To do so, we propose a simple yet effective feedback mechanism, which fuses the representations of nodes based on behavior similarity. Then, to avoid irrelevant behavior preferences introduced by the feedback mechanism, we construct a behavior self-contrast task before and after feedback, to align the node representations between the final output and the first layer of GNN. Different from the widely adopted self-discriminating task, the behavior self-contrast task avoids complex message propagation on different perturbed graphs, which is more efficient than previous methods. Extensive experiments on three public datasets demonstrate that the proposed method has distinct advantages over other CL methods on recommendation accuracy. Penghang Yu, Bing-Kun Bao, Zhiyi Tan 0002, Guanming Lu |
ACM Trans. Knowl. Discov. Data | 3 |
| 2023 | Multi-modal Context-Aware Network for Scene Graph Generation
Bing-Kun Bao, Zhiyi Tan 0002 |
ICIG (2) | 3 |
| 2023 | Multi-View Graph Convolutional Network for Multimedia RecommendationabstractMultimedia recommendation has received much attention in recent years. It models user preferences based on both behavior information and item multimodal information. Though current GCN-based methods achieve notable success, they suffer from two limitations: (1) Modality noise contamination to the item representations. Existing methods often mix modality features and behavior features in a single view (e.g., user-item view) for propagation, the noise in the modality features may be amplified and coupled with behavior features. In the end, it leads to poor feature discriminability; (2) Incomplete user preference modeling caused by equal treatment of modality features. Users often exhibit distinct modality preferences when purchasing different items. Equally fusing each modality feature ignores the relative importance among different modalities, leading to the suboptimal user preference modeling. Penghang Yu, Zhiyi Tan 0002, Guanming Lu, Bing-Kun Bao |
ACM Multimedia | 2 |
| 2023 | PRM-KGED: paper recommender model using knowledge graph embedding and deep neural network
Nimbeshaho Thierry, Bing-Kun Bao, Zafar Ali, Zhiyi Tan 0002, Ingabire Batamira Christ Chatelain, Pavlos Kefalas |
Appl. Intell. | 4 |
| 2023 | Centralized sub-critic based hierarchical-structured reinforcement learning for temporal sentence grounding
Yingyuan Zhao, Zhiyi Tan 0002, Bing-Kun Bao, Zhengzheng Tu |
Multim. Syst. | 2 |
| 2023 | Adaptive Text Denoising Network for Image Caption EditingabstractImage caption editing, which aims at editing the inaccurate descriptions of the images, is an interdisciplinary task of computer vision and natural language processing. As the task requires encoding the image and its corresponding inaccurate caption simultaneously and decoding to generate an accurate image caption, the encoder-decoder framework is widely adopted for image caption editing. However, existing methods mostly focus on the decoder, yet ignore a big challenge on the encoder: the semantic inconsistency between image and caption. To this end, we propose a novel A daptive T ext D enoising Net work (ATD-Net) to filter out noises at the word level and improve the model’s robustness at sentence level. Specifically, at the word level, we design a cross-attention mechanism called Textual Attention Mechanism (TAM), to differentiate the misdescriptive words. The TAM is designed to encode the inaccurate caption word by word based on the content of both image and caption. At the sentence level, in order to minimize the influence of misdescriptive words on the semantic of an entire caption, we introduce a Bidirectional Encoder to extract the correct semantic representation from the raw caption. The Bidirectional Encoder is able to model the global semantics of the raw caption, which enhances the robustness of the framework. We extensively evaluate our proposals on the MS-COCO image captioning dataset and prove the effectiveness of our method when compared with the state-of-the-arts. Bing-Kun Bao, Zhiyi Tan 0002, Changsheng Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | Unbiased feature enhancement framework for cross-modality person re-identification
Bairu Chen, Zhiyi Tan 0002, Xi Shao, Bing-Kun Bao |
Multim. Syst. | 3 |