EDBT 2026 Demo / reviewers in the wild / expert
Shiyin Tan
dblp:283/3406
· DBLP profile ↗
13ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0001-8316-2838ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MMPG: MoE-based Adaptive Multi-Perspective Graph Fusion for Protein Representation LearningabstractGraph Neural Networks (GNNs) have been widely adopted for Protein Representation Learning (PRL), as residue interaction networks can be naturally represented as graphs. Current GNN-based PRL methods typically rely on single-perspective graph construction strategies, which capture partial properties of residue interactions, resulting in incomplete protein representations. To address this limitation, we propose MMPG, a framework that constructs protein graphs from multiple perspectives and adaptively fuses them via Mixture of Experts (MoE) for PRL. MMPG constructs graphs from physical, chemical, and geometric perspectives to characterize different properties of residue interactions. To capture both perspective-specific features and their synergies, we develop an MoE module, which dynamically routes perspectives to specialized experts, where experts learn intrinsic features and cross-perspective interactions. We quantitatively verify that MoE automatically specializes experts in modeling distinct levels of interaction—from individual representations, to pairwise inter-perspective synergies, and ultimately to a global consensus across all perspectives. Through integrating this multi-level information, MMPG produces superior protein representations and achieves advanced performance on four different downstream protein tasks. Yusong Wang 0003, Jialun Shen, Shiyin Tan, Mingkun Xu, Changshuo Wang 0001, Zixing Song, Prayag Tiwari |
AAAI | 5 |
| 2025 | Thermal-Aware Low-Light Image Enhancement: A Real-World Benchmark and a New Light-Weight ModelabstractEnhancing images captured under low-light conditions has been a topic of research for several years. Nonetheless, existing image restoration techniques mainly concentrate on reconstructing images from RGB data, often neglecting the possibility of utilizing additional modalities. With the progress in handheld technology, capturing thermal images with mobile devices has become straightforward. Investigating the integration of thermal data into image restoration presents a valuable research opportunity. Therefore, in this paper, we propose a multimodal low-light image enhancement task based on thermal information and establish a dataset named TLIE (Thermal-aware Low-light Image Enhancement), consisting of 1,113 samples. Each sample in our dataset includes a low-light image, a normal-light image, and the corresponding thermal map. Additionally, based on TLIE dataset, we develop a multimodal approach that simultaneously processes input images and thermal map data to produce the predicted normal-light images. We compare our method with previous unimodal and multimodal state-of-the-art LIE methods, and the experimental results and detailed ablation studies prove the effectiveness of our method. Zhen Wang 0004, Yaozu Wu, Dongyuan Li, Shiyin Tan, Zhishuai Yin |
AAAI | 4 |
| 2025 | Enhancing Graph Contrastive Learning for Protein Graphs from Perspective of InvarianceabstractGraph Contrastive Learning (GCL) improves Graph Neural Network (GNN)-based protein representation learning by enhancing its generalization and robustness. Existing GCL approaches for protein representation learning rely on 2D topology, where graph augmentation is solely based on topological features, ignoring the intrinsic biological properties of proteins. Besides, 3D structure-based protein graph augmentation remains unexplored, despite proteins inherently exhibiting 3D structures. To bridge this gap, we propose novel biology-aware graph augmentation strategies from the perspective of invariance and integrate them into the protein GCL framework. Specifically, we introduce Functional Community Invariance (FCI)-based graph augmentation, which employs spectral constraints to preserve topology-driven community structures while incorporating residue-level chemical similarity as edge weights to guide edge sampling and maintain functional communities. Furthermore, we propose 3D Protein Structure Invariance (3-PSI)-based graph augmentation, leveraging dihedral angle perturbations and secondary structure rotations to retain critical 3D structural information of proteins while diversifying graph views. Extensive experiments on four different protein-related tasks demonstrate the superiority of our proposed GCL protein representation learning framework. Yusong Wang 0003, Shiyin Tan, Jialun Shen, Haobo Song, Qi Xu 0008, Prayag Tiwari, Mingkun Xu |
ICML | 2 |
| 2025 | Taming Recommendation Bias with Causal Intervention on Evolving Personal PopularityabstractPopularity bias occurs when popular items are recommended far more frequently than they should be, negatively impacting both user experience and recommendation accuracy. Existing debiasing methods mitigate popularity bias often uniformly across all users and only partially consider the time evolution of users or items. However, users have different levels of preference for item popularity, and this preference is evolving over time. To address these issues, we propose a novel method called CausalEPP (Causal Intervention on Evolving Personal Popularity) for taming recommendation bias, which accounts for the evolving personal popularity of users. Specifically, we first introduce a metric called Evolving Personal Popularity to quantify each user's preference for popular items. Then, we design a causal graph that integrates evolving personal popularity into the conformity effect, and apply deconfounded training to mitigate the popularity bias of the causal graph. During inference, we consider the evolution consistency between users and items to achieve a better recommendation. Empirical studies demonstrate that CausalEPP outperforms baseline methods in reducing popularity bias while improving recommendation accuracy. Shiyin Tan, Dongyuan Li, Renhe Jiang, Zhen Wang 0004, Xingtong Yu, Manabu Okumura |
KDD (2) | 1 |
| 2025 | Video-based Transparent Object Segmentation via Temporal Feature AggregationabstractTransparent object segmentation from a single image has been investigated for several years. However, detecting transparent areas from video has not been well explored, especially for different kinds of transparent categories besides glass, due to the scarcity of such a dataset. Therefore, in this paper, we propose the video-based transparent object segmentation task and introduce the first-of-its-kind corresponding dataset named TransVid, which contains nearly 400 videos with a total of 18,523 frames. Based on TranVid, we further propose a new method called TranSeg, in which we innovatively introduce Graph Neural Networks into the temporal segmentation task and combined with a novel Diffusion Model to make the model's segmentation results more accurate. Experimental results show that TranSeg achieves higher accuracy with fewer parameters than previous state-of-the-art models, demonstrating the effectiveness of our method. Moreover, comprehensive ablation analysis reveal several fascinating insights and suggest viable paths for further research. Zhen Wang 0004, Dongyuan Li, Yaozu Wu, Peide Zhu, Shiyin Tan, Renhe Jiang |
ACM Multimedia | 5 |
| 2025 | DyG-Mamba: Continuous State Space Modeling on Dynamic GraphsabstractDynamic graph modeling aims to uncover evolutionary patterns in real-world systems, enabling accurate social recommendation and early detection of cancer cells. Inspired by the success of recent state space models in efficiently capturing long-term dependencies, we propose DyG-Mamba by translating dynamic graph modeling into a long-term sequence modeling problem. Specifically, inspired by Ebbinghaus' forgetting curve, we treat the irregular timespans between events as control signals, allowing DyG-Mamba to dynamically adjust the forgetting of historical information. This mechanism ensures effective usage of irregular timespans, thereby improving both model effectiveness and inductive capability. In addition, inspired by Ebbinghaus' review cycle, we redefine core parameters to ensure that DyG-Mamba selectively reviews historical information and filters out noisy inputs, further enhancing the model’s robustness. Through exhaustive experiments on 12 datasets covering dynamic link prediction and node classification tasks, we show that DyG-Mamba achieves state-of-the-art performance on most datasets, while demonstrating significantly improved computational and memory efficiency. Our code is available at https://github.com/Clearloveyuan/DyG-Mamba. Dongyuan Li, Shiyin Tan, Ying Zhang 0065, Ming Jin 0005, Shirui Pan, Manabu Okumura, Renhe Jiang |
NeurIPS | 2 |
| 2025 | A Unified Retrieval Framework with Document Ranking and EDU Filtering for Multi-document SummarizationabstractIn the field of multi-document summarization (MDS), transformerbased models have demonstrated remarkable success, yet they suffer an input length limitation.Current methods apply truncation after the retrieval process to fit the context length; however, they heavily depend on manually well-crafted queries, which are impractical to create for each document set for MDS.Additionally, these methods retrieve information at a coarse granularity, leading to the inclusion of irrelevant content.To address these issues, we propose a novel retrieval-based framework that integrates query selection and document ranking and shortening into a unified process.Our approach identifies the most salient elementary discourse units (EDUs) from input documents and utilizes them as latent queries.These queries guide the document ranking by calculating relevance scores.Instead of traditional truncation, our approach filters out irrelevant EDUs to fit the context length, ensuring that only critical information is preserved for summarization.We evaluate our framework on multiple MDS datasets, demonstrating consistent improvements in ROUGE metrics while confirming its scalability and flexibility across diverse model architectures.Additionally, we validate its effectiveness through an in-depth analysis, emphasizing its ability to dynamically select appropriate queries and accurately rank documents based on their relevance scores.These results demonstrate that our framework effectively addresses context-length constraints, establishing it as a robust and reliable solution for MDS. 1 * Both authors contributed equally to this research. Shiyin Tan, Jaeeon Park, Dongyuan Li, Renhe Jiang, Manabu Okumura |
SIGIR | 1 |
| 2024 | Community-Invariant Graph Contrastive LearningabstractGraph augmentation has received great attention in recent years for graph contrastive learning (GCL) to learn well-generalized node/graph representations. However, mainstream GCL methods often favor randomly disrupting graphs for augmentation, which shows limited generalization and inevitably leads to the corruption of high-level graph information, i.e., the graph community. Moreover, current knowledge-based graph augmentation methods can only focus on either topology or node features, causing the model to lack robustness against various types of noise. To address these limitations, this research investigated the role of the graph community in graph augmentation and figured out its crucial advantage for learnable graph augmentation. Based on our observations, we propose a community-invariant GCL framework to maintain graph community structure during learnable graph augmentation. By maximizing the spectral changes, this framework unifies the constraints of both topology and feature augmentation, enhancing the model’s robustness. Empirical evidence on 21 benchmark datasets demonstrates the exclusive merits of our framework. Code is released on Github (https://github.com/ShiyinTan/CI-GCL.git). Shiyin Tan, Dongyuan Li, Renhe Jiang, Ying Zhang 0065, Manabu Okumura |
ICML | 1 |
| 2023 | Temporal and Topological Augmentation-based Cross-view Contrastive Learning Model for Temporal Link PredictionabstractWith the booming development of social media, temporal link prediction (TLP), as a core technology, has been receiving increasing attention. However, current methods are based on graph neural networks, which suffer from the over-smoothing issue and easily yield indistinguishable node representations, degrading the prediction accuracy. Besides, they lack the ability to eliminate noisy temporal information and ignore the importance of high-order neighbor information for measuring the link probability between nodes. To solve these issues, we design a cross-view graph contrastive learning (GCL) framework for TLP, called Tacl. We first design two augmented views for GCL by enhancing the temporal and topological information to obtain distinguishable node representations. Then, we learn the evolution rule of temporal networks to help constrain consistency of node representations and eliminate noise. Finally, we incorporate the high-order neighbor information to measure the link probability between nodes. Extensive experiments demonstrate the effectiveness and robustness of Tacl. Dongyuan Li, Shiyin Tan, Yusong Wang 0003, Kotaro Funakoshi, Manabu Okumura |
CIKM | 2 |
| 2022 | Temporality- and Frequency-aware Graph Contrastive Learning for Temporal NetworkabstractGraph contrastive learning (GCL) methods aim to learn more distinguishable representations by contrasting positive and negative samples. They have received increasing attention in recent years due to their wide application in recommender systems and knowledge graphs. However, almost all GCL methods are applied to static networks and can not be extended to temporal networks directly. Furthermore, recent GCL models treat low- and high-frequency nodes equally in overall training objectives, which hinders the prediction precision. To solve the aforementioned problems, in this paper, we propose a Temporality- and Frequency-aware Graph Contrastive Learning for temporal networks (TF-GCL). Specifically, to learn more diverse representations for infrequent nodes and fully explore temporal information, we first generate two augmented views from the input graph based on topological and temporal perspectives. We then design a temporality and frequency-aware objective function to maximize the agreement between node representations of the two views. Experimental results demonstrate that TF-GCL remarkably achieves more robust node representations and significantly outperforms the state-of-the-art methods on six temporal link prediction benchmark datasets. Considering the reproducibility, we release our code on Github. Shiyin Tan, Jingyi You, Dongyuan Li |
CIKM | 1 |
| 2022 | Predicting combinations of drugs by exploiting graph embedding of heterogeneous networksabstractBACKGROUND: Drug combination, offering an insight into the increased therapeutic efficacy and reduced toxicity, plays an essential role in the therapy of many complex diseases. Although significant efforts have been devoted to the identification of drugs, the identification of drug combination is still a challenge. The current algorithms assume that the independence of feature selection and drug prediction procedures, which may result in an undesirable performance. RESULTS: To address this issue, we develop a novel Semi-supervised Heterogeneous Network Embedding algorithm (called SeHNE) to predict the combination patterns of drugs by exploiting the graph embedding. Specifically, the ATC similarity of drugs, drug-target, and protein-protein interaction networks are integrated to construct the heterogeneous networks. Then, SeHNE jointly learns drug features by exploiting the topological structure of heterogeneous networks and predicting drug combination. One distinct advantage of SeHNE is that features of drugs are extracted under the guidance of classification, which improves the quality of features, thereby enhancing the performance of prediction of drugs. Experimental results demonstrate that the proposed algorithm is more accurate than state-of-the-art methods on various data, implying that the joint learning is promising for the identification of drug combination. CONCLUSIONS: The proposed model and algorithm provide an effective strategy for the prediction of combinatorial patterns of drugs, implying that the graph-based drug prediction is promising for the discovery of drugs. Shiyin Tan, Zengfa Dou, Xiaoke Ma 0001 |
BMC Bioinform. | 2 |
| 2022 | Joint multi-label learning and feature extraction for temporal link prediction
Xiaoke Ma 0001, Shiyin Tan, Xianghua Xie, Xiaoxiong Zhong, Jingjing Deng 0001 |
Pattern Recognit. | 2 |
| 2020 | SeHNE: Semi-supervised Heterogeneous Network Embedding for Drug CombinationabstractDrug combinations, offering increased therapeutic efficacy and reduced toxicity, play an important role in therapy of many complex diseases. Although great efforts have been devoted to the prediction of single drugs, the identification of drug combination is really limited. The current algorithms assume the independence of features and prediction, resulting in an undesirable performance. To address this issue, we develop a novel semisupervised heterogeneous network embedding algorithm (called SeHNE) to predict drug combinations, where ATC similarity of drugs, drug-target and protein-protein interaction (PPI) networks are integrated to construct heterogeneous network. SeHNE jointly learns features of drugs by exploiting the topological structure of heterogeneous networks, and prediction of drug combination. One typical advantage of SeHNE is that features are extracted under the guidance of classification, thereby improving the accuracy of algorithms. Experimental results demonstrate that proposed algorithm is more accurate than state-of-the-art methods on the dataset we collected, and the re-training process could improve the accuracy of classifier. Shiyin Tan, Xiaoke Ma 0001 |
BIBM | 1 |