VLDB 2026 Research / reviewers in the wild / expert
Dan Wu 0002
dblp:19/5635-2
· DBLP profile ↗
13ranked-venue papers
0as first author
10since 2021 · last 2026
0000-0001-5838-0198ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Connectivity-Guided Sparsification of 2-FWL GNNs: Preserving Full Expressivity with Improved EfficiencyabstractHigher-order Graph Neural Networks (HOGNNs) based on the 2-FWL test achieve superior expressivity by modeling 2-node and 3-node interactions, but incur cubic computational cost. Existing efficiency methods typically reduce this burden at the expense of expressivity. We propose Co-Sparsify, a connectivity-aware sparsification framework that eliminates provably redundant computations while preserving full 2-FWL expressive power. Our key insight is that 3-node interactions are expressively necessary only within biconnected components, namely, maximal subgraphs where every node pair lies on a cycle. Outside these components, structural relationships are fully captured via 2-node message passing and graph readouts, rendering higher-order modeling unnecessary. Co-Sparsify restricts 2-node message passing to connected components and 3-node interactions to biconnected components, eliminating redundant computation without approximation or sampling. We prove that Co-Sparsified GNNs match the expressivity of the 2-FWL test. Empirically, when applied to PPGN, Co-Sparsify matches or exceeds accuracy on synthetic substructure counting tasks and achieves state-of-the-art performance on real-world benchmarks (ZINC, QM9 and TUD). This study demonstrates that high expressivity and scalability are not mutually exclusive: principled, topology-guided sparsification enables powerful, efficient GNNs with theoretical guarantees. Rongqin Chen 0001, Fan Mo 0002, Pak Lon Ip, Shenghui Zhang, Dan Wu 0002, Ye Li 0002, Leong Hou U |
AAAI | 5 |
| 2025 | A Multi-scenario Attention-based Generative Model for Personalized Blood Pressure Time Series ForecastingabstractContinuous blood pressure (BP) monitoring is essential for timely diagnosis and intervention in critical care settings. However, BP varies significantly across individuals, this inter-patient variability motivates the development of personalized models tailored to each patient’s physiology. In this work, we propose a personalized BP forecasting model mainly using electrocardiogram (ECG) and photoplethysmogram (PPG) signals. This time-series model incorporates 2D representation learning to capture complex physiological relationships. Experiments are conducted on datasets collected from three diverse scenarios with BP measurements from 60 subjects total. Results demonstrate that the model achieves accurate and robust BP forecasts across scenarios within the Association for the Advancement of Medical Instrumentation (AAMI) standard criteria. This reliable early detection of abnormal fluctuations in BP is crucial for at-risk patients undergoing surgery or intensive care. The proposed model provides a valuable addition for continuous BP tracking to reduce mortality and improve prognosis. Cheng Wan 0006, Chenjie Xie, Dan Wu 0002, Ye Li 0002 |
ICASSP | 4 |
| 2025 | Enhanced Subgraph Learning in 2-FWL GNNs via Local Connectivity, Spectral, and Distance EncodingsabstractDespite the theoretical expressiveness of 2-dimensional Folklore Weisfeiler-Lehman (2-FWL) Graph Neural Networks (GNNs), a significant gap persists between their theoretical capacity and their practical performance. To bridge this gap, we identify a critical limitation in current Graph Structural Encodings (GSEs): insufficient sensitivity to subtle structural variations, particularly in local connectivity, spectral features, and distance-based patterns. We show that widely used GSEs-such as Relative Random Walk Probability (RRWP) and monomial-based methods-lack full sensitivity across spectral frequency bands and long-range distances. Moreover, they fail to capture fine-grained local connectivity, which is essential for identifying cut nodes, biconnected components, and other higher-order structures that 2-FWL GNNs theoretically encode. To address these limitations, we propose CSDGSE (Connectivity, Spectral, and Distance Graph Structural Encoding), a novel GSE framework that jointly enhances sensitivity to: (1) exact local connectivity via hierarchical graph decomposition(2) full-frequency spectral features using expressive graph polynomials (e.g., Chebyshev), and (3) full-range distance interactions. A key innovation is our scalable divide-and-conquer algorithm for computing exact local connectivity across all node pairs, enabling efficient integration into modern GSEs. Extensive experiments show that CSDGSE outperforms existing GSEs in capturing complex structural patterns, achieving state-of-the-art results on molecular property prediction benchmarks like ZINC. Our work sets a new standard for GSEs by aligning theoretical expressiveness with practical effectiveness through enhanced structural sensitivity. Rongqin Chen 0001, Yan Li 0122, Dan Wu 0002, Fan Mo 0002, Shenghui Zhang, Pak Lon Ip, Hoi Cheong Iam, Ye Li 0002, Leong Hou U |
KDD (2) | 3 |
| 2025 | Multi-Scale Spatiotemporal Dynamic Graph Neural Network for Early Prediction of Mortality Risks in Heart Failure PatientsabstractHeart Failure (HF) stands as a principal public health issue worldwide, imposing a significant burden on healthcare systems. While existing prognostic methods have achieved certain milestones in predicting the early mortality risk of HF patients, they have not fully considered the dynamic interdependencies among physiological parameters. This paper introduces a novel Multi-scale Spatiotemporal Dynamic Graph Neural Network, MSTD-GNN, which enhances the prediction capability for early mortality in HF patients by dynamically extracting spatio-temporal information of physiological parameters from ICU patient Electronic Health Records (EHRs). Our model constructs dynamic graphs to model multivariate time series data, revealing the implicit dependencies between physiological parameters and capturing the inherent dynamics of the data. We conducted experiments using the MIMIC-III and MIMIC-IV datasets. The experimental results show that, compared to existing methods, MSTD-GNN demonstrates superior performance in predicting the early mortality risk of HF patients. On the MIMIC-III and MIMIC-IV datasets, the AUC scores of MSTD-GNN reached 83.93% and 81.74%, respectively. Furthermore, through dynamic graphs, our model unveils the dynamic relationships between physiological variables across different time scales. Rongqin Chen 0001, Jifu Qu, Ye Li 0002, Dan Wu 0002 |
IEEE J. Biomed. Health Informatics | 6 |
| 2024 | ECG-LLM: Leveraging Large Language Models for Low-Quality ECG Signal RestorationabstractElectrocardiography (ECG) signals are often plagued by various types of noise, which substantially undermines the precision of subsequent analysis. This paper presents ECG-LLM, an innovative model designed to address the challenges posed by low-quality ECG signals. Leveraging the capabilities of Large Language Models (LLMs), ECG-LLM forecasts and imputes missing values in 12-lead ECG data. ECG-LLM adapts the auto-regressive characteristics of LLMs, utilizes textual markers for timestamp information, and employs a strategy of freezing Transformer layers while training new embedding and projection layers. We used ECG datasets from multiple sources for our experiments. Experimental results demonstrate that ECG-LLM significantly outperforms state-of-the-art time series forecasting models, achieving a Mean Squared Error (MSE) of 0.644 and a Mean Absolute Error (MAE) of 0.456 for a forecast length of 64. Additionally, after using ECG-LLM to restore the electrocardiogram, the disease recognition accuracy of the baseline model was improved. Our findings highlight the feasibility and effectiveness of applying LLMs to ECG signal processing, offering a new perspective for medical signal analysis and providing a potential new approach for signal preprocessing in healthcare. Code is available at https://github.com/dragonlfy/ECG-LLM. Guosheng Cui, Cheng Wan 0006, Dan Wu 0002, Ye Li 0002 |
BIBM | 4 |
| 2024 | Advancing Semi-Supervised EEG Emotion Recognition through Feature Extraction with Mixup and Large Language ModelsabstractThe scarcity of labeled EEG data presents a significant challenge in emotion recognition. To address this issue, we propose PAWS, a semi-supervised learning framework specifically designed to enhance EEG-based emotion recognition, particularly in scenarios where labeled data is extremely limited. PAWS leverages the power of Intermediate Mixup, which improves domain adaptation by generating more robust features through strategic augmentation. Additionally, PAWS integrates large language models (LLMs) for enhanced feature extraction, enabling the framework to capture complex patterns in EEG signals even with minimal labeled data. In experiments with 1%, 5%, 10%, 15%, and 20% labeled data, PAWS consistently outperformed existing methods, with the most significant performance gains observed in scenarios with very few labeled samples. Notably, with 20% labeled data, PAWS achieved an accuracy of 95.81% ± 0.52% on the SEED dataset and 83.41% ± 0.31% on the SEED-IV dataset. These results demonstrate PAWS’s effectiveness in leveraging limited labeled data to achieve superior emotion recognition performance. This framework not only advances the state-of-the-art in EEG-based emotion recognition but also provides a robust foundation for future research in semi-supervised learning. Code is available at https://github.com/dragonlfy/PAWS. Shiyi Yao, Dan Wu 0002, Ye Li 0002 |
BIBM | 4 |
| 2024 | Fast label prediction based on shrunk anchor graph for semi-supervised incomplete multiview classificationabstractExisting anchor graph-based semi-supervised classification methods can not adopt partial available labels of data to produce discriminative anchor graph, which is even challenging for incomplete multi-view data. Addressing above issues, a fast label prediction based on shrunk anchor graph (FLP-SAG) is designed for semi-supervised incomplete multi-view classification, which is capable of learning discriminative anchor graph iteratively. Firstly, in each view a similarity-based anchor graph is constructed and expanded to the size of complete data to align the multiple views. Then these pre-constructed anchor graphs are fused to get a common anchor graph, which is ready to be shrunk based on the predicted labels with high confidence scores in each iteration. To speed up the classification, an efficient two-step label prediction strategy is developed without the calculation of dense matrix inverse. Experimental results on four real world datasets comparing with several recently proposed methods demonstrate the superiority of the proposed method. Guosheng Cui, Fusheng Hao, Dan Wu 0002, Ye Li 0002 |
ICME | 3 |
| 2024 | Semi-supervised Multi-view Clustering based on NMF with Fusion RegularizationabstractMulti-view clustering has attracted significant attention and application. Nonnegative matrix factorization is one popular feature of learning technology in pattern recognition. In recent years, many semi-supervised nonnegative matrix factorization algorithms were proposed by considering label information, which has achieved outstanding performance for multi-view clustering. However, most of these existing methods have either failed to consider discriminative information effectively or included too much hyper-parameters. Addressing these issues, a semi-supervised multi-view nonnegative matrix factorization with a novel fusion regularization (FRSMNMF) is developed in this article. In this work, we uniformly constrain alignment of multiple views and discriminative information among clusters with designed fusion regularization. Meanwhile, to align the multiple views effectively, two kinds of compensating matrices are used to normalize the feature scales of different views. Additionally, we preserve the geometry structure information of labeled and unlabeled samples by introducing the graph regularization simultaneously. Due to the proposed methods, two effective optimization strategies based on multiplicative update rules are designed. Experiments implemented on six real-world datasets have demonstrated the effectiveness of our FRSMNMF comparing with several state-of-the-art unsupervised and semi-supervised approaches. Guosheng Cui, Ruxin Wang 0001, Dan Wu 0002, Ye Li 0002 |
ACM Trans. Knowl. Discov. Data | 3 |
| 2023 | TCSA: A Text-Guided Cross-View Medical Semantic Alignment Framework for Adaptive Multi-view Visual Representation Learning
Hongyang Lei, Huazhen Huang, Guosheng Cui, Ruxin Wang 0001, Dan Wu 0002, Ye Li 0002 |
ISBRA | 6 |
| 2023 | Incomplete Multiview Clustering Using Normalizing Alignment Strategy With Graph RegularizationabstractMatrix factorization has demonstrated promising performance in the incomplete multiview clustering (IMC) tasks. However, many algorithms require feature normalization operations to ensure the stability of model results, so either the convergence is unstable, or the objective function cannot fit the data well. Addressing these issues, we propose a novel IMC algorithm using a normalizing alignment strategy (IMCNAS) based on nonnegative matrix factorization. Specifically, the columns of the basis matrices are constrained into unit vector space, which integrates the feature normalization and the optimizing process, and makes the model converge fast and stable. On the other hand, this enables the model to fit the data better and produce more reasonable factorization results. Further, we develop a novel pairwise co-regularization to align incomplete multiple views more directly, without introducing a common consensus matrix like traditional centroid-based co-regularization. Graph regularization is also incorporated in the proposed model to utilize the geometrical information of data. We implement IMCNAS with a centroid-based regularization and a pairwise co-regularization respectively, and leads to two variants, i.e., IMCNAS-1 and IMCNAS-2. Both variants are optimized with multiplicative updating rules. Extensive experiments conducted on various real-world datasets comparing several state-of-the-art IMC methods verified the effectiveness of the proposed methods. The source code is available at:https://github.com/GuoshengCui/IMCNAS. Guosheng Cui, Ruxin Wang 0001, Dan Wu 0002, Ye Li 0002 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2020 | A Semi-supervised Approach for Early Identifying the Abnormal Carotid Arteries Using a Modified Variational AutoencoderabstractCarotid artery lesions could be the pathology of subclinical atherosclerosis and hence lead to the onset of stroke. Early detection of abnormal carotid artery might help to better identify individuals susceptible to stroke. Considering the carotid artery ultrasonography is time-consuming and costly, the object of this paper is to establish a model to detect the status of carotid artery for preliminary screening of stroke, according to the simple physiological examination and survey information. However, most of the previous studies were based on the linear regression or the traditional machine learning methods, those suffer from two limitations. One is the limited labeled samples, and the other one is the missing data. To address these issues, we firstly propose a semi-supervised approach based on a modified variational autoencoder (VAE) to identify the abnormal carotid arteries. In this paper, a mixture of mean and K th nearest neighbours (MKNN) and a modified VAE were used for missing data imputation. The experimental results demonstrate that the proposed method can not only handle the missing values, but also outperform four widely used supervised approaches. Therefore, we can conclude that this semi-supervised model is a promising way to identify the abnormal carotid arteries. Xiaoxiang Huang, Guosheng Cui, Dan Wu 0002, Ye Li 0002 |
BIBM | 3 |
| 2020 | Learning physical properties in complex visual scenes: An intelligent machine for perceiving blood flow dynamics from static CT angiography imaging
Zhifan Gao, Xin Wang 0045, Shanhui Sun, Dan Wu 0002, Youbing Yin, Xin Liu 0023, Heye Zhang, Victor Hugo C. de Albuquerque |
Neural Networks | 4 |
| 2015 | Motion Estimation of Common Carotid Artery Wall Using a H ∞ Filter Based Block Matching Method
Zhifan Gao, Huahua Xiong, Heye Zhang, Dan Wu 0002, Minhua Lu, Kelvin K. L. Wong, Yuan-Ting Zhang |
MICCAI (3) | 4 |