EDBT 2026 Demo / reviewers in the wild / expert
Can Wang 0004
dblp:71/4716-4
· DBLP profile ↗
27ranked-venue papers in the field
5as first author
11since 2021 · last 2025
0000-0002-2890-0057ORCID · conflict
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 19 (3 first)Information Retrieval & Web Search · 5 (1 first)Database Systems & Data Management · 3 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | OASIS: Harnessing Diffusion Adversarial Network for Ocean Salinity Imputation using Sparse Drifter TrajectoriesabstractOcean salinity plays a vital role in circulation, climate, and marine ecosystems, yet its measurement is often sparse, irregular, and noisy, especially in drifter-based datasets. Traditional approaches, such as remote sensing and optimal interpolation, rely on linearity and stationarity, and are limited by cloud cover, sensor drift, and low satellite revisit rates. While machine learning models offer flexibility, they often fail under severe sparsity and lack principled ways to incorporate physical covariates without specialized sensors. In this paper, we introduce the OceAn Salinity Imputation System, a novel diffusion adversarial framework designed to address these challenges by: (1) employing a transformer-based global dependency capturing module to learn long-range spatio-temporal correlations from sparse trajectories; (2) constructing a generative imputation model that conditions on easily observed tidal covariates to progressively refine imputed salinity fields; and (3) using a scheduler diffusion method to enhance the model's robustness. This unified architecture exploits the periodic nature of tidal signals as a proxy for unmeasured physical drivers, without the need for additional equipment. We evaluate OASIS on four benchmark datasets, including one real-world measurement from Fort Pierce Inlet and three simulated Gulf of Mexico trajectories. Results show consistent improvements over both traditional and neural baselines, achieving up to 52.5% reduction in MAE compared to Kriging. We also develop a lightweight, web-based deployment system that enables salinity imputation through interactive and batch interfaces, available at: https://github.com/yfeng77/OASIS. Bo Li 0042, Yingqi Feng, Ming Jin 0005, Xin Zheng 0008, Yufei Tang, Laurent M. Chérubin, Can Wang 0004, Alan Wee-Chung Liew, Qinghua Lu 0001, Jingwei Yao, Hong Zhang 0028, Shirui Pan, Xingquan Zhu 0001 |
CIKM | 7 |
| 2025 | Cross-Modal Sequential Point-of-Interest Recommendation with Lightweight Hybrid Fusion Strategy
Tianxing Wang 0004, Can Wang 0004, Hui Tian 0001, Hong Shen 0001 |
DaWaK | 2 |
| 2025 | Local-Aware Convolutional Modulation for Short-Term Sequential Recommendation
Tianxing Wang 0004, Can Wang 0004, Hui Tian 0001, Hong Shen 0001 |
DaWaK | 2 |
| 2025 | Test-Time GNN Model Evaluation on Dynamic GraphsabstractDynamic graph neural networks (DGNNs) have emerged as a leading paradigm for learning from dynamic graphs, which are commonly used to model real-world systems and applications. However, due to the evolving nature of dynamic graph data distributions over time, well-trained DGNNs often face significant performance uncertainty when inferring on unseen and unlabeled test graphs in practical deployment. In this case, evaluating the performance of deployed DGNNs at test time is crucial to determine whether a well-trained DGNN is suited for inference on an unseen dynamic test graph. In this work, we introduce a new research problem: DGNN model evaluation, which aims to assess the performance of a specific DGNN model trained on observed dynamic graphs by estimating its performance on unseen dynamic graphs during test time. Specifically, we propose a Dynamic Graph neural network Evaluator, dubbed DYGEvAL, toaddress this new problem. The proposed DyGEvAL involves a two-stage framework: (1) test-time dynamic graph simulation, which captures the training-test distributional differences as supervision signals and trains an evaluator; and (2) DyGEvAL development and training, which accurately estimates the performance of the well-trained DGNN model on the test-time dynamic graphs. Extensive experiments demonstrate that the proposed DyGEvAL serves as an effective evaluator for assessing various DGNN backbones across different dynamic graphs under distribution shifts. Bo Li 0042, Xin Zheng 0008, Ming Jin 0005, Can Wang 0004, Shirui Pan |
ICDM | 4 |
| 2025 | Heterogeneous network for Hierarchical Fine-Grained Domain Fake News Detection
Yue Wang 0150, Shizhong Yuan, Weimin Li 0001, Yifan Feng 0002, Fangfang Liu 0008, Can Wang 0004, Quan-Ke Pan |
Inf. Process. Manag. | 7 |
| 2025 | User-based clustering deep model for the sequential point-of-interest recommendation
Tianxing Wang 0004, Can Wang 0004, Hui Tian 0001, Alan Wee-Chung Liew |
Knowl. Inf. Syst. | 2 |
| 2023 | How Does ChatGPT Affect Fake News Detection Systems?
Bo Li 0042, Jiaxin Ju, Can Wang 0004, Shirui Pan |
ADMA (2) | 3 |
| 2023 | GeoMixer: The MLP-Based Sequential POI Recommender with Travel Routing ModellingabstractNowadays, with the rise of location-based services, the personalized sequential POI recommendation has become a pivotal element for enhancing customer experiences. Although many previous POI recommendation models have shown promising results and improvements in this area, several challenges still exist in this field. Firstly, the previous sequential recommenders do not well-utilize the geographical features that are highly affecting the user’s future choices of visits. Furthermore, the self-attention mechanism, which is a popular method used in sequential POI recommendation, has a limitation in treating the input user sequence as an unordered set. Using positional embedding is a typical way to overcome this limitation. However, the use of such embeddings may potentially restrict the model’s ability to learn meaningful patterns in user preferences among POIs. To address these challenges, we propose GeoMixer, a novel MLP-based sequential POI recommender that incorporates travel routing distance to capture geographical features and leverages Multi-layer Perceptron (MLP) architecture to model the spatial and sequential patterns in the sequential POI recommendations. By adopting MLP mixing layers, GeoMixer has the capability of memorizing the chronological order of the input POIs without the positional embedding and can emphasize the important latent features of each POI. The use of the travel routing information improves the model’s ability of capturing spatial patterns during the model learning process. Extensive experiments on real-world datasets show that GeoMixer outperforms state-of-theart methods in various metrics, highlighting the significance of incorporating travel routing distance and leveraging MLP architecture in sequential POI recommendation systems. Tianxing Wang 0004, Can Wang 0004, Hui Tian 0001, Hong Shen 0001 |
ICDM | 2 |
| 2023 | Global-Aware External Attention Deep Model for Sequential Recommendation
Tianxing Wang 0004, Can Wang 0004 |
PAKDD (3) | 2 |
| 2022 | Hyperbolic Personalized Tag Recommendation
Weibin Zhao, Aoran Zhang 0001, Lin Shang 0001, Yonghong Yu, Li Zhang 0013, Can Wang 0004, Jiajun Chen 0001, Hongzhi Yin |
DASFAA (2) | 6 |
| 2022 | An influence maximization method based on crowd emotion under an emotion-based attribute social networkabstractMost research on influence maximization focuses on the network structure features of the diffusion process but lacks the consideration of multi-dimensional characteristics. This paper proposes the attributed influence maximization based on the crowd emotion, aiming to apply the user’s emotion and group features to study the influence of multi-dimensional characteristics on information propagation. To measure the interaction effects of individual emotions, we define the user emotion power and the cluster credibility, and propose a potential influence user discovery algorithm based on the emotion aggregation mechanism to locate seed candidate sets. A two-factor information propagation model is then introduced, which considers the complexity of real networks. Experiments on real-world datasets demonstrate the effectiveness of the proposed algorithm. The results outperform the heuristic methods and are almost consistent with the greedy methods yet with improved time performance. Weimin Li 0001, Yaqiong Li, Wei Liu 0027, Can Wang 0004 |
Inf. Process. Manag. | 4 |
| 2020 | Spectrum-Guided Adversarial Disparity LearningabstractIt has been a significant challenge to portray intraclass disparity precisely in the area of activity recognition, as it requires a robust representation of the correlation between subject-specific variation for each activity class. In this work, we propose a novel end-to-end knowledge directed adversarial learning framework, which portrays the class-conditioned intraclass disparity using two competitive encoding distributions and learns the purified latent codes by denoising learned disparity. Furthermore, the domain knowledge is incorporated in an unsupervised manner to guide the optimization and further boosts the performance. The experiments on four HAR benchmark datasets demonstrate the robustness and generalization of our proposed methods over a set of state-of-the-art. We further prove the effectiveness of automatic domain knowledge incorporation in performance enhancement. Zhe Liu 0023, Lina Yao 0001, Lei Bai 0001, Xianzhi Wang 0001, Can Wang 0004 |
KDD | 5 |
| 2020 | Efficient processing of reverse nearest neighborhood queries in spatial databases
Md. Saiful Islam 0003, Bojie Shen, Can Wang 0004, David Taniar, Junhu Wang |
Inf. Syst. | 3 |
| 2019 | Top-N Hashtag Prediction via Coupling Social Influence and Homophily
Can Wang 0004, Yunwei Zhao, Chihung Chi, Willem-Jan van den Heuvel, Kwok-Yan Lam, Bela Stantic |
ADMA | 1 |
| 2019 | TOM: A Threat Operating Model for Early Warning of Cyber Security Threats
Can Wang 0004, Yunwei Zhao, Kwok-Yan Lam, Chihung Chi, Hui Tian 0001 |
ADMA | 3 |
| 2019 | DAMTRNN: A Delta Attention-Based Multi-task RNN for Intention Recognition
Weitong Chen 0001, Lin Yue, Bohan Li 0001, Can Wang 0004, Quan Z. Sheng |
ADMA | 4 |
| 2019 | A Methodology for Resolving Heterogeneity and Interdependence in Data Analytics
Yunwei Zhao, Can Wang 0004, Min Shu, Chihung Chi, Yonghong Yu |
ADMA | 3 |
| 2019 | Unfolding the Mixed and Intertwined: A Multilevel View of Topic Evolution on Twitter
Yunwei Zhao, Can Wang 0004, Willem-Jan van den Heuvel, Chihung Chi, Weimin Li 0001 |
ADMA | 2 |
| 2018 | Geographical Proximity Boosted Recommendation Algorithms for Real Estate
Yonghong Yu, Can Wang 0004, Li Zhang 0013, Rong Gao 0001, Hua Wang 0002 |
WISE (2) | 2 |
| 2018 | Coupled Clustering Ensemble by Exploring Data InterdependenceabstractClustering ensembles combine multiple partitions of data into a single clustering solution. It is an effective technique for improving the quality of clustering results. Current clustering ensemble algorithms are usually built on the pairwise agreements between clusterings that focus on the similarity via consensus functions, between data objects that induce similarity measures from partitions and re-cluster objects, and between clusters that collapse groups of clusters into meta-clusters. In most of those models, there is a strong assumption on IIDness (i.e., independent and identical distribution), which states that base clusterings perform independently of one another and all objects are also independent. In the real world, however, objects are generally likely related to each other through features that are either explicit or even implicit. There is also latent but definite relationship among intermediate base clusterings because they are derived from the same set of data. All these demand a further investigation of clustering ensembles that explores the interdependence characteristics of data. To solve this problem, a new coupled clustering ensemble (CCE) framework that works on the interdependence nature of objects and intermediate base clusterings is proposed in this article. The main idea is to model the coupling relationship between objects by aggregating the similarity of base clusterings, and the interactive relationship among objects by addressing their neighborhood domains. Once these interdependence relationships are discovered, they will act as critical supplements to clustering ensembles. We verified our proposed framework by using three types of consensus function: clustering-based, object-based, and cluster-based. Substantial experiments on multiple synthetic and real-life benchmark datasets indicate thatCCEcan effectively capture the implicit interdependence relationships among base clusterings and among objects with higher clustering accuracy, stability, and robustness compared to 14 state-of-the-art techniques, supported by statistical analysis. In addition, we show that the final clustering quality is dependent on the data characteristics (e.g., quality and consistency) of base clusterings in terms of sensitivity analysis. Finally, the applications in document clustering, as well as on the datasets with much larger size and dimensionality, further demonstrate the effectiveness, efficiency, and scalability of our proposed models. Can Wang 0004, Chihung Chi, Zhong She, Longbing Cao, Bela Stantic |
ACM Trans. Knowl. Discov. Data | 1 |
| 2013 | Coupled clustering ensemble: Incorporating coupling relationships both between base clusterings and objectsabstractClustering ensemble is a powerful approach for improving the accuracy and stability of individual (base) clustering algorithms. Most of the existing clustering ensemble methods obtain the final solutions by assuming that base clusterings perform independently with one another and all objects are independent too. However, in real-world data sources, objects are more or less associated in terms of certain coupling relationships. Base clusterings trained on the source data are complementary to one another since each of them may only capture some specific rather than full picture of the data. In this paper, we discuss the problem of explicating the dependency between base clusterings and between objects in clustering ensembles, and propose a framework for coupled clustering ensembles (CCE). CCE not only considers but also integrates the coupling relationships between base clusterings and between objects. Specifically, we involve both the intra-coupling within one base clustering (i.e., cluster label frequency distribution) and the inter-coupling between different base clusterings (i.e., cluster label co-occurrence dependency). Furthermore, we engage both the intra-coupling between two objects in terms of the base clustering aggregation and the inter-coupling among other objects in terms of neighborhood relationship. This is the first work which explicitly addresses the dependency between base clusterings and between objects, verified by the application of such couplings in three types of consensus functions: clustering-based, object-based and cluster-based. Substantial experiments on synthetic and UCI data sets demonstrate that the CCE framework can effectively capture the interactions embedded in base clusterings and objects with higher clustering accuracy and stability compared to several state-of-the-art techniques, which is also supported by statistical analysis. Can Wang 0004, Zhong She, Longbing Cao |
ICDE | 1 |
| 2013 | Efficient Mining of Contrast Patterns on Large Scale Imbalanced Real-Life Data
Jinjiu Li, Can Wang 0004, Wei Wei 0039, Chunming Liu |
PAKDD (1) | 2 |
| 2013 | A Coupled Clustering Approach for Items Recommendation
Yonghong Yu, Can Wang 0004, Yang Gao 0001, Longbing Cao, Xixi Chen |
PAKDD (2) | 2 |
| 2013 | Erratum: A Coupled Clustering Approach for Items Recommendation
Yonghong Yu, Can Wang 0004, Yang Gao 0001, Longbing Cao |
PAKDD (2) | 2 |
| 2013 | Efficient Selection of Globally Optimal Rules on Large Imbalanced Data Based on Rule Coverage Relationship AnalysisabstractRule-based anomaly and fraud detection systems often suffer from massive false alerts against a huge number of enterprise transactions. A crucial and challenging problem is to effectively select a globally optimal rule set which can capture very rare anomalies dispersed in large-scale background transactions. The existing rule selection methods which suffer significantly from complex rule interactions and overlapping in large imbalanced data, often lead to very high false positive rate. In this paper, we analyze the interactions and relationships between rules and their coverage on transactions, and propose a novel metric, Max Coverage Gain. Max Coverage Gain selects the optimal rule set by evaluating the contribution of each rule in terms of overall performance to cut out those locally significant but globally redundant rules, without any negative impact on the recall. An effective algorithm, MCGminer, is then designed with a series of built-in mechanisms and pruning strategies to handle complex rule interactions and reduce computational complexity towards identifying the globally optimal rule set. Substantial experiments on 13 UCI data sets and a real time online banking transactional database demonstrate that MCGminer achieves significant improvement on both accuracy, scalability, stability and efficiency on large imbalanced data compared to several state-of-the-art rule selection techniques. Longbing Cao, Jinjiu Li, Can Wang 0004, Philip S. Yu |
SDM | 3 |
| 2012 | CD: A Coupled Discretization Algorithm
Can Wang 0004, Mingchun Wang, Zhong She, Longbing Cao |
PAKDD (2) | 1 |
| 2011 | Coupled nominal similarity in unsupervised learningabstractThe similarity between nominal objects is not straightforward, especially in unsupervised learning. This paper proposes coupled similarity metrics for nominal objects, which consider not only intra-coupled similarity within an attribute (i.e., value frequency distribution) but also inter-coupled similarity between attributes (i.e. feature dependency aggregation). Four metrics are designed to calculate the inter-coupled similarity between two categorical values by considering their relationships with other attributes. The theoretical analysis reveals their equivalent accuracy and superior efficiency based on intersection against others, in particular for large-scale data. Substantial experiments on extensive UCI data sets verify the theoretical conclusions. In addition, experiments of clustering based on the derived dissimilarity metrics show a significant performance improvement. Can Wang 0004, Longbing Cao, Mingchun Wang, Jinjiu Li, Wei Wei 0039, Yuming Ou |
CIKM | 1 |