Wenxin Tai

dblp:284/4234 · DBLP profile ↗
← Back
15ranked-venue papers in the field
2as first author
15since 2021 · last 2026
0000-0001-7364-8324ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 7Data Mining & Knowledge Discovery · 6 (1 first)Database Systems & Data Management · 1 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 1
YearPublicationVenuePosition
2026 Tracing paths, pruning noise: Toward robust IP geolocation via topology-guided shaping and refinement
Xueting Liu 0005, Wenxin Tai, Ting Zhong, Yong Wang 0046, Kai Chen 0005, Fan Zhou 0002
Inf. Process. Manag.2
2026 Invariant learning improves out-of-distribution generalization for IP geolocation
Xueting Liu 0005, Wenxin Tai, Joojo Walker, Yong Wang 0046, Kai Chen 0005, Fan Zhou 0002
Inf. Process. Manag.3
2026 Corrigendum: Score-based Graph Learning for Urban Flow Prediction
abstract
This is a corrigendum for the article "Score-based Graph Learning for Urban Flow Prediction" published in ACM Trans. Intell. Syst. Technol. 15, 3, Article 59 (May 2024), 25 pages.
Xucheng Luo, Wenxin Tai, Kunpeng Zhang 0001, Goce Trajcevsky, Fan Zhou 0002
ACM Trans. Intell. Syst. Technol.3
2025 Extracting key insights from earnings call transcript via information-theoretic contrastive learning
Wenxin Tai, Fan Zhou 0002, Qiang Gao 0003, Ting Zhong, Kunpeng Zhang 0001
Inf. Process. Manag.2
2025 Landmark-v6: A stable IPv6 landmark representation method based on multi-feature clustering
Zhaorui Ma, Xinhao Hu, Fenlin Liu, Xiangyang Luo 0001, Wenxin Tai, Guoming Ren, Zheng Er
Inf. Process. Manag.6
2025 Information diffusion prediction via meta-knowledge learners
Zhangtao Cheng, Jienan Zhang, Xovee Xu, Wenxin Tai, Fan Zhou 0002, Goce Trajcevski, Ting Zhong
Inf. Sci.4
2024 Motif-Consistent Counterfactuals with Adversarial Refinement for Graph-level Anomaly Detection
abstract
Graph-level anomaly detection is significant in diverse domains. To improve detection performance, counterfactual graphs have been exploited to benefit the generalization capacity by learning causal relations. Most existing studies directly introduce perturbations (e.g., flipping edges) to generate counterfactual graphs, which are prone to alter the semantics of generated examples and make them off the data manifold, resulting in sub-optimal performance. To address these issues, we propose a novel approach, Motif-consistent Counterfactuals with Adversarial Refinement (MotifCAR), for graph-level anomaly detection. The model combines the motif of one graph, the core subgraph containing the identification (category) information, and the contextual subgraph (non-motif) of another graph to produce a raw counterfactual graph. However, the produced raw graph might be distorted and cannot satisfy the important counterfactual properties: Realism, Validity, Proximity and Sparsity. Towards that, we present a Generative Adversarial Network (GAN)-based graph optimizer to refine the raw counterfactual graphs. It adopts the discriminator to guide the generator to generate graphs close to realistic data, i.e., meet the property Realism. Further, we design the motif consistency to force the motif of the generated graphs to be consistent with the realistic graphs, meeting the property Validity. Also, we devise the contextual loss and connection loss to control the contextual subgraph and the newly added links to meet the properties Proximity and Sparsity. As a result, the model can generate high-quality counterfactual graphs. Experiments demonstrate the superiority of MotifCAR.
Chunjing Xiao, Shikang Pang, Wenxin Tai, Goce Trajcevski, Fan Zhou 0002
KDD3
2024 Analyzing and Mitigating Repetitions in Trip Recommendation
abstract
Trip recommendation has emerged as a highly sought-after service over the past decade. Although current studies significantly understand human intention consistency, they struggle with undesired repetitive outcomes that need resolution. We make two pivotal discoveries using statistical analyses and experimental designs: (1) The occurrence of repetitions is intricately linked to the models and decoding strategies. (2) During training and decoding, adding perturbations to logits can reduce repetition. Motivated by these observations, we introduce AR-Trip (Anti Repetition for Trip Recommendation), which incorporates a cycle-aware predictor comprising three mechanisms to avoid duplicate Points-of-Interest (POIs) and demonstrates their effectiveness in alleviating repetition. Experiments on four public datasets illustrate that AR-Trip successfully mitigates repetition issues while enhancing precision.
Wenzheng Shu, Kangqi Xu, Wenxin Tai, Ting Zhong, Yong Wang 0046, Fan Zhou 0002
SIGIR3
2024 Score-based Graph Learning for Urban Flow Prediction
abstract
Accurate urban flow prediction (UFP) is crucial for a range of smart city applications such as traffic management, urban planning, and risk assessment. To capture the intrinsic characteristics of urban flow, recent efforts have utilized spatial and temporal graph neural networks to deal with the complex dependence between the traffic in adjacent areas. However, existing graph neural network based approaches suffer from several critical drawbacks, including improper graph representation of urban traffic data, lack of semantic correlation modeling among graph nodes, and coarse-grained exploitation of external factors. To address these issues, we propose DiffUFP , a novel probabilistic graph-based framework for UFP. DiffUFP consists of two key designs: (1) a semantic region dynamic extraction method that effectively captures the underlying traffic network topology, and (2) a conditional denoising score-based adjacency matrix generator that takes spatial, temporal, and external factors into account when constructing the adjacency matrix rather than simply concatenation in existing studies. Extensive experiments conducted on real-world datasets demonstrate the superiority of DiffUFP over the state-of-the-art UFP models and the effect of the two specific modules.
Xucheng Luo, Wenxin Tai, Kunpeng Zhang 0001, Goce Trajcevski, Fan Zhou 0002
ACM Trans. Intell. Syst. Technol.3
2023 Enhancing Information Diffusion Prediction with Self-Supervised Disentangled User and Cascade Representations
abstract
Accurately predicting information diffusion is critical for a vast range of applications. Existing methods generally consider user re-sharing behaviors to be driven by a single intent, and/or assume cascade temporal influence to be unchanged, which might not be consistent with real-world scenarios. To address these issues, we propose a self-supervised disentanglement framework (DisenIDP) for information diffusion prediction. First, we construct intent-aware hypergraphs to capture users' potential intents from different perspectives, and then perform the light hypergraph convolution to adaptively activate disentangled intents. Second, we extract long-term and short-term cascade influence via independent attention-based encoders. Finally, we set a self-supervised disentanglement task to alleviate the information loss and learn better-disentanglement representations. Extensive experiments conducted on two real-world social datasets demonstrate that DisenIDP outperforms state-of-the-art models across several settings.
Zhangtao Cheng, Wenxue Ye, Leyuan Liu 0002, Wenxin Tai, Fan Zhou 0002
CIKM4
2023 Towards Trustworthy Rumor Detection with Interpretable Graph Structural Learning
abstract
The exponential growth of digital information has amplified the necessity for effective rumor detection on social media. However, existing approaches often neglect the inherent noise and uncertainty in rumor propagation, leading to obscure learning mechanisms. Moreover, current deep-learning methodologies, despite their top-tier performance, are heavily dependent on supervised learning, which is labor-intensive and inefficient. Their prediction credibility is also questionable. To tackle these issues, we present a new framework, TrustRD, for reliable rumor detection. Our framework incorporates a self-supervised learning module, designed to derive interpretable and informative representations with less reliance on large labeled data sets. A downstream model based on Bayesian networks, which is further refined with adversarial training, enhances performance while providing a quantifiable trustworthiness assessment of results. Our methods' effectiveness is confirmed through experiments on two benchmark datasets.
Leyuan Liu 0002, Zhangtao Cheng, Wenxin Tai, Fan Zhou 0002
CIKM4
2023 TrustGeo: Uncertainty-Aware Dynamic Graph Learning for Trustworthy IP Geolocation
abstract
The rising popularity of online social network services has attracted a lot of research focusing on mining various user patterns. Among them, accurate IP geolocation is essential for a plethora of location-aware applications. However, despite extensive research efforts and significant advances, the "accurate and reliable'' desideratum is yet to be achieved at a higher quality level. This work presents a graph neural network (GNN)-based model, called TrustGeo, for trustworthy street-level IP geolocation. A distinct and important aspect of TrustGeo is the incorporation of sources of uncertainty in the learning process. The results of our extensive experimental evaluations on three real-world datasets demonstrate the superiority of our framework in significantly improving the accuracy and trustworthiness of street-level IP geolocation. Our code and datasets are available at https://github.com/ICDM-UESTC/TrustGeo.
Wenxin Tai, Bin Chen 0030, Fan Zhou 0002, Ting Zhong, Goce Trajcevski, Yong Wang 0046, Kai Chen 0005
KDD1
2023 Imputation-based Time-Series Anomaly Detection with Conditional Weight-Incremental Diffusion Models
abstract
Existing anomaly detection models for time series are primarily trained with normal-point-dominant data and would become ineffective when anomalous points intensively occur in certain episodes. To solve this problem, we propose a new approach, called DiffAD, from the perspective of time series imputation. Unlike previous prediction- and reconstruction-based methods that adopt either partial or complete data as observed values for estimation, DiffAD uses a density ratio-based strategy to select normal observations flexibly that can easily adapt to the anomaly concentration scenarios. To alleviate the model bias problem in the presence of anomaly concentration, we design a new denoising diffusion-based imputation method to enhance the imputation performance of missing values with conditional weight-incremental diffusion, which can preserve the information of observed values and substantially improves data generation quality for stable anomaly detection. Besides, we customize a multi-scale state space model to capture the long-term dependencies across episodes with different anomaly patterns. Extensive experimental results on real-world datasets show that DiffAD performs better than state-of-the-art benchmarks.
Chunjing Xiao, Zehua Gou, Wenxin Tai, Kunpeng Zhang 0001, Fan Zhou 0002
KDD3
2023 RIPGeo: Robust Street-Level IP Geolocation
abstract
IP geolocation refers to the process of determining the geographic locations of Internet Protocol (IP) addresses, which is important for mobile computing and spatial data management. Despite extensive research efforts, a client-independent geolocation service with high accuracy and reliability has not yet been developed. This paper presents a graph neural network (GNN) model, dubbed RIPGeo, for robust street-level IP geolocation. Three factors that affect data quality are identified, and the importance of considering data quality in algorithm development is emphasized. Two novel self-supervised perturbational training strategies are proposed to enhance the generalization and robustness of the model. A multi-task learning framework is introduced to solve the homogenized representation problem caused by perturbational training, demonstrating much more efficiency than prevailing solutions. Theoretical analysis and experimental results demonstrate the superiority of our framework in significantly improving the accuracy and stability of street-level IP geolocation.
Wenxin Tai, Bin Chen 0030, Ting Zhong, Yong Wang 0046, Kai Chen 0005, Fan Zhou 0002
MDM1
2022 Contrastive Trajectory Learning for Tour Recommendation
abstract
The main objective of Personalized Tour Recommendation (PTR) is to generate a sequence of point-of-interest (POIs) for a particular tourist, according to the user-specific constraints such as duration time, start and end points, the number of attractions planned to visit, and so on. Previous PTR solutions are based on either heuristics for solving the orienteering problem to maximize a global reward with a specified budget or approaches attempting to learn user visiting preferences and transition patterns with the stochastic process or recurrent neural networks. However, existing learning methodologies rely on historical trips to train the model and use the next visited POI as the supervised signal, which may not fully capture the coherence of preferences and thus recommend similar trips to different users, primarily due to the data sparsity problem and long-tailed distribution of POI popularity. This work presents a novel tour recommendation model by distilling knowledge and supervision signals from the trips in a self-supervised manner. We propose Contrastive Trajectory Learning for Tour Recommendation (CTLTR), which utilizes the intrinsic POI dependencies and traveling intent to discover extra knowledge and augments the sparse data via pre-training auxiliary self-supervised objectives. CTLTR provides a principled way to characterize the inherent data correlations while tackling the implicit feedback and weak supervision problems by learning robust representations applicable for tour planning. We introduce a hierarchical recurrent encoder-decoder to identify tourists’ intentions and use the contrastive loss to discover subsequence semantics and their sequential patterns through maximizing the mutual information. Additionally, we observe that a data augmentation step as the preliminary of contrastive learning can solve the overfitting issue resulting from data sparsity. We conduct extensive experiments on a range of real-world datasets and demonstrate that our model can significantly improve the recommendation performance over the state-of-the-art baselines in terms of both recommendation accuracy and visiting orders.
Fan Zhou 0002, Xovee Xu, Wenxin Tai, Goce Trajcevski
ACM Trans. Intell. Syst. Technol.4