Xiaojun Chen 0006

dblp:20/3215-6 · DBLP profile ↗
← Back
23ranked-venue papers in the field
6as first author
7since 2021 · last 2024
—ORCID · conflict

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 8 (5 first)Information Retrieval & Web Search · 6Data Mining & Knowledge Discovery · 5 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 4
YearPublicationVenuePosition
2024 A Payment Transaction Pre-training Model for Fraud Transaction Detection
abstract
The surge in merchant fraud poses a significant threat to market order and consumer security. Effective security monitoring for merchants is crucial in safeguarding the digital life ecosystem and users' financial well-being. Detecting daily fraudulent payment transactions, a challenging task for current methods, requires efficient transformation of transactions into embeddings, especially in representing merchants based on their behavioral transactions. To address this, we propose the Grouping Sampling-based Sequence Generation (GSSG) method to generate meaningful sequences, enabling interactions among correlated transactions. We introduce Hierarchical Embedding Learning (HEL) and Hierarchical Masking pre-training (HMP) for the effective representation of hierarchical structures within flat transaction sequences. Pretrained on WeChat Pay data, our model, PTP, demonstrates superior performance in downstream fraud transaction detection, especially in few-shot learning scenarios, showcasing great potential in payment transaction scenarios.
Wenxi Huang, Zhangyi Zhao, Xiaojun Chen 0006, Qin Zhang 0011, Mark Junjie Li, Hanjing Su, Qingyao Wu
CIKM3
2024 Open-world structured sequence learning via dense target encoding
Qin Zhang 0011, Qincai Li, Haolong Xiang, Zhizhi Yu, Junyang Chen 0001, Peng Zhang 0001, Xiaojun Chen 0006
Inf. Sci.8
2023 Joint reasoning with knowledge subgraphs for Multiple Choice Question Answering
Qin Zhang 0011, Shangsi Chen, Xiaojun Chen 0006
Inf. Process. Manag.4
2022 A Dynamic Variational Framework for Open-World Node Classification in Structured Sequences
abstract
Structured sequences are a popular data representation, used to model complex data such as traffic networks. A key machine learning task for structured sequences is node classification, that is predicting the class labels of unlabeled nodes. Though many node classification models were proposed, they assume a closed world setting, that all class labels appear in the training data. But in the real-world, the presence of never-before-seen class labels in testing data can considerably degrade a classifier’s accuracy. A promising solution to this issue is to build classifiers for an open-world setting, where samples with unknown class labels are continuously observed such that training and testing data may have different class label spaces. Several approaches have been proposed for open-world learning problems in computer vision and natural language processing, but they cannot be applied directly to structured sequences due to the complexity of their non-Euclidean properties and their dynamic nature. This paper addresses this important research gap by proposing a novel Open-world Structured Sequence node Classification (OSSC) model, to learn from structured sequences in an open-world setting. OSSC captures the structural and temporal information via a GCN-based dynamic variational framework. A latent distribution sequence is learned for each node using both stochastic states and deterministic states, to capture the evolution of node attributes and topology, followed by a sampling process to generate node representations. An open-world classification loss is further adopted to ensure that node representations are sensitive to unknown classes. And a combination of Openmax and Softmax is utilized to recognize nodes from unknown classes and to classify others to one of the known classes. Experiments on real-world datasets show that the proposed OSSC method is capable of learning accurate open-world node classifiers from structured sequence data.
Qin Zhang 0011, Qincai Li, Xiaojun Chen 0006, Peng Zhang 0001, Shirui Pan, Philippe Fournier-Viger, Joshua Zhexue Huang
ICDM3
2021 Deep Self-Adaptive Hashing for Image Retrieval
abstract
Hashing technology has been widely used in image retrieval due to its computational and storage efficiency. Recently, deep unsupervised hashing methods have attracted increasing attention due to the high cost of human annotations in the real world and the superiority of deep learning technology. However, most deep unsupervised hashing methods usually pre-compute a similarity matrix to model the pairwise relationship in the pre-trained feature space. Then this similarity matrix would be used to guide hash learning, in which most of the data pairs are treated equivalently. The above process is confronted with the following defects:1) The pre-computed similarity matrix is inalterable and disconnected from the hash learning process, which cannot explore the underlying semantic information. 2) The informative data pairs may be buried by the large number of less-informative data pairs. To solve the aforementioned problems, we propose a Deep Self-Adaptive Hashing(DSAH) model to adaptively capture the semantic information with two special designs: Adaptive Neighbor Discovery(AND) and Pairwise Information Content(PIC). Firstly, we adopt the AND to initially construct a neighborhood-based similarity matrix, and then refine this initial similarity matrix with a novel update strategy to further investigate the semantic structure behind the learned representation. Secondly, we measure the priorities of data pairs with PIC and assign adaptive weights to them, which is relies on the assumption that more dissimilar data pairs contain more discriminative information for hash learning. Extensive experiments on several datasets demonstrate that the above two technologies facilitate the deep hashing model to achieve superior performance.
Qinghong Lin, Xiaojun Chen 0006, Qin Zhang 0011, Shangxuan Tian
CIKM2
2021 Learning unsupervised node representation from multi-view network
Chen Wang 0032, Xiaojun Chen 0006, Bingkun Chen, Feiping Nie 0001, Zhong Ming 0001
Inf. Sci.2
2021 Adaptive discriminant analysis for semi-supervised feature selection
Weichan Zhong, Xiaojun Chen 0006, Feiping Nie 0001, Joshua Zhexue Huang
Inf. Sci.2
2020 Multi-source Domain Adaptation for Sentiment Classification with Granger Causal Inference
abstract
In this paper, we propose a multi-source domain adaptation method with a Granger-causal objective (MDA-GC) for cross-domain sentiment classification. Specifically, for each source domain, we build an expert model by using a novel sentiment-guided capsule network, which captures the domain invariant knowledge that bridges the knowledge gap between the source and target domains. Then, an attention mechanism is devised to assign importance weights to a mixture of experts, each of which specializes in a different source domain. In addition, we propose a Granger causal objective to make the weights assigned to individual experts correlate strongly with their contributions to the decision at hand. Experimental results on a benchmark dataset demonstrate that the proposed MDA-GC model significantly outperforms the compared methods.
Min Yang 0007, Ying Shen 0001, Xiaojun Chen 0006, Chengming Li 0004
SIGIR3
2020 Semi-Supervised Feature Selection via Sparse Rescaled Linear Square Regression
abstract
With the rapid increase of the data size, it has increasing demands for selecting features by exploiting both labeled and unlabeled data. In this paper, we propose a novel semi-supervised embedded feature selection method. The new method extends the least square regression model by rescaling the regression coefficients in the least square regression with a set of scale factors, which is used for evaluating the importance of features. An iterative algorithm is proposed to optimize the new model. It has been proved that solving the new model is equivalent to solving a sparse model with a flexible and adaptable ℓ2;pnorm regularization. Moreover, the optimal solution of scale factors provides a theoretical explanation for why we can use {||w1||2, . . .,||wd||2} to evaluate the importance of features. Experimental results on eight benchmark data sets show the superior performance of the proposed method.
Xiaojun Chen 0006, Guowen Yuan, Feiping Nie 0001, Zhong Ming 0001
IEEE Trans. Knowl. Data Eng.1
2019 Structured Spectral Clustering of PurTree Data
Xiaojun Chen 0006, Yixiang Fang, Rui Mao 0001
DASFAA (2)1
2019 Exploring Communities in Large Profiled Graphs (Extended Abstract)
abstract
Given a graph G and a vertex q ∊ G, the community search (CS) problem aims to efficiently find a subgraph of G whose vertices are closely related to q. Communities are prevalent in social and biological networks, and can be used in product advertisement and social event recommendation. In this paper, we study profiled community search (PCS), where CS is performed on a profiled graph. This is a graph in which each vertex has labels arranged in a hierarchical manner. Compared with existing CS approaches, PCS can sufficiently identify vertices with semantic commonalities and thus find more high-quality diverse communities. As a naive solution for PCS is highly expensive, we have developed a tree index, which facilitates efficient and online solutions for PCS.
Yankai Chen 0001, Yixiang Fang, Reynold Cheng, Xiaojun Chen 0006, Jie Zhang 0002
ICDE5
2019 Discovering author interest evolution in order-sensitive and Semantic-aware topic modeling
Min Yang 0007, Qiang Qu 0001, Xiaojun Chen 0006, Wenting Tu, Ying Shen 0001, Jia Zhu 0003
Inf. Sci.3
2019 Exploring Communities in Large Profiled Graphs
abstract
Given a graph $G$G and a vertex $q\in G$q∈G, the community search (CS) problem aims to efficiently find a subgraph of $G$G whose vertices are closely related to $q$q. Communities are prevalent in social and biological networks, and can be used in product advertisement and social event recommendation. In this paper, we study profiled community search (PCS), where CS is performed on a profiled graph. This is a graph in which each vertex has labels arranged in a hierarchical manner. Extensive experiments show that PCS can identify communities with themes that are common to their vertices, and is more effective than existing CS approaches. As a naive solution for PCS is highly expensive, we have also developed a tree index, which facilitates efficient and online solutions for PCS.
Yankai Chen 0001, Yixiang Fang, Reynold Cheng, Xiaojun Chen 0006, Jie Zhang 0002
IEEE Trans. Knowl. Data Eng.5
2019 On Spatial-Aware Community Search
abstract
Communities are prevalent in social networks, knowledge graphs, and biological networks. Recently, the topic of community search (CS) has received plenty of attention. The CS problem aims to look for a dense subgraph that contains a query vertex. Existing CS solutions do not consider the spatial extent of a community. They can yield communities whose locations of vertices span large areas. In applications that facilitate setting social events (e.g., finding conference attendees to join a dinner), it is important to find groups of people who are physically close to each other, so it is desirable to have aspatial-aware community(or SAC), whose vertices are close structurally and spatially. Given a graph$G$and a query vertex$q$, we develop an exact solution to find the SAC containing$q$, but it cannot scale to large datasets, so we design three approximation algorithms. We further study the problem of continuous SAC search on a “dynamic spatial graph,” whose vertices’ locations change with time, and propose three fast solutions. We evaluate the solutions on both real and synthetic datasets, and the results show that SACs are better than communities returned by existing solutions. Moreover, our approximation solutions perform accurately and efficiently.
Yixiang Fang, Reynold Cheng, Xiaodong Li 0009, Siqiang Luo, Jiafeng Hu, Xiaojun Chen 0006
IEEE Trans. Knowl. Data Eng.7
2018 Spectral Clustering of Large-scale Data by Directly Solving Normalized Cut
abstract
During the past decades, many spectral clustering algorithms have been proposed. However, their high computational complexities hinder their applications on large-scale data. Moreover, most of them use a two-step approach to obtain the optimal solution, which may deviate from the solution by directly solving the original problem. In this paper, we propose a new optimization algorithm, namely Direct Normalized Cut (DNC), to directly optimize the normalized cut model. DNC has a quadratic time complexity, which is a significant reduction comparing with the cubic time complexity of the traditional spectral clustering. To cope with large-scale data, a Fast Normalized Cut (FNC) method with linear time and space complexities is proposed by extending DNC with an anchor-based strategy. In the new method, we first seek a set of anchors and then construct a representative similarity matrix by computing distances between the anchors and the whole data set. To find high quality anchors that best represent the whole data set, we propose a Balanced k-means (BKM) to partition a data set into balanced clusters and use the cluster centers as anchors. Then DNC is used to obtain the final clustering result from the representative similarity matrix. A series of experiments were conducted on both synthetic data and real-world data sets, and the experimental results show the superior performance of BKM, DNC and FNC.
Xiaojun Chen 0006, Weijun Hong, Feiping Nie 0001, Min Yang 0007, Joshua Zhexue Huang
KDD1
2018 Investigating Deep Reinforcement Learning Techniques in Personalized Dialogue Generation
abstract
In this paper, we propose a personalized dialogue generation system, which combines reinforcement learning techniques with an attention-based hierarchical recurrent encoderdecoder model. Firstly, we incorporate user-specific information into the decoder to capture user's background information and speaking style. Secondly, we employ reinforcement learning techniques to maximize future reward in dialogue, which enables our system to generate topic-coherent, informative and grammatical responses. Moreover, we propose three types of rewards to characterize good conversations. Finally, we compare the performance of the following reinforcement learning methods in dialogue generation: policy gradient, Q-learning, and actor-critic algorithms. We conduct experiments to verify the effectiveness of the proposed model on two dialogue datasets. Experimental results demonstrate that our model can generate better personalized dialogues for different users. Quantitatively, our method achieves better performance than the state-of-the-art dialogue systems in terms of BLEU score, perplexity, and human evaluation.
Min Yang 0007, Qiang Qu 0001, Kai Lei, Jia Zhu 0003, Zhou Zhao 0001, Xiaojun Chen 0006, Joshua Zhexue Huang
SDM6
2018 PurTreeClust: A Clustering Algorithm for Customer Segmentation from Massive Customer Transaction Data
abstract
Clustering of customer transaction data is an important procedure to analyze customer behaviors in retail and e-commerce companies. Note that products from companies are often organized as a product tree, in which the leaf nodes are goods to sell, and the internal nodes (except root node) could be multiple product categories. Based on this tree, we propose the “personalized product tree”, named purchase tree, to represent a customer's transaction records. So the customers' transaction data set can be compressed into a set of purchase trees. We propose a partitional clustering algorithm, named PurTreeClust, for fast clustering of purchase trees. A new distance metric is proposed to effectively compute the distance between two purchase trees. To cluster the purchase tree data, we first rank the purchase trees as candidate representative trees with a novel separate density, and then select the top k customers as the representatives of k customer groups. Finally, the clustering results are obtained by assigning each customer to the nearest representative. We also propose a gap statistic based method to evaluate the number of clusters. A series of experiments were conducted on ten real-life transaction data sets, and experimental results show the superior performance of the proposed method.
Xiaojun Chen 0006, Yixiang Fang, Min Yang 0007, Feiping Nie 0001, Zhou Zhao 0001, Joshua Zhexue Huang
IEEE Trans. Knowl. Data Eng.1
2017 Personalized Response Generation via Domain adaptation
abstract
In this paper, we propose a novel personalized response generation model via domain adaptation (PRG-DM). First, we learn the human responding style from large general data (without user-specific information). Second, we fine tune the model on a small size of personalized data to generate personalized responses with a dual learning mechanism. Moreover, we propose three new rewards to characterize good conversations that are personalized, informative and grammatical. We employ the policy gradient method to generate highly rewarded responses. Experimental results show that our model can generate better personalized responses for different users.
Min Yang 0007, Zhou Zhao 0001, Wei Zhao 0033, Xiaojun Chen 0006, Jia Zhu 0003, Lianqiang Zhou, Zigang Cao
SIGIR4
2016 PurTreeClust: A purchase tree clustering algorithm for large-scale customer transaction data
abstract
Clustering of customer transaction data is usually an important procedure to analyze customer behaviors in retail and e-commerce companies. Note that products from companies are often organized as a product tree, in which the leaf nodes are goods to sell, and the internal nodes (except root node) could be multiple product categories. Based on this tree, we present to use a “personalized product tree”, called purchase tree, to represent a customer's transaction data. The customer transaction data set can be represented as a set of purchase trees. We propose a PurTreeClust algorithm for clustering of large-scale customers from purchase trees. We define a new distance metric to effectively compute the distance between two purchase trees from the entire levels in the tree. A cover tree is then built for indexing the purchase tree data and we propose a leveled density estimation method for selecting initial cluster centers from a cover tree. PurTreeClust, a fast clustering method for clustering of large-scale purchase trees, is then presented. Last, we propose a gap statistic based method for estimating the number of clusters from the purchase tree clustering results. A series of experiments were conducted on ten large-scale transaction data sets which contain up to four million transaction records, and experimental results have verified the effectiveness and efficiency of the proposed method. We also compared our method with three clustering algorithms, e.g., spectral clustering, hierarchical agglomerative clustering and DBSCAN. The experimental results have demonstrated the superior performance of the proposed method.
Xiaojun Chen 0006, Joshua Zhexue Huang
ICDE1
2013 TW-k-Means: Automated Two-Level Variable Weighting Clustering Algorithm for Multiview Data
abstract
This paper proposes TW-k-means, an automated two-level variable weighting clustering algorithm for multiview data, which can simultaneously compute weights for views and individual variables. In this algorithm, a view weight is assigned to each view to identify the compactness of the view and a variable weight is also assigned to each variable in the view to identify the importance of the variable. Both view weights and variable weights are used in the distance function to determine the clusters of objects. In the new algorithm, two additional steps are added to the iterative k-means clustering process to automatically compute the view weights and the variable weights. We used two real-life data sets to investigate the properties of two types of weights in TW-k-means and investigated the difference between the weights of TW-k-means and the weights of the individual variable weighting method. The experiments have revealed the convergence property of the view weights in TW-k-means. We compared TW-k-means with five clustering algorithms on three real-life data sets and the results have shown that the TW-k-means algorithm significantly outperformed the other five clustering algorithms in four evaluation indices.
Xiaojun Chen 0006, Xiaofei Xu 0001, Joshua Zhexue Huang, Yunming Ye
IEEE Trans. Knowl. Data Eng.1
2012 Scalable Subspace Logistic Regression Models for High Dimensional Data
Xiaojun Chen 0006, Joshua Zhexue Huang, Shengzhong Feng
APWeb2
2012 Scalable Random Forests for Massive Data
Bingguo Li, Xiaojun Chen 0006, Mark Junjie Li, Joshua Zhexue Huang, Shengzhong Feng
PAKDD (1)2
2006 Neighborhood Density Method for Selecting Initial Cluster Centers in K-Means Clustering
Yunming Ye, Joshua Zhexue Huang, Xiaojun Chen 0006, Shuigeng Zhou, Graham J. Williams, Xiaofei Xu 0001
PAKDD3