VLDB 2026 Research / reviewers in the wild / expert
Xiaojun Chen 0006
dblp:20/3215-6
· DBLP profile ↗
87ranked-venue papers
20as first author
36since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 57 · 12 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 5 first-author · 10 since 2021Databases, data management, data science and information retrieval · 23 · 6 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TGLsta: Low-resource Textual Graph Learning with Semantic and Topological Awareness via LLMsabstractTextual Graphs (TGs) present a graph-based representation of textual data and find wide applications in real-world scenarios, such as citation networks, knowledge graphs, and social networks. While the traditional "pre-train, fine-tune" framework effectively addresses tasks requiring abundant labeled data, it falls short in scenarios with limited resource or zero-shot learning capabilities, particularly in low-resource textual graph node classification. Additionally, prevalent approaches that convert text nodes into shallow or manually engineered features fail to capture the rich semantic nuances within the text. The conventional methods often neglect the fusion of semantic and topological information, resulting in suboptimal model learning. To overcome these challenges, we proposed a novel method of low-resource textual graph node classification based on large language models, i.e., Textual graph learning with semantic and topological awareness (TGLsta), which comprehensively explores the semantic information, near neighborhood information, and the topology information in textual graphs, where these components are the most important information source contained in textual graphs. Graph prompt tuning for both zero- and few-shot textual graph node classification is further introduced. Qin Zhang 0011, Xiaochen Fan, Xiaojun Chen 0006, Shirui Pan |
AAAI | 5 |
| 2025 | Hard400: A Bilingual Code Generation Evaluation Benchmark for Large Language Models
Qin Zhang 0011, Hong Zhou 0004, Xiaojun Chen 0006, Han Liu 0002 |
ICIC (18) | 5 |
| 2025 | MRBench: A Multi-Image Reasoning Benchmark with Adaptive Knowledge RetrievalabstractMulti-image understanding is crucial in real-world applications such as social media analysis and news reporting. However, existing benchmarks fall short in evaluating models' ability to integrate external knowledge and perform cross-image reasoning. To address this gap, we introduce MRBench, a comprehensive benchmark designed to assess knowledge-based reasoning across 12 diverse domains, incorporating four types of image relations: visually similar, identical entities, attribute-associated, and independent images. Additionally, we propose Multimodal Adaptive Retrieval Reasoning (MARR), a novel framework that enables the analysis of relationships among multiple input images and adaptively determines when to terminate the retrieval process. Extensive evaluations of state-of-the-art multimodal large language models (MLLMs) show a notable gap between model and human performance. The best-performing model, Gemini 2.0, reaches 56.86% accuracy, still 20.24% below humans. Proprietary models generally surpass open-source ones, particularly on visually similar and same-entity tasks, underscoring current limits in multi-image reasoning and retrieval and positioning MRBench as a key diagnostic tool. Our benchmark is available for further research and development in this field. https://github.com/Bruce-XJChen/MRBench. Wenxi Huang, Xiaojun Chen 0006, Qin Zhang 0011, Ting Wan, Liangjie Zhang |
ACM Multimedia | 2 |
| 2025 | A Structured Bipartite Graph Learning method for ensemble clustering
Zitong Zhang 0003, Xiaojun Chen 0006, Chen Wang 0032, Ruili Wang 0001, Feiping Nie 0001 |
Pattern Recognit. | 2 |
| 2025 | Structured multi-view k-means clustering
Zitong Zhang 0003, Xiaojun Chen 0006, Chen Wang 0032, Ruili Wang 0001, Feiping Nie 0001 |
Pattern Recognit. | 2 |
| 2024 | Multi-Level Cross-Modal Alignment for Image ClusteringabstractRecently, the cross-modal pretraining model has been employed to produce meaningful pseudo-labels to supervise the training of an image clustering model. However, numerous erroneous alignments in a cross-modal pretraining model could produce poor-quality pseudo labels and degrade clustering performance. To solve the aforementioned issue, we propose a novel Multi-level Cross-modal Alignment method to improve the alignments in a cross-modal pretraining model for downstream tasks, by building a smaller but better semantic space and aligning the images and texts in three levels, i.e., instance-level, prototype-level, and semantic-level. Theoretical results show that our proposed method converges, and suggests effective means to reduce the expected clustering risk of our method. Experimental results on five benchmark datasets clearly show the superiority of our new method. Liping Qiu, Qin Zhang 0011, Xiaojun Chen 0006, Shaotian Cai |
AAAI | 3 |
| 2024 | ROG_PL: Robust Open-Set Graph Learning via Region-Based Prototype LearningabstractOpen-set graph learning is a practical task that aims to classify the known class nodes and to identify unknown class samples as unknowns. Conventional node classification methods usually perform unsatisfactorily in open-set scenarios due to the complex data they encounter, such as out-of-distribution (OOD) data and in-distribution (IND) noise. OOD data are samples that do not belong to any known classes. They are outliers if they occur in training (OOD noise), and open-set samples if they occur in testing. IND noise are training samples which are assigned incorrect labels. The existence of IND noise and OOD noise is prevalent, which usually cause the ambiguity problem, including the intra-class variety problem and the inter-class confusion problem. Thus, to explore robust open-set learning methods is necessary and difficult, and it becomes even more difficult for non-IID graph data. To this end, we propose a unified framework named ROG_PL to achieve robust open-set learning on complex noisy graph data, by introducing prototype learning. In specific, ROG_PL consists of two modules, i.e., denoising via label propagation and open-set prototype learning via regions. The first module corrects noisy labels through similarity-based label propagation and removes low-confidence samples, to solve the intra-class variety problem caused by noise. The second module learns open-set prototypes for each known class via non-overlapped regions and remains both interior and border prototypes to remedy the inter-class confusion problem. The two modules are iteratively updated under the constraints of classification loss and prototype diversity loss. To the best of our knowledge, the proposed ROG_PL is the first robust open-set node classification method for graph data with complex noise. Experimental evaluations of ROG_PL on several benchmark graph datasets demonstrate that it has good performance. Qin Zhang 0011, Jiexin Lu, Liping Qiu, Shirui Pan, Xiaojun Chen 0006, Junyang Chen 0001 |
AAAI | 6 |
| 2024 | Enhancing Noise Robustness of Retrieval-Augmented Language Models with Adaptive Adversarial TrainingabstractLarge Language Models (LLMs) exhibit substantial capabilities yet encounter challenges, including hallucination, outdated knowledge, and untraceable reasoning processes.Retrievalaugmented generation (RAG) has emerged as a promising solution, integrating knowledge from external databases to mitigate these challenges.However, inappropriate retrieved passages can potentially hinder the LLMs' capacity to generate comprehensive and high-quality responses.Prior RAG studies on the robustness of retrieval noises often confine themselves to a limited set of noise types, deviating from realworld retrieval environments and limiting practical applicability.In this study, we initially investigate retrieval noises and categorize them into three distinct types, reflecting real-world environments.We analyze the impact of these various retrieval noises on the robustness of LLMs.Subsequently, we propose a novel RAG approach known as Retrieval-augmented Adaptive Adversarial Training (RAAT).RAAT leverages adaptive adversarial training to dynamically adjust the model's training process in response to retrieval noises.Concurrently, it employs multi-task learning to ensure the model's capacity to internally recognize noisy contexts.Extensive experiments demonstrate that the LLaMA-2 7B model trained using RAAT exhibits significant improvements in F1 and EM scores under diverse noise conditions.For reproducibility, we release our code and data at: https://github.com/calubkk/RAAT. Feiteng Fang, Yuelin Bai, Shiwen Ni, Min Yang 0007, Xiaojun Chen 0006, Ruifeng Xu 0001 |
ACL (1) | 5 |
| 2024 | A Payment Transaction Pre-training Model for Fraud Transaction DetectionabstractThe surge in merchant fraud poses a significant threat to market order and consumer security. Effective security monitoring for merchants is crucial in safeguarding the digital life ecosystem and users' financial well-being. Detecting daily fraudulent payment transactions, a challenging task for current methods, requires efficient transformation of transactions into embeddings, especially in representing merchants based on their behavioral transactions. To address this, we propose the Grouping Sampling-based Sequence Generation (GSSG) method to generate meaningful sequences, enabling interactions among correlated transactions. We introduce Hierarchical Embedding Learning (HEL) and Hierarchical Masking pre-training (HMP) for the effective representation of hierarchical structures within flat transaction sequences. Pretrained on WeChat Pay data, our model, PTP, demonstrates superior performance in downstream fraud transaction detection, especially in few-shot learning scenarios, showcasing great potential in payment transaction scenarios. Wenxi Huang, Zhangyi Zhao, Xiaojun Chen 0006, Qin Zhang 0011, Mark Junjie Li, Hanjing Su, Qingyao Wu |
CIKM | 3 |
| 2024 | MoZIP: A Multilingual Benchmark to Evaluate Large Language Models in Intellectual PropertyabstractLarge language models (LLMs) have demonstrated impressive performance in various natural language processing (NLP) tasks. However, there is limited understanding of how well LLMs perform in specific domains (e.g, the intellectual property (IP) domain). In this paper, we contribute a new benchmark, the first Multilingual-oriented quiZ on Intellectual Property (MoZIP), for the evaluation of LLMs in the IP domain. The MoZIP benchmark includes three challenging tasks: IP multiple-choice quiz (IPQuiz), IP question answering (IPQA), and patent matching (PatentMatch). In addition, we also develop a new IP-oriented multilingual large language model (called MoZi), which is a BLOOMZ-based model that has been supervised fine-tuned with multilingual IP-related text data. We evaluate our proposed MoZi model and four well-known LLMs (i.e., BLOOMZ, BELLE, ChatGLM and ChatGPT) on the MoZIP benchmark. Experimental results demonstrate that MoZi outperforms BLOOMZ, BELLE and ChatGLM by a noticeable margin, while it had lower scores compared with ChatGPT. Notably, the performance of current LLMs on the MoZIP benchmark has much room for improvement, and even the most powerful ChatGPT does not reach the passing level. Our source code, data, and models are available at https://github.com/AI-for-Science/MoZi. Shiwen Ni, Minghuan Tan, Yuelin Bai, Fuqiang Niu, Min Yang 0007, Bowen Zhang 0005, Ruifeng Xu 0001, Xiaojun Chen 0006, Chengming Li 0004, Xiping Hu |
LREC/COLING | 8 |
| 2024 | Unsupervised Multiple Choices Question Answering Via Universal CorpusabstractUnsupervised question answering is a promising yet challenging task, which alleviates the burden of building large-scale annotated data in a new domain. It motivates us to study the unsupervised multiple-choice question answering (MCQA) problem. In this paper, we propose a novel framework designed to generate synthetic MCQA data barely based on contexts from the universal domain without relying on any form of manual annotation. Possible answers are extracted and used to produce related questions, then we leverage both named entities (NE) and knowledge graphs to discover plausible distractors to form complete synthetic samples. Experiments on multiple MCQA datasets demonstrate the effectiveness of our method. Qin Zhang 0011, Xiaojun Chen 0006 |
ICASSP | 3 |
| 2024 | Representation Surgery for Multi-Task Model MergingabstractMulti-task learning (MTL) compresses the information from multiple tasks into a unified backbone to improve computational efficiency and generalization. Recent work directly merges multiple independently trained models to perform MTL instead of collecting their raw data for joint training, greatly expanding the application scenarios of MTL. However, by visualizing the representation distribution of existing model merging schemes, we find that the merged model often suffers from the dilemma of representation bias. That is, there is a significant discrepancy in the representation distribution between the merged and individual models, resulting in poor performance of merged MTL. In this paper, we propose a representation surgery solution called ``Surgery" to reduce representation bias in the merged model. Specifically, Surgery is a lightweight task-specific plugin that takes the representation of the merged model as input and attempts to output the biases contained in the representation from the merged model. We then designed an unsupervised optimization objective that updates the Surgery plugin by minimizing the distance between the merged model's representation and the individual model's representation. Extensive experiments demonstrate significant MTL performance improvements when our Surgery plugin is applied to state-of-the-art (SOTA) model merging schemes. Enneng Yang, Li Shen 0008, Zhenyi Wang 0001, Guibing Guo, Xiaojun Chen 0006, Xingwei Wang 0001, Dacheng Tao |
ICML | 5 |
| 2024 | ReCoS: A Novel Benchmark for Cross-Modal Image-Text Retrieval in Complex Real-Life ScenariosabstractImage-text retrieval stands as a pivotal task within information retrieval, gaining increasing importance with the rapid advancements in Visual-Language Pretraining models. However, current benchmarks for evaluating these models face limitations, exemplified by instances such as BLIP2 achieving near-perfect performance on existing benchmarks. In response, this paper advocates for a more robust evaluation benchmark for image-text retrieval, one that embraces several essential characteristics. Firstly, a comprehensive benchmark should cover a diverse range of tasks in both perception and cognition-based retrieval. Recognizing this need, we introduce ReCoS, a novel benchmark specifically designed for cross-modal image-text retrieval in complex real-life scenarios. Unlike existing benchmarks, ReCoS encompasses 12 retrieval tasks, with a particular focus on three cognition-based tasks, providing a more holistic assessment of model capabilities. To ensure the novelty of the benchmark, we emphasize the use of original data sources, steering clear of reliance on existing publicly available datasets to minimize the risk of data leakage. Additionally, to strike a balance between the complexity of the real world and benchmark usability, ReCoS includes text descriptions that are neither overly detailed, making retrieval overly simplistic, nor under-detailed to the point where retrieval becomes impossible. Our evaluation results shed light on the challenges faced by existing methods, especially in cognition-based retrieval tasks within ReCoS. This underscores the necessity for innovative approaches in addressing the complexities of image-text retrieval in real-world scenarios. Our code and benchmark datasets are available for further research and development in this field https://github.com/Bruce-XJChen/ReCos Xiaojun Chen 0006, Jimeng Lou, Wenxi Huang, Ting Wan, Qin Zhang 0011, Min Yang 0007 |
ACM Multimedia | 1 |
| 2024 | II-Bench: An Image Implication Understanding Benchmark for Multimodal Large Language ModelsabstractThe rapid advancements in the development of multimodal large language models (MLLMs) have consistently led to new breakthroughs on various benchmarks. In response, numerous challenging and comprehensive benchmarks have been proposed to more accurately assess the capabilities of MLLMs. However, there is a dearth of exploration of the higher-order perceptual capabilities of MLLMs. To fill this gap, we propose the Image Implication understanding Benchmark, II-Bench, which aims to evaluate the model's higher-order perception of images. Through extensive experiments on II-Bench across multiple MLLMs, we have made significant findings. Initially, a substantial gap is observed between the performance of MLLMs and humans on II-Bench. The pinnacle accuracy of MLLMs attains 74.8%, whereas human accuracy averages 90%, peaking at an impressive 98%. Subsequently, MLLMs perform worse on abstract and complex images, suggesting limitations in their ability to understand high-level semantics and capture image details. Finally, it is observed that most models exhibit enhanced accuracy when image sentiment polarity hints are incorporated into the prompts. This observation underscores a notable deficiency in their inherent understanding of image sentiment. We believe that II-Bench will inspire the community to develop the next generation of MLLMs, advancing the journey towards expert artificial general intelligence (AGI). II-Bench is publicly available at https://huggingface.co/datasets/m-a-p/II-Bench. Feiteng Fang, Xeron Du, Chenhao Zhang 0005, Noah Wang, Yuelin Bai, Qixuan Zhao, Liyang Fan, Chengguang Gan, Hongquan Lin, Jiaming Li 0004, Yuansheng Ni, Haihong Wu, Yaswanth Narsupalli, Zhigang Zheng, Chengming Li 0004, Xiping Hu, Ruifeng Xu 0001, Xiaojun Chen 0006, Min Yang 0007, Ruibo Liu, Wenhao Huang 0001, Ge Zhang 0009, Shiwen Ni |
NeurIPS | 20 |
| 2024 | EGonc : Energy-based Open-Set Node Classification with substitute UnknownsabstractOpen-set Classification (OSC) is a critical requirement for safely deploying machine learning models in the open world, which aims to classify samples from known classes and reject samples from out-of-distribution (OOD).
Existing methods exploit the feature space of trained network and attempt at estimating the uncertainty in the predictions.
However, softmax-based neural networks are found to be overly confident in their predictions even on data they have never seen before and
the immense diversity of the OOD examples also makes such methods fragile.
To this end, we follow the idea of estimating the underlying density of the training data to decide whether a given input is close to the in-distribution (IND) data and adopt Energy-based models (EBMs) as density estimators.
A novel energy-based generative open-set node classification method, \textit{EGonc}, is proposed to achieve open-set graph learning.
Specifically, we generate substitute unknowns to mimic the distribution of real open-set samples firstly, based on the information of graph structures.
Then, an additional energy logit representing the virtual OOD class is learned from the residual of the feature against the principal space, and matched with the original logits by a constant scaling. This virtual logit serves as the indicator of OOD-ness.
EGonc has nice theoretical properties that guarantee an overall distinguishable margin between the detection scores for IND and OOD samples.
Comprehensive experimental evaluations of EGonc also demonstrate its superiority. Qin Zhang 0011, Zelin Shi, Shirui Pan, Junyang Chen 0001, Huisi Wu, Xiaojun Chen 0006 |
NeurIPS | 6 |
| 2024 | Partial order relation-based gene ontology embedding improves protein function predictionabstractProtein annotation has long been a challenging task in computational biology. Gene Ontology (GO) has become one of the most popular frameworks to describe protein functions and their relationships. Prediction of a protein annotation with proper GO terms demands high-quality GO term representation learning, which aims to learn a low-dimensional dense vector representation with accompanying semantic meaning for each functional label, also known as embedding. However, existing GO term embedding methods, which mainly take into account ancestral co-occurrence information, have yet to capture the full topological information in the GO-directed acyclic graph (DAG). In this study, we propose a novel GO term representation learning method, PO2Vec, to utilize the partial order relationships to improve the GO term representations. Extensive evaluations show that PO2Vec achieves better outcomes than existing embedding methods in a variety of downstream biological tasks. Based on PO2Vec, we further developed a new protein function prediction method PO2GO, which demonstrates superior performance measured in multiple metrics and annotation specificity as well as few-shot prediction capability in the benchmarks. These results suggest that the high-quality representation of GO structure is critical for diverse biological tasks including computational protein annotation. Bin Wang 0002, Yan Kou, Xiaojun Chen 0006, Yi Pan 0001, Shuangwei Hu, Zhenjiang Zech Xu |
Briefings Bioinform. | 5 |
| 2024 | Open-world structured sequence learning via dense target encoding
Qin Zhang 0011, Qincai Li, Haolong Xiang, Zhizhi Yu, Junyang Chen 0001, Peng Zhang 0001, Xiaojun Chen 0006 |
Inf. Sci. | 8 |
| 2024 | A fine-grained self-adapting prompt learning approach for few-shot learning with pre-trained language models
Xiaojun Chen 0006, Philippe Fournier-Viger, Bowen Zhang 0005, Guodong Long, Qin Zhang 0011 |
Knowl. Based Syst. | 1 |
| 2023 | Semantic-Enhanced Image ClusteringabstractImage clustering is an important and open challenging task in computer vision. Although many methods have been proposed to solve the image clustering task, they only explore images and uncover clusters according to the image features, thus being unable to distinguish visually similar but semantically different images. In this paper, we propose to investigate the task of image clustering with the help of visual-language pre-training model. Different from the zero-shot setting, in which the class names are known, we only know the number of clusters in this setting. Therefore, how to map images to a proper semantic space and how to cluster images from both image and semantic spaces are two key problems. To solve the above problems, we propose a novel image clustering method guided by the visual-language pre-training model CLIP, named Semantic-Enhanced Image Clustering (SIC). In this new method, we propose a method to map the given images to a proper semantic space first and efficient methods to generate pseudo-labels according to the relationships between images and semantics. Finally, we propose to perform clustering with consistency learning in both image space and semantic space, in a self-supervised learning fashion. The theoretical result of convergence analysis shows that our proposed method can converge at a sublinear speed. Theoretical analysis of expectation risk also shows that we can reduce the expectation risk by improving neighborhood consistency, increasing prediction confidence, or reducing neighborhood imbalance. Experimental results on five benchmark datasets clearly show the superiority of our new method. Shaotian Cai, Liping Qiu, Xiaojun Chen 0006, Qin Zhang 0011, Longteng Chen |
AAAI | 3 |
| 2023 | A Survey for Efficient Open Domain Question AnsweringabstractQin Zhang, Shangsi Chen, Dongkuan Xu, Qingqing Cao, Xiaojun Chen, Trevor Cohn, Meng Fang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Qin Zhang 0011, Shangsi Chen, Dongkuan Xu, Xiaojun Chen 0006, Trevor Cohn |
ACL (1) | 5 |
| 2023 | G2Pxy: Generative Open-Set Node Classification on Graphs with Proxy UnknownsabstractNode classification is the task of predicting the labels of unlabeled nodes in a graph. State-of-the-art methods based on graph neural networks achieve excellent performance when all labels are available during training. But in real-life, models are of ten applied on data with new classes, which can lead to massive misclassification and thus significantly degrade performance. Hence, developing open-set classification methods is crucial to determine if a given sample belongs to a known class. Existing methods for open-set node classification generally use transductive learning with part or all of the features of real unseen class nodes to help with open-set classification. In this paper, we propose a novel generative open-set node classification method, i.e., G2Pxy, which follows a stricter inductive learning setting where no information about unknown classes is available during training and validation. Two kinds of proxy unknown nodes, inter-class unknown proxies and external unknown proxies are generated via mixup to efficiently anticipate the distribution of novel classes. Using the generated proxies, a closed-set classifier can be transformed into an open-set one, by augmenting it with an extra proxy classifier. Under the constraints of both cross entropy loss and complement entropy loss, G2Pxy achieves superior effectiveness for unknown class detection and known class classification, which is validated by experiments on bench mark graph datasets. Moreover, G2Pxy does not have specific requirement on the GNN architecture and shows good generalizations. Qin Zhang 0011, Zelin Shi, Xiaojun Chen 0006, Philippe Fournier-Viger, Shirui Pan |
IJCAI | 4 |
| 2023 | Joint reasoning with knowledge subgraphs for Multiple Choice Question Answering
Qin Zhang 0011, Shangsi Chen, Xiaojun Chen 0006 |
Inf. Process. Manag. | 4 |
| 2022 | Deep Unsupervised Hashing with Latent Semantic ComponentsabstractDeep unsupervised hashing has been appreciated in the regime of image retrieval. However, most prior arts failed to detect the semantic components and their relationships behind the images, which makes them lack discriminative power. To make up the defect, we propose a novel Deep Semantic Components Hashing (DSCH), which involves a common sense that an image normally contains a bunch of semantic components with homology and co-occurrence relationships. Based on this prior, DSCH regards the semantic components as latent variables under the Expectation-Maximization framework and designs a two-step iterative algorithm with the objective of maximum likelihood of training data. Firstly, DSCH constructs a semantic component structure by uncovering the fine-grained semantics components of images with a Gaussian Mixture Modal~(GMM), where an image is represented as a mixture of multiple components, and the semantics co-occurrence are exploited. Besides, coarse-grained semantics components, are discovered by considering the homology relationships between fine-grained components, and the hierarchy organization is then constructed. Secondly, DSCH makes the images close to their semantic component centers at both fine-grained and coarse-grained levels, and also makes the images share similar semantic components close to each other. Extensive experiments on three benchmark datasets demonstrate that the proposed hierarchical semantic components indeed facilitate the hashing model to achieve superior performance. Qinghong Lin, Xiaojun Chen 0006, Qin Zhang 0011, Shaotian Cai, Hongfa Wang |
AAAI | 2 |
| 2022 | A Dynamic Variational Framework for Open-World Node Classification in Structured SequencesabstractStructured sequences are a popular data representation, used to model complex data such as traffic networks. A key machine learning task for structured sequences is node classification, that is predicting the class labels of unlabeled nodes. Though many node classification models were proposed, they assume a closed world setting, that all class labels appear in the training data. But in the real-world, the presence of never-before-seen class labels in testing data can considerably degrade a classifier’s accuracy. A promising solution to this issue is to build classifiers for an open-world setting, where samples with unknown class labels are continuously observed such that training and testing data may have different class label spaces. Several approaches have been proposed for open-world learning problems in computer vision and natural language processing, but they cannot be applied directly to structured sequences due to the complexity of their non-Euclidean properties and their dynamic nature. This paper addresses this important research gap by proposing a novel Open-world Structured Sequence node Classification (OSSC) model, to learn from structured sequences in an open-world setting. OSSC captures the structural and temporal information via a GCN-based dynamic variational framework. A latent distribution sequence is learned for each node using both stochastic states and deterministic states, to capture the evolution of node attributes and topology, followed by a sampling process to generate node representations. An open-world classification loss is further adopted to ensure that node representations are sensitive to unknown classes. And a combination of Openmax and Softmax is utilized to recognize nodes from unknown classes and to classify others to one of the known classes. Experiments on real-world datasets show that the proposed OSSC method is capable of learning accurate open-world node classifiers from structured sequence data. Qin Zhang 0011, Qincai Li, Xiaojun Chen 0006, Peng Zhang 0001, Shirui Pan, Philippe Fournier-Viger, Joshua Zhexue Huang |
ICDM | 3 |
| 2022 | Logic tensor network with massive learned knowledge for aspect-based sentiment analysis
Hu Huang 0009, Bowen Zhang 0005, Liwen Jing 0001, Xianghua Fu, Xiaojun Chen 0006, Jianyang Shi |
Knowl. Based Syst. | 5 |
| 2022 | Directly solving normalized cut for multi-view data
Chen Wang 0032, Xiaojun Chen 0006, Feiping Nie 0001, Joshua Zhexue Huang |
Pattern Recognit. | 2 |
| 2022 | Semisupervised Feature Selection via Structured Manifold LearningabstractRecently, semisupervised feature selection has gained more attention in many real applications due to the high cost of obtaining labeled data. However, existing methods cannot solve the "multimodality" problem that samples in some classes lie in several separate clusters. To solve the multimodality problem, this article proposes a new feature selection method for semisupervised task, namely, semisupervised structured manifold learning (SSML). The new method learns a new structured graph which consists of more clusters than the known classes. Meanwhile, we propose to exploit the submanifold in both labeled data and unlabeled data by consuming the nearest neighbors of each object in both labeled and unlabeled objects. An iterative optimization algorithm is proposed to solve the new model. A series of experiments was conducted on both synthetic and real-world datasets and the experimental results verify the ability of the new method to solve the multimodality problem and its superior performance compared with the state-of-the-art methods. Xiaojun Chen 0006, Renjie Chen 0004, Qingyao Wu, Feiping Nie 0001, Min Yang 0007, Rui Mao 0001 |
IEEE Trans. Cybern. | 1 |
| 2022 | Semisupervised Feature Selection With Sparse Discriminative Least Squares RegressionabstractIn big data time, selecting informative features has become an urgent need. However, due to the huge cost of obtaining enough labeled data for supervised tasks, researchers have turned their attention to semisupervised learning, which exploits both labeled and unlabeled data. In this article, we propose a sparse discriminative semisupervised feature selection (SDSSFS) method. In this method, the$\epsilon $-dragging technique for the supervised task is extended to the semisupervised task, which is used to enlarge the distance between classes in order to obtain a discriminative solution. The flexible$\ell _{2,p}$norm is implicitly used as regularization in the new model. Therefore, we can obtain a more sparse solution by setting smaller$p$. An iterative method is proposed to simultaneously learn the regression coefficients and$\epsilon $-dragging matrix and predicting the unknown class labels. Experimental results on ten real-world datasets show the superiority of our proposed method. Chen Wang 0032, Xiaojun Chen 0006, Guowen Yuan, Feiping Nie 0001, Min Yang 0007 |
IEEE Trans. Cybern. | 2 |
| 2021 | Deep Self-Adaptive Hashing for Image RetrievalabstractHashing technology has been widely used in image retrieval due to its computational and storage efficiency. Recently, deep unsupervised hashing methods have attracted increasing attention due to the high cost of human annotations in the real world and the superiority of deep learning technology. However, most deep unsupervised hashing methods usually pre-compute a similarity matrix to model the pairwise relationship in the pre-trained feature space. Then this similarity matrix would be used to guide hash learning, in which most of the data pairs are treated equivalently. The above process is confronted with the following defects:1) The pre-computed similarity matrix is inalterable and disconnected from the hash learning process, which cannot explore the underlying semantic information. 2) The informative data pairs may be buried by the large number of less-informative data pairs. To solve the aforementioned problems, we propose a Deep Self-Adaptive Hashing(DSAH) model to adaptively capture the semantic information with two special designs: Adaptive Neighbor Discovery(AND) and Pairwise Information Content(PIC). Firstly, we adopt the AND to initially construct a neighborhood-based similarity matrix, and then refine this initial similarity matrix with a novel update strategy to further investigate the semantic structure behind the learned representation. Secondly, we measure the priorities of data pairs with PIC and assign adaptive weights to them, which is relies on the assumption that more dissimilar data pairs contain more discriminative information for hash learning. Extensive experiments on several datasets demonstrate that the above two technologies facilitate the deep hashing model to achieve superior performance. Qinghong Lin, Xiaojun Chen 0006, Qin Zhang 0011, Shangxuan Tian |
CIKM | 2 |
| 2021 | Learning unsupervised node representation from multi-view network
Chen Wang 0032, Xiaojun Chen 0006, Bingkun Chen, Feiping Nie 0001, Zhong Ming 0001 |
Inf. Sci. | 2 |
| 2021 | Adaptive discriminant analysis for semi-supervised feature selection
Weichan Zhong, Xiaojun Chen 0006, Feiping Nie 0001, Joshua Zhexue Huang |
Inf. Sci. | 2 |
| 2021 | Selection of diverse features with a diverse regularization
Weichan Zhong, Xiaojun Chen 0006, Qingyao Wu, Min Yang 0007, Joshua Zhexue Huang |
Pattern Recognit. | 2 |
| 2021 | Neural Attentive Network for Cross-Domain Aspect-Level Sentiment ClassificationabstractThis work takes the lead to study the aspect-level sentiment classificationin the domain adaptation scenario. Given a document of any domains, the model needs to figure out the sentiments with respect to fine-grained aspects in the documents. Two main challenges exist in this problem. One is to build a robust document modeling across domains; the other is to mine the domain-specific aspects and make use of the sentiment lexicon. In this paper, we propose a novel approach Neural Attentive model for cross-domain Aspect-level sentiment CLassification (NAACL), which leverages the benefits of the supervised deep neural network as well as the unsupervised probabilistic generative model to strengthen the representation learning. NAACL jointly learns two tasks: (i) a domain classifier, working on documents in both the source and target domains to recognize the domain information of input texts and transfer knowledge from the source domain to the target domain. In particular, a weakly supervised Latent Dirichlet Allocation model (wsLDA) is proposed to learn the domain-specificaspectandsentiment lexiconrepresentations that are then used to calculate the aspect/lexicon-aware document representations via a multi-view attention mechanism; (ii) an aspect-level sentiment classifier, sharing the document modeling with the domain classifier. It makes use of the domain classification results and the aspect/sentiment-aware document representations to classify the aspect-level sentiment of the document in domain adaptation scenario. NAACL is evaluated on both English and Chinese datasets with the out-of-domain as well as in-domain setups. Quantitatively, the experiments demonstrate that NAACL has robust superiority over the compared methods in terms of classification accuracy and F1 score. The qualitative evaluation also shows that the proposed model is capable of reasonably paying attention to those words that are important to judge the sentiment polarity of the input text given an aspect. Min Yang 0007, Wenpeng Yin 0001, Qiang Qu 0001, Wenting Tu, Ying Shen 0001, Xiaojun Chen 0006 |
IEEE Trans. Affect. Comput. | 6 |
| 2021 | Fast Manifold Ranking With Local Bipartite GraphabstractDuring the past decades, manifold ranking has been widely applied to content-based image retrieval and shown excellent performance. However, manifold ranking is computationally expensive in both graph construction and ranking learning. Much effort has been devoted to improve its performance by introducing approximating techniques. In this paper, we propose a fast manifold ranking method, namely Local Bipartite Manifold Ranking (LBMR). Given a set of images, we first extract multiple regions from each image to form a large image descriptor matrix, and then use the anchor-based strategy to construct a local bipartite graph in which a regional k -means (RKM) is proposed to obtain high quality anchors. We propose an iterative method to directly solve the manifold ranking problem from the local bipartite graph, which monotonically decreases the objective function value in each iteration until the algorithm converges. Experimental results on several real-world image datasets demonstrate the effectiveness and efficiency of our proposed method. Xiaojun Chen 0006, Yuzhong Ye, Qingyao Wu, Feiping Nie 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | Hierarchical Human-Like Deep Neural Networks for Abstractive Text SummarizationabstractDeveloping an abstractive text summarization (ATS) system that is capable of generating concise, appropriate, and plausible summaries for the source documents is a long-term goal of artificial intelligence (AI). Recent advances in ATS are overwhelmingly contributed by deep learning techniques, which have taken the state-of-the-art of ATS to a new level. Despite the significant success of previous methods, generating high-quality and human-like abstractive summaries remains a challenge in practice. The human reading cognition, which is essential for reading comprehension and logical thinking, is still relatively new territory and underexplored in deep neural networks. In this article, we propose a novel Hierarchical Human-like deep neural network for ATS (HH-ATS), inspired by the process of how humans comprehend an article and write the corresponding summary. Specifically, HH-ATS is composed of three primary components (i.e., a knowledge-aware hierarchical attention module, a multitask learning module, and a dual discriminator generative adversarial network), which mimic the three stages of human reading cognition (i.e., rough reading, active reading, and postediting). Experimental results on two benchmark data sets (CNN/Daily Mail and Gigaword) demonstrate that HH-ATS consistently and substantially outperforms the compared methods. Min Yang 0007, Chengming Li 0004, Ying Shen 0001, Qingyao Wu, Zhou Zhao 0001, Xiaojun Chen 0006 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2021 | An Effective Hybrid Learning Model for Real-Time Event SummarizationabstractReal-time event summarization (RES) aims at extracting a handful of document updates from an overwhelming document stream as the real-time event summary that tracks and summarizes the evolving event of interest. It has been attracting much attention, especially with the growth of streaming applications. Despite the effectiveness of previous studies, obtaining relevant, nonredundant, and timely event summaries remains challenging in real-life applications. This study proposes an effective Hybrid learning model for RES (HRES), which attempts to resolve all three challenges (i.e., nonredundancy, relevance, and timeliness) of RES in a unified framework. The main idea is to: 1) exploit the factual background knowledge from the knowledge base (KB) to capture the informative knowledge and implicit information from the input document/query for better text matching; 2) design a memory network to memorize the input facts temporally from the historical document stream and avoid pushing redundant facts in subsequent timesteps; 3) leverage relevance prediction as an auxiliary task to strengthen the document modeling and help to extract relevant documents; and 4) consider both historical dependencies and future uncertainty of the real-time document stream by exploiting the reinforcement learning technique. Extensive experiments demonstrate that HRES has robust superiority over competitors and gains the state-of-the-art results. Min Yang 0007, Qiang Qu 0001, Ying Shen 0001, Zhou Zhao 0001, Xiaojun Chen 0006, Chengming Li 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2020 | A Noise Adaptive Model for Distantly Supervised Relation Extraction
Bowen Zhang 0005, Yunming Ye, Xiaojun Chen 0006, Xutao Li 0003 |
NLPCC (1) | 4 |
| 2020 | Multi-source Domain Adaptation for Sentiment Classification with Granger Causal InferenceabstractIn this paper, we propose a multi-source domain adaptation method with a Granger-causal objective (MDA-GC) for cross-domain sentiment classification. Specifically, for each source domain, we build an expert model by using a novel sentiment-guided capsule network, which captures the domain invariant knowledge that bridges the knowledge gap between the source and target domains. Then, an attention mechanism is devised to assign importance weights to a mixture of experts, each of which specializes in a different source domain. In addition, we propose a Granger causal objective to make the weights assigned to individual experts correlate strongly with their contributions to the decision at hand. Experimental results on a benchmark dataset demonstrate that the proposed MDA-GC model significantly outperforms the compared methods. Min Yang 0007, Ying Shen 0001, Xiaojun Chen 0006, Chengming Li 0004 |
SIGIR | 3 |
| 2020 | Enhanced Balanced Min Cut
Xiaojun Chen 0006, Weijun Hong, Feiping Nie 0001, Joshua Zhexue Huang, Li Shen 0001 |
Int. J. Comput. Vis. | 1 |
| 2020 | Discriminative Streaming Network Embedding
Yiyan Qi, Jiefeng Cheng, Xiaojun Chen 0006, Reynold Cheng, Albert Bifet, Pinghui Wang |
Knowl. Based Syst. | 3 |
| 2020 | A memory network based end-to-end personalized task-oriented dialogue generation
Bowen Zhang 0005, Xiaofei Xu 0001, Xutao Li 0003, Yunming Ye, Xiaojun Chen 0006, Zhongjie Wang 0003 |
Knowl. Based Syst. | 5 |
| 2020 | An Advanced Deep Generative Framework for Temporal Link Prediction in Dynamic NetworksabstractTemporal link prediction in dynamic networks has attracted increasing attention recently due to its valuable real-world applications. The primary challenge of temporal link prediction is to capture the spatial-temporal patterns and high nonlinearity of dynamic networks. Inspired by the success of image generation, we convert the dynamic network into a sequence of static images and formulate the temporal link prediction as a conditional image generation problem. We propose a novel deep generative framework, called NetworkGAN, to tackle the challenging temporal link prediction task efficiently, which simultaneously models the spatial and temporal features in the dynamic networks via deep learning techniques. The proposed NetworkGAN inherits the advantages of the graph convolutional network (GCN), the temporal matrix factorization (TMF), the long short-term memory network (LSTM), and the generative adversarial network (GAN). Specifically, an attentive GCN is first designed to automatically learn the spatial features of dynamic networks. Second, we propose a TMF enhanced attentive LSTM (TMF-LSTM) to capture the temporal dependencies and evolutionary patterns of dynamic networks, which predicts the network snapshot at next timestamp based on the network snapshots observed at previous timestamps. Furthermore, we employ a GAN framework to further refine the performance of temporal link prediction by using a discriminative model to guide the training of the deep generative model (i.e., TMF-LSTM) in an adversarial process. To verify the effectiveness of the proposed model, we conduct extensive experiments on five real-world datasets. Experimental results demonstrate the significant advantages of NetworkGAN compared to other strong competitors. Min Yang 0007, Junhao Liu 0001, Lei Chen 0072, Zhou Zhao 0001, Xiaojun Chen 0006, Ying Shen 0001 |
IEEE Trans. Cybern. | 5 |
| 2020 | Leveraging Long and Short-Term Information in Content-Aware Movie Recommendation via Adversarial TrainingabstractMovie recommendation systems provide users with ranked lists of movies based on individual's preferences and constraints. Two types of models are commonly used to generate ranking results: 1) long-term models and 2) session-based models. The long-term-based models represent the interactions between users and movies that are supposed to change slowly across time, while the session-based models encode the information of users' interests and changing dynamics of movies' attributes in short terms. In this paper, we propose the LSIC model, leveraging long and short-term information for content-aware movie recommendation using adversarial training. In the adversarial process, we train a generator as an agent of reinforcement learning which recommends the next movie to a user sequentially. We also train a discriminator which attempts to distinguish the generated list of movies from the real records. The poster information of movies is integrated to further improve the performance of movie recommendation, which is specifically essential when few ratings are available. The experiments demonstrate that the proposed model has robust superiority over competitors and achieves the state-of-the-art results. Wei Zhao 0033, Benyou Wang, Min Yang 0007, Jianbo Ye, Zhou Zhao 0001, Xiaojun Chen 0006, Ying Shen 0001 |
IEEE Trans. Cybern. | 6 |
| 2020 | An Ensemble of Generation- and Retrieval-Based Image Captioning With Dual Generator Generative Adversarial NetworkabstractImage captioning, which aims to generate a sentence to describe the key content of a query image, is an important but challenging task. Existing image captioning approaches can be categorised into two types: generation-based methods and retrieval-based methods. Retrieval-based methods describe images by retrieving pre-existing captions from a repository. Generation-based methods synthesize a new sentence that verbalizes the query image. Both ways have certain advantages but suffer from their own disadvantages. In the paper, we propose a novel EnsCaption model, which aims at enhancing an ensemble of retrieval-based and generation-based image captioning methods through a novel dual generator generative adversarial network. Specifically, EnsCaption is composed of a caption generation model that synthesizes tailored captions for the query image, a caption re-ranking model that retrieves the best-matching caption from a candidate caption pool consisting of generated captions and pre-retrieved captions, and a discriminator that learns the multi-level difference between the generated/retrieved captions and the ground-truth captions. During the adversarial training process, the caption generation model and the caption re-ranking model provide improved synthetic and retrieved candidate captions with high ranking scores from the discriminator, while the discriminator based on multi-level ranking is trained to assign low ranking scores to the generated and retrieved image captions. Our model absorbs the merits of both generation-based and retrieval-based approaches. We conduct comprehensive experiments to evaluate the performance of EnsCaption on two benchmark datasets: MSCOCO and Flickr-30K. Experimental results show that EnsCaption achieves impressive performance compared to the strong baseline methods. Min Yang 0007, Junhao Liu 0001, Ying Shen 0001, Zhou Zhao 0001, Xiaojun Chen 0006, Qingyao Wu, Chengming Li 0004 |
IEEE Trans. Image Process. | 5 |
| 2020 | Semi-Supervised Feature Selection via Sparse Rescaled Linear Square RegressionabstractWith the rapid increase of the data size, it has increasing demands for selecting features by exploiting both labeled and unlabeled data. In this paper, we propose a novel semi-supervised embedded feature selection method. The new method extends the least square regression model by rescaling the regression coefficients in the least square regression with a set of scale factors, which is used for evaluating the importance of features. An iterative algorithm is proposed to optimize the new model. It has been proved that solving the new model is equivalent to solving a sparse model with a flexible and adaptable ℓ2;pnorm regularization. Moreover, the optimal solution of scale factors provides a theoretical explanation for why we can use {||w1||2, . . .,||wd||2} to evaluate the importance of features. Experimental results on eight benchmark data sets show the superior performance of the proposed method. Xiaojun Chen 0006, Guowen Yuan, Feiping Nie 0001, Zhong Ming 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2020 | LABIN: Balanced Min Cut for Large-Scale DataabstractAlthough many spectral clustering algorithms have been proposed during the past decades, they are not scalable to large-scale data due to their high computational complexities. In this paper, we propose a novel spectral clustering method for large-scale data, namely, large-scale balanced min cut (LABIN). A new model is proposed to extend the self-balanced min-cut (SBMC) model with the anchor-based strategy and a fast spectral rotation with linear time complexity is proposed to solve the new model. Extensive experimental results show the superior performance of our proposed method in comparison with the state-of-the-art methods including SBMC. Xiaojun Chen 0006, Renjie Chen 0004, Qingyao Wu, Yixiang Fang, Feiping Nie 0001, Joshua Zhexue Huang |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2019 | Exploring Human-Like Reading Strategy for Abstractive Text SummarizationabstractThe recent artificial intelligence studies have witnessed great interest in abstractive text summarization. Although remarkable progress has been made by deep neural network based methods, generating plausible and high-quality abstractive summaries remains a challenging task. The human-like reading strategy is rarely explored in abstractive text summarization, which however is able to improve the effectiveness of the summarization by considering the process of reading comprehension and logical thinking. Motivated by the humanlike reading strategy that follows a hierarchical routine, we propose a novel Hybrid learning model for Abstractive Text Summarization (HATS). The model consists of three major components, a knowledge-based attention network, a multitask encoder-decoder network, and a generative adversarial network, which are consistent with the different stages of the human-like reading strategy. To verify the effectiveness of HATS, we conduct extensive experiments on two real-life datasets, CNN/Daily Mail and Gigaword datasets. The experimental results demonstrate that HATS achieves impressive results on both datasets. Min Yang 0007, Qiang Qu 0001, Wenting Tu, Ying Shen 0001, Zhou Zhao 0001, Xiaojun Chen 0006 |
AAAI | 6 |
| 2019 | Semi-Supervised Feature Selection with Adaptive Discriminant AnalysisabstractIn this paper, we propose a novel Adaptive Discriminant Analysis for semi-supervised feature selection, namely SADA. Instead of computing fixed similarities before performing feature selection, SADA simultaneously learns an adaptive similarity matrix S and a projection matrix W with an iterative method. In each iteration, S is computed from the projected distance with the learned W and W is computed with the learned S. Therefore, SADA can learn better projection matrix W by weakening the effect of noise features with the adaptive similarity matrix. Experimental results on 4 data sets show the superiority of SADA compared to 5 semisupervised feature selection methods. Weichan Zhong, Xiaojun Chen 0006, Guowen Yuan, Yiqin Li, Feiping Nie 0001 |
AAAI | 2 |
| 2019 | Structured Spectral Clustering of PurTree Data
Xiaojun Chen 0006, Yixiang Fang, Rui Mao 0001 |
DASFAA (2) | 1 |
| 2019 | Exploring Communities in Large Profiled Graphs (Extended Abstract)abstractGiven a graph G and a vertex q ∊ G, the community search (CS) problem aims to efficiently find a subgraph of G whose vertices are closely related to q. Communities are prevalent in social and biological networks, and can be used in product advertisement and social event recommendation. In this paper, we study profiled community search (PCS), where CS is performed on a profiled graph. This is a graph in which each vertex has labels arranged in a hierarchical manner. Compared with existing CS approaches, PCS can sufficiently identify vertices with semantic commonalities and thus find more high-quality diverse communities. As a naive solution for PCS is highly expensive, we have developed a tree index, which facilitates efficient and online solutions for PCS. Yankai Chen 0001, Yixiang Fang, Reynold Cheng, Xiaojun Chen 0006, Jie Zhang 0002 |
ICDE | 5 |
| 2019 | Knowledge-enhanced Hierarchical Attention for Community Question Answering with Multi-task and Adaptive LearningabstractIn this paper, we propose a Knowledge-enhanced Hierarchical Attention for community question answering with Multi-task learning and Adaptive learning (KHAMA). First, we propose a hierarchical attention network to fully fuse knowledge from input documents and knowledge base (KB) by exploiting the semantic compositionality of the input sequences. The external factual knowledge helps recognize background knowledge (entity mentions and their relationships) and eliminate noise information from long documents that have sophisticated syntactic and semantic structures. In addition, we build multiple CQA models with adaptive boosting and then combine these models to learn a more effective and robust CQA system. Further- more, KHAMA is a multi-task learning model. It regards CQA as the primary task and question categorization as the auxiliary task, aiming at learning a category-aware document encoder and enhance the quality of identifying essential information from long questions. Extensive experiments on two benchmarks demonstrate that KHAMA achieves substantial improvements over the compared methods. Min Yang 0007, Lei Chen 0072, Xiaojun Chen 0006, Qingyao Wu, Wei Zhou 0028, Ying Shen 0001 |
IJCAI | 3 |
| 2019 | Learning Personalized End-to-End Task-Oriented Dialogue Generation
Bowen Zhang 0005, Xiaofei Xu 0001, Xutao Li 0003, Yunming Ye, Xiaojun Chen 0006, Lianjie Sun |
NLPCC (1) | 5 |
| 2019 | Sentiment analysis through critic learning for optimizing convolutional neural networks with rules
Bowen Zhang 0005, Xiaofei Xu 0001, Xutao Li 0003, Xiaojun Chen 0006, Yunming Ye, Zhongjie Wang 0003 |
Neurocomputing | 4 |
| 2019 | Discovering author interest evolution in order-sensitive and Semantic-aware topic modeling
Min Yang 0007, Qiang Qu 0001, Xiaojun Chen 0006, Wenting Tu, Ying Shen 0001, Jia Zhu 0003 |
Inf. Sci. | 3 |
| 2019 | Subspace Weighting Co-Clustering of Gene Expression DataabstractMicroarray technology enables the collection of vast amounts of gene expression data from biological experiments. Clustering algorithms have been successfully applied to exploring the gene expression data. Since a set of genes may be only correlated to a subset of samples, it is useful to use co-clustering to recover co-clusters in the gene expression data. In this paper, we propose a novel algorithm, called Subspace Weighting Co-Clustering (SWCC), for high dimensional gene expression data. In SWCC, a gene subspace weight matrix is introduced to identify the contribution of gene objects in distinguishing different sample clusters. We design a new co-clustering objective function to recover the co-clusters in the gene expression data, in which the subspace weight matrix is introduced. An iterative algorithm is developed to solve the objective function, in which the subspace weight matrix is automatically computed during the iterative co-clustering process. Our empirical study shows encouraging results of the proposed algorithm in comparison with six state-of-the-art clustering algorithms on ten gene expression data sets. We also propose to use SWCC for gene clustering and selection. The experimental results show that the selected genes can improve the classification performance of Random Forests. Xiaojun Chen 0006, Joshua Zhexue Huang, Qingyao Wu, Min Yang 0007 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2019 | Spectral Clustering of Customer Transaction Data With a Two-Level Subspace Weighting MethodabstractFinding customer groups from transaction data is very important for retail and e-commerce companies. Recently, a "Purchase Tree" data structure is proposed to compress the customer transaction data and a local PurTree spectral clustering method is proposed to cluster the customer transaction data. However, in the PurTree distance, the node weights for the children nodes of a parent node are set as equal and the differences between different nodes are not distinguished. In this paper, we propose a two-level subspace weighting spectral clustering (TSW) algorithm for customer transaction data. In the new method, a PurTree subspace metric is proposed to measure the dissimilarity between two customers represented by two purchase trees, in which a set of level weights are introduced to distinguish the importance of different tree levels and a set of sparse node weights are introduced to distinguish the importance of different tree nodes in a purchase tree. TSW learns an adaptive similarity matrix from the local distances in order to better uncover the cluster structure buried in the customer transaction data. Simultaneously, it learns a set of level weights and a set of sparse node weights in the PurTree subspace distance. An iterative optimization algorithm is proposed to optimize the proposed model. We also present an efficient method to compute a regularization parameter in TSW. TSW was compared with six clustering algorithms on ten benchmark data sets and the experimental results show the superiority of the new method. Xiaojun Chen 0006, Wenya Sun, Zhihui Li 0001, Xizhao Wang, Yunming Ye |
IEEE Trans. Cybern. | 1 |
| 2019 | Exploring Communities in Large Profiled GraphsabstractGiven a graph $G$G and a vertex $q\in G$q∈G, the community search (CS) problem aims to efficiently find a subgraph of $G$G whose vertices are closely related to $q$q. Communities are prevalent in social and biological networks, and can be used in product advertisement and social event recommendation. In this paper, we study profiled community search (PCS), where CS is performed on a profiled graph. This is a graph in which each vertex has labels arranged in a hierarchical manner. Extensive experiments show that PCS can identify communities with themes that are common to their vertices, and is more effective than existing CS approaches. As a naive solution for PCS is highly expensive, we have also developed a tree index, which facilitates efficient and online solutions for PCS. Yankai Chen 0001, Yixiang Fang, Reynold Cheng, Xiaojun Chen 0006, Jie Zhang 0002 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2019 | On Spatial-Aware Community SearchabstractCommunities are prevalent in social networks, knowledge graphs, and biological networks. Recently, the topic of community search (CS) has received plenty of attention. The CS problem aims to look for a dense subgraph that contains a query vertex. Existing CS solutions do not consider the spatial extent of a community. They can yield communities whose locations of vertices span large areas. In applications that facilitate setting social events (e.g., finding conference attendees to join a dinner), it is important to find groups of people who are physically close to each other, so it is desirable to have aspatial-aware community(or SAC), whose vertices are close structurally and spatially. Given a graph$G$and a query vertex$q$, we develop an exact solution to find the SAC containing$q$, but it cannot scale to large datasets, so we design three approximation algorithms. We further study the problem of continuous SAC search on a “dynamic spatial graph,” whose vertices’ locations change with time, and propose three fast solutions. We evaluate the solutions on both real and synthetic datasets, and the results show that SACs are better than communities returned by existing solutions. Moreover, our approximation solutions perform accurately and efficiently. Yixiang Fang, Reynold Cheng, Xiaodong Li 0009, Siqiang Luo, Jiafeng Hu, Xiaojun Chen 0006 |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2019 | Multitask Learning for Cross-Domain Image CaptioningabstractRecent artificial intelligence research has witnessed great interest in automatically generating text descriptions of images, which are known as the image captioning task. Remarkable success has been achieved on domains where a large number of paired data in multimedia are available. Nevertheless, annotating sufficient data is labor-intensive and time-consuming, establishing significant barriers for adapting the image captioning systems to new domains. In this study, we introduc a novel Multitask Learning Algorithm for cross-Domain Image Captioning (MLADIC). MLADIC is a multitask system that simultaneously optimizes two coupled objectives via a dual learning mechanism: image captioning and text-to-image synthesis, with the hope that by leveraging the correlation of the two dual tasks, we are able to enhance the image captioning performance in the target domain. Concretely, the image captioning task is trained with an encoder-decoder model (i.e., CNN-LSTM) to generate textual descriptions of the input images. The image synthesis task employs the conditional generative adversarial network (C-GAN) to synthesize plausible images based on text descriptions. In C-GAN, a generative model $G$ synthesizes plausible images given text descriptions, and a discriminative model $D$ tries to distinguish the images in training data from the generated images by $G$. The adversarial process can eventually guide $G$ to generate plausible and high-quality images. To bridge the gap between different domains, a two-step strategy is adopted in order to transfer knowledge from the source domains to the target domains. First, we pre-train the model to learn the alignment between the neural representations of images and that of text data with the sufficient labeled source domain data. Second, we fine-tune the learned model by leveraging the limited image-text pairs and unpaired data in the target domain. We conduct extensive experiments to evaluate the performance of MLADIC by using the MSCOCO as the source domain data, and using Flickr30k and Oxford-102 as the target domain data. The results demonstrate that MLADIC achieves substantially better performance than the strong competitors for the cross-domain image captioning task. Min Yang 0007, Wei Zhao 0033, Yabing Feng, Zhou Zhao 0001, Xiaojun Chen 0006, Kai Lei |
IEEE Trans. Multim. | 6 |
| 2019 | MARES: multitask learning algorithm for Web-scale real-time event summarization
Min Yang 0007, Wenting Tu, Qiang Qu 0001, Kai Lei, Xiaojun Chen 0006, Jia Zhu 0003, Ying Shen 0001 |
World Wide Web | 5 |
| 2018 | A Stratified Feature Ranking Method for Supervised Feature SelectionabstractMost feature selection methods usually select the highest rank features which may be highly correlated with each other. In this paper, we propose a Stratified Feature Ranking (SFR) method for supervised feature selection. In the new method, a Subspace Feature Clustering (SFC) is proposed to identify feature clusters, and a stratified feature ranking method is proposed to rank the features such that the high rank features are lowly correlated. Experimental results show the superiority of SFR. Renjie Chen 0004, Xiaojun Chen 0006, Guowen Yuan, Wenya Sun, Qingyao Wu |
AAAI | 2 |
| 2018 | Discriminative Semi-Supervised Feature Selection via Rescaled Least Squares Regression-SupplementabstractIn this paper, we propose a Discriminative Semi-Supervised Feature Selection (DSSFS) method. In this method, a ε-dragging technique is introduced to the Rescaled Linear Square Regression in order to enlarge the distances between different classes. An iterative method is proposed to simultaneously learn the regression coefficients, ε-draggings matrix and predicting the unknown class labels. Experimental results show the superiority of DSSFS. Guowen Yuan, Xiaojun Chen 0006, Chen Wang 0032, Feiping Nie 0001, Liping Jing |
AAAI | 2 |
| 2018 | A Semi-Supervised Network Embedding Model for Protein Complexes DetectionabstractProtein complex is a group of associated polypeptide chains which plays essential roles in biological process. Given a graph representing protein-protein interactions (PPI) network, it is critical but non-trivial to detect protein complexes.In this paper, we propose a semi-supervised network embedding model by adopting graph convolutional networks to effectively detect densely connected subgraphs. We conduct extensive experiment on two popular PPI networks with various data sizes and densities. The experimental results show our approach achieves state-of-the-art performance. Wei Zhao 0033, Jia Zhu 0003, Min Yang 0007, Danyang Xiao, Gabriel Pui Cheong Fung, Xiaojun Chen 0006 |
AAAI | 6 |
| 2018 | PLASTIC: Prioritize Long and Short-term Information in Top-n Recommendation using Adversarial TrainingabstractRecommender systems provide users with ranked lists of items based on individual's preferences and constraints. Two types of models are commonly used to generate ranking results: long-term models and session-based models. While long-term models represent the interactions between users and items that are supposed to change slowly across time, session-based models encode the information of users' interests and changing dynamics of items' attributes in short terms. In this paper, we propose a PLASTIC model, Prioritizing Long And Short-Term Information in top-n reCommendation using adversarial training. In the adversarial process, we train a generator as an agent of reinforcement learning which recommends the next item to a user sequentially. We also train a discriminator which attempts to distinguish the generated list of items from the real list recorded. Extensive experiments show that our model exhibits significantly better performances on two widely used real-world datasets. Wei Zhao 0033, Benyou Wang, Jianbo Ye, Yongqiang Gao, Min Yang 0007, Xiaojun Chen 0006 |
IJCAI | 6 |
| 2018 | Spectral Clustering of Large-scale Data by Directly Solving Normalized CutabstractDuring the past decades, many spectral clustering algorithms have been proposed. However, their high computational complexities hinder their applications on large-scale data. Moreover, most of them use a two-step approach to obtain the optimal solution, which may deviate from the solution by directly solving the original problem. In this paper, we propose a new optimization algorithm, namely Direct Normalized Cut (DNC), to directly optimize the normalized cut model. DNC has a quadratic time complexity, which is a significant reduction comparing with the cubic time complexity of the traditional spectral clustering. To cope with large-scale data, a Fast Normalized Cut (FNC) method with linear time and space complexities is proposed by extending DNC with an anchor-based strategy. In the new method, we first seek a set of anchors and then construct a representative similarity matrix by computing distances between the anchors and the whole data set. To find high quality anchors that best represent the whole data set, we propose a Balanced k-means (BKM) to partition a data set into balanced clusters and use the cluster centers as anchors. Then DNC is used to obtain the final clustering result from the representative similarity matrix. A series of experiments were conducted on both synthetic data and real-world data sets, and the experimental results show the superior performance of BKM, DNC and FNC. Xiaojun Chen 0006, Weijun Hong, Feiping Nie 0001, Min Yang 0007, Joshua Zhexue Huang |
KDD | 1 |
| 2018 | Investigating Deep Reinforcement Learning Techniques in Personalized Dialogue GenerationabstractIn this paper, we propose a personalized dialogue generation system, which combines reinforcement learning techniques with an attention-based hierarchical recurrent encoderdecoder model. Firstly, we incorporate user-specific information into the decoder to capture user's background information and speaking style. Secondly, we employ reinforcement learning techniques to maximize future reward in dialogue, which enables our system to generate topic-coherent, informative and grammatical responses. Moreover, we propose three types of rewards to characterize good conversations. Finally, we compare the performance of the following reinforcement learning methods in dialogue generation: policy gradient, Q-learning, and actor-critic algorithms. We conduct experiments to verify the effectiveness of the proposed model on two dialogue datasets. Experimental results demonstrate that our model can generate better personalized dialogues for different users. Quantitatively, our method achieves better performance than the state-of-the-art dialogue systems in terms of BLEU score, perplexity, and human evaluation. Min Yang 0007, Qiang Qu 0001, Kai Lei, Jia Zhu 0003, Zhou Zhao 0001, Xiaojun Chen 0006, Joshua Zhexue Huang |
SDM | 6 |
| 2018 | A Topic Drift Model for authorship attribution
Min Yang 0007, Xiaojun Chen 0006, Wenting Tu, Jia Zhu 0003, Qiang Qu 0001 |
Neurocomputing | 2 |
| 2018 | Feature-enhanced attention network for target-dependent sentiment classification
Min Yang 0007, Qiang Qu 0001, Xiaojun Chen 0006, Chaoxue Guo, Ying Shen 0001, Kai Lei |
Neurocomputing | 3 |
| 2018 | Personalized response generation by Dual-learning based domain adaptation
Min Yang 0007, Wenting Tu, Qiang Qu 0001, Zhou Zhao 0001, Xiaojun Chen 0006, Jia Zhu 0003 |
Neural Networks | 5 |
| 2018 | TWCC: Automated Two-way Subspace Weighting Partitional Co-Clustering
Xiaojun Chen 0006, Min Yang 0007, Joshua Zhexue Huang, Zhong Ming 0001 |
Pattern Recognit. | 1 |
| 2018 | PurTreeClust: A Clustering Algorithm for Customer Segmentation from Massive Customer Transaction DataabstractClustering of customer transaction data is an important procedure to analyze customer behaviors in retail and e-commerce companies. Note that products from companies are often organized as a product tree, in which the leaf nodes are goods to sell, and the internal nodes (except root node) could be multiple product categories. Based on this tree, we propose the “personalized product tree”, named purchase tree, to represent a customer's transaction records. So the customers' transaction data set can be compressed into a set of purchase trees. We propose a partitional clustering algorithm, named PurTreeClust, for fast clustering of purchase trees. A new distance metric is proposed to effectively compute the distance between two purchase trees. To cluster the purchase tree data, we first rank the purchase trees as candidate representative trees with a novel separate density, and then select the top k customers as the representatives of k customer groups. Finally, the clustering results are obtained by assigning each customer to the nearest representative. We also propose a gap statistic based method to evaluate the number of clusters. A series of experiments were conducted on ten real-life transaction data sets, and experimental results show the superior performance of the proposed method. Xiaojun Chen 0006, Yixiang Fang, Min Yang 0007, Feiping Nie 0001, Zhou Zhao 0001, Joshua Zhexue Huang |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2018 | Local Adaptive Projection Framework for Feature Selection of Labeled and Unlabeled DataabstractMost feature selection methods first compute a similarity matrix by assigning a fixed value to pairs of objects in the whole data or to pairs of objects in a class or by computing the similarity between two objects from the original data. The similarity matrix is fixed as a constant in the subsequent feature selection process. However, the similarities computed from the original data may be unreliable, because they are affected by noise features. Moreover, the local structure within classes cannot be recovered if the similarities between the pairs of objects in a class are equal. In this paper, we propose a novel local adaptive projection (LAP) framework. Instead of computing fixed similarities before performing feature selection, LAP simultaneously learns an adaptive similarity matrix and a projection matrix with an iterative method. In each iteration, is computed from the projected distance with the learned and W is computed with the learned . Therefore, LAP can learn better projection matrix by weakening the effect of noise features with the adaptive similarity matrix. A supervised feature selection with LAP (SLAP) method and an unsupervised feature selection with LAP (ULAP) method are proposed. Experimental results on eight data sets show the superiority of SLAP compared with seven supervised feature selection methods and the superiority of ULAP compared with five unsupervised feature selection methods. Xiaojun Chen 0006, Guowen Yuan, Feiping Nie 0001, Xiaojun Chang, Joshua Zhexue Huang |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | Detecting Review Spammer GroupsabstractWith an increasing number of paid writers posting fake reviews to promote or demote some target entities through Internet, review spammer detection has become a crucial and challenging task. In this paper, we propose a three-phase method to address the problem of identifying review spammer groups and individual spammers, who get paid for posting fake comments. We evaluate the effectiveness and performance of the approach on a real-life online shopping review dataset from amazon.com. The experimental result shows that our model achieved comparable or better performance than previous work on spammer detection. Min Yang 0007, Xiaojun Chen 0006 |
AAAI | 3 |
| 2017 | Attention Based LSTM for Target Dependent Sentiment ClassificationabstractWe present an attention-based bidirectional LSTM approach to improve the target-dependent sentiment classification. Our method learns the alignment between the target entities and the most distinguishing features. We conduct extensive experiments on a real-life dataset. The experimental results show that our model achieves state-of-the-art results. Min Yang 0007, Wenting Tu, Xiaojun Chen 0006 |
AAAI | 5 |
| 2017 | Identifying and Tracking Sentiments and Topics from Social Media Texts during Natural DisastersabstractWe study the problem of identifying the topics and sentiments and tracking their shifts from social media texts in different geographical regions during emergencies and disasters.We propose a location-based dynamic sentiment-topic model (LDST) which can jointly model topic, sentiment, time and Geolocation information.The experimental results demonstrate that LDST performs very well at discovering topics and sentiments from social media and tracking their shifts in different geographical regions during emergencies and disasters 1 . Min Yang 0007, Jincheng Mei, Heng Ji 0001, Wei Zhao 0033, Zhou Zhao 0001, Xiaojun Chen 0006 |
EMNLP | 6 |
| 2017 | A Self-Balanced Min-Cut Algorithm for Image ClusteringabstractMany spectral clustering algorithms have been proposed and successfully applied to image data analysis such as content based image retrieval, image annotation, and image indexing. Conventional spectral clustering algorithms usually involve a two-stage process: eigendecomposition of similarity matrix and clustering assignments from eigenvectors by k-means or spectral rotation. However, the final clustering assignments obtained by the two-stage process may deviate from the assignments by directly optimize the original objective function. Moreover, most of these methods usually have very high computational complexities. In this paper, we propose a new min-cut algorithm for image clustering, which scales linearly to the data size. In the new method, a self-balanced min-cut model is proposed in which the Exclusive Lasso is implicitly introduced as a balance regularizer in order to produce balanced partition. We propose an iterative algorithm to solve the new model, which has a time complexity of O(n) where n is the number of samples. Theoretical analysis reveals that the new method can simultaneously minimize the graph cut and balance the partition across all clusters. A series of experiments were conducted on both synthetic and benchmark data sets and the experimental results show the superior performance of the new method. Xiaojun Chen 0006, Joshua Zhexue Huang, Feiping Nie 0001, Renjie Chen 0004, Qingyao Wu |
ICCV | 1 |
| 2017 | Scalable Normalized Cut with Improved Spectral RotationabstractMany spectral clustering algorithms have been proposed and successfully applied to many high-dimensional applications. However, there are still two problems that need to be solved: 1) existing methods for obtaining the final clustering assignments may deviate from the true discrete solution, and 2) most of these methods usually have very high computational complexity. In this paper, we propose a Scalable Normalized Cut method for clustering of large scale data. In the new method, an efficient method is used to construct a small representation matrix and then clustering is performed on the representation matrix. In the clustering process, an improved spectral rotation method is proposed to obtain the solution of the final clustering assignments. A series of experimental were conducted on 14 benchmark data sets and the experimental results show the superior performance of the new method. Xiaojun Chen 0006, Feiping Nie 0001, Joshua Zhexue Huang, Min Yang 0007 |
IJCAI | 1 |
| 2017 | Semi-supervised Feature Selection via Rescaled Linear RegressionabstractWith the rapid increase of complex and high-dimensional sparse data, demands for new methods to select features by exploiting both labeled and unlabeled data have increased. Least regression based feature selection methods usually learn a projection matrix and evaluate the importances of features using the projection matrix, which is lack of theoretical explanation. Moreover, these methods cannot find both global and sparse solution of the projection matrix. In this paper, we propose a novel semi-supervised feature selection method which can learn both global and sparse solution of the projection matrix. The new method extends the least square regression model by rescaling the regression coefficients in the least square regression with a set of scale factors, which are used for ranking the features. It has shown that the new model can learn global and sparse solution. Moreover, the introduction of scale factors provides a theoretical explanation for why we can use the projection matrix to rank the features. A simple yet effective algorithm with proved convergence is proposed to optimize the new model. Experimental results on eight real-life data sets show the superiority of the method. Xiaojun Chen 0006, Guowen Yuan, Feiping Nie 0001, Joshua Zhexue Huang |
IJCAI | 1 |
| 2017 | Personalized Response Generation via Domain adaptationabstractIn this paper, we propose a novel personalized response generation model via domain adaptation (PRG-DM). First, we learn the human responding style from large general data (without user-specific information). Second, we fine tune the model on a small size of personalized data to generate personalized responses with a dual learning mechanism. Moreover, we propose three new rewards to characterize good conversations that are personalized, informative and grammatical. We employ the policy gradient method to generate highly rewarded responses. Experimental results show that our model can generate better personalized responses for different users. Min Yang 0007, Zhou Zhao 0001, Wei Zhao 0033, Xiaojun Chen 0006, Jia Zhu 0003, Lianqiang Zhou, Zigang Cao |
SIGIR | 4 |
| 2016 | Online Multi-Instance Multi-Label learning for protein function predictionabstractProtein function prediction is a challenging and essential research problem in the field of computational biology. Conventionally, a protein consists of a number of structural domains and performs multiple function. By representing proteins, domains and functions by bags as well as instances and classes respectively, we are able to model the protein function prediction task as the Multi-Instance Multi-Label (MIML) learning problem. Existing MIML algorithms mainly focus on batch setting where training examples are available before learning. Such offline paradigm works well in simulation, but it may be not feasible for real-world online applications where data comes one by one or chunk by chunk. In this paper, we investigate the protein function prediction problem under a new learning framework, called Online Multi-Instance Multi-Label (OMIML) learning, where MIML protein examples arrive sequentially in an online setting, and develop two OMIML algorithms (OMIML-I and OMIML-B) to make predictions for the incoming data. In the proposed OMIML algorithms, variable-length features are constructed to represent the MIML protein examples based on an incremental vocabulary mechanism. In particular, the incremental vocabularies that OMIML-I and OMIML-B are based on consist of instances and bags, respectively. Then we seek an online prediction for each new arrived protein example by incorporating the constructed features into an online multi-label learning algorithm which is constructed by introducing an artificial label into an online multi-label ranking model. We evaluate the algorithms on the protein dataset consisting of seven real-world organisms. Experimental results have demonstrated the effectiveness of the proposed OMIML algorithms for protein function prediction. Feng Wu 0004, Qiong Liu 0006, Tianyong Hao, Xiaojun Chen 0006, Qingyao Wu |
BIBM | 4 |
| 2016 | PurTreeClust: A purchase tree clustering algorithm for large-scale customer transaction dataabstractClustering of customer transaction data is usually an important procedure to analyze customer behaviors in retail and e-commerce companies. Note that products from companies are often organized as a product tree, in which the leaf nodes are goods to sell, and the internal nodes (except root node) could be multiple product categories. Based on this tree, we present to use a “personalized product tree”, called purchase tree, to represent a customer's transaction data. The customer transaction data set can be represented as a set of purchase trees. We propose a PurTreeClust algorithm for clustering of large-scale customers from purchase trees. We define a new distance metric to effectively compute the distance between two purchase trees from the entire levels in the tree. A cover tree is then built for indexing the purchase tree data and we propose a leveled density estimation method for selecting initial cluster centers from a cover tree. PurTreeClust, a fast clustering method for clustering of large-scale purchase trees, is then presented. Last, we propose a gap statistic based method for estimating the number of clusters from the purchase tree clustering results. A series of experiments were conducted on ten large-scale transaction data sets which contain up to four million transaction records, and experimental results have verified the effectiveness and efficiency of the proposed method. We also compared our method with three clustering algorithms, e.g., spectral clustering, hierarchical agglomerative clustering and DBSCAN. The experimental results have demonstrated the superior performance of the proposed method. Xiaojun Chen 0006, Joshua Zhexue Huang |
ICDE | 1 |
| 2013 | TW-k-Means: Automated Two-Level Variable Weighting Clustering Algorithm for Multiview DataabstractThis paper proposes TW-k-means, an automated two-level variable weighting clustering algorithm for multiview data, which can simultaneously compute weights for views and individual variables. In this algorithm, a view weight is assigned to each view to identify the compactness of the view and a variable weight is also assigned to each variable in the view to identify the importance of the variable. Both view weights and variable weights are used in the distance function to determine the clusters of objects. In the new algorithm, two additional steps are added to the iterative k-means clustering process to automatically compute the view weights and the variable weights. We used two real-life data sets to investigate the properties of two types of weights in TW-k-means and investigated the difference between the weights of TW-k-means and the weights of the individual variable weighting method. The experiments have revealed the convergence property of the view weights in TW-k-means. We compared TW-k-means with five clustering algorithms on three real-life data sets and the results have shown that the TW-k-means algorithm significantly outperformed the other five clustering algorithms in four evaluation indices. Xiaojun Chen 0006, Xiaofei Xu 0001, Joshua Zhexue Huang, Yunming Ye |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2012 | Scalable Subspace Logistic Regression Models for High Dimensional Data
Xiaojun Chen 0006, Joshua Zhexue Huang, Shengzhong Feng |
APWeb | 2 |
| 2012 | Scalable Random Forests for Massive Data
Bingguo Li, Xiaojun Chen 0006, Mark Junjie Li, Joshua Zhexue Huang, Shengzhong Feng |
PAKDD (1) | 2 |
| 2012 | A feature group weighting method for subspace clustering of high-dimensional data
Xiaojun Chen 0006, Yunming Ye, Xiaofei Xu 0001, Joshua Zhexue Huang |
Pattern Recognit. | 1 |
| 2006 | MFCRank: A Web Ranking Algorithm Based on Correlation of Multiple Features
Yunming Ye, Yan Li 0040, Xiaofei Xu 0001, Joshua Zhexue Huang, Xiaojun Chen 0006 |
CICLing | 5 |
| 2006 | Neighborhood Density Method for Selecting Initial Cluster Centers in K-Means Clustering
Yunming Ye, Joshua Zhexue Huang, Xiaojun Chen 0006, Shuigeng Zhou, Graham J. Williams, Xiaofei Xu 0001 |
PAKDD | 3 |