VLDB 2026 Research / reviewers in the wild / expert
Xiaoming Jin
dblp:26/5593
· DBLP profile ↗
42ranked-venue papers
8as first author
1since 2021 · last 2023
0000-0003-1531-3803ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 2 first-authorDatabases, data management, data science and information retrieval · 19 · 5 first-authorGraphics, computer vision, multimedia, augmented reality and games · 17 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 2 first-authorSystems, architecture and hardware · 1 · 1 since 2021Computer networks · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
13 papers |
Transfer learning and domain adaptation · 38% Representation and self-supervised learning · 20% Information extraction and text analysis · 12% | |
| Databases, data mining, and information retrieval
12 papers |
Data mining · 36% Recommender systems · 28% Query processing and optimization · 13% |
Topics — the 30 heaviest of 49, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Face, body and person analysis
pedestrian attribute recognition |
0.7 | 2 | 2019 | Incremental Few-Shot Learning for Pedestrian Attribute Recognition · IJCAI 2019 Grouping Attribute Recognition for Pedestrian with Joint Recurrent Learning · IJCAI 2018 |
Natural language and speech › Information extraction and text analysis
text classification |
0.5 | 3 | 2019 | Adaptive Region Embedding for Text Classification · AAAI 2019 Short Text Classification Improved by Learning Multi-Granularity Topics · IJCAI 2011 Multi-domain active learning for text classification · KDD 2012 |
Machine learning › Transfer learning and domain adaptation
heterogeneous transfer learning |
0.4 | 1 | 2020 | Heterogeneous Transfer Learning with Weighted Instance-Correspondence Data · AAAI 2020 |
Machine learning › Transfer learning and domain adaptation › few-shot learning
few-shot class-incremental learning |
0.4 | 1 | 2019 | Incremental Few-Shot Learning for Pedestrian Attribute Recognition · IJCAI 2019 |
Machine learning › Transfer learning and domain adaptation
few-shot learning |
0.4 | 1 | 2019 | Incremental Few-Shot Learning for Pedestrian Attribute Recognition · IJCAI 2019 |
Machine learning › Representation and self-supervised learning › representation learning › embedding learning › image embedding
region embedding |
0.4 | 1 | 2019 | Adaptive Region Embedding for Text Classification · AAAI 2019 |
Machine learning › Representation and self-supervised learning › text embedding
text representation learning |
0.4 | 1 | 2019 | Adaptive Region Embedding for Text Classification · AAAI 2019 |
Information retrieval
similarity search |
0.3 | 2 | 2016 | Robust Iterative Quantization for Efficient ℓp-norm Similarity Search · IJCAI 2016 Probabilistic Similarity Query on Dimension Incomplete Data · ICDM 2009 |
Machine learning › Deep learning architectures and training › deep generative model
deep belief network |
0.3 | 1 | 2018 | Automatic Gating of Attributes in Deep Structure · IJCAI 2018 |
Computer vision › Image recognition and object detection
image classification |
0.3 | 1 | 2018 | Automatic Gating of Attributes in Deep Structure · IJCAI 2018 |
Data integration and cleaning
missing data |
0.3 | 2 | 2014 | Searching Dimension Incomplete Databases · IEEE Trans. Knowl. Data Eng. 2014 Probabilistic Similarity Query on Dimension Incomplete Data · ICDM 2009 |
Query processing and optimization
similarity query processing |
0.3 | 2 | 2014 | Searching Dimension Incomplete Databases · IEEE Trans. Knowl. Data Eng. 2014 Probabilistic Similarity Query on Dimension Incomplete Data · ICDM 2009 |
Machine learning › Transfer learning and domain adaptation › cross-domain learning
cross-domain text classification |
0.3 | 2 | 2012 | Topic Correlation Analysis for Cross-Domain Text Classification · AAAI 2012 Transfer Learning via Cluster Correspondence Inference · ICDM 2010 |
Machine learning › Efficient and distributed learning
active learning |
0.2 | 1 | 2016 | Active Learning with Cross-Class Knowledge Transfer · AAAI 2016 |
Machine learning › Transfer learning and domain adaptation › knowledge transfer
cross-category knowledge transfer |
0.2 | 1 | 2016 | Active Learning with Cross-Class Knowledge Transfer · AAAI 2016 |
Machine learning › Efficient and distributed learning › active learning
uncertainty sampling |
0.2 | 1 | 2016 | Active Learning with Cross-Class Knowledge Transfer · AAAI 2016 |
Machine learning › Transfer learning and domain adaptation › zero-shot learning
zero-shot classification |
0.2 | 1 | 2016 | Transductive Zero-Shot Recognition via Shared Model Space Learning · AAAI 2016 |
Recommender systems › interactive recommendation
active learning for recommendation |
0.2 | 1 | 2016 | Multi-Domain Active Learning for Recommendation · AAAI 2016 |
Recommender systems
cross-domain recommendation |
0.2 | 1 | 2016 | Multi-Domain Active Learning for Recommendation · AAAI 2016 |
Data mining › text mining
topic model |
0.2 | 2 | 2012 | Topic Mining over Asynchronous Text Sequences · IEEE Trans. Knowl. Data Eng. 2012 Mining common topics from multiple asynchronous text streams · WSDM 2009 |
Computer vision › Image recognition and object detection
attribute recognition |
0.2 | 1 | 2015 | Learning Predictable and Discriminative Attributes for Visual Recognition · AAAI 2015 |
Machine learning › Representation and self-supervised learning › representation learning
feature extraction |
0.2 | 1 | 2015 | Gaussian Cardinality Restricted Boltzmann Machines · AAAI 2015 |
Machine learning › Probabilistic and Bayesian machine learning › boltzmann machine
restricted boltzmann machine |
0.2 | 1 | 2015 | Gaussian Cardinality Restricted Boltzmann Machines · AAAI 2015 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
sparse coding |
0.2 | 1 | 2015 | Gaussian Cardinality Restricted Boltzmann Machines · AAAI 2015 |
Data mining
pattern mining |
0.2 | 3 | 2009 | Probabilistic Similarity Query on Dimension Incomplete Data · ICDM 2009 Efficient Discovery of Emerging Frequent Patterns in ArbitraryWindows on Data Streams · ICDE 2006 Concept Sampling: Towards Systematic Selection in Large-Scale Mixed Concepts in Machine Learning · IJCAI 2007 |
Recommender systems
collaborative filtering |
0.2 | 1 | 2013 | Celebrity Recommendation with Collaborative Social Topic Regression · IJCAI 2013 |
Recommender systems
social recommendation |
0.2 | 1 | 2013 | Celebrity Recommendation with Collaborative Social Topic Regression · IJCAI 2013 |
Machine learning › Transfer learning and domain adaptation › cross-domain learning
multi-domain learning |
0.1 | 1 | 2012 | Multi-domain active learning for text classification · KDD 2012 |
Natural language and speech › Information extraction and text analysis
topic model |
0.1 | 1 | 2012 | Topic Correlation Analysis for Cross-Domain Text Classification · AAAI 2012 |
Data mining › text mining
topic detection |
0.1 | 1 | 2012 | Topic Mining over Asynchronous Text Sequences · IEEE Trans. Knowl. Data Eng. 2012 |
Methods — techniques the papers use, named apart from their topics
instance weighting · 0.4meta-network · 0.4meta-learning · 0.4convolutional neural network · 0.4adaptive region embedding · 0.4recurrent neural network · 0.3feature extraction · 0.3deep belief network · 0.3body region proposal · 0.3attribute gating · 0.3variance-based strategy · 0.2expected entropy · 0.2generative topic model · 0.2probabilistic framework · 0.2filtering bounds · 0.2collaborative topic regression · 0.2topic modeling · 0.1non-negative matrix factorization · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | DeGTeC: A deep graph-temporal clustering framework for data-parallel job characterization in data centers
Kaizhong Chen, Lan Yi, Xiaoming Jin |
Future Gener. Comput. Syst. | 5 |
| 2020 | Heterogeneous Transfer Learning with Weighted Instance-Correspondence DataabstractInstance-correspondence (IC) data are potent resources for heterogeneous transfer learning (HeTL) due to the capability of bridging the source and the target domains at the instance-level. To this end, people tend to use machine-generated IC data, because manually establishing IC data is expensive and primitive. However, existing IC data machine generators are not perfect and always produce the data that are not of high quality, thus hampering the performance of domain adaption. In this paper, instead of improving the IC data generator, which might not be an optimal way, we accept the fact that data quality variation does exist but find a better way to use the data. Specifically, we propose a novel heterogeneous transfer learning method named Transfer Learning with Weighted Correspondence (TLWC), which utilizes IC data to adapt the source domain to the target domain. Rather than treating IC data equally, TLWC can assign solid weights to each IC data pair depending on the quality of the data. We conduct extensive experiments on HeTL datasets and the state-of-the-art results verify the effectiveness of TLWC. Xiaoming Jin, Guiguang Ding, Jungong Han, Jiyong Zhang 0001, Sicheng Zhao |
AAAI | 2 |
| 2019 | Adaptive Region Embedding for Text ClassificationabstractDeep learning models such as convolutional neural networks and recurrent networks are widely applied in text classification. In spite of their great success, most deep learning models neglect the importance of modeling context information, which is crucial to understanding texts. In this work, we propose the Adaptive Region Embedding to learn context representation to improve text classification. Specifically, a metanetwork is learned to generate a context matrix for each region, and each word interacts with its corresponding context matrix to produce the regional representation for further classification. Compared to previous models that are designed to capture context information, our model contains less parameters and is more flexible. We extensively evaluate our method on 8 benchmark datasets for text classification. The experimental results prove that our method achieves state-of-the-art performances and effectively avoids word ambiguity. Liuyu Xiang, Xiaoming Jin, Lan Yi, Guiguang Ding |
AAAI | 2 |
| 2019 | Towards Better Uncertainty Sampling: Active Learning with Multiple Views for Deep Convolutional Neural NetworkabstractConvolutional neural network (CNN) has been successfully applied to many fields, such as image classification and object detection. It relies on huge amount of data. However, labelling a large amount of data is expensive. Active learning is one of the approaches to alleviate the labelling effort. We propose a new active learning approach for CNN. Different from existing active learning algorithms for CNN, first, the active query strategy is measured from multiple views, not only the last output of CNN; second, multiple views are obtained from multiple hidden layers in CNN, not from other related data or models. We evaluate our approach on three widely used datasets: Fashion-MNIST, SVHN and CIFAR-10. Experimental results show that the proposed method outperforms baseline methods in image classification. Xiaoming Jin, Guiguang Ding, Lan Yi, Chenggang Yan 0001 |
ICME | 2 |
| 2019 | Incremental Few-Shot Learning for Pedestrian Attribute RecognitionabstractPedestrian attribute recognition has received increasing attention due to its important role in video surveillance applications. However, most existing methods are designed for a fixed set of attributes. They are unable to handle the incremental few-shot learning scenario, i.e. adapting a well-trained model to newly added attributes with scarce data, which commonly exists in the real world. In this work, we present a meta learning based method to address this issue. The core of our framework is a meta architecture capable of disentangling multiple attribute information and generalizing rapidly to new coming attributes. By conducting extensive experiments on the benchmark dataset PETA and RAP under the incremental few-shot setting, we show that our method is able to perform the task with competitive performances and low resource requirements. Liuyu Xiang, Xiaoming Jin, Guiguang Ding, Jungong Han, Leida Li |
IJCAI | 2 |
| 2019 | Image Emotion Distribution Learning with Graph Convolutional NetworksabstractRecently, with the rapid progress of techniques in visual analysis, a lot of attention has been paid to affective computing due to its wide potential applications. Traditional affective analysis mainly focus on single label image emotion classification. But a single image may invoke different emotions for different persons, even for one person. So emotion distribution learning is proposed to capture the underlying emotion distribution for images. Currently, state-of-the-art works model the distribution by deep convolutional networks equipped with distribution specific loss. However, the correlation among different emotions is ignored in these works. Some emotions usually co-appear, while some are hardly invoked at the same time. Properly modeling the correlation is important for image emotion distribution learning. Graph convolutional networks have shown great performance in capturing the underlying relationship in graph, and have been successfully applied in vision problems, such as zero-shot image classification. So, in this paper, we propose to apply graph convolutional networks for emotion distribution learning, termed EmotionGCN, which captures the correlation among emotions. The EmotionGCN can make use of correlation either mined from data, or directly from psychological models, such as Mikels' wheel. Extensive experiments are conducted on the FlickrLDL and TwitterLDL datasets, and the results on seven evaluation metrics demonstrate the superiority of the proposed method. Xiaoming Jin |
ICMR | 2 |
| 2018 | Automatic Gating of Attributes in Deep StructureabstractDeep structure has been widely applied in a large variety of fields for its excellence of representing data. Attributes are a unique type of data descriptions that have been successfully utilized in numerous tasks to enhance performance. However, to introduce attributes into deep structure is complicated and challenging, because different layers in deep structure accommodate features of different abstraction levels, while different attributes may naturally represent the data in different abstraction levels. This demands adaptively and jointly modeling of attributes and deep structure by carefully examining their relationship. Different from existing works that treat attributes straightforwardly as the same level without considering their abstraction levels, we can make better use of attributes in deep structure by properly connecting them. In this paper, we move forward along this new direction by proposing a deep structure named Attribute Gated Deep Belief Network (AG-DBN) that includes a tunable attribute-layer gating mechanism and automatically learns the best way of connecting attributes to appropriate hidden layers. Experimental results on a manually-labeled subset of ImageNet, a-Yahoo and a-Pascal data set justify the superiority of AG-DBN against several baselines including CNN model and other AG-DBN variants. Specifically, it outperforms the CNN model, VGG19, by significantly reducing the classification error from 26.70% to 13.56% on a-Pascal. Xiaoming Jin, Lan Yi, Guiguang Ding, Dou Shen |
IJCAI | 1 |
| 2018 | Grouping Attribute Recognition for Pedestrian with Joint Recurrent LearningabstractPedestrian attributes recognition is to predict attribute labels of pedestrian from surveillance images, which is a very challenging task for computer vision due to poor imaging quality and small training dataset. It is observed that semantic pedestrian attributes to be recognised tend to show semantic or visual spatial correlation. Attributes can be grouped by the correlation while previous works mostly ignore this phenomenon. Inspired by Recurrent Neural Network (RNN)'s super capability of learning context correlations, this paper proposes an end-to-end Grouping Recurrent Learning (GRL) model that takes advantage of the intra-group mutual exclusion and inter-group correlation to improve the performance of pedestrian attribute recognition. Our GRL method starts with the detection of precise body region via Body Region Proposal followed by feature extraction from detected regions. These features, along with the semantic groups, are fed into RNN for recurrent grouping attribute recognition, where intra group correlations can be learned. Extensive empirical evidence shows that our GRL model achieves state-of-the-art results, based on pedestrian attribute datasets, i.e. standard PETA and RAP datasets. Xin Zhao 0020, Liufang Sang, Guiguang Ding, Xiaoming Jin |
IJCAI | 5 |
| 2017 | Common-Specific Multimodal Learning for Deep Belief NetworkabstractMultimodal Deep Belief Network has been widely used to extract representations for multimodal data by fusing the high-level features of each data modality into common representations. Such straightforward fusion strategy can benefit the classification and information retrieval tasks. However, it may introduce noise in case the high-level features are not naturally common hence non-fusable for different modalities. Intuitively, each modality may have its own specific features and corresponding representation capabilities thus should not be simply fused. Therefore, it is more reasonable to fuse only the common features and represent the multimodal data by both the fused features and the modality-specific features. To distinguish common features from modal-specific features is a challenging task for traditional DBN models where all features are crudely mixed. This paper proposes the Common-Specific Multimodal Deep Belief Network (CSDBN) to solve the problem. CS-DBN automatically separates common features from modal-specific features and fuses only the common ones for data representation. Experimental results demonstrate the superiority of CS-DBN for classification tasks compared with the baseline approaches. Changsheng Xiang, Xiaoming Jin |
CIKM | 2 |
| 2016 | Transductive Zero-Shot Recognition via Shared Model Space LearningabstractZero-shot Recognition (ZSR) is to learn recognition models for novel classes without labeled data. It is a challenging task and has drawn considerable attention in recent years. The basic idea is to transfer knowledge from seen classes via the shared attributes. This paper focus on the transductive ZSR, i.e., we have unlabeled data for novel classes. Instead of learning models for seen and novel classes separately as in existing works, we put forward a novel joint learning approach which learns the shared model space (SMS) for models such that the knowledge can be effectively transferred between classes using the attributes. An effective algorithm is proposed for optimization. We conduct comprehensive experiments on three benchmark datasets for ZSR. The results demonstrates that the proposed SMS can significantly outperform the state-of-the-art related approaches which validates its efficacy for the ZSR task. Guiguang Ding, Xiaoming Jin, Jianmin Wang 0001 |
AAAI | 3 |
| 2016 | Active Learning with Cross-Class Knowledge TransferabstractWhen there are insufficient labeled samples for training a supervised model, we can adopt active learning to select the most informative samples for human labeling, or transfer learning to transfer knowledge from related labeled data source. Combining transfer learning with active learning has attracted much research interest in recent years. Most existing works follow the setting where the class labels in source domain are the same as the ones in target domain. In this paper, we focus on a more challenging cross-class setting where the class labels are totally different in two domains but related to each other in an intermediary attribute space, which is barely investigated before. We propose a novel and effective method that utilizes the attribute representation as the seed parameters to generate the classification models for classes. And we propose a joint learning framework that takes into account the knowledge from the related classes in source domain, and the information in the target domain. Besides, it is simple to perform uncertainty sampling, a fundamental technique for active learning, based on the framework. We conduct experiments on three benchmark datasets and the results demonstrate the efficacy of the proposed method. Guiguang Ding, Xiaoming Jin |
AAAI | 4 |
| 2016 | Multi-Domain Active Learning for RecommendationabstractRecently, active learning has been applied to recommendation to deal with data sparsity on a single domain. In this paper, we propose an active learning strategy for recommendation to alleviate the data sparsity in a multi-domain scenario. Specifically, our proposed active learning strategy simultaneously consider both specific and independent knowledge over all domains. We use the expected entropy to measure the generalization error of the domain-specific knowledge and propose a variance-based strategy to measure the generalization error of the domain-independent knowledge. The proposed active learning strategy use a unified function to effectively combine these two measurements. We compare our strategy with five state-of-the-art baselines on five different multi-domain recommendation tasks, which are constituted by three real-world data sets. The experimental results show that our strategy performs significantly better than all the baselines and reduces human labeling efforts by at least 5.6%, 8.3%, 11.8%, 12.5% and 15.4% on the five tasks, respectively. Xiaoming Jin, Lianghao Li, Guiguang Ding, Qiang Yang 0001 |
AAAI | 2 |
| 2016 | Robust Iterative Quantization for Efficient ℓp-norm Similarity Search
Guiguang Ding, Jungong Han, Xiaoming Jin |
IJCAI | 4 |
| 2015 | Learning Predictable and Discriminative Attributes for Visual RecognitionabstractUtilizing attributes for visual recognition has attracted increasingly interest because attributes can effectively bridge the semantic gap between low-level visual features and high-level semantic labels. In this paper, we propose a novel method for learning predictable and discriminative attributes. Specifically, we require the learned attributes can be reliably predicted from visual features, and discover the inherent discriminative structure of data. In addition, we propose to exploit the intra-category locality of data to overcome the intra-category variance in visual data. We conduct extensive experiments on Animals with Attributes (AwA) and Caltech256 datasets, and the results demonstrate that the proposed method achieves state-of-the-art performance. Guiguang Ding, Xiaoming Jin, Jianmin Wang 0001 |
AAAI | 3 |
| 2015 | Gaussian Cardinality Restricted Boltzmann MachinesabstractRestricted Boltzmann Machine (RBM) has been applied to a wide variety of tasks due to its advantage in feature extraction. Implementing sparsity constraint in the activated hidden units of RBM is an important improvement on RBM. The sparsity constraints in the existing methods are usually specified by users and are independent of the input data. However, the input data could be heterogeneous in content and thus naturally demand elastic and adaptive settings of the sparsity constraints. To solve this problem, we proposed a generalized model with adaptive sparsity constraint, named Gaussian Cardinality Restricted Boltzmann Machines (GC-RBM). In this model, the thresholds of hidden unit activations are decided by the input data and a given Gaussian distribution on the pre-training phase. We provide a principled method to train the GC-RBM with Gaussian prior. Experimental results on two real world data sets justify the effectiveness of the proposed method and its superiority over CaRBM in terms of classification accuracy. Xiaoming Jin, Guiguang Ding, Dou Shen |
AAAI | 2 |
| 2014 | User Interests Imbalance Exploration in Social Recommendation: A Fitness AdaptationabstractRecent years have witnessed an increasing interest in how to incorporate social network information into recommendation algorithms to enhance the user experience. In this paper, we find the phenomenon that users in the contexts of recommendation system and social network do not share the same interest space. Based on this finding, we proposed the social regulatory factor regression model (SRFRM) which could connect different interest spaces in different contexts together in an unified latent factor model. Specifically, different from the traditional social based latent factor models with strong limitation that all sides share the same feature space, the proposed method leverages the regulatory factor number on both sides to meet the fact that users and items or users in different contexts may not share the same interest space. It works by incorporating two linear transformation matrices into the matrix co-factorization framework that matrix factorization of user ratings is regularized by that of social trust network. We study a large subsets of data from epinions.com and douban.com respectively. The experimental results indicate that users in different contexts have different interest spaces and our model achieves a higher performance compared with related state-of-the-art methods. Tianchun Wang, Xiaoming Jin, Xuetao Ding |
CIKM | 2 |
| 2014 | Max-margin latent feature relational models for entity-attribute networksabstractLink prediction is a fundamental task in statistical analysis of network data. Though much research has concentrated on predicting entity-entity relationships in homogeneous networks, it has attracted increasing attentions to predict relationships in heterogeneous networks, which consist of multiple types of nodes and relational links. Existing work on heterogeneous network link prediction mainly focuses on using input features that are explicitly extracted by humans. This paper presents an approach to automatically learn latent features from partially observed heterogeneous networks, with a particular focus on entity-attribute networks (EANs), and making predictions for unseen pairs. To make the latent features discriminative, we adopt the max-margin idea under the framework of maximum entropy discrimination (MED). Our maximum entropy discrimination joint relational model (MED-JRM) can jointly predict entity-entity relationships as well as the missing attributes of entities in EANs. Experimental results on several real networks demonstrate that our model has improved performance over state-of-the-art homogeneous and heterogeneous network link prediction algorithms. Fei Xia 0005, Ning Chen 0002, Jun Zhu 0001, Aonan Zhang, Xiaoming Jin |
IJCNN | 5 |
| 2014 | Searching Dimension Incomplete DatabasesabstractSimilarity query is a fundamental problem in database, data mining and information retrieval research. Recently, querying incomplete data has attracted extensive attention as it poses new challenges to traditional querying techniques. The existing work on querying incomplete data addresses the problem where the data values on certain dimensions are unknown. However, in many real-life applications, such as data collected by a sensor network in a noisy environment, not only the data values but also the dimension information may be missing. In this work, we propose to investigate the problem of similarity search on dimension incomplete data. A probabilistic framework is developed to model this problem so that the users can find objects in the database that are similar to the query with probability guarantee. Missing dimension information poses great computational challenge, since all possible combinations of missing dimensions need to be examined when evaluating the similarity between the query and the data objects. We develop the lower and upper bounds of the probability that a data object is similar to the query. These bounds enable efficient filtering of irrelevant data objects without explicitly examining all missing dimension combinations. A probability triangle inequality is also employed to further prune the search space and speed up the query process. The proposed probabilistic framework and techniques can be applied to both whole and subsequence queries. Extensive experimental results on real-life data sets demonstrate the effectiveness and efficiency of our approach. Wei Cheng 0002, Xiaoming Jin, Jian-Tao Sun, Xuemin Lin 0001, Xiang Zhang 0001, Wei Wang 0010 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2013 | Short text classification by detecting information pathabstractShort text is becoming ubiquitous in many modern information systems. Due to the shortness and sparseness of short texts, there are less informative word co-occurrences among them, which naturally pose great difficulty for classification tasks on such data. To overcome this difficulty, this paper proposes a new way for effectively classifying the short texts. Our method is based on a key observation that there usually exists ordered subsets in short texts, which is termed ``information path'' in this work, and classification on each subset based on the classification results of some pervious subsets can yield higher overall accuracy than classifying the entire data set directly. We propose a method to detect the information path and employ it in short text classification. Different from the state-of-art methods, our method does not require any external knowledge or corpus that usually need careful fine-tuning, which makes our method easier and more robust on different data sets. Experiments on two real world data sets show the effectiveness of the proposed method and its superiority over the existing methods. Shitao Zhang, Xiaoming Jin, Dou Shen, Bin Cao 0001, Xuetao Ding |
CIKM | 2 |
| 2013 | Celebrity Recommendation with Collaborative Social Topic Regression
Xuetao Ding, Xiaoming Jin, Lianghao Li |
IJCAI | 2 |
| 2013 | Path Planning for Aviation Actuators in Deploying Regular Sensor NetworksabstractThe topology problem focusing on the coverage of sensor networks has been researched by many workers. These sensors can be dropped randomly using aircraft, or deployed by people who take them by vehicle to be positioned. For places that are hard to reach, deployment of sensors requires the use of manned or unmanned aircraft. Although the scattered sensors can be organized by a network protocol, it is more efficient if sensors are positioned regularly. This paper focuses on a path-planning algorithm that considers the performance of aircraft engines and the shortest time to position a sensor network. We propose a geometry-based path-planning algorithm (GPA) to plan a route that can navigate an aircraft to complete a mission. Simulation results show that the complexity and the memory required for the GPA are less than for the A* algorithm. The GPA can be used in path planning of aircraft to deploy regularly shaped sensor networks. Qunhua Pan, Juhui Pan, Xiaoming Jin, Keyu Shen |
MSN | 3 |
| 2012 | Topic Correlation Analysis for Cross-Domain Text ClassificationabstractCross-domain text classification aims to automatically train a precise text classifier for a target domain by using labeled text data from a related source domain. To this end, the distribution gap between different domains has to be reduced. In previous works, a certain number of shared latent features (e.g., latent topics, principal components, etc.) are extracted to represent documents from different domains, and thus reduce the distribution gap. However, only relying the shared latent features as the domain bridge may limit the amount of knowledge transferred. This limitation is more serious when the distribution gap is so large that only a small number of latent features can be shared between domains. In this paper, we propose a novel approach named Topic Correlation Analysis (TCA), which extracts both the shared and the domain-specific latent features to facilitate effective knowledge transfer. In TCA, all word features are first grouped into the shared and the domain-specific topics using a joint mixture model. Then the correlations between the two kinds of topics are inferred and used to induce a mapping between the domain-specific topics from different domains. Finally, both the shared and the mapped domain-specific topics are utilized to span a new shared feature space where the supervised knowledge can be effectively transferred. The experimental results on two real-world data sets justify the superiority of the proposed method over the stat-of-the-art baselines. Lianghao Li, Xiaoming Jin, Mingsheng Long |
AAAI | 2 |
| 2012 | Multi-domain active learning for text classificationabstractActive learning has been proven to be effective in reducing labeling efforts for supervised learning. However, existing active learning work has mainly focused on training models for a single domain. In practical applications, it is common to simultaneously train classifiers for multiple domains. For example, some merchant web sites (like Amazon.com) may need a set of classifiers to predict the sentiment polarity of product reviews collected from various domains (e.g., electronics, books, shoes). Though different domains have their own unique features, they may share some common latent features. If we apply active learning on each domain separately, some data instances selected from different domains may contain duplicate knowledge due to the common features. Therefore, how to choose the data from multiple domains to label is crucial to further reducing the human labeling efforts in multi-domain learning. In this paper, we propose a novel multi-domain active learning framework to jointly select data instances from all domains with duplicate information considered. In our solution, a shared subspace is first learned to represent common latent features of different domains. By considering the common and the domain-specific features together, the model loss reduction induced by each data instance can be decomposed into a common part and a domain-specific part. In this way, the duplicate information across domains can be encoded into the common part of model loss reduction and taken into account when querying. We compare our method with the state-of-the-art active learning approaches on several text classification tasks: sentiment classification, newsgroup classification and email spam filtering. The experiment results show that our method reduces the human labeling efforts by 33.2%, 42.9% and 68.7% on the three tasks, respectively. Lianghao Li, Xiaoming Jin, Sinno Jialin Pan, Jian-Tao Sun |
KDD | 2 |
| 2012 | Topic Mining over Asynchronous Text SequencesabstractTime stamped texts, or text sequences, are ubiquitous in real-world applications. Multiple text sequences are often related to each other by sharing common topics. The correlation among these sequences provides more meaningful and comprehensive clues for topic mining than those from each individual sequence. However, it is nontrivial to explore the correlation with the existence of asynchronism among multiple sequences, i.e., documents from different sequences about the same topic may have different time stamps. In this paper, we formally address this problem and put forward a novel algorithm based on the generative topic model. Our algorithm consists of two alternate steps: the first step extracts common topics from multiple sequences based on the adjusted time stamps provided by the second step; the second step adjusts the time stamps of the documents according to the time distribution of the topics discovered by the first step. We perform these two steps alternately and after iterations a monotonic convergence of our objective function can be guaranteed. The effectiveness and advantage of our approach were justified through extensive empirical studies on two real data sets consisting of six research paper repositories and two news article feeds, respectively. Xiang Wang 0001, Xiaoming Jin, Mengen Chen, Dou Shen |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2011 | Short Text Classification Improved by Learning Multi-Granularity Topics
Mengen Chen, Xiaoming Jin, Dou Shen |
IJCAI | 2 |
| 2010 | Transfer Learning via Cluster Correspondence InferenceabstractTransfer learning targets to leverage knowledge from one domain for tasks in a new domain. It finds abundant applications, such as text/sentiment classification. Many previous works are based on cluster analysis, which assume some common clusters shared by both domains. They mainly focus on the one-to-one cluster correspondence to bridge different domains. However, such a correspondence scheme might be too strong for real applications where each cluster in one domain corresponds to many clusters in the other domain. In this paper, we propose a Cluster Correspondence Inference (CCI) method to iteratively infer many-to-many correspondence among clusters from different domains. Specifically, word clusters and document clusters are exploited for each domain using nonnegative matrix factorization, then the word clusters from different domains are corresponded in a many-to-many scheme, with the help of shared word space as a bridge. These two steps are run iteratively and label information is transferred from source domain to target domain through the inferred cluster correspondence. Experiments on various real data sets demonstrate that our method outperforms several state-of-the-art approaches for cross-domain text classification. Mingsheng Long, Wei Cheng 0002, Xiaoming Jin, Jianmin Wang 0001, Dou Shen |
ICDM | 3 |
| 2009 | Probabilistic Similarity Query on Dimension Incomplete DataabstractRetrieving similar data has drawn many research efforts in the literature due to its importance in data mining, database and information retrieval. This problem is challenging when the data is incomplete. In previous research, data incompleteness refers to the fact that data values for some dimensions are unknown. However, in many practical applications (e.g., data collection by sensor network under bad environment), not only data values but even data dimension information may also be missing, which will make most similarity query algorithms infeasible. In this work, we propose the novel similarity query problem on dimension incomplete data and adopt a probabilistic framework to model this problem. For this problem, users can give a distance threshold and a probability threshold to specify their retrieval requirements. The distance threshold is used to specify the allowed distance between query and data objects and the probability threshold is used to require that the retrieval results satisfy the distance condition at least with the given probability. Instead of enumerating all possible cases to recover the missed dimensions, we propose an efficient approach to speed up the retrieval process by leveraging the inherent relations between query and dimension incomplete data objects. During the query process, we estimate the lower/upper bounds of the probability that the query is satisfied by a given data object, and utilize these bounds to filter irrelevant data objects efficiently. Furthermore, a probability triangle inequality is proposed to further speed up query processing. According to our experiments on real data sets, the proposed similarity query method is verified to be effective and efficient on dimension incomplete data. Wei Cheng 0002, Xiaoming Jin, Jian-Tao Sun |
ICDM | 2 |
| 2009 | Mining common topics from multiple asynchronous text streamsabstractText streams are becoming more and more ubiquitous, in the forms of news feeds, weblog archives and so on, which result in a large volume of data. An effective way to explore the semantic as well as temporal information in text streams is topic mining, which can further facilitate other knowledge discovery procedures. In many applications, we are facing multiple text streams which are related to each other and share common topics. The correlation among these streams can provide more meaningful and comprehensive clues for topic mining than those from each individual stream. However, it is nontrivial to explore the correlation with the existence of asynchronism among multiple streams, i.e. documents from different streams about the same topic may have different timestamps, which remains unsolved in the context of topic mining. In this paper, we formally address this problem and put forward a novel algorithm based on the generative topic model. Our algorithm consists of two alternate steps: the first step extracts common topics from multiple streams based on the adjusted timestamps by the second step; the second step adjusts the timestamps of the documents according to the time distribution of the discovered topics by the first step. We perform these two steps alternately and a monotone convergence of our objective function is guaranteed. The effectiveness and advantage of our approach were justified by extensive empirical studies on two real data sets consisting of six research paper streams and two news article streams, respectively. Xiang Wang 0002, Xiaoming Jin, Dou Shen |
WSDM | 3 |
| 2007 | Similarity Search over Incomplete Symbolic Sequences
Xiaoming Jin |
DEXA | 2 |
| 2007 | Concept Sampling: Towards Systematic Selection in Large-Scale Mixed Concepts in Machine Learning
Yi Zhang 0010, Xiaoming Jin |
IJCAI | 2 |
| 2006 | Understanding and Enhancing the Folding-In Method in Latent Semantic Indexing
Xiang Wang 0002, Xiaoming Jin |
DEXA | 2 |
| 2006 | Efficient Discovery of Emerging Frequent Patterns in ArbitraryWindows on Data StreamsabstractThis paper proposes an effective data mining technique for finding useful patterns in streaming sequences. At present, typical approaches to this problem are to search for patterns in a fixed-size window sliding through the stream of data being collected. The practical values of such approaches are limited in that, in typical application scenarios, the patterns are emerging and it is difficult, if not impossible, to determine a priori a suitable window size within which useful patterns may exist. It is therefore desirable to devise techniques that can identify useful patterns with arbitrary window sizes. Attempts to this problem are challenging, however, because it requires a highly efficient searching in a substantially bigger solution space. This paper presents a new method which includes firstly a pruning strategy to reduce the search space and secondly a mining strategy that adopts a dynamic index structure to allow efficient discovery of emerging patterns in a streaming sequence. Experimental results on real data and synthetic data show that the proposed method outperforms other existing schemes both in computational efficiency and effectiveness in finding useful patterns. Xiaoming Jin, Xinqiang Zuo, Kwok-Yan Lam, Jianmin Wang 0001, Jia-Guang Sun 0001 |
ICDE | 1 |
| 2006 | A Simple Approximation for Dynamic Time Warping Search in Large Time Series Database
Xiaoming Jin |
IDEAL | 2 |
| 2006 | Indexing and Mining of Graph Database Based on Interconnected Subgraph
Haichuan Shang, Xiaoming Jin |
IDEAL | 2 |
| 2006 | A Multi-Hierarchical Representation for Similarity Measurement of Time Series
Xinqiang Zuo, Xiaoming Jin |
PAKDD | 2 |
| 2005 | Watermarking Spatial Trajectory Database
Xiaoming Jin, Jianmin Wang 0001, Deyi Li |
DASFAA | 1 |
| 2005 | Accurate Symbolization of Time Series
Xinqiang Zuo, Xiaoming Jin |
PAKDD | 2 |
| 2004 | Dynamic Symbolization of Streaming Time Series
Xiaoming Jin, Jianmin Wang 0001, Jia-Guang Sun 0001 |
IDEAL | 1 |
| 2002 | Indexing and Mining of the Local Patterns in Sequence Database
Xiaoming Jin, Yuchang Lu, Chunyi Shi |
IDEAL | 1 |
| 2002 | Similarity measure based on partial information of time seriesabstractSimilarity measure of time series is an important subroutine in many KDD applications. Previous similarity models mainly focus on the prominent series behaviors by considering the whole information of time series. In this paper, we address the problem: which portion of information is more suitable for similarity measure for the data collected from a certain field. We propose a model for the retrieval and representation of the partial information in time series data, and a methodology for evaluating the similarity measurements based on partial information. The methodology is to retrieve various portions of information from the raw data and represent it in a concise form, then cluster the time series using the partial information and evaluate the similarity measurements through comparing the results with a standard classification. Experiments on data set from stock market give some interesting observations and justify the usefulness of our approach. Xiaoming Jin, Yuchang Lu, Chunyi Shi |
KDD | 1 |
| 2002 | Distribution Discovery: Local Analysis of Temporal Rules
Xiaoming Jin, Yuchang Lu, Chunyi Shi |
PAKDD | 1 |
| 2001 | Micro Similarity Queries in Time Series Database
Xiaoming Jin, Yuchang Lu, Chunyi Shi |
PAKDD | 1 |