EDBT 2026 Demo / reviewers in the wild / expert
Zhongzhi Shi
dblp:52/1213 · also Zhong-Zhi Shi
· DBLP profile ↗
174ranked-venue papers
9as first author
4since 2021 · last 2024
0000-0002-3280-1676ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 95 · 4 first-author · 3 since 2021Databases, data management, data science and information retrieval · 48 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 23 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 3 first-authorSystems, architecture and hardware · 6 · 1 first-authorComputer networks · 5 · 1 since 2021Software engineering, systems software and programming languages · 5Human-computer interaction and ubiquitous computing · 5Theory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
12 papers |
Information extraction and text analysis · 27% Transfer learning and domain adaptation · 18% Representation and self-supervised learning · 16% | |
| Databases, data mining, and information retrieval
8 papers |
Recommender systems · 53% Data mining · 44% Knowledge graphs · 2% | |
| Theoretical computer science
5 papers |
Information theory · 45% Algorithmic game theory and mechanism design · 31% Mathematical optimization · 21% | |
| Computer graphics and multimedia
2 papers |
Multimedia analysis and retrieval · 50% Image and video processing · 50% |
Topics — the 30 heaviest of 56, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis
text classification |
0.3 | 2 | 2013 | Triplex transfer learning: exploiting both shared and distinct concepts for text classification · WSDM 2013 Concept Learning for Cross-Domain Text Classification: A General Probabilistic Framework · IJCAI 2013 |
Recommender systems
social recommendation |
0.3 | 1 | 2018 | Leveraging Hypergraph Random Walk Tag Expansion and User Social Relation for Microblog Recommendation · ICDM 2018 |
Recommender systems › side information integration
tag-based recommendation |
0.3 | 1 | 2018 | Leveraging Hypergraph Random Walk Tag Expansion and User Social Relation for Microblog Recommendation · ICDM 2018 |
Machine learning › Transfer learning and domain adaptation › cross-domain learning
cross-domain text classification |
0.3 | 2 | 2013 | Concept Learning for Cross-Domain Text Classification: A General Probabilistic Framework · IJCAI 2013 Mining Distinction and Commonality across Multiple Domains Using Generative Model for Text Classification · IEEE Trans. Knowl. Data Eng. 2012 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference |
0.2 | 1 | 2015 | Bayesian Maximum Margin Principal Component Analysis · AAAI 2015 |
Machine learning › Representation and self-supervised learning › representation learning
dimensionality reduction |
0.2 | 1 | 2015 | Bayesian Maximum Margin Principal Component Analysis · AAAI 2015 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
principal component analysis |
0.2 | 1 | 2015 | Bayesian Maximum Margin Principal Component Analysis · AAAI 2015 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference |
0.2 | 1 | 2015 | Bayesian Maximum Margin Principal Component Analysis · AAAI 2015 |
Natural language and speech › Information extraction and text analysis
topic model |
0.2 | 2 | 2014 | Mining Distinction and Commonality across Multiple Domains Using Generative Model for Text Classification · IEEE Trans. Knowl. Data Eng. 2012 Ratable Aspects over Sentiments: Predicting Ratings for Unrated Reviews · ICDM 2014 |
Natural language and speech › Information extraction and text analysis › sentiment analysis › aspect-based sentiment analysis
aspect extraction |
0.2 | 1 | 2014 | Ratable Aspects over Sentiments: Predicting Ratings for Unrated Reviews · ICDM 2014 |
Data mining › text mining › sentiment analysis
review mining |
0.2 | 1 | 2014 | Ratable Aspects over Sentiments: Predicting Ratings for Unrated Reviews · ICDM 2014 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
concept learning |
0.2 | 1 | 2013 | Concept Learning for Cross-Domain Text Classification: A General Probabilistic Framework · IJCAI 2013 |
Algorithmic game theory and mechanism design › pricing
pricing mechanism |
0.2 | 1 | 2013 | Optimal Pricing for Improving Efficiency of Taxi Systems · IJCAI 2013 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › latent structure discovery
latent semantic learning |
0.1 | 1 | 2012 | Multi-task Semi-supervised Semantic Feature Learning for Classification · ICDM 2012 |
Natural language and speech › Information extraction and text analysis › topic model
probabilistic latent semantic analysis |
0.1 | 1 | 2012 | Mining Distinction and Commonality across Multiple Domains Using Generative Model for Text Classification · IEEE Trans. Knowl. Data Eng. 2012 |
Recommender systems › collaborative filtering
matrix factorization |
0.1 | 1 | 2012 | Multi-task Semi-supervised Semantic Feature Learning for Classification · ICDM 2012 |
Recommender systems › collaborative filtering › matrix factorization
nonnegative matrix tri-factorization |
0.1 | 1 | 2012 | Multi-task Semi-supervised Semantic Feature Learning for Classification · ICDM 2012 |
Multimedia analysis and retrieval
cross-modal retrieval |
0.1 | 1 | 2012 | Cross-media knowledge discovery · KDD 2012 |
Machine learning › Representation and self-supervised learning › representation learning › embedding learning
probabilistic embedding |
0.1 | 1 | 2011 | Combining Supervised and Unsupervised Models via Unconstrained Probabilistic Embedding · IJCAI 2011 |
Machine learning › Learning paradigms › multi-label classification
classifier chain |
0.1 | 1 | 2010 | Sorted label classifier chains for learning images with multi-label · ACM Multimedia 2010 |
Machine learning › Learning paradigms
multi-label classification |
0.1 | 1 | 2010 | Sorted label classifier chains for learning images with multi-label · ACM Multimedia 2010 |
Machine learning › Transfer learning and domain adaptation › domain adaptation
multi-source domain adaptation |
0.1 | 1 | 2010 | Cross-Domain Learning from Multiple Sources: A Consensus Regularization Perspective · IEEE Trans. Knowl. Data Eng. 2010 |
Data mining › text mining › topic modeling
latent dirichlet allocation |
0.1 | 1 | 2010 | D-LDA: A Topic Modeling Approach without Constraint Generation for Semi-defined Classification · ICDM 2010 |
Data mining › text mining › text classification › weakly supervised classification
semi-supervised classification |
0.1 | 1 | 2010 | D-LDA: A Topic Modeling Approach without Constraint Generation for Semi-defined Classification · ICDM 2010 |
Data mining › text mining
topic modeling |
0.1 | 1 | 2010 | D-LDA: A Topic Modeling Approach without Constraint Generation for Semi-defined Classification · ICDM 2010 |
Image and video processing › pattern detection
curve detection |
0.1 | 1 | 2010 | Nonparametric Curve Extraction Based on Ant Colony System · AAAI 2010 |
Machine learning › Graph learning › limited supervision › multi-view semi-supervised learning
co-training |
0.1 | 1 | 2009 | Coboost learning of visual categories with 1st and 2nd order features from Google images · ACM Multimedia 2009 |
Computer vision › Image recognition and object detection › object recognition › category recognition
object category modeling |
0.1 | 1 | 2009 | Coboost learning of visual categories with 1st and 2nd order features from Google images · ACM Multimedia 2009 |
Machine learning › Learning paradigms
semi-supervised learning |
0.1 | 1 | 2009 | Coboost learning of visual categories with 1st and 2nd order features from Google images · ACM Multimedia 2009 |
Natural language and speech › Information extraction and text analysis
web information extraction |
0.1 | 1 | 2009 | Towards combining web classification and web information extraction: a case study · KDD 2009 |
Methods — techniques the papers use, named apart from their topics
generative modeling · 0.4discriminative modeling · 0.4cognitive modeling · 0.4semi-supervised learning · 0.4latent dirichlet allocation · 0.4generative topic modeling · 0.4tag expansion · 0.3hypergraph random walk · 0.3alternating iterative optimization · 0.3mean-field variational inference · 0.2max-margin learning · 0.2hough transform · 0.2data augmentation · 0.2ant colony system · 0.2probabilistic framework · 0.2non-negative matrix tri-factorization · 0.2manifold regularization · 0.2double latent variable model · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Hybrid particle swarm optimization with adaptive learning strategy
Lanyu Wang, Dongping Tian, Xiaorui Gou, Zhongzhi Shi |
Soft Comput. | 4 |
| 2023 | Multiagent Reinforcement Learning With Heterogeneous Graph Attention NetworkabstractMost recent research on multiagent reinforcement learning (MARL) has explored how to deploy cooperative policies for homogeneous agents. However, realistic multiagent environments may contain heterogeneous agents that have different attributes or tasks. The heterogeneity of the agents and the diversity of relationships cause the learning of policy excessively tough. To tackle this difficulty, we present a novel method that employs a heterogeneous graph attention network to model the relationships between heterogeneous agents. The proposed method can generate an integrated feature representation for each agent by hierarchically aggregating latent feature information of neighbor agents, with the importance of the agent level and the relationship level being entirely considered. The method is agnostic to specific MARL methods and can be flexibly integrated with diverse value decomposition methods. We conduct experiments in predator-prey and StarCraft Multiagent Challenge (SMAC) environments, and the empirical results demonstrate that the performance of our method is superior to existing methods in several heterogeneous scenarios. Wei Du 0010, Shifei Ding, Chenglong Zhang 0001, Zhongzhi Shi |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | Matching images and texts with multi-head attention network for cross-media hashing retrieval
Zhixin Li 0001, Xiumin Xie, Huifang Ma, Zhongzhi Shi |
Eng. Appl. Artif. Intell. | 5 |
| 2021 | Integrating Scene Semantic Knowledge into Image CaptioningabstractMost existing image captioning methods use only the visual information of the image to guide the generation of captions, lack the guidance of effective scene semantic information, and the current visual attention mechanism cannot adjust the focus intensity on the image. In this article, we first propose an improved visual attention model. At each timestep, we calculated the focus intensity coefficient of the attention mechanism through the context information of the model, then automatically adjusted the focus intensity of the attention mechanism through the coefficient to extract more accurate visual information. In addition, we represented the scene semantic knowledge of the image through topic words related to the image scene, then added them to the language model. We used the attention mechanism to determine the visual information and scene semantic information that the model pays attention to at each timestep and combined them to enable the model to generate more accurate and scene-specific captions. Finally, we evaluated our model on Microsoft COCO (MSCOCO) and Flickr30k standard datasets. The experimental results show that our approach generates more accurate captions and outperforms many recent advanced models in various evaluation metrics. Haiyang Wei, Zhixin Li 0001, Feicheng Huang, Canlong Zhang, Huifang Ma, Zhongzhi Shi |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2020 | Multi-view RBM with posterior consistency and domain adaptation
Nan Zhang 0014, Shifei Ding, Tongfeng Sun, Hongmei Liao, Zhongzhi Shi |
Inf. Sci. | 6 |
| 2020 | On rule acquisition methods for data classification in heterogeneous incomplete decision systems
Zuqiang Meng, Zhongzhi Shi |
Knowl. Based Syst. | 2 |
| 2019 | A novel density peaks clustering with sensitivity of local density and density-adaptive metric
Mingjing Du 0001, Shifei Ding, Yu Xue 0004, Zhongzhi Shi |
Knowl. Inf. Syst. | 4 |
| 2019 | A multiway p-spectral clustering algorithm
Shifei Ding, Qiankun Hu, Hongjie Jia, Zhongzhi Shi |
Knowl. Based Syst. | 5 |
| 2018 | Leveraging Hypergraph Random Walk Tag Expansion and User Social Relation for Microblog RecommendationabstractRecommending valuable contents for microblog users is an important way to improve users' experiences. As high quality descriptors of user semantics, tags have always been used to represent users' interests or attributes. In this work, we propose a microblog recommendation approach via hypergraph random walk tag expansion and user social relation. More specifically, microblogs are considered as hyperedges and terms are taken as hypervertexs for each user, and the weighting strategies for both hyperedges and hypervertexs are established. Random walk is performed on the weighted hypergraph to obtain a number of terms as tags for users. And then the tag similarity matrix and the user-tag matrix can be constructed based on tag probability correlations and weight of each tag. Besides, the significance of user social relation is also considered for recommendation. Moreover, an iterative updating scheme is developed to get the final user-tag matrix for computing the similarities between microblogs and users. Experimental results show that the algorithm is effective in microblog recommendation. Huifang Ma, Weizhong Zhao, Zhongzhi Shi |
ICDM | 5 |
| 2018 | An improved density peaks clustering algorithm with fast finding cluster centers
Xiao Xu 0006, Shifei Ding, Zhongzhi Shi |
Knowl. Based Syst. | 3 |
| 2017 | Automatic image annotation based on Gaussian mixture model considering cross-modal correlations
Dongping Tian, Zhongzhi Shi |
J. Vis. Commun. Image Represent. | 2 |
| 2016 | Denoising Laplacian multi-layer extreme learning machine
Nan Zhang 0014, Shifei Ding, Zhongzhi Shi |
Neurocomputing | 3 |
| 2016 | Recent advances in Support Vector Machines
Shifei Ding, Zhongzhi Shi, Dacheng Tao, Bo An 0001 |
Neurocomputing | 2 |
| 2016 | Multi-feature fusion deep networks
Gang Ma 0001, Bo Zhang 0022, Zhongzhi Shi |
Neurocomputing | 4 |
| 2016 | On quick attribute reduction in decision-theoretic rough set models
Zuqiang Meng, Zhongzhi Shi |
Inf. Sci. | 2 |
| 2016 | Image matting in the perception granular deep learning
Hong Hu 0001, Liang Pang 0001, Zhongzhi Shi |
Knowl. Based Syst. | 3 |
| 2016 | On efficient methods of computing attribute-value blocks in incomplete decision systems
Zuqiang Meng, Qiuling Gan, Zhongzhi Shi |
Knowl. Based Syst. | 3 |
| 2015 | Bayesian Maximum Margin Principal Component AnalysisabstractSupervised dimensionality reduction has shown great advantages in finding predictive subspaces. Previous methods rarely consider the popular maximum margin principle and are prone to overfitting to usually small training data, especially for those under the maximum likelihood framework. In this paper, we present a posterior-regularized Bayesian approach to combine Principal Component Analysis (PCA) with the max-margin learning. Based on the data augmentation idea for max-margin learning and the probabilistic interpretation of PCA, our method can automatically infer the weight and penalty parameter of max-margin learning machine, while finding the most appropriate PCA subspace simultaneously under the Bayesian framework. We develop a fast mean-field variational inference algorithm to approximate the posterior. Experimental results on various classification tasks show that our method outperforms a number of competitors. Changying Du, Shandian Zhe, Fuzhen Zhuang, Yuan Qi 0001, Qing He 0003, Zhongzhi Shi |
AAAI | 6 |
| 2015 | An Environment Visual Awareness Approach in Cognitive Model ABGPabstractABGP is a special cognitive model, which consists of awareness, beliefs, goals and plans. As most agent architectures, ABGP agents obtain knowledge from the natural scenes only through single preestablished rules as well, don't directly capture the natural scenes information like human visual. Inspired by the biological visual cortex (V1) and the higher brain areas perceiving visual features, we propose a novel deep network model convolutional generative stochastic model (CGSM) used to visual feature representation, and firstly introduce it into the awareness module of the cognitive model ABGP to construct a state-of-the-art cognitive model ABGP-CGSM. For the novel cognitive model ABGP-CGSM, we construct a rat-robot maze search simulation platform to show the validity recognizing natural scenes. According to the simulation results on the noise and noiseless natural scenes, the rat-robot implemented by ABGP-CGSM has an excellent success rate when passing through the maze. The simulation shows that the ABGP-CGSM model proposed in our work can directly enhance the capability of communication between agent and natural scenes, improve the ability to cognize the real world as human being and conduct the agent to plan independently its path in terms of the visual information from the natural scenes. Gang Ma 0001, Bo Zhang 0003, Baoyuan Qi, Zhongzhi Shi |
ICTAI | 5 |
| 2015 | A new solution algorithm for solving rule-sets based bilevel decision problemsabstractSummary Bilevel decision addresses compromises between two interacting decision entities within a given hierarchical complex system under distributed environments. Bilevel programming typically solves bilevel decision problems. However, formulation of objectives and constraints in mathematical functions is required, which are difficult, and sometimes impossible, in real‐world situations because of various uncertainties. Our study develops a rule‐set based bilevel decision approach, which models a bilevel decision problem by creating, transforming and reducing related rule sets. This study develops a new rule‐sets based solution algorithm to obtain an optimal solution from the bilevel decision problem described by rule sets. A case study and a set of experiments illustrate both functions and the effectiveness of the developed algorithm in solving a bilevel decision problem. Copyright © 2012 John Wiley & Sons, Ltd. Jie Lu 0001, Zheng Zheng 0001, Guangquan Zhang 0001, Qing He 0003, Zhongzhi Shi |
Concurr. Comput. Pract. Exp. | 5 |
| 2015 | Parallel sampling from big data with uncertainty distribution
Qing He 0003, Fuzhen Zhuang, Tianfeng Shang, Zhongzhi Shi |
Fuzzy Sets Syst. | 5 |
| 2015 | Learning deep representations via extreme learning machines
Wenchao Yu, Fuzhen Zhuang, Qing He 0003, Zhongzhi Shi |
Neurocomputing | 4 |
| 2015 | QPLSA: Utilizing quad-tuples for aspect identification and rating
Wenjuan Luo, Fuzhen Zhuang, Weizhong Zhao, Qing He 0003, Zhongzhi Shi |
Inf. Process. Manag. | 5 |
| 2014 | Ratable Aspects over Sentiments: Predicting Ratings for Unrated ReviewsabstractMost existing rat able aspect generating methods for aspect mining focus on identifying and rating aspects of reviews with overall ratings, while huge amount of unrated reviews are beyond their ability. This drawback motivates the research problem in this paper: predicting aspect ratings and overall ratings for unrated reviews. To solve this problem, we novelly propose a topic model based on Latent Dirichlet Allocation with indirect supervision. Compared with the previous bag-of-words representation of review documents, we utilize the quad-tuples of (head, modifier, rating, entity) to explicitly model the associations between modifiers and ratings. Specifically, our solution for aspect mining in unrated reviews is decomposed into three steps. Firstly, rat able aspects are generated over sentiments from training reviews with overall ratings. Afterwards, inference of aspect identification and rating for unrated reviews are provided. Finally, overall ratings are predicted for unrated reviews. Under this framework, aspect and sentiment associations are captured in the form of joint probabilities through a generative process. The effectiveness of our approach is testified on a real-world dataset crawled from Trip Advisor http://www.tripadvisor.com/, and extensive experiments show that our method significantly outperforms state-of-the-art methods. Wenjuan Luo, Fuzhen Zhuang, Xiaohu Cheng, Qing He 0003, Zhongzhi Shi |
ICDM | 5 |
| 2014 | Cognitive memory systems in consciousness and memory modelabstractMemory is a fundamental component in human brain and plays very important roles for all mental processes. The analysis of memory systems through cognitive architectures can be performed at the computational, or functional level, on the basis of empirical data. In this paper we discuss memory systems in the extended Consciousness and Memory Model (CAM) The knowledge representations used in CAM for working memory, semantic memory, episodic memory and procedural memory are introduced. It will be explained how, in CAM, all of these knowledge types are represented in dynamic decription logic (DDL), a formal logic with the capability for description and reasoning regarding dynamic application domains characterized by actions. Zhongzhi Shi |
IJCNN | 1 |
| 2014 | Balanced Seed Selection for Budgeted Influence Maximization in Social Networks
Fuzhen Zhuang, Qing He 0003, Zhongzhi Shi |
PAKDD (1) | 4 |
| 2014 | Transfer Learning with Multiple Sources via Consensus Regularized Autoencoders
Fuzhen Zhuang, Xiaohu Cheng, Sinno Jialin Pan, Wenchao Yu, Qing He 0003, Zhongzhi Shi |
ECML/PKDD (3) | 6 |
| 2014 | Scalable bootstrap clustering for massive dataabstractThe bootstrap provides a simple and powerful means of improving the accuracy of clustering. However, for today's increasingly large datasets, the computation of bootstrap-based quantities can be prohibitively demanding. In this paper we introduce the Bag of Little Bootstraps Clustering (BLBC), a new procedure which utilizes the Bag of Little Bootstraps technique to obtain a robust, computationally efficient means of clustering for massive data. Moreover, BLBC is suited to implementation on modern parallel and distributed computing architectures which are often used to process large datasets. We investigate empirically the performance characteristics of BLBC and compare to the performances of existing methods via experiments on simulated data and real data. The results show that BLBC has a significantly more favorable computational profile than the bootstrap based clustering while maintaining good statistical correctness. Fuzhen Zhuang, Xiang Ao 0001, Qing He 0003, Zhongzhi Shi |
SNPD | 5 |
| 2014 | Perception granular computing in visual haze-free task
Hong Hu 0001, Liang Pang 0001, Dongping Tian, Zhongzhi Shi |
Expert Syst. Appl. | 4 |
| 2014 | Track on Intelligent Computing and Applications: Selected papers from the 2012 International Workshop on Information, Intelligence and Computing (IWIIC 2012)
Shifei Ding, Zhongzhi Shi |
Neurocomputing | 2 |
| 2014 | Clustering in extreme learning machine feature space
Qing He 0003, Xin Jin 0004, Changying Du, Fuzhen Zhuang, Zhongzhi Shi |
Neurocomputing | 5 |
| 2014 | Combining supervised and unsupervised models via unconstrained probabilistic embedding
Xiang Ao 0001, Ping Luo 0001, Xudong Ma, Fuzhen Zhuang, Qing He 0003, Zhongzhi Shi, Zhiyong Shen |
Inf. Sci. | 6 |
| 2014 | Pessimistic rough set based decisions: A multigranulation fusion strategy
Shunyong Li, Jiye Liang, Zhongzhi Shi, Feng Wang 0038 |
Inf. Sci. | 4 |
| 2014 | Interaction relationships of caches in agent-based HD video surveillance: Discovery and utilization
Wenjia Niu, Gang Li 0009, Endong Tong, Xinghua Yang, Liang Chang 0003, Zhongzhi Shi, Song Ci |
J. Netw. Comput. Appl. | 6 |
| 2014 | Bloom filter-based workflow management to enable QoS guarantee in wireless sensor networks
Endong Tong, Wenjia Niu, Gang Li 0009, Ding Tang, Liang Chang 0003, Zhongzhi Shi, Song Ci |
J. Netw. Comput. Appl. | 6 |
| 2014 | A novel self-adaptive extreme learning machine based on affinity propagation for radial basis function neural network
Shifei Ding, Gang Ma 0001, Zhongzhi Shi |
Neural Comput. Appl. | 3 |
| 2014 | Wavelet twin support vector machine
Shifei Ding, Fulin Wu, Zhongzhi Shi |
Neural Comput. Appl. | 3 |
| 2014 | A Rough RBF Neural Network Based on Weighted Regularized Extreme Learning Machine
Shifei Ding, Gang Ma 0001, Zhongzhi Shi |
Neural Process. Lett. | 3 |
| 2014 | Triplex Transfer Learning: Exploiting Both Shared and Distinct Concepts for Text ClassificationabstractTransfer learning focuses on the learning scenarios when the test data from target domains and the training data from source domains are drawn from similar but different data distributions with respect to the raw features. Along this line, some recent studies revealed that the high-level concepts, such as word clusters, could help model the differences of data distributions, and thus are more appropriate for classification. In other words, these methods assume that all the data domains have the same set of shared concepts, which are used as the bridge for knowledge transfer. However, in addition to these shared concepts, each domain may have its own distinct concepts. In light of this, we systemically analyze the high-level concepts, and propose a general transfer learning framework based on nonnegative matrix trifactorization, which allows to explore both shared and distinct concepts among all the domains simultaneously. Since this model provides more flexibility in fitting the data, it can lead to better classification accuracy. Moreover, we propose to regularize the manifold structure in the target domains to improve the prediction performances. To solve the proposed optimization problem, we also develop an iterative algorithm and theoretically analyze its convergence properties. Finally, extensive experiments show that the proposed model can outperform the baseline methods with a significant margin. In particular, we show that our method works much better for the more challenging tasks when there are distinct concepts in the data. Fuzhen Zhuang, Ping Luo 0001, Changying Du, Qing He 0003, Zhongzhi Shi, Hui Xiong 0001 |
IEEE Trans. Cybern. | 5 |
| 2014 | A co-boost framework for learning object categories from Google Images with 1st and 2nd order features
Xi Liu 0009, Zhi-Ping Shi 0002, Zhongzhi Shi |
Vis. Comput. | 3 |
| 2013 | A New Similarity Measure Based on Preference Sequences for Collaborative Filtering
Tianfeng Shang, Qing He 0003, Fuzhen Zhuang, Zhongzhi Shi |
APWeb | 4 |
| 2013 | Classification of big velocity data via cross-domain Canonical Correlation AnalysisabstractMany classification techniques work well only under a common assumption that the training and test data are drawn from the same feature space and the same distribution. However, big velocity data usually show disobedience of this assumption. For example, in the field of web-document classification, new document is continuously emerging every day. Transfer learning aims at leveraging the knowledge in labeled source domains to predict the unlabeled data in a target domain, where the distributions are different in domains. As one of the important research directions of transfer learning, one kind of approaches focus on the correspondence between pivot features and all the other specific features from different domains, to extract some relevant features that may reduce the difference between the domains, have attracted wide attention and study. However, the limitation caused by the vague meanings in different domains prevents these algorithms from further improvement. To tackle this problem, we propose a cross-domain canonical correlation analysis algorithm called CD-CCA by applying Canonical Correlation Analysis (CCA) to transfer learning. CD-CCA can learn a semantic space of multi-view correspondences from different domains respectively and transfer the knowledge by dimensionality reduction in a multi-view way. Experimental results on the 144×6 classification problems in 20Newsgroups, show that CD-CCA can significantly improve the prediction accuracy. Bo Zhang 0003, Zhongzhi Shi |
IEEE BigData | 2 |
| 2013 | Automatic image annotation via local sparse codingabstractSparse coding is an active research topic in machine learning and signal processing community. In this paper, we propose a novel local sparse model for multi-label image annotation. Existing feature descriptors and extraction algorithms pay less attention to semantic information and extracted feature dimension usually is high, which leads to heavy computation. Noise and redundant information often reduce the performance of sparse model. To address these issues, we combine label and visual information for feature selection while most previous work only utilizes labels and ignores visual information itself. First of all, we make use of label sets to seek images neighbor relations and generate the Gaussian kernel matrix over these neighbor images, then use LLP(Local Learning Projection) algorithm to get minimal local estimation error. After that, for each query image, we find its K nearest neighbors in the transformed space and use these neighbors to reconstruct it via sparse coding. Moreover, during coding, we penalize the corresponding reconstruction coefficients to implicitly reflect the neighbor relations. Finally, propagating tags from training data to test data. Image annotation experiments on the Corel5k dataset show the performance of our approach is comparable to several state-of-the-art algorithms. Dongping Tian, Hong Hu 0001, Zhongzhi Shi |
ICASSP | 5 |
| 2013 | Employing PLSA model and max-bisection for refining image annotationabstractWe present a new method for refining image annotation by fusing probabilistic latent semantic analysis (PLSA) with max-bisection (MB). We first construct a PLSA model with asymmetric modalities to estimate the posterior probabilities of each annotating keyword for an image, and then a label similarity graph is built by a weighted linear combination of label similarity and visual similarity. Followed by the rank-two relaxation heuristics over the constructed label graph is employed to further mine the correlation of the keywords so as to capture the refining annotation, which plays a critical role in semantic based image retrieval. The novelty of our method mainly lies in two aspects: exploiting PLSA to accomplish the initial semantic annotation task and implementing max-bisection based on the rank-two relaxation algorithm over the weighted label graph to refine the candidate annotations generated by the PLSA. We evaluate our method on the standard Corel dataset and the experimental results are competitive to several state-of-the-art approaches. Dongping Tian, Zhongzhi Shi |
ICIP | 4 |
| 2013 | Optimal Pricing for Improving Efficiency of Taxi Systems
Jiarui Gan, Bo An 0001, Haizhong Wang, Zhongzhi Shi |
IJCAI | 5 |
| 2013 | Concept Learning for Cross-Domain Text Classification: A General Probabilistic Framework
Fuzhen Zhuang, Ping Luo 0001, Peifeng Yin, Qing He 0003, Zhongzhi Shi |
IJCAI | 5 |
| 2013 | Extreme Learning Machine combining matrix factorization for collaborative filteringabstractCollaborative Filtering (CF) is one of the most popular techniques for information filtering in recommendation systems. Currently, there are many linear and nonlinear regression algorithms for CF. However, to our knowledge, these regression algorithms may not give satisfactory results in some practical applications. In this paper, Extreme Learning Machine (ELM), which is famous with its fast speed and good performance in generalization, is firstly employed to build a nonlinear regression model for CF, namely ELM for CF (ELMCF) algorithm. Then by combining ELM and Weighted Nonnegative Matrix Tri-Factorization (WNMTF), which can alleviate the data sparsity problem of the user-item matrix, a new nonlinear regression model is proposed, namely Extreme Learning Machine Combining Matrix Factorization for Collaborative Filtering (CELMCF) algorithm, to construct regression based CF algorithms and improve the performance of recommendation systems. Experiments are conducted on several benchmark datasets from different application domains. Experimental results show that the proposed CELMCF algorithm outperforms some state-of-the-art regression based CF algorithms (including ELMCF algorithm, Linear Regression for CF (LRCF) algorithm and Memory based CF (MemCF) algorithm) more efficiently with the competitive effectiveness. Tianfeng Shang, Qing He 0003, Fuzhen Zhuang, Zhongzhi Shi |
IJCNN | 4 |
| 2013 | Exploit Spatial Relationships among Pixels for Saliency Region Detection Using Topic Model
Guang Jiang, Xi Liu 0009, Jinpeng Yue, Zhongzhi Shi |
MMM (1) | 4 |
| 2013 | Refining Image Annotation by Integrating PLSA with Random Walk Model
Dongping Tian, Zhongzhi Shi |
MMM (1) | 3 |
| 2013 | Shared Structure Learning for Multiple Tasks with Multiple Views
Xin Jin 0004, Fuzhen Zhuang, Shuhui Wang, Qing He 0003, Zhongzhi Shi |
ECML/PKDD (2) | 5 |
| 2013 | Embedding with Autoencoder Regularization
Wenchao Yu, Guangxiang Zeng, Ping Luo 0001, Fuzhen Zhuang, Qing He 0003, Zhongzhi Shi |
ECML/PKDD (3) | 6 |
| 2013 | Triplex transfer learning: exploiting both shared and distinct concepts for text classificationabstractTransfer learning focuses on the learning scenarios when the test data from target domains and the training data from source domains are drawn from similar but different data distributions with respect to the raw features. Along this line, some recent studies revealed that the high-level concepts, such as word clusters, could help model the differences of data distributions, and thus are more appropriate for classification. In other words, these methods assume that all the data domains have the same set of shared concepts, which are used as the bridge for knowledge transfer. However, in addition to these shared concepts, each domain may have its own distinct concepts. In light of this, we systemically analyze the high-level concepts, and propose a general transfer learning framework based on nonnegative matrix trifactorization, which allows to explore both shared and distinct concepts among all the domains simultaneously. Since this model provides more flexibility in fitting the data, it can lead to better classification accuracy. Moreover, we propose to regularize the manifold structure in the target domains to improve the prediction performances. To solve the proposed optimization problem, we also develop an iterative algorithm and theoretically analyze its convergence properties. Finally, extensive experiments show that the proposed model can outperform the baseline methods with a significant margin. In particular, we show that our method works much better for the more challenging tasks when there are distinct concepts in the data. Fuzhen Zhuang, Ping Luo 0001, Changying Du, Qing He 0003, Zhongzhi Shi |
WSDM | 5 |
| 2013 | Learning semantic concepts from image database with hybrid generative/discriminative approach
Zhixin Li 0001, Zhongzhi Shi, Weizhong Zhao, Zhenjun Tang |
Eng. Appl. Artif. Intell. | 2 |
| 2013 | Parallel extreme learning machine for regression based on MapReduce
Qing He 0003, Tianfeng Shang, Fuzhen Zhuang, Zhongzhi Shi |
Neurocomputing | 4 |
| 2013 | Primal least squares twin support vector regressionabstractThe training algorithm of classical twin support vector regression (TSVR) can be attributed to the solution of a pair of quadratic programming problems (QPPs) with inequality constraints in the dual space. However, this solution is affected by time and memory constraints when dealing with large datasets. In this paper, we present a least squares version for TSVR in the primal space, termed primal least squares TSVR (PLSTSVR). By introducing the least squares method, the inequality constraints of TSVR are transformed into equality constraints. Furthermore, we attempt to directly solve the two QPPs with equality constraints in the primal space instead of the dual space; thus, we need only to solve two systems of linear equations instead of two QPPs. Experimental results on artificial and benchmark datasets show that PLSTSVR has comparable accuracy to TSVR but with considerably less computational time. We further investigate its validity in predicting the opening price of stock. Huajuan Huang, Shifei Ding, Zhongzhi Shi |
J. Zhejiang Univ. Sci. C | 3 |
| 2013 | A nonnegative matrix factorization framework for semi-supervised document clustering with dual constraints
Huifang Ma, Weizhong Zhao, Zhongzhi Shi |
Knowl. Inf. Syst. | 3 |
| 2013 | Semantic trajectory-based event detection and event pattern mining
Gang Li 0009, Guang Jiang, Zhongzhi Shi |
Knowl. Inf. Syst. | 4 |
| 2013 | Exploiting relevance, coverage, and novelty for query-focused multi-document summarization
Wenjuan Luo, Fuzhen Zhuang, Qing He 0003, Zhongzhi Shi |
Knowl. Based Syst. | 4 |
| 2012 | Multi-task Semi-supervised Semantic Feature Learning for ClassificationabstractMulti-task learning has proven to be useful to boost the learning of multiple related but different tasks. Meanwhile, latent semantic models such as LSA and LDA are popular and effective methods to extract discriminative semantic features of high dimensional dyadic data. In this paper, we present a method to combine these two techniques together by introducing a new matrix tri-factorization based formulation for semi-supervised latent semantic learning, which can incorporate labeled information into traditional unsupervised learning of latent semantics. Our inspiration for multi-task semantic feature learning comes from two facts, i.e., 1) multiple tasks generally share a set of common latent semantics, and 2) a semantic usually has a stable indication of categories no matter which task it is from. Thus to make multiple tasks learn from each other we wish to share the associations between categories and those common semantics among tasks. Along this line, we propose a novel joint Nonnegative matrix tri-factorization framework with the aforesaid associations shared among tasks in the form of a semantic-category relation matrix. Our new formulation for multi-task learning can simultaneously learn (1) discriminative semantic features of each task, (2) predictive structure and categories of unlabeled data in each task, (3) common semantics shared among tasks and specific semantics exclusive to each task. We give alternating iterative algorithm to optimize our objective and theoretically show its convergence. Finally extensive experiments on text data along with the comparison with various baselines and three state-of-the-art multi-task learning algorithms demonstrate the effectiveness of our method. Changying Du, Fuzhen Zhuang, Qing He 0003, Zhongzhi Shi |
ICDM | 4 |
| 2012 | Fast Time Series Classification Based on Infrequent ShapeletsabstractTime series shapelets are small and local time series subsequences which are in some sense maximally representative of a class. E.Keogh uses distance of the shapelet to classify objects. Even though shapelet classification can be interpretable and more accurate than many state-of-the-art classifiers, there is one main limitation of shapelets, i.e. shapelet classification training process is offline, and uses subsequence early abandon and admissible entropy pruning strategies, the time to compute is still significant. In this work, we address the later problem by introducing a novel algorithm that finds time series shapelet in significantly less time than the current methods by extracting infrequent time series shapelet candidates. Subsequences that are distinguishable are usually infrequent compared to other subsequences. The algorithm called ISDT (Infrequent Shapelet Decision Tree) uses infrequent shapelet candidates extracting to find shapelet. Experiments demonstrate the efficiency of ISDT algorithm on several benchmark time series datasets. The result shows that ISDT significantly outperforms the current shapelet algorithm. Qing He 0003, Zhi Dong, Fuzhen Zhuang, Tianfeng Shang, Zhongzhi Shi |
ICMLA (1) | 5 |
| 2012 | Parallel Decision Tree with Application to Water Quality Data Analysis
Qing He 0003, Zhi Dong, Fuzhen Zhuang, Tianfeng Shang, Zhongzhi Shi |
ISNN (2) | 5 |
| 2012 | Cross-media knowledge discoveryabstractIn this talk I introduce cloud computing based cross-media knowledge discovery. We propose a framework for cross-media semantic understanding which contains discriminative modeling, generative modeling and cognitive modeling. In cognitive modeling a new model entitled CAM is proposed which is suitable for cross-media semantic understanding. We develop an agent-aid model for load balance in cloud computing environment. For quality of service we present a utility function to evaluate the cloud performance. A Cross-Media Intelligent Retrieval System (CMIRS), which is managed by ontology-based knowledge system KMSphere, will be illustrated. Finally, the directions for further researches on cloud computing based cross-media knowledge discovery will be pointed out and discussed. Zhongzhi Shi |
KDD | 1 |
| 2012 | Quad-tuple PLSA: Incorporating Entity and Its Rating in Aspect Identification
Wenjuan Luo, Fuzhen Zhuang, Qing He 0003, Zhongzhi Shi |
PAKDD (1) | 4 |
| 2012 | Parallel Implementation of Apriori Algorithm Based on MapReduceabstractSearching frequent patterns in transactional databases is considered as one of the most important data mining problems and Apriori is one of the typical algorithms for this task. Developing fast and efficient algorithms that can handle large volumes of data becomes a challenging task due to the large databases. In this paper, we implement a parallel Apriori algorithm based on MapReduce, which is a framework for processing huge datasets on certain kinds of distributable problems using a large number of computers (nodes). The experimental results demonstrate that the proposed algorithm can scale well and efficiently process large datasets on commodity hardware. Qing He 0003, Zhongzhi Shi |
SNPD | 4 |
| 2012 | Extended rough set-based attribute reduction in inconsistent incomplete decision systems
Zuqiang Meng, Zhongzhi Shi |
Inf. Sci. | 2 |
| 2012 | Multi-view learning via probabilistic latent semantic analysis
Fuzhen Zhuang, George Karypis, Xia Ning, Qing He 0003, Zhongzhi Shi |
Inf. Sci. | 5 |
| 2012 | A Family of Dynamic Description Logics for Representing and Reasoning About Actions
Liang Chang 0003, Zhongzhi Shi, Tianlong Gu, Lingzhong Zhao |
J. Autom. Reason. | 2 |
| 2012 | Optimizing radial basis function neural network based on rough sets and affinity propagation clustering algorithmabstractA novel method based on rough sets (RS) and the affinity propagation (AP) clustering algorithm is developed to optimize a radial basis function neural network (RBFNN). First, attribute reduction (AR) based on RS theory, as a preprocessor of RBFNN, is presented to eliminate noise and redundant attributes of datasets while determining the number of neurons in the input layer of RBFNN. Second, an AP clustering algorithm is proposed to search for the centers and their widths without a priori knowledge about the number of clusters. These parameters are transferred to the RBF units of RBFNN as the centers and widths of the RBF function. Then the weights connecting the hidden layer and output layer are evaluated and adjusted using the least square method (LSM) according to the output of the RBF units and desired output. Experimental results show that the proposed method has a more powerful generalization capability than conventional methods for an RBFNN. Xinzheng Xu, Shifei Ding, Zhongzhi Shi, Hong Zhu 0005 |
J. Zhejiang Univ. Sci. C | 3 |
| 2012 | Effective semi-supervised document clustering via active learning with instance-level constraints
Weizhong Zhao, Qing He 0003, Huifang Ma, Zhongzhi Shi |
Knowl. Inf. Syst. | 4 |
| 2012 | Extracting discriminative features for CBIR
Zhi-Ping Shi 0002, Xi Liu 0009, Qingyong Li, Qing He 0003, Zhongzhi Shi |
Multim. Tools Appl. | 5 |
| 2012 | A feature binding computational model for multi-class object categorization and recognition
Xishun Wang, Xi Liu 0009, Zhongzhi Shi, Hongjian Sui |
Neural Comput. Appl. | 3 |
| 2012 | Mining Distinction and Commonality across Multiple Domains Using Generative Model for Text ClassificationabstractThe distribution difference among multiple domains has been exploited for cross-domain text categorization in recent years. Along this line, we show two new observations in this study. First, the data distribution difference is often due to the fact that different domains use different index words to express the same concept. Second, the association between the conceptual feature and the document class can be stable across domains. These two observations actually indicate the distinction and commonality across domains. Inspired by the above observations, we propose a generative statistical model, named Collaborative Dual-PLSA (CD-PLSA), to simultaneously capture both the domain distinction and commonality among multiple domains. Different from Probabilistic Latent Semantic Analysis (PLSA) with only one latent variable, the proposed model has two latent factors y and z, corresponding to word concept and document class, respectively. The shared commonality intertwines with the distinctions over multiple domains, and is also used as the bridge for knowledge transformation. An Expectation Maximization (EM) algorithm is developed to solve the CD-PLSA model, and further its distributed version is exploited to avoid uploading all the raw data to a centralized location and help to mitigate privacy concerns. After the training phase with all the data from multiple domains we propose to refine the immediate outputs using only the corresponding local data. In summary, we propose a two-phase method for cross-domain text classification, the first phase for collaborative training with all the data, and the second step for local refinement. Finally, we conduct extensive experiments over hundreds of classification tasks with multiple source domains and multiple target domains to validate the superiority of the proposed method over existing state-of-the-art methods of supervised and transfer learning. It is noted to mention that as shown by the experimental results CD-PLSA for the collaborative training is more tolerant of distribution differences, and the local refinement also gains significant improvement in terms of classification accuracy. Fuzhen Zhuang, Ping Luo 0001, Zhiyong Shen, Qing He 0003, Yuhong Xiong, Zhongzhi Shi, Hui Xiong 0001 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2012 | Erratum to "Mining Distinction and Commonality across Multiple Domains Using Generative Model for Text Classification"abstractThe distribution difference among multiple domains has been exploited for cross-domain text categorization in recent years. Along this line, we show two new observations in this study. First, the data distribution difference is often due to the fact that different domains use different index words to express the same concept. Second, the association between the conceptual feature and the document class can be stable across domains. These two observations actually indicate the distinction and commonality across domains. Inspired by the above observations, we propose a generative statistical model, named Collaborative Dual-PLSA (CD-PLSA), to simultaneously capture both the domain distinction and commonality among multiple domains. Different from Probabilistic Latent Semantic Analysis (PLSA) with only one latent variable, the proposed model has two latent factors y and z, corresponding to word concept and document class, respectively. The shared commonality intertwines with the distinctions over multiple domains, and is also used as the bridge for knowledge transformation. An Expectation Maximization (EM) algorithm is developed to solve the CD-PLSA model, and further its distributed version is exploited to avoid uploading all the raw data to a centralized location and help to mitigate privacy concerns. After the training phase with all the data from multiple domains we propose to refine the immediate outputs using only the corresponding local data. In summary, we propose a two-phase method for cross-domain text classification, the first phase for collaborative training with all the data, and the second step for local refinement. Finally, we conduct extensive experiments over hundreds of classification tasks with multiple source domains and multiple target domains to validate the superiority of the proposed method over existing state-of-the-art methods of supervised and transfer learning. It is noted to mention that as shown by the experimental results CD-PLSA for the collaborative training is more tolerant of distribution differences, and the local refinement also gains significant improvement in terms of classification accuracy. Fuzhen Zhuang, Ping Luo 0001, Zhiyong Shen, Qing He 0003, Yuhong Xiong, Zhongzhi Shi, Hui Xiong 0001 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2011 | Parallel Outlier Detection Using KD-Tree Based on MapReduceabstractDistributed and Parallel algorithms have attracted a vast amount of interest and research in recent decades, to handle large-scale data set in real-world applications. In this paper, we focus on a parallel implementation of KD-Tree based outlier detection method to deal with large-scale data set. As one of the state-of-the-art outlier detection methods, KD-Tree based has been approved to be an effective algorithm. However, it still cannot process large-scale data set efficiently due to its serial implementation. Based on the current and powerful parallel programming framework -- MapReduce, we propose to implement the parallel KD-Tree based outlier detection algorithm (e.g., PKDTree for short). Experimental results demonstrate the efficiency of PKDTree according to the evaluation criterions of scale up, speedup and size up. Qing He 0003, Fuzhen Zhuang, Zhongzhi Shi |
CloudCom | 5 |
| 2011 | Combining Supervised and Unsupervised Models via Unconstrained Probabilistic Embedding
Xudong Ma, Ping Luo 0001, Fuzhen Zhuang, Qing He 0003, Zhongzhi Shi, Zhiyong Shen |
IJCAI | 5 |
| 2011 | An approach for adaptive associative classification
Kun Yue, Wenjia Niu, Zhongzhi Shi |
Expert Syst. Appl. | 4 |
| 2011 | A parallel incremental extreme SVM classifier
Qing He 0003, Changying Du, Fuzhen Zhuang, Zhongzhi Shi |
Neurocomputing | 5 |
| 2011 | CARSA: A context-aware reasoning-based service agent model for AI planning of web service composition
Wenjia Niu, Gang Li 0009, Hui Tang 0001, Zhongzhi Shi |
J. Netw. Comput. Appl. | 5 |
| 2011 | Multi-granularity context model for dynamic Web service composition
Wenjia Niu, Gang Li 0009, Zhijun Zhao, Hui Tang 0001, Zhongzhi Shi |
J. Netw. Comput. Appl. | 5 |
| 2011 | Modeling continuous visual features for semantic image annotation and retrieval
Zhixin Li 0001, Zhi-Ping Shi 0002, Xi Liu 0009, Zhongzhi Shi |
Pattern Recognit. Lett. | 4 |
| 2011 | CHSMST: a clustering algorithm based on hyper surface and minimum spanning tree
Qing He 0003, Weizhong Zhao, Zhongzhi Shi |
Soft Comput. | 3 |
| 2010 | Nonparametric Curve Extraction Based on Ant Colony SystemabstractCurve extraction is an important and basic technique in image processing and computer vision. Due to the complexity of the images and the limitation of segmentation algorithms, there are always a large number of noisy pixels in the segmented binary images. In this paper, we present an approach based on ant colony system (ACS) to detect nonparametric curves from a binary image containing discontinuous curves and noisy points. Compared with the well-known Hough transform (HT) method, the ACS-based curve extraction approach can deal with both regular and irregular curves without knowing their shapes in advance. The proposed approach has many characteristics such as faster convergence, implicit parallelism and strong ability to deal with highly-noised images. Moreover, our approach can extract multiple curves from an image, which is impossible for the previous genetic algorithm based approach. Experimental results show that the proposed ACS-based approach is effective and efficient. Qing Tan, Qing He 0003, Zhongzhi Shi |
AAAI | 3 |
| 2010 | Combining the Missing Link: An Incremental Topic Model of Document Content and HyperlinkabstractThe content and structure of linked information such as sets of web pages or research paper archives are dynamic and keep on changing. Even though different methods are proposed to exploit both the link structure and the content information, no existing approach can effectively deal with this evolution. We propose a novel joint model, called Link-IPLSI, to combine texts and links in a topic modeling framework incrementally. The model takes advantage of a novel link updating technique that can cope with dynamic changes of online document streams in a faster and scalable way. Furthermore, an adaptive asymmetric learning method is adopted to freely control the assignment of weights to terms and citations. Experimental results on two different sources of online information demonstrate the time saving strength of our method and indicate that our model leads to systematic improvements in the quality of classification and link prediction. Huifang Ma, Zhixin Li 0001, Zhongzhi Shi |
APWeb | 3 |
| 2010 | Collaborative Dual-PLSA: mining distinction and commonality across multiple domains for text classificationabstractThe distribution difference among multiple data domains has been considered for the cross-domain text classification problem. In this study, we show two new observations along this line. First, the data distribution difference may come from the fact that different domains use different key words to express the same concept. Second, the association between this conceptual feature and the document class may be stable across domains. These two issues are actually the distinction and commonality across data domains. Inspired by the above observations, we propose a generative statistical model, named Collaborative Dual-PLSA (CD-PLSA), to simultaneously capture both the domain distinction and commonality among multiple domains. Different from Probabilistic Latent Semantic Analysis (PLSA) with only one latent variable, the proposed model has two latent factors y and z, corresponding to word concept and document class respectively. The shared commonality intertwines with the distinctions over multiple domains, and is also used as the bridge for knowledge transformation. We exploit an Expectation Maximization (EM) algorithm to learn this model, and also propose its distributed version to handle the situation where the data domains are geographically separated from each other. Finally, we conduct extensive experiments over hundreds of classification tasks with multiple source domains and multiple target domains to validate the superiority of the proposed CD-PLSA model over existing state-of-the-art methods of supervised and transfer learning. In particular, we show that CD-PLSA is more tolerant of distribution differences. Fuzhen Zhuang, Ping Luo 0001, Zhiyong Shen, Qing He 0003, Yuhong Xiong, Zhongzhi Shi, Hui Xiong 0001 |
CIKM | 6 |
| 2010 | Automatic image annotation with continuous PLSAabstractAutomatic image annotation has become an important and challenging problem due to the existence of semantic gap. In this paper, we firstly extend probabilistic latent semantic analysis (PLSA) to model continuous quantity. In addition, corresponding Expectation-Maximization (EM) algorithm is derived to determine the model parameters. Furthermore, in order to deal with the data of different modalities in terms of their characteristics, we present a semantic annotation model which employs continuous PLSA and standard PLSA to model visual features and textual words respectively. The model learns the correlation between these two modalities by an asymmetric learning approach and then it can predict semantic annotation for unseen images. We compare our approach with several state-of-the-art approaches on a standard Corel dataset. The experiment results show that our approach performs more effectively and accurately. Zhixin Li 0001, Zhi-Ping Shi 0002, Xi Liu 0009, Zhongzhi Shi |
ICASSP | 4 |
| 2010 | A novel sparse coding model based on structural similarityabstractUnderstanding and modeling the function of the neurons and neural systems are primary goal of systems neuroscience. Sparse coding theory demonstrates that the neurons in primary visual cortex form a sparse representation of natural scenes in the viewpoint of statistics. In this paper, we propose a novel sparse coding model based on structural similarity (SS_SC) for natural image feature extraction. The advantage for our model is to be able to preserve structural information from a scene, which human visual perception is highly adapted for. Using the proposed sparse coding model, the validity of image feature extraction is testified. Furthermore, compared with standard sparse coding (SC) model, the experimental results show that the quality of reconstructed images obtained by our method outperforms the SC method. Zhi-Ping Shi 0002, Xi Liu 0009, Zhongzhi Shi |
ICASSP | 4 |
| 2010 | D-LDA: A Topic Modeling Approach without Constraint Generation for Semi-defined ClassificationabstractWe study what we call semi-defined classification, which deals with the categorization tasks where the taxonomy of the data is not well defined in advance. It is motivated by the real-world applications, where the unlabeled data may also come from some other unknown classes besides the known classes for the labeled data. Given the unlabeled data, our goal is to not only identify the instances belonging to the known classes, but also cluster the remaining data into other meaningful groups. It differs from traditional semi-supervised clustering in the sense that in semi-supervised clustering the supervision knowledge is too far from being representative of a target classification, while in semi-defined classification the labeled data may be enough to supervise the learning on the known classes. In this paper we propose the model of Double-latent-layered LDA (D-LDA for short) for this problem. Compared with LDA with only one latent variable y for word topics, D-LDA contains another latent variable z for (known and unknown) document classes. With this double latent layers consisting of y and z and the dependency between them, D-LDA directly injects the class labels into z to supervise the exploiting of word topics in y. Thus, the semi-supervised learning in D-LDA does not need the generation of pair wise constraints, which is required in most of the previous semi-supervised clustering approaches. We present the experimental results on ten different data sets for semi-defined classification. Our results are either comparable to (on one data sets), or significantly better (on the other nine data set) than the six compared methods, including the state-of-the-art semi-supervised clustering methods. Fuzhen Zhuang, Ping Luo 0001, Zhiyong Shen, Qing He 0003, Yuhong Xiong, Zhongzhi Shi |
ICDM | 6 |
| 2010 | Local Bayesian Based Rejection Method for HSC Ensemble
Qing He 0003, Wenjuan Luo, Fuzhen Zhuang, Zhongzhi Shi |
ISNN (1) | 4 |
| 2010 | An Ontology-Based Semantic Web Service Space Organization and Management Model
Zhongzhi Shi |
KSEM | 2 |
| 2010 | Sorted label classifier chains for learning images with multi-labelabstractIn the real world, images always have several visual objects instead of only one, which makes it difficult for conventional object recognition methods to deal with them. In this paper, we present a topologically sorted classifier chain method for learning images with multi-label. We first provide a means of generating a topo-logically sorted label chain ordering by employing a topological sort algorithm and then apply the chain ordering to the classifier chain model proposed by [1] to classify multi-label images. Our method can capture the correlations between labels very effectively due to the sorted label chain ordering and the advantages brought by classifier chain method. We evaluate the proposed method on Corel dataset and demonstrate the micro and macro F1 measures superior to the state-of-the-art methods. Xi Liu 0009, Zhi-Ping Shi 0002, Zhixin Li 0001, Xishun Wang, Zhongzhi Shi |
ACM Multimedia | 5 |
| 2010 | Orthogonal Nonnegative Matrix Tri-factorization for Semi-supervised Document Co-clustering
Huifang Ma, Weizhong Zhao, Qing Tan, Zhongzhi Shi |
PAKDD (2) | 4 |
| 2010 | Mining Hot Clusters of Similar Anomalies for System Management
David Zhang 0001, Zhongzhi Shi |
PRICAI | 3 |
| 2010 | Exploiting Associations between Word Clusters and Document Classes for Cross-Domain Text CategorizationabstractCross-domain text categorization targets on adapting the knowledge learnt from a labeled source-domain to an unlabeled target-domain, where the documents from the source and target domains are drawn from different distributions. However, in spite of the different distributions in raw word features, the associations between word clusters (conceptual features) and document classes may remain stable across different domains. In this paper, we exploit these unchanged associations as the bridge of knowledge transformation from the source domain to the target domain by the nonnegative matrix tri-factorization. Specifically, we formulate a joint optimization framework of the two matrix tri-factorizations for the source and target domain data respectively, in which the associations between word clusters and document classes are shared between them. Then, we give an iterative algorithm for this optimization and theoretically show its convergence. The comprehensive experiments show the effectiveness of this method. In particular, we show that the proposed method can deal with some difficult scenarios where baseline methods usually do not perform well. Fuzhen Zhuang, Ping Luo 0001, Hui Xiong 0001, Qing He 0003, Yuhong Xiong, Zhongzhi Shi |
SDM | 6 |
| 2010 | Similarity-Based Bayesian Learning from Semi-structured Log Files for Fault Diagnosis of Web ServicesabstractWith the rapid development of XML language which has good flexibility and interoperability, more and more log files of software running information are represented in XML format, especially for Web services. Fault diagnosis by analyzing semi-structured and XML like log files is becoming an important issue in this area. For most related learning methods, there is a basic assumption that training data should be in identical structure, which does not hold in many situations in practice. In order to learn from training data in different structures, we propose a similarity-based Bayesian learning approach for fault diagnosis in this paper. Our method is to first estimate similarity degrees of structural elements from different log files. Then the basic structure of combined Bayesian network (CBN) is constructed, and the similarity-based learning algorithm is used to compute probabilities in CBN. Finally, test log data can be classified into possible fault categories based on the generated CBN. Experimental results show our approach outperforms other learning approaches on those training datasets which have different structures. Zhongzhi Shi, Wenjia Niu, Kunrong Chen, Xinghua Yang |
Web Intelligence | 2 |
| 2010 | Fusing semantic aspects for image annotation and retrieval
Zhixin Li 0001, Zhi-Ping Shi 0002, Xi Liu 0009, Zhongzhi Shi |
J. Vis. Commun. Image Represent. | 5 |
| 2010 | Cross-Domain Learning from Multiple Sources: A Consensus Regularization PerspectiveabstractClassification across different domains studies how to adapt a learning model from one domain to another domain which shares similar data characteristics. While there are a number of existing works along this line, many of them are only focused on learning from a single source domain to a target domain. In particular, a remaining challenge is how to apply the knowledge learned from multiple source domains to a target domain. Indeed, data from multiple source domains can be semantically related, but have different data distributions. It is not clear how to exploit the distribution differences among multiple source domains to boost the learning performance in a target domain. To that end, in this paper, we propose a consensus regularization framework for learning from multiple source domains to a target domain. In this framework, a local classifier is trained by considering both local data available in one source domain and the prediction consensus with the classifiers learned from other source domains. Moreover, we provide a theoretical analysis as well as an empirical study of the proposed consensus regularization framework. The experimental results on text categorization and image classification problems show the effectiveness of this consensus regularization learning method. Finally, to deal with the situation that the multiple source domains are geographically distributed, we also develop the distributed version of the proposed algorithm, which avoids the need to upload all the data to a centralized location and helps to mitigate privacy concerns. Fuzhen Zhuang, Ping Luo 0001, Hui Xiong 0001, Yuhong Xiong, Qing He 0003, Zhongzhi Shi |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2009 | Automated Design of Logic Circuits with a Increasable Evolution ApproachabstractSince the scalability of logic circuits is becoming larger and more complex, the auto-design is becoming more and more difficult. In order to improve automatic design and performance evaluation of logic circuits in efficiency and capability of optimization, multiobjective simulated annealing (MSA) based increasable evolution approach is designed to evolve logic circuits automatically with an extended matrix encoding method, which can be able to reflect the potential performance of a circuit and reduce the risk of deleting a circuit with a good developing potential during evolution is devised. In the process of evolution, each individual is renewedly associated to a corresponding objective in terms of a novel adaptive evaluation method at each generation. In experiments, complicated arithmetic circuits are designed to assess the performance of MSA against other algorithms. Results indicate that the proposed method could design logic circuits efficiently. Naixue Xiong, Athanasios V. Vasilakos, Zhongzhi Shi |
HPCC | 5 |
| 2009 | Modeling latent aspects for automatic image annotationabstractIn this paper, we present an approach based on probabilistic latent semantic analysis (PLSA) to accomplish the tasks of automatic image annotation. In order to model training data precisely, we represent an image as a bag of visual words and employ two PLSA models to capture semantic information from visual and textual modalities respectively. Furthermore, an adaptive learning approach is proposed to combine the aspects learned from both modalities. For each image document, distribution over aspects is fused by different weight in terms of the entropy of its feature distribution. Consequently, the two models are linked with the same distribution over aspects. This structure can predict semantic annotation for an unseen image because it associates visual and textual modalities properly. We compare our approach with several previous approaches on a standard Corel dataset. The experiment results show that our approach performs more effectively and accurately. Zhixin Li 0001, Zhi-Ping Shi 0002, Zhongzhi Shi |
ICIP | 4 |
| 2009 | Filter object categories using CoBoost with 1ST and 2ND order featuresabstractWe describe a method for filtering object category from a large number of noisy images. This problem is particularly difficult due to the greater variation within object categories and lack of labeled object images. Our method deals with it by combining a co-training algorithm CoBoost with two features - 1stand 2ndorder features, which define bag of words representation and spatial relationship between local features respectively. We iteratively train two boosting classifiers based on the 1stand 2ndorder features, during which each classifier provides labeled data for the other classifier. It is effective because the 1stand 2ndorder features make up an independent and redundant feature split. We evaluate our method on Berg dataset and demonstrate the precision comparative to the state-of-the-art. Xi Liu 0009, Zhi-Ping Shi 0002, Zhongzhi Shi |
ICIP | 3 |
| 2009 | Learning image semantics with latent aspect modelabstractAutomatic image annotation has become an important and challenging problem due to the existence of semantic gap. In this paper, we present an approach based on probabilistic latent semantic analysis (PLSA) to accomplish the tasks of semantic image annotation and retrieval. In order to model training images precisely, we employ two PLSA models to capture semantic information from visual and textual modalities respectively. Then an adaptive asymmetric learning approach is proposed to fuse aspects which are learned from both modalities. For each image document, the weight of each modality is determined by its contribution to the content of the image. Consequently, the two models are linked with the same distribution over aspects. This structure can predict semantic annotation for an unseen image because it associates visual and textual modalities properly. Finally, we compare our approach with several previous approaches on a standard Corel dataset. The experiment results show that our approach performs more effective and accurate. Zhixin Li 0001, Xi Liu 0009, Zhi-Ping Shi 0002, Zhongzhi Shi |
ICME | 4 |
| 2009 | Filter object categories: employing visual consistency and semisupervised approachabstractWe describe a method for filtering object category from a large number of noisy images. This problem is particularly difficult due to the greater variation within object categories and only a few labeled object images available. Our method deals with it by using visual consistency and semi-supervised approach. The images of one category often share some visual consistency so that the most irrelevant images can be first removed. Among the left images, a voting method is used to obtain more object exemplars with the initial object exemplars manually selected by users. Finally with all the obtained exemplars and those unlabeled images, we create a semi-supervised classifier to rank all the images. We evaluate our method on Berg dataset and demonstrate the precision comparative to the state-of-the-art. Besides, we collect five more categories from Google images to show the effectiveness of the method. Xi Liu 0009, Zhixin Li 0001, Zhi-Ping Shi 0002, Zhongzhi Shi |
ICME | 4 |
| 2009 | An Efficient Coding Model for Image Representation
Zhi-Ping Shi 0002, Zhixin Li 0001, Zhongzhi Shi |
ICONIP (1) | 4 |
| 2009 | Rule Extraction and Reduction for Hyper Surface Classification
Qing He 0003, Zhongzhi Shi |
ISNN (2) | 3 |
| 2009 | Towards combining web classification and web information extraction: a case studyabstractWeb content analysis often has two sequential and separate steps: Web Classification to identify the target Web pages, and Web Information Extraction to extract the metadata contained in the target Web pages. This decoupled strategy is highly ineffective since the errors in Web classification will be propagated to Web information extraction and eventually accumulate to a high level. In this paper we study the mutual dependencies between these two steps and propose to combine them by using a model of Conditional Random Fields (CRFs). This model can be used to simultaneously recognize the target Web pages and extract the corresponding metadata. Systematic experiments in our project OfCourse for online course search show that this model significantly improves the F1 value for both of the two steps. We believe that our model can be easily generalized to many Web applications. Ping Luo 0001, Yuhong Xiong, Zhongzhi Shi |
KDD | 5 |
| 2009 | Coboost learning of visual categories with 1st and 2nd order features from Google imagesabstractConventional object recognition techniques rely heavily on manually annotated image datasets to achieve good performances. However, collecting high quality datasets is really laborious. In this paper, we propose a semi-supervised framework for learning visual categories from Google Images. The 1st and 2nd order features, which define bag of words representation and spatial relationship between local features respectively, make up an independent and redundant feature split. We then integrate a cotraining algorithm CoBoost with these two features. We create two boosting classifiers based on the 1st and 2nd order features respectively in the training, during which one classifier provides labels for the other. Besides, the 2nd order features are generated dynamically rather than extracted exhaustively to avoid high computation. An active learning technique is also introduced to further improve the performance. We evaluate our method on the benchmark datasets, showing results competitive with the state-of-the-art unsupervised approaches and some supervised techniques. Xi Liu 0009, Zhi-Ping Shi 0002, Zhixin Li 0001, Zhongzhi Shi |
ACM Multimedia | 4 |
| 2009 | Reasoning about Web Services with Local Closed World AssumptionabstractThis paper presents a formalism for representing and reasoning about Web services with local closed world assumption (LCWA) on the basis of $\mathcal{ALCO@K}$. In our formalism, the knowledge about the states of the world is encoded in $\mathcal {ALCO@}$-ABoxes; atomic services are represented in terms of their preconditions (epistemic queries to the knowledge base) and effects (possibly negated $\mathcal{ALCO@}$-assertions involving only atomic concepts); and composite services are built up with action constructors in dynamic logics. We also summarize some reasoning tasks and develop a calculus for them. Our formalism also enjoys \emph{introspection}. The main features of our proposal (i.e., dynamic reasoning, local closed world assumption and introspection) make it more philosophically satisfying and much closer towards a practical formalism for agents with incomplete knowledge in the Web full of static information and dynamic processing. Hong Hu 0001, Zhongzhi Shi |
Web Intelligence | 3 |
| 2009 | Active Learning of Instance-Level Constraints for Semi-supervised Document ClusteringabstractThis paper presents a framework that actively selects informative documents pairs for semi-supervised document clustering. The semi-supervised document clustering algorithm is a Constrained DBSCAN (Cons-DBSCAN), which incorporates instance-level constraints to guide the clustering process in DBSCAN. By obtaining user feedbacks, our proposed active learning algorithm can get informative instance level constraints to aid clustering process. Experimental results show that Cons-DBSCAN with the proposed active learning approach can provide an appealing clustering performance. Weizhong Zhao, Qing He 0003, Huifang Ma, Zhongzhi Shi |
Web Intelligence | 4 |
| 2009 | A fast approach to attribute reduction in incomplete decision systems with tolerance relation-based rough sets
Zuqiang Meng, Zhongzhi Shi |
Inf. Sci. | 2 |
| 2009 | A dominance tree and its application in evolutionary multi-objective optimization
Chuan Shi 0001, Zhenyu Yan 0001, Kevin Lü 0001, Zhongzhi Shi, Bai Wang 0001 |
Inf. Sci. | 4 |
| 2009 | An index and retrieval framework integrating perceptive features and semantics for multimedia databases
Zhi-Ping Shi 0002, Qing He 0003, Zhongzhi Shi |
Multim. Tools Appl. | 3 |
| 2009 | Information-Theoretic Distance Measures for Clustering Validation: Generalization and NormalizationabstractThis paper studies the generalization and normalization issues of information-theoretic distance measures for clustering validation. Along this line, we first introduce a uniform representation of distance measures, defined as quasi-distance, which is induced based on a general form of conditional entropy. The quasi-distance possesses three properties: symmetry, the triangle law, and the minimum reachable. These properties ensure that the quasi-distance naturally lends itself as the external measure for clustering validation. In addition, we observe that the ranges of the distance measures are different when they apply for clustering validation on different data sets. Therefore, when comparing the performances of clustering algorithms on different data sets, distance normalization is required to equalize ranges of the distance measures. A critical challenge for distance normalization is to obtain the ranges of a distance measure when a data set is provided. To that end, we theoretically analyze the computation of the maximum value of a distance measure for a data set. Finally, we compare the performances of the partition clustering algorithm K-means on various real-world data sets. The experiments show that the normalized distance measures have better performance than the original distance measures when comparing clusterings of different data sets. Also, the normalized Shannon distance has the best performance among four distance measures under study. Ping Luo 0001, Hui Xiong 0001, Guoxing Zhan, Junjie Wu 0002, Zhongzhi Shi |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2008 | On modeling cognitive process with granular computingabstractExploring mechanism of cognitive process is a valuable means of investigating artificial intelligence. But human beingpsilas perception of cognitive process is very limited so far, and related researches are required to be further deepened. This paper focuses on modeling cognitive process based on theory of granular computing(GrC). First, fundamental principle of GrC is used to analyze cognitive process and establish its model. Secondly, a tolerance relation on vector space(ldquorv_spacerdquo for short) is defined by using a distance fucntion dis, and a tolerance granular space model is established, in which the problem of approximate representation of unknown concept(abstraction of perceptive information) is discussed and analyzed. The analyses show that a given concept, sometimes, is difficult to be learned(i.e., difficult to be accurately represented), so appropriate granular world is required. The solution to this problem involves relationship and transformation between granular worlds. At last, such relationship and transformation are investigated, and many related properties and theorems are deduced, solving the facing problem to some extent. The theory and approach of representation of a concept in a granular world and transformation between different granular worlds are made of support cognitive processpsila granular model. Zuqiang Meng, Zhongzhi Shi |
FUZZ-IEEE | 2 |
| 2008 | A Survey on Statistical Pattern Feature Extraction
Shifei Ding, Weikuan Jia, Chunyang Su, Fengxiang Jin, Zhongzhi Shi |
ICIC (2) | 5 |
| 2008 | A DDL-Based Model for Web Service Composition in Context-Aware EnvironmentabstractDynamic description logic (DDL) is among the few emerging service composition solutions through logical reasoning. To overcome low efficiency and lacking context-aware support of DDL reasoning, we propose a new DDL-based service composition model, which supports context-based service pre-filtering over DDL reasoning space. The pre-filtering runs under the BPEL workflow and a distributed reasoning algorithm need to reasoning different contexts after pre-filtering. Wenjia Niu, Zhongzhi Shi, Changlin Wan, Liang Chang 0003 |
ICWS | 2 |
| 2008 | Neural Network Research Progress and Applications in Forecast
Shifei Ding, Weikuan Jia, Chunyang Su, Zhongzhi Shi |
ISNN (2) | 5 |
| 2008 | Nonlinear Complex Neural Circuits Analysis and Design by q-Value Weighted Bounded Operator
Hong Hu 0001, Zhongzhi Shi |
ISNN (1) | 2 |
| 2008 | Extreme Support Vector Machine Classifier
Qiuge Liu, Qing He 0003, Zhongzhi Shi |
PAKDD | 3 |
| 2008 | Semantic Filtering for DDL-Based Service Composition
Wenjia Niu, Zhongzhi Shi, Liang Chang 0003 |
PRICAI | 2 |
| 2008 | Focused Crawling with Heterogeneous Semantic InformationabstractFocused crawlers selectively retrieve Web documents that are relevant to a predefined set of topics. To intelligently make predictions and decisions about relevant URLs and web pages, different topic models have been introduced to represent topic-specific knowledge. Yet it is difficult to support semantic interoperability among different models. Moreover, some manually specified additional semantic information, such as semantic markups and social annotations, could not be effectively used to improve crawling. This paper proposes to boost focused crawling with four kinds of semantic models and semantic information, including thesauruses, categories, ontologies, and folksonomies. A statistical semantic association model is proposed to integrate different semantic models, represent heterogeneous semantic information, and support semantic relevance computation. A focused crawling framework is developed which adopts both keyword based contents and different kinds of additional information for relevance prediction and ranking. Experiments show that the proposed model and framework effectively integrates heterogeneous semantic information for focused crawling. Zhongzhi Shi |
Web Intelligence | 3 |
| 2008 | Minimal Consistent Subset for Hyper Surface Classification MethodabstractHyper Surface Classification (HSC), which is based on Jordan Curve Theorem in Topology, has proven to be a simple and effective method for classifying a larger database in our previous work. To select a representative subset from the original sample set, the Minimal Consistent Subset (MCS) of HSC is studied in this paper. For HSC method, one of the most important features of MCS is that it has the same classification model as the entire sample dataset, and can totally reflect its classification ability. From this point of view, MCS is the best way of sampling from the original dataset for HSC. Furthermore, because of the minimum property of MCS, every single deletion or multiple deletions from it will lead to a reduction in generalization ability, which can be exactly predicted by the proposed formula in this paper. Qing He 0003, Xiu-Rong Zhao, Zhongzhi Shi |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2007 | Supervised Information Feature Compression Algorithm Based on Divergence Criterion
Shifei Ding, Wei Ning, Fengxiang Jin, Shixiong Xia, Zhongzhi Shi |
ICIC (2) | 5 |
| 2007 | Agent Grid Collaborative Schema Based on Policy
Maoguang Wang, Zhongzhi Shi, Shixiong Xia, Shifei Ding |
ICIC (3) | 2 |
| 2007 | Incremental Nonlinear Proximal Support Vector Machine
Qiuge Liu, Qing He 0003, Zhongzhi Shi |
ISNN (3) | 3 |
| 2007 | Distributed classification in peer-to-peer networksabstractThis work studies the problem of distributed classification in peer-to-peer(P2P) networks. While there has been a significant amount of work in distributed classification, most of existing algorithms are not designed for P2P networks. Indeed, as server-less and router-less systems, P2P networks impose several challenges for distributed classification: (1) it is not practical to have global synchronization in large-scale P2P networks; (2)there are frequent topology changes caused by frequent failure and recovery of peers; and (3) there are frequent on-the-fly data updates on each peer. Ping Luo 0001, Hui Xiong 0001, Kevin Lü 0001, Zhongzhi Shi |
KDD | 4 |
| 2007 | A Dynamic Description Logic for Representation and Reasoning About Actions
Liang Chang 0003, Zhongzhi Shi |
KSEM | 3 |
| 2007 | Multi-Agent Based Web Search with Heterogeneous Semantics
Zhongzhi Shi |
PRIMA | 2 |
| 2007 | Research on Business Intelligence in enterprise computing environmentabstractBusiness intelligence (BI) is the process of gathering enough of the right information in the right manner at the right time, and delivering the right results to the right people for decision-making purposes so that it can continue to yield real business benefits, or have a positive impact on business strategy, tactics, and operations in the enterprises. This paper was intended as a short introduction to the study of business intelligence in enterprise computing environment. In addition, the conclusions point out the challenges to broad and deep deployment of business intelligence systems, and provide the proposals of making business intelligence more effective. Zhongzhi Shi, Qing He 0003, Maoguang Wang |
SMC | 3 |
| 2007 | Distributed computing environment: Approaches and applicationsabstractIn information systems in particular business Intelligence systems, for data produced at physically distributed locations most traditional data mining approaches require data to be transmitted to a single location for centralized processing and mining. However, the continual transmission of a large number of data to a central location must be impractical and expensive. Thus, distributed and parallel data mining algorithms and applications were rapidly developed. The paper surveys the-state-of-the art in approaches and applications of distributed computing environment. The goal is to summarize a brief introduction to this field with pointers for further exploration. Zhongzhi Shi, Maoguang Wang, Wenjuan Wu |
SMC | 3 |
| 2007 | MSMiner - a developing platform for OLAP
Zhongzhi Shi, Youping Huang, Qing He 0003, Shaohui Liu, Liangxi Qin, Ziyan Jia, Jiayou Li, Huijing Huang |
Decis. Support Syst. | 1 |
| 2007 | Distributed data mining in grid computing environments
Ping Luo 0001, Kevin Lü 0001, Zhongzhi Shi, Qing He 0003 |
Future Gener. Comput. Syst. | 3 |
| 2007 | Distributed data mining on Agent Grid: Issues, platform and development toolkit
Jiewen Luo, Maoguang Wang, Zhongzhi Shi |
Future Gener. Comput. Syst. | 4 |
| 2007 | A revisit of fast greedy heuristics for mapping a class of independent tasks onto heterogeneous computing systems
Ping Luo 0001, Kevin Lü 0001, Zhongzhi Shi |
J. Parallel Distributed Comput. | 3 |
| 2007 | Context optimization of AI planning for semantic Web services composition
Lirong Qiu, Liang Chang 0003, Zhongzhi Shi |
Serv. Oriented Comput. Appl. | 4 |
| 2007 | Classification based on dimension transposition for high dimension data
Qing He 0003, Xiu-Rong Zhao, Zhongzhi Shi |
Soft Comput. | 3 |
| 2007 | On Defining Partition Entropy by InequalitiesabstractPartition entropyis the numerical metric of uncertainty within a partition of a finite set, whileconditional entropymeasures the degree of difficulty in predicting a decision partition when a condition partition is provided. Since two direct methods exist for defining conditional entropy based on its partition entropy, the inequality postulates of monotonicity, which conditional entropy satisfies, are actually additional constraints on its entropy. Thus, in this paper partition entropy is defined as a function of probability distribution, satisfying all the inequalities of not only partition entropy itself but also its conditional counterpart. These inequality postulates formalize the intuitive understandings of uncertainty contained in partitions of finite sets. We study the relationships between these inequalities, and reduce the redundancies among them. According to two different definitions of conditional entropy from its partition entropy, the convenient and unified checking conditions for any partition entropy are presented, respectively. These properties generalize and illuminate the common nature of all partition entropies. Ping Luo 0001, Guoxing Zhan, Qing He 0003, Zhongzhi Shi, Kevin Lü 0001 |
IEEE Trans. Inf. Theory | 4 |
| 2006 | Semantic Web Services Composition Using AI Planning of Description LogicsabstractWeb services composition techniques are gaining momentum as the opportunity to establish reusable and versatile inter-operability applications. The purpose of semantic Web services is to use semantic specification to automate the discovery, invocation, and composition Web services. Description logics is the formalized foundation of semantic Web services and provides well-defined semantics. And many researchers propose their composition approach based on planning techniques. We propose our service composition methods based on description logics and AI planning technologies. Our algorithm for services composition uses backward-chaining search method to find potential candidate services. And we propose a DAG-based method to generate the planning process and filtering the inappropriate services during the DAG generation process. We test our approach on a simple, yet realistic example, and the preliminary results demonstrate that our implementation provides a practical solution Lirong Qiu, Changlin Wan, Zhongzhi Shi |
APSCC | 4 |
| 2006 | Supervised Feature Extraction Algorithm Based on Continuous Divergence Criterion
Shifei Ding, Zhongzhi Shi, Fengxiang Jin |
ICIC (2) | 2 |
| 2006 | Divergence-Based Supervised Information Feature Compression Algorithm
Shifei Ding, Zhongzhi Shi |
ISNN (1) | 2 |
| 2006 | Stock Time Series Forecasting Using Support Vector Machines Employing Analyst Recommendations
Chuan Shi 0001, Sulan Zhang, Zhongzhi Shi |
ISNN (2) | 4 |
| 2006 | HyperSurface Classifiers Ensemble for High Dimensional Data Sets
Xiu-Rong Zhao, Qing He 0003, Zhongzhi Shi |
ISNN (1) | 3 |
| 2006 | Eliminate Redundancy in Parallel Search: A Multi-agent Coordination Approach
Jiewen Luo, Zhongzhi Shi |
PRICAI | 2 |
| 2006 | An Improved Multiobjective Evolutionary Algorithm Based on Dominating Tree
Chuan Shi 0001, Qingyong Li, Zhongzhi Shi |
PRICAI | 4 |
| 2006 | Description Logic Based Composition of Web Services
Lirong Qiu, Zhongzhi Shi |
PRIMA | 5 |
| 2006 | Agent Grid Collaborative Environment
Zhongzhi Shi |
PRIMA | 1 |
| 2006 | Techniques, Process, and Enterprise Solutions of Business IntelligenceabstractBusiness Intelligence (BI) has been viewed as sets of powerful tools and approaches to improving business executive decision-making, business operations, and increasing the value of the enterprise. The technology categories of BI mainly encompass Data Warehousing, OLAP, and Data Mining. This article reviews the concept of Business Intelligence and provides a survey, from a comprehensive point of view, on the BI technical framework, process, and enterprise solutions. In addition, the conclusions point out the possible reasons for the difficulties of broad deployment of enterprise BI, and the proposals of constructing a better BI system. Zhongzhi Shi, Maoguang Wang, Wenjuan Wu |
SMC | 3 |
| 2006 | A heterogeneous computing system for data mining workflows in multi-agent environmentsabstractAbstract:The computing‐intensive data mining (DM) process calls for the support of a heterogeneous computing system, which consists of multiple computers with different configurations connected by a high‐speed large‐area network for increased computational power and resources. The DM process can be described as a multi‐phase pipeline process, and in each phase there could be many optional methods. This makes the workflow for DM very complex and it can be modeled only by a directed acyclic graph (DAG). A heterogeneous computing system needs an effective and efficient scheduling framework, which orchestrates all the computing hardware to perform multiple competitive DM workflows. Motivated by the need for a practical solution of the scheduling problem for the DM workflow, this paper proposes a dynamic DAG scheduling algorithm according to the characteristics of an execution time estimation model for DM jobs. Based on an approximate estimation of job execution time, this algorithm first maps DM jobs to machines in adecentralizedanddiligent(defined in this paper) manner. Then the performance of this initial mapping can be improved throughjob migrationswhen necessary. The scheduling heuristic used considers the factors of both theminimal completion timecriterion and thecritical pathin a DAG. We implement this system in an established multi‐agent system environment, in which the reuse of existing DM algorithms is achieved by encapsulating them into agents. The system evaluation and its usage in oil well logging analysis are also discussed. Ping Luo 0001, Kevin Lü 0001, Qing He 0003, Zhongzhi Shi |
Expert Syst. J. Knowl. Eng. | 5 |
| 2006 | Progress and Challenge of Artificial Intelligence
Zhongzhi Shi, Nanning Zheng 0001 |
J. Comput. Sci. Technol. | 1 |
| 2006 | Perceptual learning and abstraction in machine learning: an application to autonomous roboticsabstractThis paper deals with the possible benefits of perceptual learning in artificial intelligence. On the one hand, perceptual learning is more and more studied in neurobiology and is now considered as an essential part of any living system. In fact, perceptual learning and cognitive learning are both necessary for learning and often depend on each other. On the other hand, many works in machine learning are concerned with "abstraction" in order to reduce the amount of complexity related to some learning tasks. In the abstraction framework, perceptual learning can be seen as a specific process that learns how to transform the data before the traditional learning task itself takes place. In this paper, we argue that biologically inspired perceptual learning mechanisms could be used to build efficient low-level abstraction operators that deal with real-world data. To illustrate this, we present an application where perceptual-learning-inspired metaoperators are used to perform an abstraction on an autonomous robot visual perception. The goal of this work is to enable the robot to learn how to identify objects it encounters in its environment. Nicolas Bredèche, Zhongzhi Shi, Jean-Daniel Zucker |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2005 | An Improved Relative Criterion Using BP Algorithm
Zhongzhi Shi |
ISNN (1) | 3 |
| 2005 | An Interval-Based Knowledge Model and Query Language for Temporal Information
Zhongzhi Shi, Xiaoxiao He, Lirong Qiu, Jiewen Luo |
PRIMA | 2 |
| 2005 | Multi-agent Cooperation: A Description Logic View
Jiewen Luo, Zhongzhi Shi, Maoguang Wang |
PRIMA | 2 |
| 2005 | Dynamic Interaction Protocol Load in Multi-Agent System Collaboration
Maoguang Wang, Zhongzhi Shi, Wenpin Jiao |
PRIMA | 2 |
| 2005 | Ontology-Driven Knowledge Management on the GridabstractThe combination of large data set size, geographic distribution of resources and users, and sophisticated applications on data and information needs robust infrastructures. In this paper, knowledge management sphere (KMSphere) is proposed and developed to explore important aspects of service-oriented and ontology-driven knowledge management on the grid. The main idea of KMSphere is to integrate ontologies with a service-oriented grid, build a knowledge space on top of databases, and then organize, utilize, and manage the knowledge resources in that space. Zhongzhi Shi, Lirong Qiu |
Web Intelligence | 2 |
| 2005 | Learning from Ontologies for Common Meaningful StructuresabstractWe put forward a hypothesis that there exist common meaningful structures among ontologies whose domains are analogous to each other The initial motivation of our hypothesis is to make full use of the structural information in existing ontologies, in order to benefit the domain of ontology. To verify the hypothesis we give a precise definition of the candidate of the common meaningful structure called MICISO (maximum isomorphic common induced sub-ontology). Based on the hypothesis and the definition we present a novel data mining problem called MICISO mining, whose aim is learning from ontologies to find out MICISOs and further recommend the common meaningful structures. We also provide an algorithm for MICISO mining, based on which we have developed a practical tool for mining and checking such structures. With the tool, the algorithm is implemented with quite a few pairs of existing ontologies, and the interesting meaningful results support our hypothesis. Thus we consider that the hypothesis is preliminarily verified. We suppose that our work sparks a novel promising thinking for the domain of ontology -to study existing ontologies for useful things. Guojie Li, Zhongzhi Shi |
Web Intelligence | 3 |
| 2005 | A logical foundation for the semantic Web
Zhongzhi Shi, Mingkai Dong 0003 |
Sci. China Ser. F Inf. Sci. | 1 |
| 2004 | A Knowledge-Based Data Model and Query Algebra for the Next-Generation Web
Qiujian Sheng, Zhongzhi Shi |
APWeb | 2 |
| 2004 | Some discussions about MOGAs: individual relations, non-dominated set, and application on automatic negotiationabstractThis paper studies the relations of individuals in evolutionary populations, and then investigates some features of the relations. The goal is to find efficient methods to construct the non-dominated set. It is proved that the individuals can be sorted by quick sort. To demonstrate the efficiency of our new method, we propose a multi-objective genetic algorithm (MOGA) based on quick sort, which is called QKMOGA. We apply QKMOGA on automatic negotiation for agents. A simple negotiation model between two agents is described, and the negotiation protocols are constructed with QKMOGA. Two experimental results show that the performance is satisfactory on the diversity and efficiency of the solutions. Jinhua Zheng, Charles Ling 0001, Zhongzhi Shi |
IEEE Congress on Evolutionary Computation | 3 |
| 2004 | Rough Set Based Image Texture Recognition Algorithm
Zheng Zheng 0001, Hong Hu 0001, Zhongzhi Shi |
KES | 3 |
| 2004 | Association-Rule Based Information Source Selection
Hui Yang 0004, Minjie Zhang 0001, Zhongzhi Shi |
PRICAI | 3 |
| 2004 | The Information Entropy, Rough Entropy And Knowledge Granulation In Rough Set TheoryabstractRough set theory is a relatively new mathematical tool for use in computer applications in circumstances which are characterized by vagueness and uncertainty. In this paper, we introduce the concepts of information entropy, rough entropy and knowledge granulation in rough set theory, and establish the relationships among those concepts. These results will be very helpful for understanding the essence of concept approximation and establishing granular computing in rough set theory. Jiye Liang, Zhongzhi Shi |
Int. J. Uncertain. Fuzziness Knowl. Based Syst. | 2 |
| 2004 | Dynamic Service Matchmaking in Intelligent Web
Zhongzhi Shi, Hanjun Zhang, Mingkai Dong 0003 |
J. Web Eng. | 2 |
| 2003 | A Teamwork Protocol for Multi-agent System
Qiujian Sheng, Zhi-Kun Zhao, Shaohui Liu, Zhongzhi Shi |
PRIMA | 4 |
| 2002 | CSIM: a document clustering algorithm based on swarm intelligenceabstractThis paper presents a document clustering algorithm based on swarm intelligence and k-means: CSIM. First, a document clustering algorithm based on swarm intelligence is employed. It is derived from a basic model interpreting ant colony organization of cemeteries. Swarm intelligence for flexibility, self-organization and robustness has been applied in a variety of areas. Taking advantage of these traits, good initial clusters are obtained in the first phase in CSIM. We then combine it with the classical k-means clustering method by using the clusters as initial centers. CSIM inherits the prominent properties of both swarm intelligence and k-means. It also offsets the weakness of those two techniques. Experimental results show the good performance of the hybrid document clustering algorithm. Bin Wu 0001, Shaohui Liu, Zhongzhi Shi |
IEEE Congress on Evolutionary Computation | 4 |
| 2002 | The Multi-class Classification Method in Large Database Based on Hyper Surface
Quing He, Zhongzhi Shi, Li-An Ren |
ICMLA | 2 |
| 2002 | Innovating Web Page Classification Through Reducing Noise
Xiaoli Li 0001, Zhongzhi Shi |
J. Comput. Sci. Technol. | 2 |
| 2000 | A Document Classifier Based on Word Semantic Association
Xiaoli Li 0001, Jimin Liu, Zhongzhi Shi |
PRICAI | 3 |
| 2000 | Integrating object-oriented analysis with action logic for model buildingabstractDecision models play an important role in decision-making, and supporting model-building is one of the most important functions of model management in decision support systems. As concepts at different abstract levels have to be used in the process of model-building, representing these concepts in a coherent way has been recognized as a key research topic. In this paper, a model-building framework is proposed which integrates object-oriented analysis with action logic as the representation tool. This model-building framework can provide representations for concepts at different abstract levels and can describe the process of abstracting decision models from decision situations or problems represented in lower abstract level concepts. Qijia Tian, Jian Ma 0008, Duanning Zhou, Zhongzhi Shi |
SMC | 4 |
| 1999 | A multimedia synchronization model based on timed Petri net
Zhongzhi Shi |
J. Comput. Sci. Technol. | 2 |
| 1999 | RAO logic for multiagent framework
Zhongzhi Shi, Qijia Tian |
J. Comput. Sci. Technol. | 1 |
| 1997 | Applying case-based reasoning to engine oil design
Zhongzhi Shi |
Artif. Intell. Eng. | 1 |
| 1995 | Minimal model semantics for sorted constraint representation
Lejian Liao, Zhongzhi Shi |
J. Comput. Sci. Technol. | 2 |
| 1995 | A necessary condition about the optimum partition on a finite set of samples and its application to clustering analysis
Shiwei Ye, Zhongzhi Shi |
J. Comput. Sci. Technol. | 2 |
| 1990 | Attribute theory in learning systems
Zhongzhi Shi, Jianchao Han |
Future Gener. Comput. Syst. | 1 |
| 1984 | FORMANAGER: An Office Forms Management SystemabstractThe form has become an important abstraction for data management in an office application environment.Structured office forms present data to users in an easily understood and easily manipulated manner.In this paper we classify forms systems in terms of three dimensions: data structuring, user interfaces, and programming interfaces.Current forms systems are analyzed under these dimensions.We have designed a comprehensive forms management system, FORMANAGER, that includes facilities for form specification, form processing, and form control.The system transforms data from a relational database into a hierarchical data structure which defines the form.The design and algorithms for implementation of the system are described, and future extensions to enhance the capabilities of forms systems are proposed. Alan R. Hevner, Zhongzhi Shi, Dawei Luo |
ACM Trans. Inf. Syst. | 3 |