VLDB 2026 Research / reviewers in the wild / expert
Yong Shi 0001
dblp:84/5467-1
· DBLP profile ↗
140ranked-venue papers
31as first author
39since 2021 · last 2027
0000-0001-7974-1079ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 97 · 20 first-author · 27 since 2021Databases, data management, data science and information retrieval · 32 · 8 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-authorSystems, architecture and hardware · 6 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 first-authorTheory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Multi-objective learning with multi-gradient descent for training sparse and interpretable neural networks
Yongjie Feng, Peng Zhang 0001, Hong Yang 0003, Byron J. Gao, Yong Shi 0001 |
Inf. Sci. | 5 |
| 2026 | Multi-label financial statement fraud detection based on long short-term memory and multilayer perceptron hybrid model
Zhensong Chen 0001, Yanxin Liu, Yong Shi 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | Multi-graph fusion guided robust adaptive learning for subspace clustering
Jianyu Miao, Xiaochan Zhang, Tiejun Yang, Yingjie Tian 0001, Yong Shi 0001 |
Expert Syst. Appl. | 6 |
| 2026 | Positive-unlabeled learning for anomaly detection based on dual-branch generative adversarial networks
Qinyan Wei, Aihua Li, Juxiang Hu, Che Han, Yuxue Chi, Yong Shi 0001 |
Knowl. Based Syst. | 6 |
| 2026 | Region-Prompt-Guided Anomaly Detection With Entropy-Based Consistency ModelingabstractVisual industrial anomaly detection has evolved from one-class modeling to more challenging multi-class settings, where diverse categories and complex visual patterns must be jointly handled. Existing approaches often assume that anomalies lie far from normal samples in feature or spatial space. However, this assumption frequently fails due to two key issues: cross-class semantic confusion, where normal structures of one category are misclassified as anomalies in another, and pixel similarity failure, where anomalous regions visually blend into normal backgrounds. To address these challenges, we propose RPGAD (Region-Prompt Guided Anomaly Detection), an information-theoretic framework that models anomalies as semantic predictive instability, reflected in the joint responses of dual paths. RPGAD integrates two components: 1) DPENet (Dual-Path regional Energy evaluation Network), which compares region-level responses across normal-only and mixed paths through an entropy-guided energy formulation to generate robust region prompts; and 2) RDNet (Reverse Distillation Network), which selectively reconstructs prompted regions and employs a Prototype-Contrastive Optimal Transport (PCOT) loss to enhance inter-class separability and local feature aggregation. Experiments on five anomaly detection benchmarks - MVTecAD, VisA, BTAD, MPDD, and Real-IAD - demonstrate the effectiveness of RPGAD. At $256 \times 256$ resolution, RPGAD achieves strong overall performance, with mAD of 87.8%, 80.3%, 85.2%, 86.2%, and 77.7% on five benchmarks, and pixel-level AP and F1-max gains of up to 12.2 and 10.4 points over strong baselines. These results confirm that RPGAD provides accurate and robust multi-class anomaly detection and localization in complex visual scenarios. Yong Shi 0001, Zhiquan Qi |
IEEE Trans. Image Process. | 1 |
| 2025 | Inter-graph and Intra-graph: Utilizing global financial markets and constituent stocks for stock index prediction
Yong Shi 0001, Yunong Wang |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | Robust sparse orthogonal basis clustering for unsupervised feature selection
Jianyu Miao, Tiejun Yang, Yingjie Tian 0001, Yong Shi 0001 |
Expert Syst. Appl. | 5 |
| 2025 | ExGAT: Context extended graph attention neural network
Pei Quan, Lei Zheng 0011, Wen Zhang 0001, Yang Xiao 0017, Lingfeng Niu, Yong Shi 0001 |
Neural Networks | 6 |
| 2025 | A universal network strategy for lightspeed computation of entropy-regularized optimal transport
Yong Shi 0001, Lei Zheng 0011, Pei Quan, Yang Xiao 0017, Lingfeng Niu |
Neural Networks | 1 |
| 2025 | Self-Supervised Random Forest on Transformed Distribution for Anomaly DetectionabstractAnomaly detection, the task of differentiating abnormal data points from normal ones, presents a significant challenge in the realm of machine learning. Numerous strategies have been proposed to tackle this task, with classification-based methods, specifically those utilizing a self-supervised approach via random affine transformations (RATs), demonstrating remarkable performance on both image and non-image data. However, these methods encounter a notable bottleneck, the overlap of constructed labeled datasets across categories, which hampers the subsequent classifiers' ability to detect anomalies. Consequently, the creation of an effective data distribution becomes the pivotal factor for success. In this article, we introduce a model called "self-supervised forest (sForest)," which leverages the random Fourier transform (RFT) and random orthogonal rotations to craft a controlled data distribution. Our model utilizes the RFT to map input data into a new feature space. With this transformed data, we create a self-labeled training dataset using random orthogonal rotations. We theoretically prove that the data distribution formulated by our methodology is more stable compared to one derived from RATs. We then use the self-labeled dataset in a random forest (RF) classifier to distinguish between normal and anomalous data points. Comprehensive experiments conducted on both real and artificial datasets illustrate that sForest outperforms other anomaly detection methods, including distance-based, kernel-based, forest-based, and network-based benchmarks. Hanyuan Hang, Shumin Ma, Xin Shen 0003, Yong Shi 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Explicit unsupervised feature selection based on structured graph and locally linear embedding
Jianyu Miao, Tiejun Yang, Yingjie Tian 0001, Yong Shi 0001 |
Expert Syst. Appl. | 6 |
| 2024 | Sparse optimization guided pruning for neural networks
Yong Shi 0001, Anda Tang, Lingfeng Niu, Ruizhi Zhou |
Neurocomputing | 1 |
| 2024 | Wasserstein distance regularized graph neural networks
Yong Shi 0001, Lei Zheng 0011, Pei Quan, Lingfeng Niu |
Inf. Sci. | 1 |
| 2024 | Black-box attacks on dynamic graphs via adversarial topology perturbations
Haicheng Tao, Jie Cao 0001, Lei Chen 0079, Hong-Liang Sun, Yong Shi 0001, Xingquan Zhu 0001 |
Neural Networks | 5 |
| 2023 | Federated learning with ℓ1 regularization
Yong Shi 0001, Yuanying Zhang, Peng Zhang 0001, Yang Xiao 0017, Lingfeng Niu |
Pattern Recognit. Lett. | 1 |
| 2023 | GAN-CL: Generative Adversarial Networks for Learning From Complementary LabelsabstractLearning from complementary labels (CLs) is a useful learning paradigm, where the CL specifies the classes that the instance does not belong to, instead of providing the ground truth as in the ordinary supervised learning scenario. In general, although it is less laborious and more efficient to collect CLs compared with ordinary labels, the less informative signal in the complementary supervision is less helpful to learn competent feature representation. Consequently, the final classifier's performance greatly deteriorates. In this article, we leverage generative adversarial networks (GANs) to derive an algorithm GAN-CL to effectively learn from CLs. In addition to the role in original GAN, the discriminator also serves as a normal classifier in GAN-CL, with the objective constructed partly with the complementary information. To further prove the effectiveness of our schema, we study the global optimality of both generator and discriminator for the GAN-CL under mild assumptions. We conduct extensive experiments on benchmark image datasets using deep models, to demonstrate the compelling improvements, compared with state-of-the-art CL learning approaches. Hanyuan Hang, Bo Wang 0049, Biao Li 0005, Yingjie Tian 0001, Yong Shi 0001 |
IEEE Trans. Cybern. | 7 |
| 2023 | Graph Influence NetworkabstractDue to the extraordinary abilities in extracting complex patterns, graph neural networks (GNNs) have demonstrated strong performances and received increasing attention in recent years. Despite their prominent achievements, recent GNNs do not pay enough attention to discriminate nodes when determining the information sources. Some of them select information sources from all or part of neighbors without distinction, and others merely distinguish nodes according to either graph structures or node features. To solve this problem, we propose the concept of the Influence Set and design a novel general GNN framework called the graph influence network (GINN), which discriminates neighbors by evaluating their influences on targets. In GINN, both topological structures and node features of the graph are utilized to find the most influential nodes. More specifically, given a target node, we first construct its influence set from the corresponding neighbors based on the local graph structure. To this aim, the pairwise influence comparison relations are extracted from the paths and a HodgeRank-based algorithm with analytical expression is devised to estimate the neighbors' structure influences. Then, after determining the influence set, the feature influences of nodes in the set are measured by the attention mechanism, and some task-irrelevant ones are further dislodged. Finally, only neighbor nodes that have high accessibility in structure and strong task relevance in features are chosen as the information sources. Extensive experiments on several datasets demonstrate that our model achieves state-of-the-art performances over several baselines and prove the effectiveness of discriminating neighbors in graph representation learning. Yong Shi 0001, Pei Quan, Yang Xiao 0017, Minglong Lei, Lingfeng Niu |
IEEE Trans. Cybern. | 1 |
| 2023 | LLP-GAN: A GAN-Based Algorithm for Learning From Label ProportionsabstractLearning from label proportions (LLP) is a widespread and important learning paradigm: only the bag-level proportional information of the grouped training instances is available for the classification task, instead of the instance-level labels in the fully supervised scenario. As a result, LLP is a typical weakly supervised learning protocol and commonly exists in privacy protection circumstances due to the sensitivity in label information for real-world applications. In general, it is less laborious and more efficient to collect label proportions as the bag-level supervised information than the instance-level one. However, the hint for learning the discriminative feature representation is also limited as a less informative signal directly associated with the labels is provided, thus deteriorating the performance of the final instance-level classifier. In this article, delving into the label proportions, we bypass this weak supervision by leveraging generative adversarial networks (GANs) to derive an effective algorithm LLP-GAN. Endowed with an end-to-end structure, LLP-GAN performs approximation in the light of an adversarial learning mechanism without imposing restricted assumptions on distribution. Accordingly, the final instance-level classifier can be directly induced upon the discriminator with minor modification. Under mild assumptions, we give the explicit generative representation and prove the global optimality for LLP-GAN. In addition, compared with existing methods, our work empowers LLP solvers with desirable scalability inheriting from deep models. Extensive experiments on benchmark datasets and a real-world application demonstrate the vivid advantages of the proposed approach. Bo Wang 0049, Hanyuan Hang, Zhiquan Qi, Yingjie Tian 0001, Yong Shi 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2022 | A polynomial-time algorithm for simple undirected graph isomorphismabstractIn the author list, "Ferry Sansoto" should be Ferry Susanto.• To reflect more accurately the contribution of the article, the title should be changed to "A permutation and equinumerosity based polynomial-time algorithm for simple undirected graph isomorphism."• In the abstract, the "Pythagorean Triples Theorem" should be removed.• In the abstract, "squared sums of elements" should be "nth power sums."• In Section 2.2, "and the sum of the individual squared elements.By checking two sums," should be ", the sum of the individual squared elements and until the sum of the nth power of the nth element in the array.By checking these sums,"• In Section 2.2, "For both vertex and edge arrays of row/column sum based on the vertex and edge adjacency matrices, if and only if one array is a permutation of another one, the corresponding two graphs are isomorphic."should be "For both the vertex and edge arrays of row/column sum based on the vertex and edge adjacency matrices, if and only if one array is a permutation of another one and the corresponding edge and vertex's adjacent relationship has been preserved, the corresponding two graphs are isomorphic." Jing He 0004, Guangyan Huang, Jie Cao 0001, Zhiwang Zhang, Hui Zheng 0001, Peng Zhang 0063, Roozbeh Zarei, Ferry Susanto, Ruchuan Wang 0001, Yimu Ji 0001, Weibei Fan, Zhijun Xie, Xiancheng Wang, Mengjiao Guo, Chihung Chi, Jiekui Zhang, Youtao Li, Xiaojun Chen 0001, Yong Shi 0001, André Van Zundert |
Concurr. Comput. Pract. Exp. | 19 |
| 2022 | An explainable multi-sparsity multi-kernel nonconvex optimization least-squares classifier method via ADMM
Zhiwang Zhang, Jing He 0004, Jie Cao 0001, Xingsen Li, Kai Zhang 0074, Pingjiang Wang, Yong Shi 0001 |
Neural Comput. Appl. | 8 |
| 2022 | Self-supervised knowledge distillation for complementary label learning
Biao Li 0005, Minglong Lei, Yong Shi 0001 |
Neural Networks | 4 |
| 2022 | SELF-LLP: Self-supervised learning from label proportions with self-ensemble
Zhiquan Qi, Bo Wang 0049, Yingjie Tian 0001, Yong Shi 0001 |
Pattern Recognit. | 5 |
| 2022 | Graph regularized locally linear embedding for unsupervised feature selection
Jianyu Miao, Tiejun Yang, Xuan Fei, Lingfeng Niu, Yong Shi 0001 |
Pattern Recognit. | 6 |
| 2022 | Learning deep feature correspondence for unsupervised anomaly detection and segmentation
Jie Yang 0064, Yong Shi 0001, Zhiquan Qi |
Pattern Recognit. | 2 |
| 2022 | DigGCN: Learning Compact Graph Convolutional Networks via Diffusion AggregationabstractRecent interests in graph neural networks (GNNs) have received increasing concerns due to their superior ability in the network embedding field. The GNNs typically follow a message passing scheme and represent nodes by aggregating features from neighbors. However, the current aggregation methods assume that the network structure is static and define the local receptive fields under visible connections, which consequently fails to consider latent or high-order structures. Besides, the aggregation methods are known to have a depth dilemma due to the over-smoothness issues. To solve the above shortcomings, we present in this article a compact graph convolutional network framework which defines the graph receptive fields based on diffusion paths and explicitly compresses the neural networks with sparsity regularization. The proposed model seeks to learn from invisible connections and recover the latent proximity. First, we infer the high-order proximity and construct diffusion paths by diffusion samplings. Compared with random walk samplings, the diffusion samplings are based on regions instead of paths. The network inference then obtains accurate weights that can be leveraged to build small but informative receptive fields with salient neighbors. Second, to utilize the deep information while avoiding overfitting, we propose learning a lightweight model by introducing a nonconvex regularizer. Numerical comparisons with the existing network embedding methods under unsupervised feature learning and supervised classification show the effectiveness of our model. Minglong Lei, Pei Quan, Rongrong Ma, Yong Shi 0001, Lingfeng Niu |
IEEE Trans. Cybern. | 4 |
| 2022 | Fuzzy-Based Concept Learning Method: Exploiting Data With Fuzzy Conceptual ClusteringabstractConcepts have been adopted in concept-cognitive learning (CCL) and conceptual clustering for concept classification and concept discovery. However, the standard CCL algorithms are incapable of tackling continuous data directly, and some standard conceptual clustering methods mainly focus on the attribute information, ignoring the object information that is also important to improve clustering analysis and concept classification ability. Therefore, in this article, we present a novel concept learning method, called the fuzzy-based concept learning model (FCLM), to address these two issues by exploiting concept hierarchical relations in concept lattices. More specifically, we first show some new related notions for FCLM based on a regular fuzzy formal decision context; among these notions, the object-oriented and attribute-oriented fuzzy concept similarities are used to achieve the concept similarity measure in concept lattices. Moreover, a novel fuzzy concept learning framework is designed, and its corresponding learning algorithms are developed. Finally, we conduct some experiments on various real-world datasets to demonstrate that the proposed method can achieve the state-of-the-art classification performance among similarity-based learning methods. In addition, we further verify the effectiveness of our method in concept discovery on the MNIST dataset. Yunlong Mi, Yong Shi 0001, Jinhai Li 0001, Mengyu Yan |
IEEE Trans. Cybern. | 2 |
| 2022 | Measuring the Network Vulnerability Based on Markov CriticalityabstractVulnerability assessment—a critical issue for networks—attempts to foresee unexpected destructive events or hostile attacks in the whole system. In this article, we consider a new Markov global connectivity metric—Kemeny constant, and take its derivative called Markov criticality to identify critical links. Markov criticality allows us to find links that are most influential on the derivative of Kemeny constant. Thus, we can utilize it to identity a critical link ( i , j ) from node i to node j , such that removing it leads to a minimization of networks’ global connectivity, i.e., the Kemeny constant. Furthermore, we also define a novel vulnerability index to measure the average speed by which we can disconnect a specified ratio of links with network decomposition. Our method is of high efficiency, which can be easily employed to calculate the Markov criticality in real-life networks. Comprehensive experiments on several synthetic and real-life networks have demonstrated our method’s better performance by comparing it with state-of-the-art baseline approaches. Hui-Jia Li, Lin Wang 0012, Zhan Bu, Jie Cao 0001, Yong Shi 0001 |
ACM Trans. Knowl. Discov. Data | 5 |
| 2022 | Optimal Estimation of Low-Rank Factors via Feature Level Data Fusion of Multiplex Signal SystemsabstractThe design of fusion engines is a subject of great importance in a variety of fields. In this paper, we focus on the problem of linear fusion at the feature level for multiple signal matrices with noises, with the features being extremal eigenvectors. When given multiple similarity matrices, the objective is to find an estimate of the latent signal eigenspace. The concentration result for the inner product of features from different matrix samples is developed, utilizing the random matrix theory. Based on of the theoretical results, we proposed an efficient algorithm,EigFuse, to solve the constrained data-driven optimization problem with different level of noises. Our method is of high efficiency by comparing it with state-of-the-art baseline approaches with multiple noise levels. Comprehensive experiments on several synthetic as well as real-life networks demonstrate our method’s superior performance. Hui-Jia Li, Zhen Wang 0004, Jie Cao 0001, Jian Pei 0001, Yong Shi 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2022 | Semi-Supervised Concept Learning by Concept-Cognitive Learning and Concept SpaceabstractIn human concept learning, people can naturally combine a handful of labeled data with abundant unlabeled data when they make classification decisions, which is also known as semi-supervised learning (SSL) in machine learning. Especially, human concept learning not only is a static process in human cognition but also can vary gradually with dynamic environments. Nevertheless, the classical SSL algorithms must be redesigned to accommodate newly input data. In this sense, concept-cognitive learning may be a good choice, as it can implement dynamic processes by imitating human cognitive processes. Meanwhile, numerous SSL methods were designed based on the feature vector information of instances, while ignoring concept structural information that is a very important process in human knowledge organization. Based on this idea, a novel SSL method, named semi-supervised concept learning method (S2CL), is proposed for dynamic SSL by employing concept spaces, in which knowledge is represented by hierarchical concept structures. Moreover, to make full use of the global and local conceptual information, we further propose an extended version of S2CL (namely,$\text{S2CL}^{\alpha }$) for concept learning. More specifically, to effectively exploit the unlabeled data, this paper first shows some new related theories for S2CL (or$\text{S2CL}^{\alpha }$) based on a regular formal decision context; then a novel SSL framework is designed, and its corresponding algorithm is developed. Finally, we conduct some experiments on various datasets to demonstrate the effectiveness of our methods, which include concept classification and incremental learning under a large quantity of unlabeled data. Yunlong Mi, Yong Shi 0001, Jinhai Li 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2021 | A polynomial-time algorithm for simple undirected graph isomorphismabstractSummary The graph isomorphism problem is to determine two finite graphs that are isomorphic which is not known with a polynomial‐time solution. This paper solves the simple undirected graph isomorphism problem with an algorithmic approach as NP=P and proposes a polynomial‐time solution to check if two simple undirected graphs are isomorphic or not. Three new representation methods of a graph as vertex/edge adjacency matrix and triple tuple are proposed. A duality of edge and vertex and a reflexivity between vertex adjacency matrix and edge adjacency matrix were first introduced to present the core idea. Beyond this, the mathematical approval is based on an equivalence between permutation and bijection. Because only addition and multiplication operations satisfy the commutative law, we propose a permutation theorem to check fast whether one of two sets of arrays is a permutation of another or not. The permutation theorem was mathematically approved by Integer Factorization Theory, Pythagorean Triples Theorem, and Fundamental Theorem of Arithmetic. For each of two n ‐ary arrays, the linear and squared sums of elements were respectively calculated to produce the results. Jing He 0004, Jinjun Chen, Guangyan Huang, Jie Cao 0001, Zhiwang Zhang, Hui Zheng 0001, Peng Zhang 0063, Roozbeh Zarei, Ferry Sansoto, Ruchuan Wang 0001, Yimu Ji 0001, Weibei Fan, Zhijun Xie, Xiancheng Wang, Mengjiao Guo, Chihung Chi, Paulo A. de Souza, Jiekui Zhang, Youtao Li, Xiaojun Chen 0001, Yong Shi 0001, David G. Green, Taraporewalla Kersi, André Van Zundert |
Concurr. Comput. Pract. Exp. | 21 |
| 2021 | Stock movement prediction with sentiment analysis based on deep learning networksabstractAbstract With the development of Internet and big data, it is more convenient for investors to share opinions or have a discuss with others via the web, which creates massive unstructured data. These data reflect investors' emotions and their investment intentions, and it will further affect the movement of the stock market. Although researchers have been attempted to use sentiment information to predict the market, the sentiment features used are driven by outdated emotion extraction systems. In this article, we proposed a new sentiment analysis system with deep neural networks for stock comments and applied estimated sentiment information to the stock movement forecasting. The empirical results showed that our deep sentiment classification method achieved a 9% improvement over the logistic regression algorithm, and provided an accurate sentiment extractor for the next predicting step. In addition our new hybrid features that mix stock trading data and sentiment information achieved 1.25% improvement among 150 Chinese stocks in the testing dataset. For American stocks, the sentiment information would reduced the predicting results. We found that emotion features extracted from comments are indeed effective for stocks with a higher price to book value and a lower beta risk value in China. Yong Shi 0001, Yuanchun Zheng, Kun Guo 0001, Xinyue Ren |
Concurr. Comput. Pract. Exp. | 1 |
| 2021 | RGSR: A two-step lossy JPG image super-resolution based on noise reduction
Biao Li 0005, Yong Shi 0001, Bo Wang 0049, Zhiquan Qi |
Neurocomputing | 2 |
| 2021 | Unsupervised anomaly segmentation via deep feature reconstruction
Yong Shi 0001, Jie Yang 0064, Zhiquan Qi |
Neurocomputing | 1 |
| 2021 | Improved incremental local outlier detection for data streams based on the landmark window model
Aihua Li, Weijia Xu, Zhidong Liu, Yong Shi 0001 |
Knowl. Inf. Syst. | 4 |
| 2021 | Distant Supervision Relation Extraction via adaptive dependency-path and additional knowledge graph supervision
Yong Shi 0001, Yang Xiao 0017, Pei Quan, Minglong Lei, Lingfeng Niu |
Neural Networks | 1 |
| 2021 | Document-level relation extraction via graph transformer networks and temporal convolutional networks
Yong Shi 0001, Yang Xiao 0017, Pei Quan, Minglong Lei, Lingfeng Niu |
Pattern Recognit. Lett. | 1 |
| 2021 | SentiVec: Learning Sentiment-Context Vector via Kernel Optimization Function for Sentiment AnalysisabstractDeep learning-based sentiment analysis (SA) methods have drawn more attention in recent years, which calls for more precise word embedding methods. This article proposes SentiVec, a kernel optimization function system for sentiment word embedding, which is based on two phases. The first phase is a supervised learning method, and the second phase consists of two unsupervised updating models, object-word-to-surrounding-words reward model (O2SR) and context-to-object-word reward model (C2OR). SentiVec is aimed at: 1) integrating the statistical information and sentiment orientation into sentiment word vectors and 2) propagating and updating the semantic information to all the word representations in a corpus. Extensive experimental results show that the optimal sentiment vectors successfully extract the features in terms of semantic and sentiment information, which makes it outperform the baseline methods on word similarity, word analogy, and SA tasks. Wei Li 0076, Yong Shi 0001, Kun Guo 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Concept-Cognitive Learning Model for Incremental Concept LearningabstractConcept-cognitive learning (CCL) is an emerging field of concerning incremental concept learning and dynamic knowledge processing in the context of dynamic environments. Although CCL has been widely researched in theory, the existing studies of CCL have one problem: the concepts obtained by CCL systems do not have generalization ability. In the meantime, the existing incremental algorithms still face some challenges that: 1) classifiers have to adapt gradually and 2) the previously acquired knowledge should be efficiently utilized. To address these problems, based on the advantage that CCL can naturally integrate new data into itself for enhancing flexibility of concept learning, we first propose a new CCL model (CCLM) to extend the classical methods of CCL, which is not only a new classifier but also good at incremental learning. Unlike the existing CCL systems, the theory of CCLM is mainly based on a formal decision context rather than a formal context. In learning concepts from dynamic environments, we show that CCLM can naturally incorporate new data into itself with a sufficient theoretical guarantee for incremental learning. For classification task and knowledge storage, our results on various data sets demonstrate that CCLM can simultaneously: 1) achieve the state-of-the-art static and dynamic classification task and 2) directly accomplish preservation of previously acquired knowledge (or concepts) under dynamic environments. Yong Shi 0001, Yunlong Mi, Jinhai Li 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2021 | A Construction of Robust Representations for Small Data Sets Using Broad Learning SystemabstractFeature processing is an important step for modeling and can improve the accuracy of machine learning models. Feature extraction methods can effectively extract features from high-dimensional data sets and enhance the accuracy of tasks. However, the performance of feature extraction methods is not stable in low-dimensional data sets. This article extends the broad learning system (BLS) to a framework for constructing robust representations in low-dimensional and small data sets. First, the BLS changed from a supervised prediction method to an ensemble feature extraction method. Second, feature extraction methods instead of random mapping are used to generate mapped features. Third, deep representations, called enhancement features, are learned from the ensemble mapped features. Fourth, data for generating mapped features and enhancement features can be randomly selected. The ensemble of mapped features and enhancement features can provide robust representations to enhance the performance of downstream tasks. A label-based autoencoder (LA) is embedded in the BLS framework as an example to show the effectiveness of the framework. A random LA (RLA) is presented to generate more different features. The experimental results show that the BLS framework can construct robust representations and significantly promote the performance of machine learning models. Peiwu Dong, Yong Shi 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2020 | Learning to Incorporate Structure Knowledge for Image InpaintingabstractThis paper develops a multi-task learning framework that attempts to incorporate the image structure knowledge to assist image inpainting, which is not well explored in previous works. The primary idea is to train a shared generator to simultaneously complete the corrupted image and corresponding structures — edge and gradient, thus implicitly encouraging the generator to exploit relevant structure knowledge while inpainting. In the meantime, we also introduce a structure embedding scheme to explicitly embed the learned structure features into the inpainting process, thus to provide possible preconditions for image completion. Specifically, a novel pyramid structure loss is proposed to supervise structure learning and embedding. Moreover, an attention mechanism is developed to further exploit the recurrent structures and patterns in the image to refine the generated structures and contents. Through multi-task learning, structure embedding besides with attention, our framework takes advantage of the structure knowledge and outperforms several state-of-the-art methods on benchmark datasets quantitatively and qualitatively. Jie Yang 0064, Zhiquan Qi, Yong Shi 0001 |
AAAI | 3 |
| 2020 | Ensemble learning with label proportions for bankruptcy prediction
Zhensong Chen 0001, Wei Chen 0061, Yong Shi 0001 |
Expert Syst. Appl. | 3 |
| 2020 | Deep learning from label proportions with labeled samples
Yong Shi 0001, Bo Wang 0049, Zhiquan Qi, Yingjie Tian 0001 |
Neural Networks | 1 |
| 2020 | Parallel RMCLP Classification Algorithm and Its Application on the Medical DataabstractTo make better use of the cloud computing technology, and to overcome the computing and storage requirements which increase rapidly with the number of training samples, in this paper, a new parallel algorithm is proposed-Parallel Regularized Multiple-Criteria Linear Programming (PRMCLP) algorithm-The RMCLP model is converted into a unconstrained optimization problem, and then, in the parallel version, it is split into several tasks, where each part is mapped and computed on a separate processor. This approach enables us to obtain efficiently the final optimization solution of the whole classification problem. At last, we apply this algorithm to Medical data classification. All experiments show that our method and approach greatly increases the training speed of RMCLP in the parallel case. Zhiquan Qi, Yingjie Tian 0001, Yong Shi 0001, Vassil Alexandrov 0001 |
IEEE Trans. Cloud Comput. | 3 |
| 2020 | s-LWSR: Super Lightweight Super-Resolution NetworkabstractIn recent years, deep-based models have achieved great success in the field of single image super-resolution (SISR), where tremendous parameters are always needed to obtain a satisfying performance. However, the high computational complexity extremely limits its applications to some mobile devices that possess less computing and storage resources. To address this problem, in this paper, we propose a flexibly adjustable super lightweight SR network: s-LWSR. Firstly, in order to efficiently abstract features from the low resolution image, we design a high-efficient U-shape based block, where an information pool is constructed to mix multi-level information from the first half part of the pipeline. Secondly, a compression mechanism based on depth-wise separable convolution is employed to further reduce the numbers of parameters with negligible performance degradation. Thirdly, by revealing the specific role of activation in deep models, we remove several activation layers in our SR model to retain more information, thus leading to the final performance improvement. Extensive experiments show that our s-LWSR, with limited parameters and operations, can achieve similar performance compared with other cumbersome DL-SR methods. Biao Li 0005, Bo Wang 0049, Zhiquan Qi, Yong Shi 0001 |
IEEE Trans. Image Process. | 5 |
| 2020 | Graph K-means Based on Leader Identification, Dynamic Game, and Opinion DynamicsabstractWith the explosion of social media networks, many modern applications are concerning about people's connections, which leads to the so-called social computing. An elusive question is to study how opinion communities form and evolve in real-world networks with great individual diversity and complex human connections. In this scenario, the classic K-means technique and its extended versions could not be directly applied, as they largely ignore the relationship among interactive objects. On the other side, traditional community detection approaches in statistical physics would be neither adequate nor fair: they only consider the network topological structure but ignore the heterogeneous-objects' attributive information. To this end, we attempt to model a realistic social media network as a discrete-time dynamical system, where the opinion matrix and the community structure could mutually affect each other. In this paper, community detection in social media networks is naturally formulated as a multi-objective optimization problem (MOOP), i.e., finding a set of densely connected components with similar opinion vectors. We propose a novel and powerful graph K-means framework, which is composed of three coupled phases in each discrete-time period. Specifically, the first phase uses a fast heuristic approach to identify those opinion leaders who have relatively high local reputation; the second phase adopts a novel dynamic game model to find the locally Pareto-optimal community structure; and the final phase employs a robust opinion dynamics model to simulate the evolution of the opinion matrix. We conduct a series of comprehensive experiments on real-world benchmark networks to validate the performance of GK-means through comparisons with the state-of-the-art graph clustering technologies. Zhan Bu, Hui-Jia Li, Chengcui Zhang, Jie Cao 0001, Aihua Li, Yong Shi 0001 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2020 | Classifying With Adaptive Hyper-Spheres: An Incremental Classifier Based on Competitive LearningabstractNowadays, datasets are always dynamic and patterns in them are changing. Instances with different labels are intertwined and often linearly inseparable, which bring new challenges to traditional learning algorithms. This paper proposes adaptive hyper-sphere (AdaHS), an adaptive incremental classifier, and its kernelized version: Nys-AdaHS. The classifier incorporates competitive training with a border zone. With adaptive hidden layer and tunable radii of hyper-spheres, AdaHS has strong capability of local learning like instance-based algorithms, but free from slow searching speed and excessive memory consumption. The experiments showed that AdaHS is robust, adaptive, and highly accurate. It is especially suitable for dynamic data in which patterns are changing, decision borders are complicated, and instances with the same label can be spherically clustered. Gang Kou, Yi Peng 0001, Yong Shi 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2020 | Recommender system for marketing optimization
Yong Shi 0001, Zhengxin Chen, Wikil Kwak |
World Wide Web | 2 |
| 2019 | Learning from Label Proportions with Generative Adversarial NetworksabstractIn this paper, we leverage generative adversarial networks (GANs) to derive an effective algorithm LLP-GAN for learning from label proportions (LLP), where only the bag-level proportional information in labels is available. Endowed with end-to-end structure, LLP-GAN performs approximation in the light of an adversarial learning mechanism, without imposing restricted assumptions on distribution. Accordingly, we can directly induce the final instance-level classifier upon the discriminator. Under mild assumptions, we give the explicit generative representation and prove the global optimality for LLP-GAN. Additionally, compared with existing methods, our work empowers LLP solver with capable scalability inheriting from deep models. Several experiments on benchmark datasets demonstrate vivid advantages of the proposed approach. Bo Wang 0049, Zhiquan Qi, Yingjie Tian 0001, Yong Shi 0001 |
NeurIPS | 5 |
| 2019 | Public blockchain evaluation using entropy and TOPSIS
Yong Shi 0001, Peiwu Dong |
Expert Syst. Appl. | 2 |
| 2019 | Concurrent concept-cognitive learning model for classification
Yong Shi 0001, Yunlong Mi, Jinhai Li 0001 |
Inf. Sci. | 1 |
| 2019 | Fast kernel extreme learning machine for ordinal regression
Yong Shi 0001, Peijia Li, Jianyu Miao, Lingfeng Niu |
Knowl. Based Syst. | 1 |
| 2019 | Feature selection with MCP $$^2$$ 2 regularization
Yong Shi 0001, Jianyu Miao, Lingfeng Niu |
Neural Comput. Appl. | 1 |
| 2019 | Constrained matrix factorization for semi-weakly learning with label proportions
Zhensong Chen 0001, Yong Shi 0001, Zhiquan Qi |
Pattern Recognit. | 2 |
| 2019 | Diffusion network embedding
Yong Shi 0001, Minglong Lei, Hong Yang 0003, Lingfeng Niu |
Pattern Recognit. | 1 |
| 2018 | City Brain, a New Architecture of Smart City Based on the Internet BrainabstractIn the ten years after the Smart City was put forward, there are still problems like the unclear concept, lack of top-down design and information island. With the further development of the Internet, the brain-like architecture of the Internet is becoming clearer and clearer. As a product of a combination of city buildings and the Internet, the Smart City will also have a new architecture, and the City Brain thus appears. Based on the Internet Brain, this paper describes how to construct the Smart City in the form of brain-like tissue, and how to evaluate the construction level of the Smart City (City IQ) relying on the Big SNS (city neural networks) and city cloud reflex arcs. Liu Feng, Fangyao Liu, Yong Shi 0001 |
CSCWD | 3 |
| 2018 | Multi-View Fusion Through Cross-Modal RetrievalabstractCross-modal retrieval, which takes text queries to retrieve relevant images or vice versa, has drawn much attention in recent years. This topic exhibits dual-heterogeneity: heterogeneity of different modalities and heterogeneous features obtained from multiple views. To address this issue, we propose an effective multi-view fusion method for cross-modal retrieval based on tensor modeling (CMTM) for cross-modal retrieval from the full-order feature interactions within the multimodal data. In order to facilitate integration of heterogeneous features from multiple views, we adopt the tensor structure to model the full-order interactions among the multi-view features effectively. Besides, a tensor factorization is applied to derive model parameters. Extensive experiments demonstrate the effectiveness of CMTM on cross-modal retrieval. Limeng Cui, Zhensong Chen 0001, Jiawei Zhang 0001, Lifang He 0001, Yong Shi 0001, Philip S. Yu |
ICIP | 5 |
| 2018 | Multi-view Collective Tensor Decomposition for Cross-modal HashingabstractMultimedia data available in various disciplines are usually heterogeneous, containing representations in multi-views, where the cross-modal search techniques become necessary and useful. It is a challenging problem due to the heterogeneity of data with multiple modalities, multi-views in each modality and the diverse data categories. In this paper, we propose a novel multi-view cross-modal hashing method named Multi-view Collective Tensor Decomposition (MCTD) to fuse these data effectively, which can exploit the complementary feature extracted from multi-modality multi-view while simultaneously discovering multiple separated subspaces by leveraging the data categories as supervision information. Our contributions are summarized as follows: 1) we exploit tensor modeling to get better representation of the complementary features and redefine a latent representation space; 2) a block-diagonal loss is proposed to explicitly pursue a more discriminative latent tensor space by exploring supervision information; 3) we propose a new feature projection method to characterize the data and to generate the latent representation for incoming new queries. An optimization algorithm is proposed to solve the objective function designed for MCTD, which works under an iterative updating procedure. Experimental results prove the state-of-the-art precision of MCTD compared with competing methods. Limeng Cui, Zhensong Chen 0001, Jiawei Zhang 0001, Lifang He 0001, Yong Shi 0001, Philip S. Yu |
ICMR | 5 |
| 2018 | The Applications of Stochastic Models in Network Embedding: A SurveyabstractNetwork embedding is a promising topic that maps the vertices to the latent space while keeps the structural proximity in the original space. The network embedding task is difficult since the network vertices have no specific time or space orders. Models that used to extract information from images and texts with regular space or time structures can not be directly applied in network heading. The key feature of network embedding methods should be further exploited. Previous network embedding reviews mainly focus on the models and algorithms used in different methods. In this survey, we review the network embedding works in the stochastic perspective either in data side or model side. Roughly, the network embedding methods fall into three main categories: matrix based methods, random walk based methods and aggregated based methods. We focus on the applications of stochastic models in solving the challenges of network embedding in data processing and modeling following the line of the three categories. Minglong Lei, Yong Shi 0001, Lingfeng Niu |
WI | 2 |
| 2018 | Inverse Convolutional Neural Networks for Learning from Label ProportionsabstractLearning from label proportions (LLP) is a new kind of learning problem which has attracted wide interest in the field of machine learning. Different from the well-known supervised learning, the training data of LLP is in form of bags and only the proportion of each class in each bag is available. Actually, many modern applications can be abstracted to this problem such as modeling voting behaviors and spam filtering. In this paper, we propose an end-to-end LLP model based on convolutional neural network called IDLLP, which employs the the idea of inverting a classifier calibration process to learn a classifier from bag probabilities. Firstly, convolutional neural network regression is used to estimate the values obtained by inverting the probability of each bag. Secondly, stochastic gradient descent based on batch is adapt to train the model, where the batch size depends on the bag size. At last, experiments demonstrate that our algorithm can obtain the best accuracies on image data compared with several recently developed methods. Yong Shi 0001, Zhiquan Qi |
WI | 1 |
| 2018 | Sparse feature kernel multi-criteria linear programming classifier
Zhiwang Zhang, Guangxia Gao, Yong Shi 0001 |
Neurocomputing | 4 |
| 2018 | DWWP: Domain-specific new words detection and word propagation system for sentiment analysis in the tourism domain
Wei Li 0076, Kun Guo 0001, Yong Shi 0001, Yuanchun Zheng |
Knowl. Based Syst. | 3 |
| 2018 | Learning from label proportions on high-dimensional data
Yong Shi 0001, Zhiquan Qi, Bo Wang 0049 |
Neural Networks | 1 |
| 2018 | Large-scale linear nonparallel SVMs
Dalian Liu, Dewei Li 0002, Yong Shi 0001, Yingjie Tian 0001 |
Soft Comput. | 3 |
| 2018 | Adaboost-LLP: A Boosting Method for Learning With Label ProportionsabstractHow to solve the classification problem with only label proportions has recently drawn increasing attention in the machine learning field. In this paper, we propose an ensemble learning strategy to deal with the learning problem with label proportions (LLP). In detail, we first give a loss function based on different weights for LLP, and then construct the corresponding weak classifier, at the same time, estimate its conditional probabilities by a standard logistic function. At last, by introducing the maximum likelihood estimation, we propose a new anyboost learning system for LLP (called Adaboost-LLP). Unlike traditional methods, our method does not make any restrictive assumptions on training set; at the same time, compared with alter- SVM, Adaboost-LLP exploits more extra weight information and uses multiple weak classifiers that can be solved efficiently to combine a strong classifier. All experiments show that our method outperforms the existing methods in both accuracy and training time. Zhiquan Qi, Yingjie Tian 0001, Lingfeng Niu, Yong Shi 0001, Peng Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2018 | Feature Selection With ℓ2, 1-2 RegularizationabstractFeature selection aims to select a subset of features from high-dimensional data according to a predefined selecting criterion. Sparse learning has been proven to be a powerful technique in feature selection. Sparse regularizer, as a key component of sparse learning, has been studied for several years. Although convex regularizers have been used in many works, there are some cases where nonconvex regularizers outperform convex regularizers. To make the process of selecting relevant features more effective, we propose a novel nonconvex sparse metric on matrices as the sparsity regularization in this paper. The new nonconvex regularizer could be written as the difference of the $\ell _{2,1}$ norm and the Frobenius ( $\ell _{2,2}$ ) norm, which is named the $\ell _{2,1-2}$ . To find the solution of the resulting nonconvex formula, we design an iterative algorithm in the framework of ConCave-Convex Procedure (CCCP) and prove its strong global convergence. An adopted alternating direction method of multipliers is embedded to solve the sequence of convex subproblems in CCCP efficiently. Using the scaled cluster indictors of data points as pseudolabels, we also apply $\ell _{2,1-2}$ to the unsupervised case. To the best of our knowledge, it is the first work considering nonconvex regularization for matrices in the unsupervised learning scenario. Numerical experiments are performed on real-world data sets to demonstrate the effectiveness of the proposed method. Yong Shi 0001, Jianyu Miao, Peng Zhang 0001, Lingfeng Niu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2018 | An interview with Professor Raj Reddy on Web Intelligence (WI) and Computational Social Science (CSS)
Ning Zhong 0001, Jiming Liu 0001, Yong Shi 0001, Yiyu Yao |
Web Intell. | 3 |
| 2017 | Inverse extreme learning machine for learning with label proportionsabstractIn large-scale learning problem, the scalability of learning algorithms is usually the key factor affecting the algorithm practical performance, which is determined by both the time complexity of the learning algorithms and the amount of supervision information (i.e., labeled data). Learning with label proportions (LLP) is a new kind of machine learning problem which has drawn much attention in recent years. Different from the well-known supervised learning, LLP can estimate a classifier from groups of weakly labeled data, where only the positive/negative class proportions of each group are known. Due to its weak requirements for the input data, LLP presents a variety of real-world applications in almost all the fields involving anonymous data, like computer vision, fraud detection and spam filtering. However, even through the required labeled data is of a very small amount, LLP still suffers from the long execution time a lot due to the high time complexity of the learning algorithm itself. In this paper, we propose a very fast learning method based on inversing output scaling process and extreme learning machine, namely Inverse Extreme Learning Machine (IELM), to address the above issues. IELM can speed up the training process by order of magnitudes for large datasets, while achieving highly competitive classification accuracy with the existing methods at the same time. Extensive experiments demonstrate the significant speedup of the proposed method. We also demonstrate the feasibility of IELM with a case study in real-world setting: modeling image attributes based on ImageNet Object Attributes dataset. Limeng Cui, Jiawei Zhang 0001, Zhensong Chen 0001, Yong Shi 0001, Philip S. Yu |
IEEE BigData | 4 |
| 2017 | Augmented SVM with ordinal partitioning for text classificationabstractOrdinal regression has received increasing interest in the past years. It aims to classify patterns by an ordinal scale. With the the explosive growth of data, the method of SVM with ordinal partitioning called SVMOP highlights its advantages due to its convenience of dealing with large scale data. However, the method of SVMOP for ordinal regression has not been exploited much. As we know, the costs should be different when dealing with mislabeled samples and how to use them plays a dominant role in model building. However, L2-loss which could enlarge the cost sensitivity has not been applied into SVM ordinal partition yet. In this paper, we propose the method of SVMOP with L2-loss for ordinal regression. Numerical results show that our approach outperforms the method of SVMOP with L1-loss and other ordianl regression models. Yong Shi 0001, Peijia Li, Lingfeng Niu |
WI | 1 |
| 2017 | A study on error correction of multiple criteria and multiple constraint levels linear programming based classificationabstractIn credit card client classification problem, reducing misclassified rate is regarded as a key issue. Unfortunately, existing machine learning methods cannot be successfully applied to this problem. This paper introduces a classification model based on multiple criteria and multiple constraint levels linear programming (MC2LP), which equips two intervals of cutoff in the model. Two parallel hyperplanes are employed to indicate the relative positions between the points and hyperplanes. Then, we discuss the correctness of new model in error correction. Matrix representations of relevant models are also offered. Finally, compared to original MCLP, known MC2LP, Logistic Regression (LR) and Support Vector Machine (SVM), the propsed model shows superiority in solving two types of error related data mining problem. Bo Wang 0049, Yong Shi 0001 |
WI | 2 |
| 2017 | Ramp loss K-Support Vector Classification-Regression; a robust and sparse multi-class approach to the intrusion detection problem
Seyed Mojtaba Hosseini Bamakan, Yong Shi 0001 |
Knowl. Based Syst. | 3 |
| 2017 | Learning with label proportions based on nonparallel support vector machines
Zhensong Chen 0001, Zhiquan Qi, Bo Wang 0049, Limeng Cui, Yong Shi 0001 |
Knowl. Based Syst. | 6 |
| 2017 | A novel clustering-based image segmentation via density peaks algorithm with mid-level feature
Yong Shi 0001, Zhensong Chen 0001, Zhiquan Qi, Limeng Cui |
Neural Comput. Appl. | 1 |
| 2017 | Nonparallel Support Vector Ordinal RegressionabstractOrdinal regression is a supervised learning problem where training samples are labeled by an ordinal scale. The ordering relation and nonmetric property of the label set distinguish it from the multiclass classification and metric regression. To better exploit the inherent structure in the label and benefit from the hidden information in data distribution, we propose a novel ordinal regression model, which is named as nonparallel support vector ordinal regression (NPSVOR) to emphasis the utilization of nonparallel proximal hyperplanes. The new model constructs a hyperplane for each rank such that the patterns of this rank lie in the close proximity while maintaining clear separation with the other ranks. Since the learning of hyperplanes can be carried out independently, NPSVOR can be trained in parallel. Furthermore, we design an efficient solver at the same time for training the hyperplanes in NPSVOR based on the alternating direction method of multipliers. Extensive experimentation demonstrates that NPSVOR yields a large and statistically significant improvement in terms of generalization performance and training speed against nine baselines. Yong Shi 0001, Lingfeng Niu, Yingjie Tian 0001 |
IEEE Trans. Cybern. | 2 |
| 2016 | An effective intrusion detection framework based on MCLP/SVM optimized by time-varying chaos particle swarm optimization
Seyed Mojtaba Hosseini Bamakan, Yingjie Tian 0001, Yong Shi 0001 |
Neurocomputing | 4 |
| 2016 | A divide-and-combine method for large scale nonparallel support vector machines
Yingjie Tian 0001, Xuchan Ju, Yong Shi 0001 |
Neural Networks | 3 |
| 2016 | Automatic Road Crack Detection Using Random Structured ForestsabstractCracks are a growing threat to road conditions and have drawn much attention to the construction of intelligent transportation systems. However, as the key part of an intelligent transportation system, automatic road crack detection has been challenged because of the intense inhomogeneity along the cracks, the topology complexity of cracks, the inference of noises with similar texture to the cracks, and so on. In this paper, we propose CrackForest, a novel road crack detection framework based on random structured forests, to address these issues. Our contributions are shown as follows: 1) apply the integral channel features to redefine the tokens that constitute a crack and get better representation of the cracks with intensity inhomogeneity; 2) introduce random structured forests to generate a high-performance crack detector, which can identify arbitrarily complex cracks; and 3) propose a new crack descriptor to characterize cracks and discern them from noises effectively. In addition, our method is faster and easier to parallel. Experimental results prove the state-of-the-art detection precision of CrackForest compared with competing methods. Yong Shi 0001, Limeng Cui, Zhiquan Qi, Zhensong Chen 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2016 | Fast and Accurate Mining the Community Structure: Integrating Center Locating and Membership OptimizationabstractMining communities or clusters in networks is valuable in analyzing, designing, and optimizing many natural and engineering complex systems, e.g., protein networks, power grid, and transportation systems. Most of the existing techniques view the community mining problem as an optimization problem based on a given quality function(e.g., modularity), however none of them are grounded with a systematic theory to identify the central nodes in the network. Moreover, how to reconcile the mining efficiency and the community quality still remains an open problem. In this paper, we attempt to address the above challenges by introducing a novel algorithm. First, a kernel function with a tunable influence factor is proposed to measure the leadership of each node, those nodes with highest local leadership can be viewed as the candidate central nodes. Then, we use a discrete-time dynamical system to describe the dynamical assignment of community membership; and formulate the serval conditions to guarantee the convergence of each node's dynamic trajectory, by which the hierarchical community structure of the network can be revealed. The proposed dynamical system is independent of the quality function used, so could also be applied in other community mining models. Our algorithm is highly efficient: the computational complexity analysis shows that the execution time is nearly linearly dependent on the number of nodes in sparse networks. We finally give demonstrative applications of the algorithm to a set of synthetic benchmark networks and also real-world networks to verify the algorithmic performance. Hui-Jia Li, Zhan Bu, Aihua Li, Zhidong Liu, Yong Shi 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2016 | Advertisement clicking prediction by using multiple criteria mathematical programming
Yong Shi 0001, Heeseok Lee, Heung Kee Kim |
World Wide Web | 2 |
| 2015 | Kernel based simple regularized multiple criteria linear program for binary classification and regressionabstractHandling data classification and regression problems through linear hyperplane is a naive and simple idea. In this paper, inspired by the idea of multiple criteria linear programs (MCLP) and multiple criteria quadratic programs (MCQP), we proposed a novel method for binary classification and regres sion problem. There are two main advantages for the proposed approach. One is that both of these two models guarantee the existence of feasible solutions when the model parameters were chosen properly. The other is that nonlinear patterns could be handled and captured by introducing kernel function into MCLP framework with a more natural way than previous work. Various classical approaches and datasets were evaluated in our experiments, and the result on both toy and real world data demonstrate the correctness and effectiveness of our proposed methods. Yong Shi 0001, Lingfeng Niu |
Intell. Data Anal. | 2 |
| 2015 | Node-coupling clustering approaches for link prediction
Fenhua Li, Jing He 0004, Guangyan Huang, Yanchun Zhang, Yong Shi 0001, Rui Zhou 0001 |
Knowl. Based Syst. | 5 |
| 2015 | Ramp loss nonparallel support vector machine for pattern classification
Dalian Liu, Yong Shi 0001, Yingjie Tian 0001 |
Knowl. Based Syst. | 2 |
| 2015 | Successive Overrelaxation for Laplacian Support Vector MachineabstractSemisupervised learning (SSL) problem, which makes use of both a large amount of cheap unlabeled data and a few unlabeled data for training, in the last few years, has attracted amounts of attention in machine learning and data mining. Exploiting the manifold regularization (MR), Belkin et al. proposed a new semisupervised classification algorithm: Laplacian support vector machines (LapSVMs), and have shown the state-of-the-art performance in SSL field. To further improve the LapSVMs, we proposed a fast Laplacian SVM (FLapSVM) solver for classification. Compared with the standard LapSVM, our method has several improved advantages as follows: 1) FLapSVM does not need to deal with the extra matrix and burden the computations related to the variable switching, which make it more suitable for large scale problems; 2) FLapSVM’s dual problem has the same elegant formulation as that of standard SVMs. This means that the kernel trick can be applied directly into the optimization model; and 3) FLapSVM can be effectively solved by successive overrelaxation technology, which converges linearly to a solution and can process very large data sets that need not reside in memory. In practice, combining the strategies of random scheduling of subproblem and two stopping conditions, the computing speed of FLapSVM is rigidly quicker to that of LapSVM and it is a valid alternative to PLapSVM. Zhiquan Qi, Yingjie Tian 0001, Yong Shi 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2014 | A new classification model using privileged information and its application
Zhiquan Qi, Yingjie Tian 0001, Yong Shi 0001 |
Neurocomputing | 3 |
| 2014 | Regularized multiple-criteria linear programming with universum and its application
Zhiquan Qi, Yingjie Tian 0001, Yong Shi 0001 |
Neural Comput. Appl. | 3 |
| 2014 | Self-Universum support vector machine
Dalian Liu, Yingjie Tian 0001, Rongfang Bie, Yong Shi 0001 |
Pers. Ubiquitous Comput. | 4 |
| 2014 | Nonparallel Support Vector Machines for Pattern ClassificationabstractWe propose a novel nonparallel classifier, called nonparallel support vector machine (NPSVM), for binary classification. Our NPSVM that is fully different from the existing nonparallel classifiers, such as the generalized eigenvalue proximal support vector machine (GEPSVM) and the twin support vector machine (TWSVM), has several incomparable advantages: 1) two primal problems are constructed implementing the structural risk minimization principle; 2) the dual problems of these two primal problems have the same advantages as that of the standard SVMs, so that the kernel trick can be applied directly, while existing TWSVMs have to construct another two primal problems for nonlinear cases based on the approximate kernel-generated surfaces, furthermore, their nonlinear problems cannot degenerate to the linear case even the linear kernel is used; 3) the dual problems have the same elegant formulation with that of standard SVMs and can certainly be solved efficiently by sequential minimization optimization algorithm, while existing GEPSVM or TWSVMs are not suitable for large scale problems; 4) it has the inherent sparseness as standard SVMs; 5) existing TWSVMs are only the special cases of the NPSVM when the parameters of which are appropriately chosen. Experimental results on lots of datasets show the effectiveness of our method in both sparseness and classification accuracy, and therefore, confirm the above conclusion further. In some sense, our NPSVM is a new starting point of nonparallel classifiers. Yingjie Tian 0001, Zhiquan Qi, Xuchan Ju, Yong Shi 0001, Xiaohui Liu 0001 |
IEEE Trans. Cybern. | 4 |
| 2013 | Spatial distance join based feature selection
Yong Shi 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2013 | Extending twin support vector machine classifier for multi-category classification problemsabstractTwin support vector machine classifier (TWSVM) was proposed by Jayadeva et al., which was used for binary classification problems. TWSVM not only overcomes the difficulties in handling the problem of exemplar unbalance in binary classification proble Juanying Xie, Kate S. Hone, Weixin Xie, Xinbo Gao 0001, Yong Shi 0001, Xiaohui Liu 0001 |
Intell. Data Anal. | 5 |
| 2013 | Structural twin support vector machine for classification
Zhiquan Qi, Yingjie Tian 0001, Yong Shi 0001 |
Knowl. Based Syst. | 3 |
| 2013 | Efficient railway tracks detection and turnouts recognition method using HOG features
Zhiquan Qi, Yingjie Tian 0001, Yong Shi 0001 |
Neural Comput. Appl. | 3 |
| 2013 | Multi-instance classification based on regularized multiple criteria linear programming
Zhiquan Qi, Yingjie Tian 0001, Yong Shi 0001 |
Neural Comput. Appl. | 3 |
| 2013 | Robust twin support vector machine for pattern classification
Zhiquan Qi, Yingjie Tian 0001, Yong Shi 0001 |
Pattern Recognit. | 3 |
| 2013 | The analytic hierarchy process: task scheduling and resource allocation in cloud computing environment
Daji Ergu, Gang Kou, Yi Peng 0001, Yong Shi 0001 |
J. Supercomput. | 4 |
| 2013 | Parallel data mining techniques on Graphics Processing Unit with Compute Unified Device Architecture (CUDA)
Liheng Jian, Ying Liu 0039, Shenshen Liang, Weidong Yi, Yong Shi 0001 |
J. Supercomput. | 6 |
| 2012 | China's national personal credit scoring system: a real-life intelligent knowledge applicationabstractCredit Reference Centre (CRC) of People's Bank of China (PBC) has built a big data: the largest personal credit database in the world with 800 million people's accounts collected from all commercial banks in China since 2003. From June 2006 to Sept 2009, Research Centre on Fictitious Economy and Data Science, Chinese Academy of Sciences (CASFEDS) and CRC jointly developed China's National Personal Credit Scoring System, known as "China Score", which is a unique and advanced KDD application under intelligent knowledge management on this big data. The system will be eventually serving all 1.3 billion population of China for their daily financial activities, such as bank accounts, credit card application, mortgage, personal loans, etc. It can become one of the most influential events of KDD techniques to human kind. This talk will introduce the key components of China Score project that includes objectives, modeling process, KDD techniques used in the projects, intelligent knowledge management and experience of the project development. In addition, the talk will also outline a number of policy recommendations based on China Score project which has been potentially impacting Chinese Government on its strategic decision making for China's economic developments. Yong Shi 0001 |
KDD | 1 |
| 2012 | A framework for application-driven classification of data streams
Peng Zhang 0001, Byron J. Gao, Ping Liu 0001, Yong Shi 0001, Li Guo 0001 |
Neurocomputing | 4 |
| 2012 | Data mining for software trustworthiness
Gang Kou, Yong Shi 0001, Guozhu Dong |
Inf. Sci. | 2 |
| 2012 | Distributed data possession checking for securing multiple replicas in geographically-dispersed clouds
Jing He 0004, Yanchun Zhang, Guangyan Huang, Yong Shi 0001, Jie Cao 0001 |
J. Comput. Syst. Sci. | 4 |
| 2012 | Training the max-margin sequence model with the relaxed slack variables
Lingfeng Niu, Jianmin Wu, Yong Shi 0001 |
Neural Networks | 3 |
| 2012 | Laplacian twin support vector machine for semi-supervised classification
Zhiquan Qi, Yingjie Tian 0001, Yong Shi 0001 |
Neural Networks | 3 |
| 2012 | Twin support vector machine with Universum data
Zhiquan Qi, Yingjie Tian 0001, Yong Shi 0001 |
Neural Networks | 3 |
| 2011 | Rigorous assessment and integration of the sequence and structure based features to predict hot spotsabstractBACKGROUND: Systematic mutagenesis studies have shown that only a few interface residues termed hot spots contribute significantly to the binding free energy of protein-protein interactions. Therefore, hot spots prediction becomes increasingly important for well understanding the essence of proteins interactions and helping narrow down the search space for drug design. Currently many computational methods have been developed by proposing different features. However comparative assessment of these features and furthermore effective and accurate methods are still in pressing need. RESULTS: In this study, we first comprehensively collect the features to discriminate hot spots and non-hot spots and analyze their distributions. We find that hot spots have lower relASA and larger relative change in ASA, suggesting hot spots tend to be protected from bulk solvent. In addition, hot spots have more contacts including hydrogen bonds, salt bridges, and atomic contacts, which favor complexes formation. Interestingly, we find that conservation score and sequence entropy are not significantly different between hot spots and non-hot spots in Ab+ dataset (all complexes). While in Ab- dataset (antigen-antibody complexes are excluded), there are significant differences in two features between hot pots and non-hot spots. Secondly, we explore the predictive ability for each feature and the combinations of features by support vector machines (SVMs). The results indicate that sequence-based feature outperforms other combinations of features with reasonable accuracy, with a precision of 0.69, a recall of 0.68, an F1 score of 0.68, and an AUC of 0.68 on independent test set. Compared with other machine learning methods and two energy-based approaches, our approach achieves the best performance. Moreover, we demonstrate the applicability of our method to predict hot spots of two protein complexes. CONCLUSION: Experimental results show that support vector machine classifiers are quite effective in predicting hot spots based on sequence features. Hot spots cannot be fully predicted through simple analysis based on physicochemical characteristics, but there is reason to believe that integration of features and machine learning methods can remarkably improve the predictive performance for hot spots. Ruoying Chen, Sixiao Yang, Yingjie Tian 0001, Yong Shi 0001 |
BMC Bioinform. | 7 |
| 2011 | Multiple criteria decision making and decision support systems - Guest editor's introduction
Gang Kou, Yong Shi 0001, Shou-Yang Wang |
Decis. Support Syst. | 2 |
| 2011 | Robust ensemble learning for mining noisy data streams
Peng Zhang 0001, Xingquan Zhu 0001, Yong Shi 0001, Li Guo 0001, Xindong Wu 0001 |
Decis. Support Syst. | 3 |
| 2011 | Multiple-kernel SVM based multiple-task oriented data mining system for gene expression data analysis
Zhen-Yu Chen 0001, Jianping Li 0001, Liwei Wei, Weixuan Xu, Yong Shi 0001 |
Expert Syst. Appl. | 5 |
| 2011 | A weighted Lq adaptive least squares support vector machine classifiers - Robust and sparse approximation
Jingli Liu, Jianping Li 0001, Weixuan Xu, Yong Shi 0001 |
Expert Syst. Appl. | 4 |
| 2011 | Credit card churn forecasting by logistic regression and decision tree
Guangli Nie, Wei Rowe, Lingling Zhang 0001, Yingjie Tian 0001, Yong Shi 0001 |
Expert Syst. Appl. | 5 |
| 2011 | Credit risk evaluation with kernel-based affine subspace nearest points learning method
Xiaofei Zhou 0002, Wenhan Jiang, Yong Shi 0001, Yingjie Tian 0001 |
Expert Syst. Appl. | 3 |
| 2010 | Robust state estimation for discrete-time stochastic neural networks with probabilistic measurement delays
Zidong Wang 0001, Yurong Liu, Xiaohui Liu 0001, Yong Shi 0001 |
Neurocomputing | 4 |
| 2010 | Kernel subclass convex hull sample selection method for SVM on face recognition
Xiaofei Zhou 0002, Wenhan Jiang, Yingjie Tian 0001, Yong Shi 0001 |
Neurocomputing | 4 |
| 2010 | Multiple criteria optimization-based data mining methods and applications: a systematic survey
Yong Shi 0001 |
Knowl. Inf. Syst. | 1 |
| 2010 | Domain-Driven Classification Based on Multiple Criteria and Multiple Constraint-Level Programming for Intelligent Credit ScoringabstractExtracting knowledge from the transaction records and the personal data of credit card holders has great profit potential for the banking industry. The challenge is to detect/predict bankrupts and to keep and recruit the profitable customers. However, grouping and targeting credit card customers by traditional data-driven mining often does not directly meet the needs of the banking industry, because data-driven mining automatically generates classification outputs that are imprecise, meaningless, and beyond users' control. In this paper, we provide a novel domain-driven classification method that takes advantage of multiple criteria and multiple constraint-level programming for intelligent credit scoring. The method involves credit scoring to produce a set of customers' scores that allows the classification results actionable and controllable by human interaction during the scoring process. Domain knowledge and experts' experience parameters are built into the criteria and constraint functions of mathematical programming and the human and machine conversation is employed to generate an efficient and precise solution. Experiments based on various data sets validated the effectiveness and efficiency of the proposed methods. Jing He 0004, Yanchun Zhang, Yong Shi 0001, Guangyan Huang |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2010 | Active Learning From Stream Data Using Optimal Weight Classifier EnsembleabstractIn this paper, we propose a new research problem on active learning from data streams, where data volumes grow continuously, and labeling all data is considered expensive and impractical. The objective is to label a small portion of stream data from which a model is derived to predict future instances as accurately as possible. To tackle the technical challenges raised by the dynamic nature of the stream data, i.e., increasing data volumes and evolving decision concepts, we propose a classifier-ensemble-based active learning framework that selectively labels instances from data streams to build a classifier ensemble. We argue that a classifier ensemble's variance directly corresponds to its error rate, and reducing a classifier ensemble's variance is equivalent to improving its prediction accuracy. Because of this, one should label instances toward the minimization of the variance of the underlying classifier ensemble. Accordingly, we introduce a minimum-variance (MV) principle to guide the instance labeling process for data streams. In addition, we derive an optimal-weight calculation method to determine the weight values for the classifier ensemble. The MV principle and the optimal weighting module are combined to build an active learning framework for data streams. Experimental results on synthetic and real-world data demonstrate the performance of the proposed work in comparison with other approaches. Xingquan Zhu 0001, Peng Zhang 0001, Xiaodong Lin 0004, Yong Shi 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 4 |
| 2010 | Multiple criteria programming models for VIP E-Mail behavior analysisabstractExcessive lose of customer account is becoming a major headache for VIP E-Mail hosting companies. Analysis of what kind of customer is more prone to lose and finding the appropriate measures to sustain those customers has become urgent needs. Recentl Peng Zhang 0001, Xingquan Zhu 0001, Zhiwang Zhang, Yong Shi 0001 |
Web Intell. Agent Syst. | 4 |
| 2009 | A New Kernel-Based Classification AlgorithmabstractA new kernel-based learning algorithm called kernel affine subspace nearest point (KASNP) approach is proposed in this paper. Inspired by the geometrical explanation of support vector machines (SVMs) and its nearest point problem in convex hulls, we extend the convex hull of each class to its corresponding affine subspace in high dimensional space induced by kernel. In two class affine subspaces, KASNP finds the nearest points and then constructs a separating hyperplane, which bisects the line segment joining them. The nearest point problem of KASNP is only an unconstrained optimal problem whose solution can be directly computed. Compared with SVM, KASNP avoids solving convex quadratic programming. Experiments on two-spiral dataset, two UCI credit datasets, and face recognition datasets show that our proposed KASNP is effective for data classification. Xiaofei Zhou 0002, Wenhan Jiang, Yingjie Tian 0001, Peng Zhang 0001, Guangli Nie, Yong Shi 0001 |
ICDM | 6 |
| 2009 | An Aggregate Ensemble for Mining Concept Drifting Data Streams with Noise
Peng Zhang 0001, Xingquan Zhu 0001, Yong Shi 0001, Xindong Wu 0001 |
PAKDD | 3 |
| 2009 | Regularized multiple criteria linear programs for classification
Yong Shi 0001, Yingjie Tian 0001, Xiaojun Chen 0001, Peng Zhang 0001 |
Sci. China Ser. F Inf. Sci. | 1 |
| 2009 | Decision analysis of data mining project based on Bayesian risk
Guangli Nie, Lingling Zhang 0001, Ying Liu 0039, Xiuyu Zheng, Yong Shi 0001 |
Expert Syst. Appl. | 5 |
| 2009 | A rough set-based multiple criteria linear programming approach for the medical diagnosis and prognosis
Zhiwang Zhang, Yong Shi 0001, Guangxia Gao |
Expert Syst. Appl. | 2 |
| 2009 | A class of classification and regression methods by multiobjective programming
Dongling Zhang, Yong Shi 0001, Yingjie Tian 0001, Meihong Zhu |
Frontiers Comput. Sci. China | 2 |
| 2009 | Multiple criteria mathematical programming for multi-class classification and application in network intrusion detection
Gang Kou, Yi Peng 0001, Zhengxin Chen, Yong Shi 0001 |
Inf. Sci. | 4 |
| 2008 | A Family of Optimization Based Data Mining Methods
Yong Shi 0001, Nian Yan, Zhenxing Chen |
APWeb | 1 |
| 2008 | Cleansing Noisy Data StreamsabstractIn this paper, we identify a new research problem on cleansing noisy data streams which contain incorrectly labeled training examples. The objective is to accurately identify and remove mislabeled data, such that the prediction models built from the cleansed streams can be more accurate than the ones trained from the raw noisy streams. For this purpose, we first use bias-variance decomposition to derive a maximum variance margin (MVM) principle for stream data cleansing. Following this principle, we further propose a local and global filtering (LgF) framework to combine the strength of local noise filtering (within one single data chunk) and global noise filtering (across a number of adjacent data chunks) to identify erroneous data. Experimental results on six data streams (including two real-world data streams) demonstrate that LgF significantly outperforms simple methods in identifying noisy examples. Xingquan Zhu 0001, Peng Zhang 0001, Xindong Wu 0001, Dan He 0001, Chengqi Zhang, Yong Shi 0001 |
ICDM | 6 |
| 2008 | Categorizing and mining concept drifting data streamsabstractMining concept drifting data streams is a defining challenge for data mining research. Recent years have seen a large body of work on detecting changes and building prediction models from stream data, with a vague understanding on the types of the concept drifting and the impact of different types of concept drifting on the mining algorithms. In this paper, we first categorize concept drifting into two scenarios: Loose Concept Drifting (LCD) and Rigorous Concept Drifting (RCD), and then propose solutions to handle each of them separately. For LCD data streams, because concepts in adjacent data chunks are sufficiently close to each other, we apply kernel mean matching (KMM) method to minimize the discrepancy of the data chunks in the kernel space. Such a minimization process will produce weighted instances to build classifier ensemble and handle concept drifting data streams. For RCD data streams, because genuine concepts in adjacent data chunks may randomly and rapidly change, we propose a new Optimal Weights Adjustment (OWA) method to determine the optimum weight values for classifiers trained from the most recent (up-to-date) data chunk, such that those classifiers can form an accurate classifier ensemble to predict instances in the yet-to-come data chunk. Experiments on synthetic and real-world datasets will show that weighted instance approach is preferable when the concept drifting is mainly caused by the changing of the class prior probability; whereas the weighted classifier approach is preferable when the concept drifting is mainly triggered by the changing of the conditional probability. Peng Zhang 0001, Xingquan Zhu 0001, Yong Shi 0001 |
KDD | 3 |
| 2008 | A Multi-criteria Convex Quadratic Programming model for credit data analysis
Yi Peng 0001, Gang Kou, Yong Shi 0001, Zhengxin Chen |
Decis. Support Syst. | 3 |
| 2007 | From similarity retrieval to cluster analysis: The case of R*-treesabstractData mining is concerned with important aspects related to both database techniques and AI/machine learning mechanisms, and provides an excellent opportunity for exploring the interesting relationship between retrieval and inference/reasoning, a fundamental issue concerning the nature of data mining. In the data mining context, this relationship can be restated as connection and differences between data retrieval and data mining. In this paper we explore this relationship by examining time series data indexed through R*-trees, and study the issues of (1) retrieval of data similar to a given query (which is a plain data retrieval task), and (2) clustering of the data based on similarity (which is a data mining task). Along the way of examination of our central theme, we also report new algorithms and new results related to these two issues. We have developed a software package consisting of a similarity analysis tool and two implemented clustering algorithms: KMeans-R and Hierarchy-R. A sketch of experimental results is also provided Jiaxiong Pi, Yong Shi 0001, Zhengxin Chen |
CIDM | 2 |
| 2007 | Succinct Matrix Approximation and Efficient k-NN ClassificationabstractThis work reveals that instead of the polynomial bounds in previous literatures there exists a sharper bound of exponential form for the L2norm of an arbitrary shaped random matrix. Based on the newly elaborated bound, a nonuniform sampling method is presented to succinctly approximate a matrix with a sparse binary one, and thus relieves the computation loads ofk-NN classifier in both time and storage. The method is also pass-efficient because sampling and quantizing are combined together in a single step and the whole process can be completed within one pass over the input matrix. In the evaluations on compression ratio and reconstruction error, the sampling method exhibits impressive capability in providing succinct and tight approximations for the input matrices. The most significant finding in the classification experiment is that thek-NN classifier based on the approximation can even outperform the standard one. This provides another strong evidence for the claim that our method is especially capable in capturing intrinsic characteristics. Yong Shi 0001 |
ICDM | 2 |
| 2007 | Active Learning from Data StreamsabstractIn this paper, we address a new research problem on active learning from data streams where data volumes grow continuously and labeling all data is considered expensive and impractical. The objective is to label a small portion of stream data from which a model is derived to predict newly arrived instances as accurate as possible. In order to tackle the challenges raised by data streams' dynamic nature, we propose a classifier ensembling based active learning framework which selectively labels instances from data streams to build an accurate classifier. A minimal variance principle is introduced to guide instance labeling from data streams. In addition, a weight updating rule is derived to ensure that our instance labeling process can adaptively adjust to dynamic drifting concepts in the data. Experimental results on synthetic and real-world data demonstrate the performances of the proposed efforts in comparison with other simple approaches. Xingquan Zhu 0001, Peng Zhang 0001, Xiaodong Lin 0004, Yong Shi 0001 |
ICDM | 4 |
| 2007 | An Empirical Study of the Noise Impact on Cost-Sensitive Learning
Xingquan Zhu 0001, Xindong Wu 0001, Taghi M. Khoshgoftaar, Yong Shi 0001 |
IJCAI | 4 |
| 2007 | A Multi-criteria Decision Support System of Water Resource Allocation Scenarios
Jing He 0004, Yanchun Zhang, Yong Shi 0001 |
KSEM | 3 |
| 2007 | Developing Mining-Grid Centric E-Finance Portals for Risk Management and Decision MakingabstractE-finance industry is rapidly transforming and evolving toward more dynamic, flexible and intelligent solutions. This paper describes a model with dynamic multilevel workflows corresponding to a multilayer Grid architecture. The mining-grid is used for multiaspect analysis in building e-finance portals on the Wisdom Web. The application and research demonstrate that mining-grid centric design is effective for developing intelligent risk management and decision making financial systems. This paper concentrates on how to develop a mining-grid centric e-finance portal (MGCFP), not only for supplying effective online financial services for both retail and corporate customers, but also for intelligent credit risk management and decision making for financial enterprises and partners. Jia Hu 0002, Ning Zhong 0001, Yong Shi 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2007 | Predicting the distance between antibody's interface residue and antigen to recognize antigen types by support vector machine
Yong Shi 0001, Wei Yin 0001, Yajun Guo |
Neural Comput. Appl. | 1 |
| 2006 | Evaluation of Cluster Analysis Algorithms Enhanced by Using R*-TreesabstractR* tree is a useful data structure for handling spatial data. However, although objects stored in the same R* tree leaf node enjoys spatial proximity, it is well-known that R* trees cannot be used directly for cluster analysis. Nevertheless, R* tree’s indexing feature can be used to assist existing cluster analysis methods, thus enhancing their performance or cluster quality. In this paper, we explore how to use R* trees to improve well-known Kmeans and hierarchical clustering methods. Based on R*Tree’s feature of indexing Minimum Bounding Box (MBB) according to spatial proximity, we extend R*-Tree’s application to cluster analysis of time series. Two improved algorithms, KMeans-R and Hierarchy-R, are proposed. The performance of these two methods is evaluated against K-Means and K-Means with sampling technique (KMeans-S) using similarity-oriented, supervised measures of cluster validity. Rand Index (RI), Adjusted Rand Index (ARI) and Information Gain (IG) are used as evaluation measures in our experiments. Compared with K-Means and KMeans-S, the clustering results from different data sets have shown that KMeans-R and Hierarchy-R have achieved better clustering quality. Jiaxiong Pi, Yong Shi 0001, Zhengxin Chen |
AICCSA | 2 |
| 2006 | Nonlinear Classification by Linear Programming with Signed Fuzzy MeasuresabstractLinear programming (LP) based models provide good solutions to classification problem especially when the data is linearly separable. The assumption of LP classification models is: the contributions from all attributes towards the classification model are the sum of contributions of each attribute. This assumption leads to a weakness of LP classification models when data is linearly inseparable. The concept of signed fuzzy measure is introduced and utilized in LP approach in order to enhance the classification power through capturing all possible interactions among any two or more attributes. The use of the Choquet integral with respect to a signed fuzzy measure on LP model is able to separate the data that is finearly inseparable. Nian Yan, Zhenyuan Wang, Yong Shi 0001, Zhengxin Chen |
FUZZ-IEEE | 3 |
| 2005 | A Shrinking-Based Clustering Approach for Multidimensional DataabstractExisting data analysis techniques have difficulty in handling multidimensional data. Multidimensional data has been a challenge for data analysis because of the inherent sparsity of the points. In this paper, we first present a novel data preprocessing technique called shrinking which optimizes the inherent characteristic of distribution of data. This data reorganization concept can be applied in many fields such as pattern recognition, data clustering, and signal processing. Then, as an important application of the data shrinking preprocessing, we propose a shrinking-based approach for multidimensional data analysis which consists of three steps: data shrinking, cluster detection, and cluster evaluation and selection. The process of data shrinking moves data points along the direction of the density gradient, thus generating condensed, widely-separated clusters. Following data shrinking, clusters are detected by finding the connected components of dense cells (and evaluated by their compactness). The data-shrinking and cluster-detection steps are conducted on a sequence of grids with different cell sizes. The clusters detected at these scales are compared by a cluster-wise evaluation measurement, and the best clusters are selected as the final result. The experimental results show that this approach can effectively and efficiently detect clusters in both low and high-dimensional spaces. Yong Shi 0001, Yuqing Song 0002, Aidong Zhang 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2002 | Duality in fuzzy multi-criteria and multi-constraint level linear programming: a parametric approach
Yihua Zhong, Yong Shi 0001 |
Fuzzy Sets Syst. | 2 |
| 1999 | Managing user performance for a corporate network
Heeseok Lee, Sufi M. Nazem, Yong Shi 0001, Justin Stolen |
Inf. Manag. | 4 |
| 1996 | A consensus ranking for information system requirementsabstractIn allocating scarce resources to a new information system (IS), a non‐trivial task becomes the determination of a best priority ranking of the IS’s intangible information requirements. Given a set of individual users’ rankings of the information requirements, illustrates, through a real‐world case study, a streamlined consensus priority ranking (SCPR) method based on a concept of minimizing the disagreement (distance) between individual rankings. Compared to a traditional weighted ranking method, the SCPR method is easy to understand, systematic and requires no weighting methodology. Thus, the SCPR method can help the system development team make efficient decisions when allocating resources for a new information system. Yong Shi 0001, Pamela Specht, Justin Stolen |
Inf. Manag. Comput. Secur. | 1 |
| 1994 | A computer-aided system for linear production designs
Yong Shi 0001, Po Lung Yu, Changqing Zhang 0004, Dazhi Zhang |
Decis. Support Syst. | 1 |
| 1994 | Allocating data files over a wide area network: Goal setting and compromise design
Heeseok Lee, Yong Shi 0001, Justin Stolen |
Inf. Manag. | 2 |