Jinmao Wei 0001

dblp:99/3777-1 · also Jin-Mao Wei 0001 · DBLP profile ↗
← Back
47ranked-venue papers
2as first author
16since 2021 · last 2026
0000-0003-0809-6687ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 9 since 2021Databases, data management, data science and information retrieval · 11 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Systems, architecture and hardware · 2Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Feature Selection via Dynamic Feature Graph
Mykola Pechenizkiy, Jinmao Wei 0001, Jian Liu 0040
IEEE Trans. Knowl. Data Eng.3
2025 Predicting Drug-Protein Interactions Based on Similarity Reconstruction and Adaptive Combination Algorithm
abstract
Applying computational methods to predict drug-protein interactions can accelerate drug discovery. An ideal computational method should identify false negative samples within the covered space of its dataset, known as introspection, and exhibit good generalizability and scalability beyond the covered space. However, current computational methods have struggled to achieve both objectives simultaneously. In this paper, we propose CombDPI, which leverages similarity reconstruction and an adaptive combination algorithm. It comprises two parts. The first part focuses on predicting false negative samples within the dataset's covered space. CombDPI assesses similarity from various perspectives, enhances reliability by reconstructing similarity relationships, and utilizes an adaptive combination algorithm to complete the predictions. The second part of CombDPI is dedicated to predicting interactions for novel drugs or proteins. In this part, the model utilizes the similarity between novel drugs or proteins and known drugs or proteins to create their respective representations. By leveraging these representations for prediction, CombDPI becomes less dependent on the prediction of other drug-protein pairs, thereby greatly enhancing its generalizability and scalability. Experimental results on three benchmark datasets, with both in-space and out-space settings, demonstrate that CombDPI outperforms existing interaction prediction methods. Furthermore, case studies highlight CombDPI's ability to discover potential drug-protein interactions.
Renhong Cheng, Jinmao Wei 0001
IEEE Trans. Comput. Biol. Bioinform.3
2025 PCLT-PPI: Predicting Multi-Type Interactions Between Proteins Based on Point Cloud Structure and Local Topology Preservation
abstract
Protein-protein interactions (PPIs) play a crucial role in cellular biochemical reactions. Computationally mining PPI can help us better understand cellular regulatory mechanisms. Most existing methods focus on the linear structure of proteins, ignoring the influence of native spatial structure on their properties. Furthermore, when neural networks are used to learn protein embeddings, the nonlinear transformations may change the topological relationships between proteins. To address the above issues, we propose a PPI prediction method based on protein point cloud structure and local topology preservation, naming it PCLT-PPI. It extracts structural features from protein point cloud structures and relational features through graph neural networks. Throughout the process, PCLT-PPI maintains the local topology of proteins in their origin and embedding spaces. Experimental results show that, under three test set partition modes (Random, BFS, DFS) and four evaluation metrics (F1, AUC, AUPR, Hamming Loss), PCLT-PPI performs better than several state-of-the-art PPI prediction methods, especially when predicting protein PPIs that are not visible during training, exhibiting stronger robustness and higher generalization ability. The results also demonstrate that point cloud structure and local topology preservation can improve PPI prediction performance, which may provide a reference for subsequent related research.
Yurui Hou, Jinmao Wei 0001, Jian Liu 0040
IEEE J. Biomed. Health Informatics4
2024 Structure-inclusive similarity based directed GNN: a method that can control information flow to predict drug-target binding affinity
abstract
MOTIVATION: Exploring the association between drugs and targets is essential for drug discovery and repurposing. Comparing with the traditional methods that regard the exploration as a binary classification task, predicting the drug-target binding affinity can provide more specific information. Many studies work based on the assumption that similar drugs may interact with the same target. These methods constructed a symmetric graph according to the undirected drug similarity or target similarity. Although these similarities can measure the difference between two molecules, it is unable to analyze the inclusion relationship of their substructure. For example, if drug A contains all the substructures of drug B, then in the message-passing mechanism of the graph neural network, drug A should acquire all the properties of drug B, while drug B should only obtain some of the properties of A. RESULTS: To this end, we proposed a structure-inclusive similarity (SIS) which measures the similarity of two drugs by considering the inclusion relationship of their substructures. Based on SIS, we constructed a drug graph and a target graph, respectively, and predicted the binding affinities between drugs and targets by a graph convolutional network-based model. Experimental results show that considering the inclusion relationship of the substructure of two molecules can effectively improve the accuracy of the prediction model. The performance of our SIS-based prediction method outperforms several state-of-the-art methods for drug-target binding affinity prediction. The case studies demonstrate that our model is a practical tool to predict the binding affinity between drugs and targets. AVAILABILITY AND IMPLEMENTATION: Source codes and data are available at https://github.com/HuangStomach/SISDTA.
Jipeng Huang, Chang Sun 0002, Rong Tang 0004, Jinmao Wei 0001
Bioinform.7
2024 Unsupervised feature selection by learning exponential weights
Jun Wang 0023, Zhichen Gu, Jinmao Wei 0001, Jian Liu 0040
Pattern Recognit.4
2023 Dynamic Feed-Forward LSTM
Chengkai Piao, Jinmao Wei 0001
KSEM (1)3
2023 Word-Context Attention for Text Representation
Chengkai Piao, Yapeng Zhu, Jinmao Wei 0001, Jian Liu 0040
Neural Process. Lett.4
2023 A Deep Neural Network-Based Co-Coding Method to Predict Drug-Protein Interactions by Analyzing the Feature Consistency Between Drugs and Proteins
abstract
Exploring drug-protein interactions (DPIs) through computational methods can effectively reduce the workload and the cost of DPI identification. Previous works try to predict DPIs by integrating and analyzing the unique features of drugs and proteins. They cannot adequately analyze the consistency between the drug features and the protein features due to their different semantics. However, the consistency of their features, such as the correlation originating from their sharing diseases, may reveal some potential DPIs. Here we propose a deep neural network-based co-coding method (DNNCC for short) to predict novel DPIs. DNNCC projects the original features of drugs and proteins to a common embedding space through a co-coding strategy. In this way, the embedding features of drugs and proteins have the same semantics. Therefore, the prediction module can discover the unknown DPIs by exploring the feature consistency between drugs and proteins. The experimental results indicate that the performance of DNNCC is significantly superior to five state-of-the-art DPI prediction methods under several evaluation metrics. The superiority of integrating and analyzing the common features of drugs and proteins is proved by the ablation experiments. The novel DPIs predicted by DNNCC verify that DNNCC is a powerful prior tool that can effectively discover potential DPIs.
Chang Sun 0002, Rong Tang 0004, Jipeng Huang, Jinmao Wei 0001, Jian Liu 0040
IEEE ACM Trans. Comput. Biol. Bioinform.4
2023 Predicting Drug-Protein Interactions by Self-Adaptively Adjusting the Topological Structure of the Heterogeneous Network
abstract
Many powerful computational methods based on graph neural networks (GNNs) have been proposed to predict drug-protein interactions (DPIs). It can effectively reduce laboratory workload and the cost of drug discovery and drug repurposing. However, many clinical functions of drugs and proteins are unknown due to their unobserved indications. Therefore, it is difficult to establish a reliable drug-protein heterogeneous network that can describe the relationships between drugs and proteins based on the available information. To solve this problem, we propose a DPI prediction method that can self-adaptively adjust the topological structure of the heterogeneous networks, and name it SATS. SATS establishes a representation learning module based on graph attention network to carry out the drug-protein heterogeneous network. It can self-adaptively learn the relationships among the nodes based on their attributes and adjust the topological structure of the network according to the training loss of the model. Finally, SATS predicts the interaction propensity between drugs and proteins based on their embeddings. The experimental results show that SATS can effectively improve the topological structure of the network. The performance of SATS outperforms several state-of-the-art DPI prediction methods under various evaluation metrics. These prove that SATS is useful to deal with incomplete data and unreliable networks. The case studies on the top section of the prediction results further demonstrate that SATS is powerful for discovering novel DPIs.
Rong Tang 0004, Chang Sun 0002, Jipeng Huang, Jinmao Wei 0001, Jian Liu 0040
IEEE J. Biomed. Health Informatics5
2022 Multi-variable AUC for sifting complementary features and its biomedical application
abstract
Although sifting functional genes has been discussed for years, traditional selection methods tend to be ineffective in capturing potential specific genes. First, typical methods focus on finding features (genes) relevant to class while irrelevant to each other. However, the features that can offer rich discriminative information are more likely to be the complementary ones. Next, almost all existing methods assess feature relations in pairs, yielding an inaccurate local estimation and lacking a global exploration. In this paper, we introduce multi-variable Area Under the receiver operating characteristic Curve (AUC) to globally evaluate the complementarity among features by employing Area Above the receiver operating characteristic Curve (AAC). Due to AAC, the class-relevant information newly provided by a candidate feature and that preserved by the selected features can be achieved beyond pairwise computation. Furthermore, we propose an AAC-based feature selection algorithm, named Multi-variable AUC-based Combined Features Complementarity, to screen discriminative complementary feature combinations. Extensive experiments on public datasets demonstrate the effectiveness of the proposed approach. Besides, we provide a gene set about prostate cancer and discuss its potential biological significance from the machine learning aspect and based on the existing biomedical findings of some individual genes.
Keyu Du, Jun Wang 0023, Jinmao Wei 0001, Jian Liu 0040
Briefings Bioinform.4
2022 Drug-Protein interaction prediction by correcting the effect of incomplete information in heterogeneous information
abstract
MOTIVATION: Large-scale heterogeneous data provide diverse perspectives for predicting drug-protein interactions (DPIs). However, the available information on molecular interactions and clinical associations related to drugs or proteins is incomplete because there may be unproven interactions and associations. This incomplete information in the available data is presented in the form of non-interaction and non-correlation, which may mislead the prediction model. Existing methods fuse incomplete and complete information without considering their integrity, so the negative effects of incomplete information still exist. RESULTS: We develop a network-based DPI prediction method named BRWCP, which uses the complete information network to correct the prediction results acquired by the incomplete information network. By integrating relevant heterogeneous information that may be incomplete, the feature similarities of drugs and proteins are obtained. Combining the feature similarities and known DPIs, an incomplete information-based drug-protein heterogeneous network is constructed. Then, a bidirectional random walk with pruning algorithm is adopted in this heterogeneous network to predict potential DPIs. Next, the predicted DPIs are combined with the chemical fingerprint similarity of drugs and amino acid sequence similarity of proteins to construct the complete information network. The bidirectional random walk with pruning algorithm is applied in the new network to obtain the final prediction results until it converges. Experimental results show that BRWCP is superior to several state-of-the-art DPI prediction methods, and case studies further confirm its ability to tap potential DPIs. AVAILABILITY AND IMPLEMENTATION: The code and data used in BRWCP are available at https://github.com/lyfdomain/BRWCP. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Chang Sun 0002, Jinmao Wei 0001, Jian Liu 0040
Bioinform.3
2022 Multi-objective data enhancement for deep learning-based ultrasound analysis
abstract
Recently, Deep Learning based automatic generation of treatment recommendation has been attracting much attention. However, medical datasets are usually small, which may lead to over-fitting and inferior performances of deep learning models. In this paper, we propose multi-objective data enhancement method to indirectly scale up the medical data to avoid over-fitting and generate high quantity treatment recommendations. Specifically, we define a main and several auxiliary tasks on the same dataset and train a specific model for each of these tasks to learn different aspects of knowledge in limited data scale. Meanwhile, a Soft Parameter Sharing method is exploited to share learned knowledge among models. By sharing the knowledge learned by auxiliary tasks to the main task, the proposed method can take different semantic distributions into account during the training process of the main task. We collected an ultrasound dataset of thyroid nodules that contains Findings, Impressions and Treatment Recommendations labeled by professional doctors. We conducted various experiments on the dataset to validate the proposed method and justified its better performance than existing methods.
Chengkai Piao, Mengyue Lv, Rongyan Zhou, Jinmao Wei 0001, Jian Liu 0040
BMC Bioinform.6
2021 Recommending irregular regions using graph attentive networks
Hengpeng Xu, Jun Wang 0023, Jinmao Wei 0001
Ad Hoc Networks3
2021 Autoencoder-based drug-target interaction prediction by preserving the consistency of chemical properties and functions of drugs
abstract
MOTIVATION: Exploring the potential drug-target interactions (DTIs) is a key step in drug discovery and repurposing. In recent years, predicting the probable DTIs through computational methods has gradually become a research hot spot. However, most of the previous studies failed to judiciously take into account the consistency between the chemical properties of drug and its functions. The changes of these relationships may lead to a severely negative effect on the prediction of DTIs. RESULTS: We propose an autoencoder-based method, AEFS, under spatial consistency constraints to predict DTIs. A heterogeneous network is established to integrate the information of drugs, proteins and diseases. The original drug features are projected to an embedding (protein) space by a multi-layer encoder, and further projected into label (disease) space by a decoder. In this process, the clinical information of drugs is introduced to assist the DTI prediction. By maintaining the distribution of drug correlation in the original feature, embedding and label space, AEFS keeps the consistency between chemical properties and functions of drugs. Experimental comparisons indicate that AEFS is more robust for imbalanced data and of significantly superior performance in DTI prediction. Case studies further confirm its ability to mine the latent DTIs. AVAILABILITY AND IMPLEMENTATION: The code of AEFS is available at https://github.com/JackieSun818/AEFS. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Chang Sun 0002, Yangkun Cao, Jinmao Wei 0001, Jian Liu 0040
Bioinform.3
2021 Multi-class feature selection by exploring reliable class correlation
Jinmao Wei 0001, Jian Liu 0040
Knowl. Based Syst.3
2021 Unsupervised Cross-View Feature Selection on incomplete data
Yuanyuan Xu 0002, Jun Wang 0023, Jinmao Wei 0001, Jian Liu 0040, Lina Yao 0001, Wenjie Zhang 0001
Knowl. Based Syst.4
2020 Document Summarization with VHTM: Variational Hierarchical Topic-Aware Mechanism
abstract
Automatic text summarization focuses on distilling summary information from texts. This research field has been considerably explored over the past decades because of its significant role in many natural language processing tasks; however, two challenging issues block its further development: (1) how to yield a summarization model embedding topic inference rather than extending with a pre-trained one and (2) how to merge the latent topics into diverse granularity levels. In this study, we propose a variational hierarchical model to holistically address both issues, dubbed VHTM. Different from the previous work assisted by a pre-trained single-grained topic model, VHTM is the first attempt to jointly accomplish summarization with topic inference via variational encoder-decoder and merge topics into multi-grained levels through topic embedding and attention. Comprehensive experiments validate the superior performance of VHTM compared with the baselines, accompanying with semantically consistent topics.
Xiyan Fu, Jun Wang 0023, Jinghan Zhang 0004, Jinmao Wei 0001, Zhenglu Yang
AAAI4
2020 To Avoid the Pitfall of Missing Labels in Feature Selection: A Generative Model Gives the Answer
abstract
In multi-label learning, instances have a large number of noisy and irrelevant features, and each instance is associated with a set of class labels wherein label information is generally incomplete. These missing labels possess two sides like a coin; people cannot predict whether their provided information for feature selection is favorable (relevant) or not (irrelevant) during tossing. Existing approaches either superficially consider the missing labels as negative or indiscreetly impute them with some predicted values, which may either overestimate unobserved labels or introduce new noises in selecting discriminative features. To avoid the pitfall of missing labels, a novel unified framework of selecting discriminative features and modeling incomplete label matrix is proposed from a generative point of view in this paper. Concretely, we relax Smoothness Assumption to infer the label observability, which can reveal the positions of unobserved labels, and employ the spike-and-slab prior to perform feature selection by excluding unobserved labels. Using a data-augmentation strategy leads to full local conjugacy in our model, facilitating simple and efficient Expectation Maximization (EM) algorithm for inference. Quantitative and qualitative experimental results demonstrate the superiority of the proposed approach under various evaluation metrics.
Yuanyuan Xu 0002, Jun Wang 0023, Jinmao Wei 0001
AAAI3
2020 Flexible Parameter Sharing Networks
Chengkai Piao, Jinmao Wei 0001, Yapeng Zhu, Hengpeng Xu
NLPCC (1)2
2019 Revealing Semantic Structures of Texts: Multi-grained Framework for Automatic Mind-map Generation
abstract
A mind-map is a diagram used to represent ideas linked to and arranged around a central concept. It’s easier to visually access the knowledge and ideas by converting a text to a mind-map. However, highlighting the semantic skeleton of an article remains a challenge. The key issue is to detect the relations amongst concepts beyond intra-sentence. In this paper, we propose a multi-grained framework for automatic mind-map generation. That is, a novel neural network is taken to detect the relations at first, which employs multi-hop self-attention and gated recurrence network to reveal the directed semantic relations via sentences. A recursive algorithm is then designed to select the most salient sentences to constitute the hierarchy. The human-like mind-map is automatically constructed with the key phrases in the salient sentences. Promising results have been achieved on the comparison with manual mind-maps. The case studies demonstrate that the generated mind-maps reveal the underlying semantic structures of the articles.
Jinmao Wei 0001, Zhong Su
IJCAI3
2019 Probabilistic Margin-Aware Multi-Label Feature Selection by Preserving Spatial Consistency
abstract
Multi-label feature selection focuses on constructing a reduced feature space for discriminating multi-label instances. In consideration of the complex structures of label and feature spaces, a critical issue that explicitly determines selection performance is how to induce consistent information from both spaces to steer feature selection. Existing approaches tackle this issue in various spatial-aware views, without sufficient consideration of the negative effects of irrelevant features and imbalanced neighbors on inferring space structure. Inspired by the superiority of margin theory in assessing reliable space structure, we approach multi-label feature selection in the learning framework of preserving label-feature space consistency through probabilistic margin in this paper. In contrast to existing approaches, our model assesses the weighted margin based on the probabilistic nearest neighbors, and preserves consistent margin information in label and feature spaces. In this manner, label-feature space consistency is elegantly achieved, which conduces to effectively capturing discriminative features suitable for multi-label learning tasks and eliminating noisy features. Experimental results on multi-label data sets demonstrate the encouraging performance of the proposed model.
Shuai An 0003, Jun Wang 0023, Jinmao Wei 0001, Jianhua Ruan
IJCNN4
2019 A general framework for learning prosodic-enhanced representation of rap lyrics
Hongru Liang, Haozheng Wang, Qian Li 0016, Jun Wang 0023, Guandong Xu, Jinmao Wei 0001, Zhenglu Yang
World Wide Web7
2018 Semi-Supervised Multi-Label Feature Selection by Preserving Feature-Label Space Consistency
abstract
Semi-supervised learning and multi-label learning pose different challenges for feature selection, which is one of the core techniques for dimension reduction, and the exploration of reducing feature space for multi-label learning with incomplete label information is far from satisfactory. Existing feature selection approaches devote attention to either of two issues, namely, alleviating negative effects of imperfectly predicted labels and quantitatively evaluating label correlations, exclusively for semi-supervised or multi-label scenarios. A unified framework to extract label correlation information with incomplete prior knowledge and embed this information in feature selection however, is rarely touched. In this paper, we propose a space consistency-based feature selection model to address this issue. Specifically, correlation information in feature space is learned based on the probabilistic neighborhood similarities, and correlation information in label space is optimized by preserving feature-label space consistency. This mechanism contributes to appropriately extracting label information in semi-supervised multi-label learning scenario and effectively employing this information to select discriminative features. An extensive experimental evaluation on real-world data shows the superiority of the proposed approach under various evaluation metrics.
Yuanyuan Xu 0002, Jun Wang 0023, Shuai An 0003, Jinmao Wei 0001, Jianhua Ruan
CIKM4
2018 A Multi-Attention based Neural Network with External Knowledge for Story Ending Predicting Task
abstract
Enabling a mechanism to understand a temporal story and predict its ending is an interesting issue that has attracted considerable attention, as in case of the ROC Story Cloze Task (SCT). In this paper, we develop a multi-attention-based neural network (MANN) with well-designed optimizations, like Highway Network, and concatenated features with embedding representations into the hierarchical neural network model. Considering the particulars of the specific task, we thoughtfully extend MANN with external knowledge resources, exceeding state-of-the-art results obviously. Furthermore, we develop a thorough understanding of our model through a careful hand analysis on a subset of the stories. We identify what traits of MANN contribute to its outperformance and how external knowledge is obtained in such an ending prediction task.
Qian Li 0016, Jinmao Wei 0001, Yanhui Gu, Adam Jatowt, Zhenglu Yang
COLING3
2018 JTAV: Jointly Learning Social Media Content Representation by Fusing Textual, Acoustic, and Visual Features
abstract
Learning social media content is the basis of many real-world applications, including information retrieval and recommendation systems, among others. In contrast with previous works that focus mainly on single modal or bi-modal learning, we propose to learn social media content by fusing jointly textual, acoustic, and visual information (JTAV). Effective strategies are proposed to extract fine-grained features of each modality, that is, attBiGRU and DCRNN. We also introduce cross-modal fusion and attentive pooling techniques to integrate multi-modal information comprehensively. Extensive experimental evaluation conducted on real-world datasets demonstrate our proposed model outperforms the state-of-the-art approaches by a large margin.
Hongru Liang, Haozheng Wang, Jun Wang 0023, Shaodi You, Zhe Sun 0009, Jinmao Wei 0001, Zhenglu Yang
COLING6
2018 Probabilistic Topic and Role Model for Information Diffusion in Social Network
Hengpeng Xu, Jinmao Wei 0001, Zhenglu Yang, Jianhua Ruan, Jun Wang 0023
PAKDD (2)2
2018 ANNC: AUC-Based Feature Selection by Maximizing Nearest Neighbor Complementarity
Xuemeng Jiang, Jun Wang 0023, Jinmao Wei 0001, Jianhua Ruan
PRICAI (1)3
2018 HAVAE: Learning Prosodic-Enhanced Representations of Rap Lyrics
Hongru Liang, Qian Li 0016, Haozheng Wang, Jun Wang 0023, Zhe Sun 0009, Jinmao Wei 0001, Zhenglu Yang
PRICAI (1)7
2018 Unsupervised learning of semantic representation for documents with the law of total probability
abstract
Abstract The semantic information of documents needs to be represented because it is the basis for many applications, such as document summarization, web search, and text analysis. Although many studies have explored this problem by enriching document vectors with the relatedness of the words involved, the performance remains far from satisfactory because the physical boundaries of documents hinder the evaluation of the relatedness between words. To address this problem, we propose an effective approach to further infer the implicit relatedness between words via their common related words. To avoid overestimation of the implicit relatedness, we restrict the inference in terms of the marginal probabilities of the words based on the law of total probability. The proposed method measures the relatedness between words, which is confirmed theoretically and experimentally. Thorough evaluation on real datasets illustrates that significant improvement on document clustering has been achieved with the proposed method compared with state-of-the-art methods.
Jinmao Wei 0001, Zhenglu Yang
Nat. Lang. Eng.2
2018 Local-Nearest-Neighbors-Based Feature Weighting for Gene Selection
abstract
Selecting functional genes is essential for analyzing microarray data. Among many available feature (gene) selection approaches, the ones on the basis of the large margin nearest neighbor receive more attention due to their low computational costs and high accuracies in analyzing the high-dimensional data. Yet, there still exist some problems that hamper the existing approaches in sifting real target genes, including selecting erroneous nearest neighbors, high sensitivity to irrelevant genes, and inappropriate evaluation criteria. Previous pioneer works have partly addressed some of the problems, but none of them are capable of solving these problems simultaneously. In this paper, we propose a new local-nearest-neighbors-based feature weighting approach to alleviate the above problems. The proposed approach is based on the trick of locally minimizing the within-class distances and maximizing the between-class distances with the nearest neighbors rule. We further define a feature weight vector, and construct it by minimizing the cost function with a regularization term. The proposed approach can be applied naturally to the multi-class problems and does not require extra modification. Experimental results on the UCI and the open microarray data sets validate the effectiveness and efficiency of the new approach.
Shuai An 0003, Jun Wang 0023, Jinmao Wei 0001
IEEE ACM Trans. Comput. Biol. Bioinform.3
2017 Unsupervised Feature Selection with Joint Clustering Analysis
abstract
Unsupervised feature selection has raised considerable interests in the past decade, due to its remarkable performance in reducing dimensionality without any prior class information. Preserving reliable locality information and achieving excellent cluster separation are two critical issues for unsupervised feature selection. However, existing methods cannot tackle two issues simultaneously. To address the problems, we propose a novel unsupervised approach that integrates sparse feature selection and robust joint clustering analysis. The joint clustering analysis seamlessly unifies the spectral clustering and the orthogonal basis clustering. Specifically, a probabilistic neighborhood graph is utilized to preserve reliable locality information in the spectral clustering, and an orthogonal basis matrix is incorporated to achieve excellent cluster separation in the orthogonal basis clustering. A compact and effective iterative algorithm is designed to optimize the proposed selection framework. Extensive experiments on both synthetic data and real-world data validate the effectiveness of our approach under various evaluation indices.
Shuai An 0003, Jun Wang 0023, Jinmao Wei 0001, Zhenglu Yang
CIKM3
2017 AVC: Selecting discriminative features on basis of AUC by maximizing variable complementarity
abstract
BACKGROUND: The Receiver Operator Characteristic (ROC) curve is well-known in evaluating classification performance in biomedical field. Owing to its superiority in dealing with imbalanced and cost-sensitive data, the ROC curve has been exploited as a popular metric to evaluate and find out disease-related genes (features). The existing ROC-based feature selection approaches are simple and effective in evaluating individual features. However, these approaches may fail to find real target feature subset due to their lack of effective means to reduce the redundancy between features, which is essential in machine learning. RESULTS: In this paper, we propose to assess feature complementarity by a trick of measuring the distances between the misclassified instances and their nearest misses on the dimensions of pairwise features. If a misclassified instance and its nearest miss on one feature dimension are far apart on another feature dimension, the two features are regarded as complementary to each other. Subsequently, we propose a novel filter feature selection approach on the basis of the ROC analysis. The new approach employs an efficient heuristic search strategy to select optimal features with highest complementarities. The experimental results on a broad range of microarray data sets validate that the classifiers built on the feature subset selected by our approach can get the minimal balanced error rate with a small amount of significant features. CONCLUSIONS: Compared with other ROC-based feature selection approaches, our new approach can select fewer features and effectively improve the classification performance.
Jun Wang 0023, Jinmao Wei 0001
BMC Bioinform.3
2017 Discrimination Structure Complementarity-Based Feature Selection
abstract
Feature selection is crucial, particularly for processing high‐dimensional data. Existing selection methods generally compute a discriminant value for a feature with respect to class variable to indicate its classification ability. However, a scalar value can hardly reveal the multifaceted classification abilities of a feature for different subproblems of a complicated multiclass problem. In view of this, we propose to select features based on discrimination structure complementarity. To this end, the classification abilities of a feature for different subproblems are evaluated individually. Consequently, a discrimination structure vector can be obtained to indicate if the feature is discriminative respectively for different subproblems. Based on discrimination structure, indispensable and dispensable features (ID‐features for short) are defined. In selection process, the ID‐features, which are complementary in discrimination structure to the selected ones, are selected. The proposed method tries to equally treat all subproblems and hence can avoid falling into the pitfall that the discriminative features for difficult subproblems are prone to be covered by the features for easy ones in multi‐class classification. Two algorithms are developed and compared with several feature selection methods using some open data sets. Experimental results demonstrate the effectiveness of the proposed method.
Jinmao Wei 0001, Zhenglu Yang
Comput. Intell.2
2017 Feature selection based on measurement of ability to classify subproblems
Jinmao Wei 0001
Neurocomputing2
2017 Feature Selection by Maximizing Independent Classification Information
abstract
Feature selection approaches based on mutual information can be roughly categorized into two groups. The first group minimizes the redundancy of features between each other. The second group maximizes the new classification information of features providing for the selected subset. A critical issue is that large new information does not signify little redundancy, and vice versa. Features with large new information but with high redundancy may be selected by the second group, and features with low redundancy but with little relevance with classes may be highly scored by the first group. Existing approaches fail to balance the importance of both terms. As such, a new information term denoted as Independent Classification Information is proposed in this paper. It assembles the newly provided information and the preserved information negatively correlated with the redundant information. Redundancy and new information are properly unified and equally treated in the new term. This strategy helps find the predictive features providing large new information and little redundancy. Moreover, independent classification information is proved as a loose upper bound of the total classification information of feature subset. Its maximization is conducive to achieve a high global discriminative performance. Comprehensive experiments demonstrate the effectiveness of the new approach.
Jun Wang 0023, Jinmao Wei 0001, Zhenglu Yang, Shu-Qin Wang
IEEE Trans. Knowl. Data Eng.2
2016 Feature Selection via Vectorizing Feature's Discriminative Information
Jun Wang 0023, Hengpeng Xu, Jinmao Wei 0001
APWeb (1)3
2016 Supervised Feature Selection by Preserving Class Correlation
abstract
Feature selection is an effective technique for dimension reduction, which assesses the importance of features and constructs an optimal feature subspace suitable for recognition task. Two recognition scenarios, i.e., single-label learning and multi-label learning, pose different challenges for feature selection. For the single-label task, how to accurately measure and reduce feature redundancy is crucial. For the multi-label task, how to effectively exploit class correlation information during selection is critical. However, both issues cannot be simultaneously resolved by any existing selection methods. In this paper, we propose effective supervised feature selection techniques to address the problems. The original class correlation information in the reduced feature space is preserved, and meanwhile the feature redundancy for classification is alleviated. To the best of our knowledge, this study is the first attempt to accomplish both recognition tasks in a unified framework. Comprehensive experimental evaluations on artificial, single-label, and multi-label data sets demonstrate the effectiveness of the new approach.
Jun Wang 0023, Jinmao Wei 0001, Zhenglu Yang
CIKM2
2016 Joint Probability Consistent Relation Analysis for Document Representation
Jinmao Wei 0001, Zhenglu Yang
DASFAA (1)2
2015 Extended Strategies for Document Clustering with Word Co-occurrences
Jinmao Wei 0001, Zhenglu Yang
APWeb2
2015 Feature Selection Method Based on Feature's Classification Bias and Performance
Jun Wang 0023, Jinmao Wei 0001
ICA3PP (2)2
2015 Enriching Document Representation with the Deviations of Word Co-occurrence Frequencies
Jinmao Wei 0001, Zhenglu Yang
ICA3PP (2)2
2015 Context Vector Model for Document Representation: A Computational Study
abstract
To tackle the sparse data problem of the bag-of-words model for document representation, the Context Vector Model (CVM) has been proposed to enrich a document with the relatedness of all the words in a corpus to the document. The nature of CVM is the combination of word vectors, wherefore the representation method for words is essential for CVM. A computational study is performed in this paper to compare the effects of the newly proposed word representation methods embedded in CVM. The experimental results demonstrate that some of the newly proposed word representation methods significantly improve the performance of CVM, for they estimate the relatedness between words better.
Jinmao Wei 0001, Hengpeng Xu
NLPCC2
2014 A novel feature measure for fuzzy clustering algorithm on microarray data
abstract
Fuzzy clustering algorithm is employed in gene microarray analysis to discover the strength of the association between genes and different clusters. Gene-based fuzzy clustering algorithm just employs all instances' values of a certain gene as this gene's features. In some sense, the original feature vector can hardly provide comprehensive discriminative information of the gene. In this paper, a novel feature vector by the proposed measure for each gene is employed in fuzzy clustering algorithm. The proposed feature vector can provide information about the influence of a given gene for the overall shape of clusters. By analysis and experiment upon microarray data sets, the performance of the fuzzy clustering algorithm based on proposed feature vector is compared with that of some classical clustering algorithms. The results demonstrate that the fuzzy clustering algorithm based on proposed feature vector is capable of obtaining better clusters than other contrast algorithms. The results by classifiers based on different clustering algorithms demonstrate that the proposed feature vector can get the same or better accuracy than the original feature vector.
Jinmao Wei 0001
FUZZ-IEEE2
2014 Gene Selection Based on Supervised Vector Representation of Genes
Jinmao Wei 0001
PRICAI4
2012 PAC Learnability of Rough Hypercuboid Classifier
Jinmao Wei 0001
ICIC (2)2
2010 A novel measure for evaluating classifiers
Jinmao Wei 0001, Xiao-Jie Yuan, Qinghua Hu, Shu-Qin Wang
Expert Syst. Appl.1
2010 Ensemble Rough Hypercuboid Approach for Classifying Cancers
abstract
Cancer classification is the critical basis for patient-tailored therapy. Conventional histological analysis tends to be unreliable because different tumors may have similar appearance. The advances in microarray technology make individualized therapy possible. Various machine learning methods can be employed to classify cancer tissue samples based on microarray data. However, few methods can be elegantly adopted for generating accurate and reliable as well as biologically interpretable rules. In this paper, we introduce an approach for classifying cancers based on the principle of minimal rough fringe. For training rough hypercuboid classifiers from gene expression data sets, the method dynamically evaluates all available genes and sifts the genes with the smallest implicit regions as the dimensions of implicit hypercuboids. An unseen object is predicted to be a certain class if it falls within the corresponding class hypercuboid. Based upon the method, ensemble rough hypercuboid classifiers are subsequently constructed. Experimental results on some open cancer gene expression data sets show that the proposed method is capable of generating accurate and interpretable rules compared with some other machine learning methods. Hence, it is a feasible way of classifying cancer tissues in biomedical applications.
Jinmao Wei 0001, Shu-Qin Wang, Xiao-Jie Yuan
IEEE Trans. Knowl. Data Eng.1