VLDB 2026 Research / reviewers in the wild / expert
Jian-Ping Mei
dblp:36/8017 · also Jianping Mei
· DBLP profile ↗
40ranked-venue papers
24as first author
19since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 20 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 3 since 2021Security and privacy · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Defense Against Model Stealing Based on Account-Aware Distribution DiscrepancyabstractMalicious users attempt to replicate commercial models functionally at low cost by training a clone model with query responses. It is challenging to timely prevent such model-stealing attacks to achieve strong protection and maintain utility. In this paper, we propose a novel non-parametric detector called Account-aware Distribution Discrepancy (ADD) to recognize queries from malicious users by leveraging account-wise local dependency. We formulate each class as a Multivariate Normal distribution (MVN) in the feature space and measure the malicious score as the sum of weighted class-wise distribution discrepancy. The ADD detector is combined with random-based prediction poisoning to yield a plug-and-play defense module named D-ADD for image classification models. Results of extensive experimental studies show that D-ADD achieves strong defense against different types of attacks with little interference in serving benign users for both soft and hard-label settings. Jian-Ping Mei, Xuyun Zhang, Tiantian Zhu 0001 |
AAAI | 1 |
| 2025 | Source-free domain adaptation with aligned transfer and self-supervised learning
Yetao Weng, Jian-Ping Mei, Chengyang Hu |
Appl. Intell. | 2 |
| 2025 | PDCleaner: A multi-view collaborative data compression method for provenance graph-based APT detection systems
Jiaobo Jin, Tiantian Zhu 0001, Qixuan Yuan, Tieming Chen, Mingqi Lv, Chenbin Zheng, Jian-Ping Mei |
Comput. Secur. | 7 |
| 2025 | MIRDETECTOR: Applying malicious intent representation for enhanced APT anomaly detection
Tiantian Zhu 0001, Tieming Chen, Mingqi Lv, Jian-Ping Mei, Zhengqiu Weng, Lili Shi |
Comput. Secur. | 6 |
| 2025 | Soft-label generator based on classifier weights
Xinkai Chu, Jian-Ping Mei, Rui Yan 0005 |
Neurocomputing | 2 |
| 2025 | Semi-supervised structured nonnegative matrix factorization for anchor graph embedding
Jian-Ping Mei, Yuanjian Mo |
Neurocomputing | 2 |
| 2025 | VulnTrace: Tracking and Detecting Code Vulnerabilities with Historical Commits and Semantic EmbeddingsabstractOpen source software has evolved into a fundamental element of the contemporary information sector; however, security threats within its supply chain are persistently rising. Within the collaborative development framework of open source, the introduction of malicious code can lead to significant security vulnerabilities. Conventional methods for detecting these vulnerabilities, which rely on machine learning, face challenges such as a lack of sufficient datasets, inadequate deep semantic understanding, and limitations to single-vulnerability detection. To address these challenges, we introduce a novel approach named VulnTrace, which analyzes historical records of submissions in open source projects to construct a high-quality dataset of vulnerabilities with accurate labels. VulnTrace employs Word2Vec alongside Abstract Syntax Tree (AST) technologies to capture both the semantic and structural details of code segments and utilizes a Transformer model for precise vulnerability identification, thereby enhancing accuracy and interpretability in detection. Experimental results indicate that VulnTrace achieves approximately 93% accuracy, 95% precision, 83% recall and an F1 score of 88% in vulnerability detection tasks, significantly reducing false positives and demonstrating remarkable robustness. Qijie Song, Jiaobo Jin, Tiantian Zhu 0001, Tieming Chen, Mingqi Lv, Licheng Pan, Jian-Ping Mei |
Int. J. Softw. Eng. Knowl. Eng. | 7 |
| 2025 | ThreatCog: An adaptive and lightweight mobile user authentication system with enhanced motion sensory signals
Tiantian Zhu 0001, Jian-Ping Mei, Xue Leng, Xiangyang Zheng, Zhengqiu Weng |
J. Inf. Secur. Appl. | 6 |
| 2025 | Self-supervised learning from images: No negative pairs, no cluster-balancing
Jian-Ping Mei, Shixiang Wang, Miaoqi Yu |
Pattern Recognit. | 1 |
| 2024 | Query-Efficient Stealing Attacks Against Image Encoders
Jian-Ping Mei, Yuhao Guan, Chunlong Lu, Mingqi Lv |
PRICAI (4) | 1 |
| 2024 | Federated Prompt Tuning: When is it Necessary?
Jian-Ping Mei, Chunlong Lu, Yuhao Guan, Mingqi Lv |
PRICAI (2) | 1 |
| 2024 | Semi-supervised nonnegative matrix factorization with label propagation and constraint propagation
Yuanjian Mo, Xiangli Li 0003, Jian-Ping Mei |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | Dual semi-supervised hypergraph regular multi-view NMF with anchor graph embedding
Jian-Ping Mei, Xiangli Li 0001, Yuanjian Mo |
Knowl. Based Syst. | 1 |
| 2024 | Output Regularization With Cluster-Based Soft TargetsabstractWhile supervised learning of over-parameterized neural networks achieved state-of-the-art performance in image classification, it tends to over-fit the labeled training samples to give inferior generalization ability. Output regularization deals with over-fitting by using soft targets as additional training signals. Although clustering is one of the most fundamental data analysis tools for discovering general-purpose and data-driven structures, it has been ignored in existing output regularization approaches. In this article, we leverage this underlying structural information by proposing Cluster-based soft targets for Output Regularization (CluOReg). This approach provides a unified way for simultaneous clustering in embedding space and neural classifier training with cluster-based soft targets via output regularization. By explicitly calculating a class relationship matrix in the cluster space, we obtain classwise soft targets shared by all samples in each class. Results of image classification experiments under various settings on a number of benchmark datasets are provided. Without resorting to external models or designed data augmentation, we get consistent and significant reductions in classification error compared with other approaches, demonstrating that cluster-based soft targets effectively complement the ground-truth label. Jian-Ping Mei, Wenhao Qiu, Defang Chen 0001, Rui Yan 0005 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Data-Locked Knowledge Distillation with Ready-To-Use SubstitutesabstractKnowledge Distillation (KD) becomes a popular way of knowledge transfer from a pre-trained cumbersome teacher model to a compact student model by training the student to mimic the teacher's outputs. However, the original training data of the teacher model, which are supposed to be used for distilling, are often inaccessible during the student training phase due to privacy and intellectual property issues. Several studies reported encouraging results for such kind of data-locked KD problem by collecting substitute inputs via synthesization or seeking ready-to-use samples from publicly available data. This paper aims to take a step further with the more cost-efficient selection-based method. Specifically, we propose a new In-Distribution Detection (IDD) criterion to measure the quality of a candidate sample based on both model confidence and class-relationship consistency. Immediate evaluation of IDD performance with several benchmark image datasets verifies the effectiveness of the designed measurement compared to existing ones, showing its potential for substitute selection from open resources. Experimental results on data-locked KD demonstrate that our approach consistently improves the previous selection-based approach and even outperforms the latest synthesization-based ones given a proper candidate set. Jian-Ping Mei, Xinkai Chu |
IJCNN | 1 |
| 2023 | SemCKD: Semantic Calibration for Cross-Layer Knowledge DistillationabstractKnowledge distillation is a technique to enhance the generalization ability of a student model by exploiting outputs from a teacher model. Recently, feature-map based variants explore knowledge transfer between manually assigned teacher-student pairs in intermediate layers for further improvement. However, layer semantics may vary in different neural networks, resulting in performance degeneration due to negative regularization from semantic mismatch in manual layer associations. To address this issue, we propose semantic calibration for cross-layer knowledge distillation (SemCKD), which automatically assigns proper target layers of the teacher model for each student layer with an attention mechanism. With a learned attention distribution, each student layer distills knowledge contained in multiple teacher layers rather than a specific intermediate layer for appropriate cross-layer supervision. We further provide theoretical analysis of the association weights and conduct extensive experiments to demonstrate the effectiveness of our approach. On average, SemCKD improves the student Top-1 classification accuracy by 4.27% across twelve different teacher-student model combinations on CIFAR-100. Code is available athttps://github.com/DefangChen/SemCKD. Can Wang 0001, Defang Chen 0001, Jian-Ping Mei, Chun Chen 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | Knowledge Distillation with the Reused Teacher ClassifierabstractKnowledge distillation aims to compress a powerful yet cumbersome teacher model into a lightweight student model without much sacrifice of performance. For this purpose, various approaches have been proposed over the past few years, generally with elaborately designed knowledge rep-resentations, which in turn increase the difficulty of model development and interpretation. In contrast, we empirically show that a simple knowledge distillation technique is enough to significantly narrow down the teacher-student performance gap. We directly reuse the discriminative classifier from the pre-trained teacher model for student inference and train a student encoder through feature alignment with a single ℓ2loss. In this way, the student model is able to achieve exactly the same performance as the teacher model provided that their extracted features are perfectly aligned. An additional projector is developed to help the student encoder match with the teacher classifier, which renders our technique applicable to various teacher and student architectures. Extensive experiments demonstrate that our technique achieves state-of-the-art results at the modest cost of compression ratio due to the added projector. Defang Chen 0001, Jian-Ping Mei, Hailin Zhang 0005, Can Wang 0001, Chun Chen 0001 |
CVPR | 2 |
| 2022 | TaskDrop: A competitive baseline for continual learning of sentiment classification
Jian-Ping Mei, Yilun Zhen, Qianwei Zhou, Rui Yan 0005 |
Neural Networks | 1 |
| 2021 | Cross-Layer Distillation with Semantic CalibrationabstractRecently proposed knowledge distillation approaches based on feature-map transfer validate that intermediate layers of a teacher model can serve as effective targets for training a student model to obtain better generalization ability. Existing studies mainly focus on particular representation forms for knowledge transfer between manually specified pairs of teacher-student intermediate layers. However, semantics of intermediate layers may vary in different networks and manual association of layers might lead to negative regularization caused by semantic mismatch between certain teacher-student layer pairs. To address this problem, we propose Semantic Calibration for Cross-layer Knowledge Distillation (SemCKD), which automatically assigns proper target layers of the teacher model for each student layer with an attention mechanism. With a learned attention distribution, each student layer distills knowledge contained in multiple layers rather than a single fixed intermediate layer from the teacher model for appropriate cross-layer supervision in training. Consistent improvements over state-of-the-art approaches are observed in extensive experiments with various network architectures for teacher and student models, demonstrating the effectiveness and flexibility of the proposed attention based soft layer association mechanism for cross-layer distillation. Defang Chen 0001, Jian-Ping Mei, Can Wang 0001, Zhe Wang 0001, Chun Chen 0001 |
AAAI | 2 |
| 2020 | Online Knowledge Distillation with Diverse PeersabstractDistillation is an effective knowledge-transfer technique that uses predicted distributions of a powerful teacher model as soft targets to train a less-parameterized student model. A pre-trained high capacity teacher, however, is not always available. Recently proposed online variants use the aggregated intermediate predictions of multiple student models as targets to train each student model. Although group-derived targets give a good recipe for teacher-free distillation, group members are homogenized quickly with simple aggregation functions, leading to early saturated solutions. In this work, we propose Online Knowledge Distillation with Diverse peers (OKDDip), which performs two-level distillation during training with multiple auxiliary peers and one group leader. In the first-level distillation, each auxiliary peer holds an individual set of aggregation weights generated with an attention-based mechanism to derive its own targets from predictions of other auxiliary peers. Learning from distinct target distributions helps to boost peer diversity for effectiveness of group-based distillation. The second-level distillation is performed to transfer the knowledge in the ensemble of auxiliary peers further to the group leader, i.e., the model used for inference. Experimental results show that the proposed framework consistently gives better performance than state-of-the-art approaches without sacrificing training or inference complexity, demonstrating the effectiveness of the proposed two-level distillation framework. Defang Chen 0001, Jian-Ping Mei, Can Wang 0001, Chun Chen 0001 |
AAAI | 2 |
| 2019 | Clustering for heterogeneous information networks with extended star-structure
Jian-Ping Mei, Huajiang Lv, Lianghuai Yang |
Data Min. Knowl. Discov. | 1 |
| 2019 | Radar emitter identification with bispectrum and hierarchical extreme learning machine
Ru Cao, Jiuwen Cao, Jian-Ping Mei, Chun Yin, Xuegang Huang |
Multim. Tools Appl. | 3 |
| 2019 | Semisupervised Fuzzy Clustering With Partition Information of SubsetsabstractPairwise constraint is a type of side information that is widely considered in existing semisupervised clustering approaches. In this paper, we explore a new form of supervision for clustering. We consider the partition results of a number of subsets as additional information to assist clustering. Compared to the pairwise constraint, which only involves the “must-link” or “cannot-link” relationship of two objects, the partition of a subset of objects provides information about the group structure of more objects and hence can possibly serve as a more effective form of supervision for clustering. In this paper, we instantiate the idea of clustering with subset partitions under the fuzzy clustering framework for document categorization. The proposed fuzzy clustering approach is formulated to learn from the partition of subsets and has the ability to handle high-dimensional document data. Specifically, the partition results of subsets are collectively transformed into pairwise relationships, based on which a penalty term is constructed and incorporated into a cosine-distance-based fuzzy c-means approach. The experimental results on benchmark data sets demonstrate the effectiveness of the proposed approach for a semisupervised document clustering. Jian-Ping Mei |
IEEE Trans. Fuzzy Syst. | 1 |
| 2017 | A social influence based trust model for recommender systemsabstractTrustworthy computing has recently attracted significant interest from researchers in several fields including multi-agent systems, social network analysis, and recommender systems. As an additional dimension of information to past rating history, trust has been shown to be helpful for improving th e accuracy of recommendations. Studies on the relationship between trust and rating behaviors may provide insights into the formation of trust in the context of online community, and lead to possible indicators for the effective use of trust in recommendations. In this paper, we study people's trust and rating behavior with the Epinions dataset. Epinions.com is a popular product review website allowing users to rate various categories of products, and establish a list of trustworthy users. We perform correlation analysis of activeness and trustworthiness defined by the number of ratings and the number of trustors to derive findings that can help the design of new decision support mechanisms in trust-based recommender systems. We then propose a trustee-influence based trust model where a trustee's activeness or trustworthiness is used to determine trust relationships. This trust model is incorporated into a memory-based and matrix factorization recommender systems to support online purchasing decision-making. Experimental results demonstrate the effectiveness of the proposed trust model for recommendation. Jian-Ping Mei, Han Yu 0001, Zhiqi Shen 0001, Chunyan Miao |
Intell. Data Anal. | 1 |
| 2017 | Large Scale Document Categorization With Fuzzy ClusteringabstractClustering documents into coherent categories is a very useful and important step for document processing and understanding. The introducing of fuzzy set theory into clustering provides a favorable mechanism to capture overlapping among document clusters. Document dataset is commonly represented as a collection of high-dimensional vectors, which may not be able to fit into memory entirely, when the dataset is large and with a very high dimensionality. However, most of the existing fuzzy clustering approaches deal with small and static datasets. Some of them may have a good scalability but they are only effective for low dimensional data. The study presented in this paper is about new efforts on fuzzy clustering of large-scale and high-dimensional data-especially suitable for document categorization. To consider both large scale and high dimensionality into the problem formulation, our key idea is to incorporate document-tailored fuzzy clustering into a scheme, which is effective for dealing with a large-scale problem. We first identified three representative schemes in fuzzy clustering for handling large-scale data, namely sampling extension, single pass, and divide ensemble. The limitation of fuzzy C-means (FCM)-based approaches for a large document clustering are then investigated. Based on the study, we propose new approaches by incorporating each of hyperspherical FCM and fuzzy coclustering with the three scale-up schemes, respectively. This enables our new approaches to maintain effectiveness for high-dimensional data with an extended scalability. Extensive experimental studies with real-world large document datasets have been conducted and the results demonstrate that the proposed approaches perform consistently better over existing ones in document categorization. Jian-Ping Mei, Yangtao Wang, Lihui Chen 0001, Chunyan Miao |
IEEE Trans. Fuzzy Syst. | 1 |
| 2016 | Hyperspherical Fuzzy clustering for online document categorizationabstractFor document data which are typically represented as high dimensional and sparse vectors, cosine distance based Hyperspherical Fuzzy C-Means(HFCM) has been shown to be more effective than classic Euclidean distance based fuzzy c-means(FCM) for document categorization. The existing HFCM approach assumes a static dataset and performs clustering in a batch mode. This design makes HFCM no more suitable when the dataset keeps increasing in its size or is too large to be loaded wholly. In this paper, we work on fuzzy clustering approaches for online document clustering. Specifically, we propose to perform hyperspherical fuzzy c-means in an online manner based on the stochastic gradient method. In this Online Hyperspherical Fuzzy C-Means(OnHFCM), documents are assumed to come one by one and centroids of clusters are updated immediately when each document is arrived. Such a formulation allows OnHFCM to be applicable to large datasets which may be incrementally changing in size. Two variants of OnHFCM with different objective functions are presented in this paper. In addition to the classic online algorithm which processes documents one by one, mini-batch algorithm that handles a small batch of documents at a time is also provided. Experimental results on real-world benchmarks demonstrate that the proposed OnHFCM algorithms achieved better performance than existing ones in online document categorization. Jian-Ping Mei, Yangtao Wang |
FUZZ-IEEE | 1 |
| 2014 | Incremental fuzzy clustering for document categorizationabstractIncremental clustering has been proposed to handle large datasets which can not fit into memory entirely. Single pass fuzzy c-means (SpFCM) and Online fuzzy c-means (OFCM) are two representative incremental fuzzy clustering methods. Both of them extend the scalability of fuzzy c-means (FCM) by processing the dataset chunk by chunk. However, due to the data sparsity and high-dimensionality, SpFCM and OFCM fail to produce reasonable results for document data. In this study, we work on clustering approaches that take care of both the large-scale and high-dimensionality issues. Specifically, we propose two methods for incrementally clustering of document data. The first method is a modification of the existing FCM-based incremental clustering with a step to normalize the centroids in each iteration, while the other method is incremental clustering, i.e., Single-Pass or Online, with weighted fuzzy co-clustering. We use several benchmark document datasets for experimental study. The experimental results show that the proposed approaches achieved significant improvements over existing SpFCM and OFCM in document clustering. Jian-Ping Mei, Yangtao Wang, Lihui Chen 0001, Chunyan Miao |
FUZZ-IEEE | 1 |
| 2014 | Stochastic gradient descent based fuzzy clustering for large dataabstractData is growing at an unprecedented rate in commercial and scientific areas. Clustering algorithms for large data which require small memory consumption and scalability become increasingly important under this circumstance. In this paper, we propose a new clustering approach called stochastic gradient based fuzzy clustering(SGFC) which achieves the optimization based on stochastic approximation to handle such kind of large data. We derive an adaptive learning rate which can be updated incrementally and maintained automatically in gradient descent approach employed in SGFC. Moreover, SGFC is extended to a mini-batch SGFC to reduce the stochastic noise. Additionally, multi-pass SGFC is also proposed to improve the clustering performance. Experiments have been conducted on synthetic data to show the effectiveness of our derived adaptive learning rate. Experimental studies have been also conducted on several large benchmark datasets including real world image and document datasets. Compared with existing fuzzy clustering approaches for large data, the mini-batch SGFC shows comparable or better accuracy with significant less time consumption. These results demonstrate the great potential of SGFC for large data analysis. Yangtao Wang, Lihui Chen 0001, Jian-Ping Mei |
FUZZ-IEEE | 3 |
| 2014 | A Social Trust Model Considering Trustees' Influence
Jian-Ping Mei, Han Yu 0001, Yong Liu 0020, Zhiqi Shen 0001, Chunyan Miao |
PRIMA | 1 |
| 2014 | Proximity-based k-partitions clustering with ranking for document categorization and analysis
Jian-Ping Mei, Lihui Chen 0001 |
Expert Syst. Appl. | 1 |
| 2014 | Incremental Fuzzy Clustering With Multiple Medoids for Large DataabstractAs an important technique of data analysis, clustering plays an important role in finding the underlying pattern structure embedded in unlabeled data. Clustering algorithms that need to store all the data into the memory for analysis become infeasible when the dataset is too large to be stored. To handle such large data, incremental clustering approaches are proposed. The key idea behind these approaches is to find representatives (centroids or medoids) to represent each cluster in each data chunk, which is a packet of the data, and final data analysis is carried out based on those identified representatives from all the chunks. In this paper, we propose a new incremental clustering approach called incremental multiple medoids-based fuzzy clustering (IMMFC) to handle complex patterns that are not compact and well separated. We would like to investigate whether IMMFC is a good alternative to capturing the underlying data structure more accurately. IMMFC not only facilitates the selection of multiple medoids for each cluster in a data chunk, but also has the mechanism to make use of relationships among those identified medoids as side information to help the final data clustering process. The detailed problem formulation, updating rules derivation, and the in-depth analysis of the proposed IMMFC are provided. Experimental studies on several large datasets that include real world malware datasets have been conducted. IMMFC outperforms existing incremental fuzzy clustering approaches in terms of clustering accuracy and robustness to the order of data. These results demonstrate the great potential of IMMFC for large-data analysis. Yangtao Wang, Lihui Chen 0001, Jian-Ping Mei |
IEEE Trans. Fuzzy Syst. | 3 |
| 2013 | Drug-target interaction prediction by learning from local information and neighborsabstractMOTIVATION: In silico methods provide efficient ways to predict possible interactions between drugs and targets. Supervised learning approach, bipartite local model (BLM), has recently been shown to be effective in prediction of drug-target interactions. However, for drug-candidate compounds or target-candidate proteins that currently have no known interactions available, its pure 'local' model is not able to be learned and hence BLM may fail to make correct prediction when involving such kind of new candidates. RESULTS: We present a simple procedure called neighbor-based interaction-profile inferring (NII) and integrate it into the existing BLM method to handle the new candidate problem. Specifically, the inferred interaction profile is treated as label information and is used for model learning of new candidates. This functionality is particularly important in practice to find targets for new drug-candidate compounds and identify targeting drugs for new target-candidate proteins. Consistent good performance of the new BLM-NII approach has been observed in the experiment for the prediction of interactions between drugs and four categories of target proteins. Especially for nuclear receptors, BLM-NII achieves the most significant improvement as this dataset contains many drugs/targets with no interactions in the cross-validation. This demonstrates the effectiveness of the NII strategy and also shows the great potential of BLM-NII for prediction of compound-protein interactions. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jian-Ping Mei, Chee Keong Kwoh 0001, Peng Yang 0010, Xiaoli Li 0001, Jie Zheng 0002 |
Bioinform. | 1 |
| 2013 | LinkFCM: Relation integrated fuzzy c-means
Jian-Ping Mei, Lihui Chen 0001 |
Pattern Recognit. | 1 |
| 2012 | Positive-unlabeled learning for disease gene identificationabstractBACKGROUND: Identifying disease genes from human genome is an important but challenging task in biomedical research. Machine learning methods can be applied to discover new disease genes based on the known ones. Existing machine learning methods typically use the known disease genes as the positive training set P and the unknown genes as the negative training set N (non-disease gene set does not exist) to build classifiers to identify new disease genes from the unknown genes. However, such kind of classifiers is actually built from a noisy negative set N as there can be unknown disease genes in N itself. As a result, the classifiers do not perform as well as they could be. RESULT: Instead of treating the unknown genes as negative examples in N, we treat them as an unlabeled set U. We design a novel positive-unlabeled (PU) learning algorithm PUDI (PU learning for disease gene identification) to build a classifier using P and U. We first partition U into four sets, namely, reliable negative set RN, likely positive set LP, likely negative set LN and weak negative set WN. The weighted support vector machines are then used to build a multi-level classifier based on the four training sets and positive training set P to identify disease genes. Our experimental results demonstrate that our proposed PUDI algorithm outperformed the existing methods significantly. CONCLUSION: The proposed PUDI algorithm is able to identify disease genes more accurately by treating the unknown data more appropriately as unlabeled set U instead of negative set N. Given that many machine learning problems in biomedical research do involve positive and unlabeled data instead of negative data, it is possible that the machine learning methods for these problems can be further improved by adopting PU learning methods, as we have done here for disease gene identification. AVAILABILITY AND IMPLEMENTATION: The executable program and data are available at http://www1.i2r.a-star.edu.sg/~xlli/PUDI/PUDI.html. Peng Yang 0010, Xiaoli Li 0001, Jian-Ping Mei, Chee Keong Kwoh 0001, See-Kiong Ng |
Bioinform. | 3 |
| 2012 | SumCR: A new subtopic-based extractive approach for text summarization
Jian-Ping Mei, Lihui Chen 0001 |
Knowl. Inf. Syst. | 1 |
| 2012 | A Fuzzy Approach for Multitype Relational Data ClusteringabstractMining interrelated data among multiple types of objects or entities is important in many real-world applications. Despite extensive study on fuzzy clustering of vector space data, very limited exploration has been made on fuzzy clustering of relational data that involve several object types. In this paper, we propose a new fuzzy clustering approach for multitype relational data (FC-MR). In FC-MR, different types of objects are clustered simultaneously. An object is assigned a large membership with respect to a cluster if its related objects in this cluster have high rankings. In each cluster, an object tends to have a high ranking if its related objects have large memberships in this cluster. The FC-MR approach is formulated to deal with multitype relational data with various structures. The objective function of FC-MR is locally optimized by an efficient iterative algorithm, which updates the fuzzy membership matrix and the ranking matrix of one type at once while keeping those of other types constant. We also discuss the simplified FC-MR for multitype relational data with two special structures, namely, star-structure and extended star-structure. Experimental studies are conducted on benchmark document datasets to illustrate how the proposed approach can be applied flexibly under different scenarios in real-world applications. The experimental results demonstrate the feasibility and effectiveness of the new approach compared with existing ones. Jian-Ping Mei, Lihui Chen 0001 |
IEEE Trans. Fuzzy Syst. | 1 |
| 2011 | Fuzzy clustering approach for star-structured multi-type relational dataabstractRecently, mining interrelated data among multiple types of objects attracts a lot of attention due to its importance in many real-world applications. Despite of extensive study on fuzzy clustering of vector space data and homogeneous relational data, very limited exploration has been made on fuzzy clustering of relational data involving several object types. In this paper, we propose FC-SMR, a fuzzy approach for clustering star-structured multi-type relational data, where the central type is related to multiple attribute types. In FC-SMR, objects of the central type are clustered based on the rankings of objects of different attribute types. We formulate the clustering problem as a constrained maximization problem and give an efficient algorithm for finding local solutions of the defined objective function. Experimental studies conducted on real-world document data show the effectiveness of the new approach. Jian-Ping Mei, Lihui Chen 0001 |
FUZZ-IEEE | 1 |
| 2011 | Fuzzy relational clustering around medoids: A unified view
Jian-Ping Mei, Lihui Chen 0001 |
Fuzzy Sets Syst. | 1 |
| 2010 | Fuzzy clustering with weighted medoids for relational data
Jian-Ping Mei, Lihui Chen 0001 |
Pattern Recognit. | 1 |
| 2008 | Entropy-based robust fuzzy clustering of relational dataabstractRelational data clustering algorithms are proposed to deal with the data represented as the similarity or dissimilarity between each pair of objects. Fuzzy clustering of relational data (FRC) is a recently proposed approach that can handle non-Euclidean distance relational data. Unfortunately, negative values may appear in the clustering process of FRC. Another related algorithm A-P (assignment prototype) applies two different memberships and obtains a more stable minimization procedure. However, the fixed exponent m and sensitivity to initialization make A-P less feasible to some data sets. In this paper, we propose a new entropy-based fuzzy clustering for relational data (EFRC). EFRC and its robust version R-EFRC make use of two types of memberships called partitioning and ranking. Experiments on typical relational data sets and 2-D noisy data sets show that the new algorithm can produce meaningful clustering results and is robust to noise. Jian-Ping Mei, Lihui Chen 0001 |
SMC | 1 |