VLDB 2026 Research / reviewers in the wild / expert
Hongwei Tao
dblp:32/7498
· DBLP profile ↗
15ranked-venue papers
6as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Software engineering, systems software and programming languages · 5 · 4 first-author · 5 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cross-project defect prediction based on transfer graph convolutional network
Hongwei Tao, Zhenhao Geng, Yongheng Xie |
Empir. Softw. Eng. | 3 |
| 2026 | CR-GAC: Cross-modal Recombination via Graph-Attention Collaborative Optimization for multimodal sentiment analysis
Haoran Chen 0004, Zuhe Li, Yushan Pan, Hongwei Tao, Huaiguang Wu, Yunyang Wang, Chenguang Yang 0001 |
Expert Syst. Appl. | 5 |
| 2026 | Software defect prediction based on graph code semantics
Hongwei Tao, Zhenhao Geng, Xiaoxu Niu, Qiaoling Cao |
Expert Syst. Appl. | 1 |
| 2025 | Prediction of incompatible bug numbers between versions of java open-source software based on deep fusion features
Xiaoxu Niu, Hongwei Tao, Qiaoling Cao, Jianxun Wang 0010, Zhenhao Geng |
Expert Syst. Appl. | 2 |
| 2025 | Software aging oriented trustworthiness measurement based on weighted Boltzmann entropy
Hongwei Tao, Han Liu 0012, Xiaoxu Niu, Licheng Ding, Yixiang Chen 0001, Qiaoling Cao |
Inf. Softw. Technol. | 1 |
| 2025 | Deep residual PLSR model with manifold optimization and Gaussian filter for enhanced image classification
Haoran Chen 0004, Wenjun Song, Hongwei Tao, Zuhe Li |
Vis. Comput. | 5 |
| 2024 | Hybrid density-based adaptive weighted collaborative representation for imbalanced learning
Junwei Jin 0001, Hongwei Tao, Chuang Han, C. L. Philip Chen |
Appl. Intell. | 4 |
| 2024 | Software Defect Prediction Method Based on Clustering Ensemble LearningabstractThe technique of software defect prediction aims to assess and predict potential defects in software projects and has made significant progress in recent years within software development. In previous studies, this technique largely relied on supervised learning methods, requiring a substantial amount of labeled historical defect data to train the models. However, obtaining these labeled data often demands significant time and resources. In contrast, software defect prediction based on unsupervised learning does not depend on known labeled data, eliminating the need for large‐scale data labeling, thereby saving considerable time and resources while providing a more flexible solution for ensuring software quality. This paper conducts software defect prediction using unsupervised learning methods on data from 16 projects across two public datasets (PROMISE and NASA). During the feature selection step, a chi‐squared sparse feature selection method is proposed. This feature selection strategy combines chi‐squared tests with sparse principal component analysis (SPCA). Specifically, the chi‐squared test is first used to filter out the most statistically significant features, and then the SPCA is applied to reduce the dimensionality of these significant features. In the clustering step, the dot product matrix and Pearson correlation coefficient (PCC) matrix are used to construct weighted adjacency matrices, and a clustering overlap method is proposed. This method integrates spectral clustering, Newman clustering, fluid clustering, and Clauset–Newman–Moore (CNM) clustering through ensemble learning. Experimental results indicate that, in the absence of labeled data, using the chi‐squared sparse method for feature selection demonstrates superior performance, and the proposed clustering overlap method outperforms or is comparable to the effectiveness of the four baseline clustering methods. Hongwei Tao, Qiaoling Cao, Xiaoxu Niu, Zhenhao Geng, Songtao Shang |
IET Softw. | 1 |
| 2024 | Cross-Project Defect Prediction Using Transfer Learning with Long Short-Term Memory NetworksabstractWith the increasing number of software projects, within‐project defect prediction (WPDP) has already been unable to meet the demand, and cross‐project defect prediction (CPDP) is playing an increasingly significant role in the area of software engineering. The classic CPDP methods mainly concentrated on applying metric features to predict defects. However, these approaches failed to consider the rich semantic information, which usually contains the relationship between software defects and context. Since traditional methods are unable to exploit this characteristic, their performance is often unsatisfactory. In this paper, a transfer long short‐term memory (TLSTM) network model is first proposed. Transfer semantic features are extracted by adding a transfer learning algorithm to the long short‐term memory (LSTM) network. Then, the traditional metric features and semantic features are combined for CPDP. First, the abstract syntax trees (AST) are generated based on the source codes. Second, the AST node contents are converted into integer vectors as inputs to the TLSTM model. Then, the semantic features of the program can be extracted by TLSTM. On the other hand, transferable metric features are extracted by transfer component analysis (TCA). Finally, the semantic features and metric features are combined and input into the logical regression (LR) classifier for training. The presented TLSTM model performs better on the f ‐measure indicator than other machine and deep learning models, according to the outcomes of several open‐source projects of the PROMISE repository. The TLSTM model built with a single feature achieves 0.7% and 2.1% improvement on Log4j‐1.2 and Xalan‐2.7, respectively. When using combined features to train the prediction model, we call this model a transfer long short‐term memory for defect prediction (DPTLSTM). DPTLSTM achieves a 2.9% and 5% improvement on Synapse‐1.2 and Xerces‐1.4.4, respectively. Both prove the superiority of the proposed model on the CPDP task. This is because LSTM capture long‐term dependencies in sequence data and extract features that contain source code structure and context information. It can be concluded that: (1) the TLSTM model has the advantage of preserving information, which can better retain the semantic features related to software defects; (2) compared with the CPDP model trained with traditional metric features, the performance of the model can validly enhance by combining semantic features and metric features. Hongwei Tao, Lianyou Fu, Qiaoling Cao, Xiaoxu Niu, Songtao Shang, Yang Xian |
IET Softw. | 1 |
| 2024 | Density-Based Discriminative Nonnegative Representation Model for Imbalanced ClassificationabstractAbstract Representation-based methods have found widespread applications in various classification tasks. However, these methods cannot deal effectively with imbalanced data scenarios. They tend to neglect the importance of minority samples, resulting in bias toward the majority class. To address this limitation, we propose a density-based discriminative nonnegative representation approach for imbalanced classification tasks. First, a new class-specific regularization term is incorporated into the framework of a nonnegative representation based classifier (NRC) to reduce the correlation between classes and improve the discrimination ability of the NRC. Second, a weight matrix is generated based on the hybrid density information of each sample’s neighbors and the decision boundary, which can assign larger weights to minority samples and thus reduce the preference for the majority class. Furthermore, the resulting model can be efficiently optimized through the alternating direction method of multipliers. Extensive experimental results demonstrate that our proposed method is superior to numerous state-of-the-art imbalanced learning methods. Junwei Jin 0001, Hongwei Tao, Jiaofen Nan, Huaiguang Wu, C. L. Philip Chen |
Neural Process. Lett. | 4 |
| 2024 | PDRLRR: A novel low-rank representation with projection distance regularization via manifold optimization for clustering
Haoran Chen 0004, Hongwei Tao, Zuhe Li, Boyue Wang |
Pattern Recognit. | 3 |
| 2024 | A comparative study of software defect binomial classification prediction models based on machine learning
Hongwei Tao, Xiaoxu Niu, Lianyou Fu, Qiaoling Cao, Songtao Shang, Yang Xian |
Softw. Qual. J. | 1 |
| 2023 | Low-rank Representation with Adaptive Dimensionality Reduction via Manifold Optimization for ClusteringabstractThe dimensionality reduction techniques are often used to reduce data dimensionality for computational efficiency or other purposes in existing low-rank representation (LRR)-based methods. However, the two steps of dimensionality reduction and learning low-rank representation coefficients are implemented in an independent way; thus, the adaptability of representation coefficients to the original data space may not be guaranteed. This article proposes a novel model, i.e., low-rank representation with adaptive dimensionality reduction (LRRARD) via manifold optimization for clustering, where dimensionality reduction and learning low-rank representation coefficients are integrated into a unified framework. This model introduces a low-dimensional projection matrix to find the projection that best fits the original data space. And the low-dimensional projection matrix and the low-rank representation coefficients interact with each other to simultaneously obtain the best projection matrix and representation coefficients. In addition, a manifold optimization method is employed to obtain the optimal projection matrix, which is an unconstrained optimization method in a constrained search space. The experimental results on several real datasets demonstrate the superiority of our proposed method. Haoran Chen 0004, Hongwei Tao, Zuhe Li, Xiao Wang 0009 |
ACM Trans. Knowl. Discov. Data | 3 |
| 2022 | Theoretical and empirical validation of software trustworthiness measure based on the decomposition of attributesabstractFrom the perspective of attribute decomposition, there are a variety of software trustworthiness metric models. However, little attention has been paid to using more rigorous methods and to performing theoretical validation. Axiomatic methods formalise the empirical understanding of software attributes through defining ideal metric properties. They can offer precise terms for the software attributes' quantification. We have utilised them to assess software trustworthiness on the basis of attribute decomposition, presented four properties, constructed a software trustworthiness measure (STMBDA for short). In this paper, we extend the set of properties, introduce two new properties, namely non-negativity and proportionality, and perfect substitutability and expectability. We verify the theoretical rationality of STMBDA by demonstrating that it conforms to the new property set and the empirical validity by evaluating the trustworthiness of 23 spacecraft software. The validation results show that STMBDA is able to effectively assess the spacecraft software trustworthiness and identify weaknesses in the development process. Hongwei Tao, Yixiang Chen 0001, Hengyang Wu |
Connect. Sci. | 1 |
| 2020 | Demagnetization Diagnosing in PMSM Based on SIDDTW Under Nonstationary ConditionsabstractDemagnetization, as one of the most frequent faults, has great influence on the performance of (permanent magnet synchronous motor) PMSM. However, the motor usually runs in nonstationary conditions, that brings great challenge to the effective diagnosis of demagnetization fault. This paper presents a new methodology of Shift-invariant Dictionary of Dynamic Time Warping (SIDDTW) to diagnose the demagnetization fault under nonstationary conditions. Firstly, according to the characteristics of current signal under demagnetization fault, the shift-invariant dictionary is constructed. Then, Matching Pursuit (MP) is used to represent the current signals that collected from the running process of PMSM, and then the sparse coefficient series are obtained. Finally, the Dynamic Time Warping (DTW) method is used to calculate the sparse coefficient series distance between the test data and the database which build in the training process. In this step, the nearest distance is matched, and corresponding operation state is recognized as the final diagnosis result. The results show that the presented method has good adaptability when dealing with nonstationary conditions both on the Simulink platform and the real-time simulation platform. Tao Peng 0010, Zhiwen Chen 0001, Chao Yang 0017, Hongwei Tao |
IECON | 5 |