VLDB 2026 Research / reviewers in the wild / expert
Ming Zheng
dblp:03/6858
· DBLP profile ↗
18ranked-venue papers
10as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 5 since 2021Systems, architecture and hardware · 5 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fed-DiTTab: Diffusion transformer for tabular data generation in federated learning
Yajun Pi, Ming Zheng, Fanhao Ma |
Neural Networks | 2 |
| 2026 | Federated Recommendation Model Based on Personalized Attention and Privacy-Preserving Dynamic GraphabstractGraph Neural Networks (GNNs) have been widely adopted in recommendation systems. When integrated into a federated learning framework, GNNs can enhance the model’s expressive capability. However, challenges arise in personalized representation and graph expansion due to the heterogeneity and locality of user data in federated recommendation systems. To address these challenges, we propose a federated recommendation model based on personalized attention and privacy-preserving dynamic graphs. The method first matches neighbor users for each selected client. Subsequently, it counts the interaction frequencies of items for both local and neighbor users to construct personalized weights, which captures the unique characteristics of different users. Additionally, we designs a method for constructing privacy-preserving dynamic graphs. In each round of federated training, the selected client adds pseudo-interaction items to its own interaction subgraph, perturbing the real interactions. After completing local training, the noisy interaction subgraph is incorporated into the global graph to capture higher-order connectivity information among users while safeguarding their interaction privacy. We conduct extensive experiments on three benchmark datasets, and the results demonstrate that the proposed PADG method achieves superior performance while effectively protecting privacy. Xiaoyao Zheng, Shukai Ye, Ming Zheng, Liangmin Guo, Qingying Yu, Yonglong Luo |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2025 | Multi-Scale Dual-Domain Attention Network for Traffic Flow PredictionabstractIn modern urban traffic management, accurate traffic flow prediction contributes to travel decision optimization, signal control, emergency response and resource scheduling, which improves road efficiency, reduces congestion and promotes sustainable urban planning. However, traffic flow data are characterized by complex spatio-temporal dependence, multi-scale periodicity, and highly dynamic changes. Existing studies mostly focus on single spatio-temporal domain modeling and ignore frequency domain information fusion, which makes it difficult to comprehensively capture the potential laws of traffic flow. To this end, a multi-scale spatio-temporal fusion Transformer prediction model is proposed, which systematically integrates frequency-domain analysis with spatio-temporal dependent modeling. The model contains three parts: (1) spatial-temporal adaptive neighbor selection algorithm, which dynamically supplements topological information based on spatio-temporal correlation to enhance the efficiency of inter-subdivisional information transfer; (2) frequency-domain feature coupling module, which fuses the frequency-domain and spatial-domain features by fast Fourier transform to enhance the ability of temporal pattern sensing; and (3) spatio-temporal-frequency-domain dual-attention encoder, which combines the linear-attention mechanism to efficiently capture the longrange spatio-temporal dependencies. Experimental results on several real traffic datasets show that the model significantly outperforms existing methods in terms of mean absolute error, root mean square error and mean absolute percentage error, and demonstrates stronger robustness in complex spatio-temporal patterns and abnormal fluctuation scenarios. Chenhui Wei, Chuanming Chen, Ming Zheng, Tianjiao Ni, Qingying Yu |
ICPADS | 4 |
| 2025 | FLO: Focal Loss-Based Oversampling for High-Dimensional Imbalanced Biomedical Data Classification
Ming Zheng, Qiran Zhou, Fulong Chen 0002 |
ICPADS | 1 |
| 2025 | Optimal Sampling Rate Selection for Parallel Hybrid Sampling Framework of Imbalanced Data ClassificationabstractABSTRACT Imbalanced data classification is one of the challenges in the field of data mining and machine learning. At present, the main method to solve imbalanced data classification issues from the data level is resampling. Hybrid sampling is widely used because it can avoid the problem of overfitting or mistakenly deleting the most useful samples when using oversampling or undersampling alone. However, the current hybrid sampling methods are mostly implemented in serial, which has the problems of excessive time cost and mutual influence. Meanwhile, few studies consider the automatic determination of the oversampling rate and undersampling rate in hybrid sampling methods but use the default sampling rate. Therefore, this study proposes a novel optimal sampling rate selection for a parallel hybrid sampling framework of imbalanced data classification. At the same time, we improved the differential evolution algorithm to optimize the sampling rate for the parallel hybrid sampling framework to obtain more stable classification performance. Experiments show that the SRDE specifically designed for the parallel hybrid sampling framework is superior to other heuristic algorithms on imbalanced data classification, and the results of statistical test experiments also verified this result. Zhuo Zhao, Ming Zheng, Fanhao Ma |
Concurr. Comput. Pract. Exp. | 2 |
| 2025 | UFIDSF: An undersampling approach based on feature importance and double side filter for imbalanced data classification
Ming Zheng, Liangchen Hu, Qingying Yu, Xiaoyao Zheng |
Future Gener. Comput. Syst. | 1 |
| 2025 | Resampling approach for imbalanced data classification based on class instance density per feature value intervals
Ming Zheng |
Inf. Sci. | 2 |
| 2025 | Fed-UGI: Federated Undersampling Learning Framework With Gini Impurity for Imbalanced Network Intrusion DetectionabstractIn the modern interconnected world, the popularization of networks and the rapid development of information technology led to the increasing security risks and threats in network systems. The existing intrusion detection system is constantly challenged by various malicious intrusion attacks. Machine learning algorithms have been widely used in intrusion detection. However, the model training requires the support of a sufficient high-quality samples, especially attack traffic data. Network intrusion detection datasets may not be shared between organizations due to data security and some privacy policy concerns. The federated learning framework is an optimal approach to address this issue, in which organizations collaborate to train a global model shared by multiple parties while keeping the data local to the client, guaranteeing the data privacy and security of all parties. However, there is a problem of class imbalance in the network traffic data owned by the organizations, which seriously affects the detection performance of the model and leads to a high consumption of model training time. Therefore, this study proposed a novel federated undersampling learning framework with Gini impurity, namely Fed-UGI. The framework is based on the hash-based block undersampling method to rebalance the client, which can solve the influence of imbalanced training data on the model detection performance and improve the model training efficiency. Moreover, the client weighted aggregation strategy based on Local Gini impurity can further optimize the effect of global model aggregation and reduce the impact of the dispersion degree and information difference in client data on model aggregation. In addition, extensive experiments on intrusion detection datasets show that compared to SOTA methods, the proposed Fed-UGI method has a good detection effect on the three metrics of F1-score, G-mean and AUC, the training time of the model is reduced by 51.76%-92.58%, especially in highly class imbalance situation. Ming Zheng, Ying Hu 0006, Xiaoyao Zheng, Yonglong Luo |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2025 | Fed-OLF: Federated Oversampling Learning Framework for Imbalanced Software Defect Prediction Under Privacy ProtectionabstractSoftware defect prediction technology can discover potential errors or hidden defects by establishing prediction models before the use of products in the field of software engineering, so as to reduce subsequent problems and improve software quality and security. However, building predictive models requires enough software defect dataset support, especially defect samples. Due to the involvement of confidential information from various organizations or enterprises, software defect data cannot be shared and effectively utilized. Therefore, to achieve collaborative training of multiparty shared software defect prediction models while keeping the data local to various organizations, we made the federated learning framework for the issue of software defect prediction. Meanwhile, the nondefect and defect instances in software defect datasets are usually imbalanced, which can seriously affect the software defect prediction performance of the model. Therefore, this study designs a novel federated oversampling learning framework Fed-OLF. First, the TabDiT method based on deep generative model is proposed in Fed-OLF to expand and rebalance the local imbalanced software defect dataset of each client with a certain degree of privacy protection. Second, a parameter aggregation strategy based on local information entropy is proposed in Fed-OLF to further optimize the parameter aggregation effect of the global shared model, thereby achieving better model performance. We conduct extensive experiments on the PROMISE dataset and the NASA Promise repository, and experimental results on the PROMISE dataset and the NASA Promise repository show that, the proposed Fed-OLF exhibits better predictive performance under the F1-score, G-mean, and AUC metrics when compared with the advanced baseline methods. In addition, we verify that both the TabDiT method and the parameter aggregation strategy based on local information entropy in Fed-OLF are useful, and the combination of them can more effectively improve model performance. Ming Zheng, Rui Zhu 0009, Xuan Zhang 0002, Zhi Jin 0001 |
IEEE Trans. Reliab. | 2 |
| 2024 | Exploratory parallel hybrid sampling framework for imbalanced data classification
Ming Zheng, Zhuo Zhao, Wanggen Li, Tong Li 0004 |
Eng. Appl. Artif. Intell. | 1 |
| 2024 | Which standard classification algorithm has more stable performance for imbalanced network traffic data?
Ming Zheng, Qingying Yu, Liangmin Guo, Fulong Chen 0002 |
Soft Comput. | 1 |
| 2023 | Neural Frailty Machine: Beyond proportional hazard assumption in neural survival regressionsabstractWe present neural frailty machine (NFM), a powerful and flexible neural modeling framework for survival regressions. The NFM framework utilizes the classical idea of multiplicative frailty in survival analysis as a principled way of extending the proportional hazard assumption, at the same time being able to leverage the strong approximation power of neural architectures for handling nonlinear covariate dependence. Two concrete models are derived under the framework that extends neural proportional hazard models and nonparametric hazard regression models. Both models allow efficient training under the likelihood objective. Theoretically, for both proposed models, we establish statistical guarantees of neural function approximation with respect to nonparametric components via characterizing their rate of convergence. Empirically, we provide synthetic experiments that verify our theoretical statements. We also conduct experimental evaluations over $6$ benchmark datasets of different scales, showing that the proposed NFM models achieve predictive performance comparable to or sometimes surpassing state-of-the-art survival models. Our code is publicly availabel at https://github.com/Rorschach1989/nfm Jiawei Qiao, Mingzhe Wu, Wen Yu 0002, Ming Zheng, Tengfei Liu 0007, Weiqiang Wang 0002 |
NeurIPS | 5 |
| 2021 | UFFDFR: Undersampling framework with denoising, fuzzy c-means clustering, and representative sample selection for imbalanced data classification
Ming Zheng, Tong Li 0004, Xiaoyao Zheng, Qingying Yu, Chuanming Chen, Changlong Lv |
Inf. Sci. | 1 |
| 2021 | An automatic sampling ratio detection method based on genetic algorithm for imbalanced data classification
Ming Zheng, Tong Li 0004, Taochun Wang, Biao Jie, Mingjing Tang, Changlong Lv |
Knowl. Based Syst. | 1 |
| 2020 | Conditional Wasserstein generative adversarial network-gradient penalty-based approach to alleviating imbalanced data classification
Ming Zheng, Tong Li 0004, Rui Zhu 0009, Yahui Tang, Mingjing Tang, Leilei Lin, Zifei Ma |
Inf. Sci. | 1 |
| 2020 | A data-driven risk measurement model of software developer turnover
Zifei Ma, Ruiyin Li, Tong Li 0004, Rui Zhu 0009, Mingjing Tang, Ming Zheng |
Soft Comput. | 8 |
| 2011 | Anonymous Communication over Invisible Mix Rings
Ming Zheng, Hai-Xin Duan |
ICA3PP (1) | 1 |
| 2007 | 2D NMR metabonomic analysis: a novel method for automated peak alignmentabstractMOTIVATION: Comparative metabolic profiling by nuclear magnetic resonance (NMR) is showing increasing promise for identifying inter-individual differences to drug response. Two dimensional (2D) (1)H (13)C NMR can reduce spectral overlap, a common problem of 1D (1)H NMR. However, the peak alignment tools for 1D NMR spectra are not well suited for 2D NMR. An automated and statistically robust method for aligning 2D NMR peaks is required to enable comparative metabonomic analysis using 2D NMR. RESULTS: A novel statistical method was developed to align NMR peaks that represent the same chemical groups across multiple 2D NMR spectra. The degree of local pattern match among peaks in different spectra is assessed using a similarity measure, and a heuristic algorithm maximizes the similarity measure for peaks across the whole spectrum. This peak alignment method was used to align peaks in 2D NMR spectra of endogenous metabolites in liver extracts obtained from four inbred mouse strains in the study of acetaminophen-induced liver toxicity. This automated alignment method was validated by manual examination of the top 50 peaks as ranked by signal intensity. Manual inspection of 1872 peaks in 39 different spectra demonstrated that the automated algorithm correctly aligned 1810 (96.7%) peaks. AVAILABILITY: Algorithm is available upon request. Ming Zheng, Yanzhou Liu 0002, Joseph Pease, Jonathan Usuka, Guochun Liao, Gary Peltz |
Bioinform. | 1 |