VLDB 2026 Research / reviewers in the wild / expert
Guancheng Lin
dblp:323/8302
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2026
0009-0009-8339-313XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GaitDG: A Single-Source Domain Generalization Framework for Cross-Domain Gait RecognitionabstractIn recent years, significant advances in gait recognition have been seen, with many methods reporting high accuracy on certain datasets. However, domain shifts, such as distribution inconsistencies in viewpoint or clothing, can severely degrade the performance of these models on unseen target domains, hindering the widespread application of gait recognition. Some unsupervised domain adaptation (UDA) methods have been proposed to address this problem. However, these approaches require continual updates with target domain data, which is often difficult to obtain due to privacy concerns and deployment complexity. This paper presents GaitDG, a single-source domain generalization framework designed to enhance the generalization ability of gait recognition models for unseen target domains, requiring training on only one source domain without accessing target domain data. During training, GaitDG employs adversarial training to disentangle domain-specific and identity-specific features, enabling the discovery of latent sub-domains and the extraction of domain-invariant features. Furthermore, GaitDG supports the integration of data augmentation to diversify the source domain data. We also introduce a data augmentation method, Segmentation Model Transfer (SMT), to mitigate recognition performance degradation caused by variations in segmentation models. As a model-agnostic approach, GaitDG can directly enhance the cross-domain recognition performance of gait recognition models without altering their structure. Comprehensive experiments on widely used gait datasets demonstrate that GaitDG significantly improves the cross-domain recognition performance of several state-of-the-art gait recognition models. Guancheng Lin, Man Zhou 0004, Lianmiao Wang, Qin Liu 0003, Yueyue Dai, Fue Zeng |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2025 | Less Is More: Unlocking Semi-Supervised Deep Learning for Vulnerability DetectionabstractDeep learning has demonstrated its effectiveness in software vulnerability detection, but acquiring a large number of labeled code snippets for training deep learning models is challenging due to labor-intensive annotation. With limited labeled data, complex deep learning models often suffer from overfitting and poor performance. To address this limitation, semi-supervised deep learning offers a promising approach by annotating unlabeled code snippets with pseudo-labels and utilizing limited labeled data together as training sets to train vulnerability detection models. However, applying semi-supervised deep learning for accurate vulnerability detection comes with several challenges. One challenge lies in how to select correctly pseudo-labeled code snippets as training data, while another involves mitigating the impact of potentially incorrectly pseudo-labeled training code snippets during model training. To address these challenges, we propose the semi-supervised vulnerability detection (SSVD) approach. SSVD leverages the information gain of model parameters as the certainty of the correctness of pseudo-labels and prioritizes high-certainty pseudo-labeled code snippets as training data. Additionally, it incorporates the proposed noise-robust triplet loss to maximize the separation between vulnerable and non-vulnerable code snippets to better propagate labels from labeled code snippets to nearby unlabeled snippets and utilizes the proposed noise-robust cross-entropy loss for gradient clipping to mitigate the error accumulation caused by incorrect pseudo-labels. We evaluate SSVD with nine semi-supervised approaches on four widely-used public vulnerability datasets. The results demonstrate that SSVD outperforms the baselines with an average of 29.82% improvement in terms of F1-score and 56.72% in terms of MCC. In addition, SSVD trained on a certain proportion of labeled data can outperform or closely match the performance of fully supervised LineVul and ReVeal vulnerability detection models trained on 100% labeled data in most scenarios. This indicates that SSVD can effectively learn from limited labeled data to enhance vulnerability detection performance, thereby reducing the effort required for labeling a large number of code snippets. Xiao Yu 0008, Guancheng Lin, Xing Hu 0008, Jacky W. Keung, Xin Xia 0001 |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2024 | Enhancing Deep Learning Vulnerability Detection through Imbalance Loss Functions: An Empirical StudyabstractSoftware Vulnerability Detection (VD) is crucial in software engineering, and Deep Learning (DL) has demonstrated effective in this domain. However, the class imbalance issue, where non-vulnerable code snippets vastly outnumber vulnerable ones, hinders the performance of DL-based Vulnerability Detection (DLVD) models. Recent studies have explored data resampling methods to address this, but these methods often lead to data distribution alterations, resulting in information loss, model overfitting, and reduced interpretability. Imbalance loss functions have thus emerged as viable alternatives. To comprehensively evaluate the effectiveness of imbalance loss functions in DLVD, we investigate six imbalance loss functions and Cross-Entropy Loss (the default for LineVul and ReVeal models) on two DLVD models across three public VD datasets, using three evaluation metrics and the Scott-Knott Effect Size Difference test. Our findings provide valuable insights into selecting loss functions and data resampling methods in DLVD. First, the DLVD model LineVul outperforms ReVeal across all datasets. Second, Label Distribution-Aware Margin loss and Random Under-Sampling generally yield the best Precision and Recall, respectively. Third, to avoid information loss and maintain interpretability, we recommend Logit Adjustment Loss (LALoss) due to its high Recall and superior F1 metric performance. Based on these findings, we suggest employing LineVul with LALoss for VD, as it detects more vulnerable code snippets (higher Recall) while providing comprehensive performance (higher F1). Yanzhong He, Guancheng Lin, Jacky W. Keung |
Internetware | 2 |
| 2024 | Revisiting Code Smell Severity Prioritization using learning to rank techniques
Guancheng Lin, Peilin Song, Xin Wang 0114 |
Expert Syst. Appl. | 2 |
| 2024 | Improving effort-aware defect prediction by directly learning to rank software modules
Xiao Yu 0008, Jiqing Rao, Lei Liu 0062, Guancheng Lin, Jacky W. Keung, Junwei Zhou 0002, Jianwen Xiang |
Inf. Softw. Technol. | 4 |
| 2023 | Revisiting "code smell severity classification using machine learning techniques"abstractIn the context of limited maintenance resources, predicting the severity of code smells is more practically useful than simply detecting them. Fontana et al. first empirically investigated some classification algorithms and some regression algorithms, for severity prediction. Their results showed that random forest and decision tree performed well on Mean Absolute Error (MAE), Mean Squared Error (MSE), and Spearman and Kendall rank correlation coefficients. However, they did not consider the issue of imbalanced data distribution in the severity dataset, and used inappropriate performance evaluation metrics. Therefore, we revisit the effectiveness of 10 classification methods and 11 regression methods, for code severity prediction using Cumulative Lift Chart (CLC) and Severity@20% as the primary performance metrics and Accuracy as the secondary performance indicator. The results show that the Gradient Boosting Regression (GBR) method performs the best in terms of these metrics. Lei Liu 0062, Peixin Yang, Kuan Zou, Guancheng Lin, Jianwen Xiang |
COMPSAC | 6 |
| 2023 | Quaternion fractional-order weighted generalized Laguerre-Fourier moments and moment invariants for color image analysis
Guancheng Lin, Wenqiang Xi |
Signal Process. Image Commun. | 3 |