Chaoyu Gong

dblp:258/4466 · DBLP profile ↗
← Back
19ranked-venue papers
12as first author
17since 2021 · last 2026
0000-0002-5540-5350ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 6 first-author · 10 since 2021Databases, data management, data science and information retrieval · 6 · 5 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 When Deepfake Detection Meets Graph Neural Network: A Unified and Lightweight Framework
abstract
The proliferation of generative video models has made detecting AI-generated and manipulated videos an urgent challenge. Existing detection approaches often fail to generalize across diverse manipulation types due to their reliance on isolated spatial, temporal, or spectral information, and typically require large models to perform well. This paper introduces SSTGNN, a lightweight Spatial-Spectral-Temporal Graph Neural Network framework that represents videos as structured graphs, enabling joint reasoning over spatial inconsistencies, temporal artifacts, and spectral distortions. SSTGNN incorporates learnable spectral filters and spatial-temporal differential modeling into a unified graph-based architecture, capturing subtle manipulation traces more effectively. Extensive experiments on diverse benchmark datasets demonstrate that SSTGNN not only achieves superior performance in both in-domain and cross-domain settings, but also offers strong efficiency and resource allocation. Remarkably, SSTGNN accomplishes these results with up to 42× fewer parameters than state-of-the-art models, making it highly lightweight and resource-friendly for real-world deployment.
Haoyu Liu 0001, Chaoyu Gong, Mengke He, Jiate Li, Kai Han 0001, Siqiang Luo
KDD (1)2
2026 Incremental open-set evidential reasoning for uncertainty-aware fault diagnosis in rotating machinery
Haiyue Fu, Chaoyu Gong, Zhi-gang Su
Knowl. Based Syst.3
2025 Unsupervised Learning for Class Distribution Mismatch
abstract
Class distribution mismatch (CDM) refers to the discrepancy between class distributions in training data and target tasks. Previous methods address this by designing classifiers to categorize classes known during training, while grouping unknown or new classes into an "other" category. However, they focus on semi-supervised scenarios and heavily rely on labeled data, limiting their applicability and performance. To address this, we propose Unsupervised Learning for Class Distribution Mismatch (UCDM), which constructs positive-negative pairs from unlabeled data for classifier training. Our approach randomly samples images and uses a diffusion model to add or erase semantic classes, synthesizing diverse training pairs. Additionally, we introduce a confidence-based labeling mechanism that iteratively assigns pseudo-labels to valuable real-world data and incorporates them into the training process. Extensive experiments on three datasets demonstrate UCDM’s superiority over previous semi-supervised methods. Specifically, with a 60\% mismatch proportion on Tiny-ImageNet dataset, our approach, without relying on labeled data, surpasses OpenMatch (with 40 labels per class) by 35.1%, 63.7%, and 72.5% in classifying known, unknown, and new classes.
Pan Du 0002, Wangbo Zhao, Xinai Lu, Zhikai Li, Chaoyu Gong, Suyun Zhao, Hong Chen 0001, Cuiping Li 0001, Kai Wang 0036, Yang You 0001
ICML6
2025 Sparse MeZO: Less Parameters for Better Performance in Zeroth-Order LLM Fine-Tuning
abstract
While fine-tuning large language models (LLMs) for specific tasks often yields impressive results, it comes at the cost of memory inefficiency due to back-propagation in gradient-based training. Memory-efficient Zeroth-order (MeZO) optimizers, recently proposed to address this issue, only require forward passes during training, making them more memory-friendly. However, compared with exact gradients, ZO-based gradients usually exhibit an estimation error, which can significantly hurt the optimization process, leading to slower convergence and suboptimal solutions. In addition, we find that the estimation error will hurt more when adding to large weights instead of small weights. Based on this observation, this paper introduces Sparse MeZO, a novel memory-efficient zeroth-order optimization approach that applies ZO only to a carefully chosen subset of parameters. We propose a simple yet effective parameter selection scheme that yields significant performance gains with Sparse-MeZO. Additionally, we develop a memory-optimized implementation for sparse masking, ensuring the algorithm requires only inference-level memory consumption, allowing Sparse-MeZO to fine-tune LLaMA-30b on a single A100 GPU. Experimental results illustrate that Sparse-MeZO consistently improves both performance and convergence speed over MeZO without any overhead. For example, it achieves a 9% absolute accuracy improvement and 3.5x speedup over MeZO on the RTE task.
Yong Liu 0020, Chaoyu Gong, Minhao Cheng, Cho-Jui Hsieh, Yang You 0001
NeurIPS3
2025 KMT-PLL: K-Means Cross-Attention Transformer for Partial Label Learning
abstract
Partial label learning (PLL) studies the problem of learning instance classification with a set of candidate labels and only one is correct. While recent works have demonstrated that the Vision Transformer (ViT) has achieved good results when training from clean data, its applications to PLL remain limited and challenging. To address this issue, we rethink the relationship between instances and object queries to propose K-means cross-attention transformer for PLL (KMT-PLL), which can continuously learn cluster centers and be used for downstream disambiguation tasks. More specifically, K-means cross-attention as a clustering process can effectively learn the cluster centers to represent label classes. The purpose of this operation is to make the similarity between instances and labels measurable, which can effectively detect noise labels. Furthermore, we propose a new corrected cross entropy formulation, which can assign weights to candidate labels according to the instance-to-label relevance to guide the training of the instance classifier. As the training goes on, the ground-truth label is progressively identified, and the refined labels and cluster centers in turn help to improve the classifier. Simulation results demonstrate the advantage of the KMT-PLL and its suitability for PLL.
Jinfu Fan, Linqing Huang, Chaoyu Gong, Yang You 0001, Min Gan, Zhongjie Wang 0004
IEEE Trans. Neural Networks Learn. Syst.3
2025 Time-Series Clustering With Dynamic Feature Mining in the Framework of Belief Function Theory
abstract
Clustering time-series data has gained abundant popularity and has been widely used in diverse scientific areas. However, few studies have systematically addressed the ambiguity and uncertainty contained in time-series clustering, which often leads to degraded clustering performance due to the high variability of time-series curves. Such ambiguity and uncertainty can be explained as the random dynamic changes of individual time series and are reflected in the clustering memberships of time-series data. Focusing on these issues, this article proposes a novel algorithm to tackle ambiguity and uncertainty in time-series clustering by capturing their dynamic features under the framework of belief function theory. Specifically, it employs an evidential Markov model for each time series to formalize dynamic features as a transition mass matrix, and an evidential clustering algorithm then derives a credal partition to group the data. Ablation studies validate the effectiveness of each component, and experiments on 128 UCR datasets demonstrate the strong performance of the proposed algorithm.
Yunshu Shi, Chaoyu Gong, Zhi-gang Su
IEEE Trans. Syst. Man Cybern. Syst.2
2024 Self-filling evidential clustering for partial multi-view data
Chaoyu Gong, Yang You 0001
Expert Syst. Appl.1
2024 Scalable Evidential K-Nearest Neighbor Classification on Big Data
abstract
TheK-Nearest Neighbor (K-NN) algorithm has garnered widespread utilization in real-world scenarios, due to its exceptional interpretability that other classification algorithms may not have. The evidential K-NN (EK-NN) algorithm builds upon the same nearest neighbor search procedure as K-NN, and provides more informative classification outcomes. However, EK-NN is not practical for big data because it is computationally complex. First, the search forKnearest neighbors of test samples from$n$training samples requires$O(n^{2})$operations. Additionally, estimating parameters involves performing complicated matrix calculations that increase in scale as the dataset becomes larger. To address these issues, we propose two scalable EK-NN classifiers, Global Exact EK-NN and Local Approximate EK-NN, under the distributed Spark framework. Along with the Local Approximate EK-NN, a new distributed gradient descent algorithm is developed to learn parameters. Data parallelism is used to reduce negative impacts caused by data distribution differences. Experimental results show that Our algorithms are able to achieve state-of-the-art scaling efficiency and accuracy on large datasets with more than 10 million samples.
Chaoyu Gong, James Demmel, Yang You 0001
IEEE Trans. Big Data1
2024 Distributed and Joint Evidential K-Nearest Neighbor Classification
abstract
The performance ofK-nearest neighbor (K-NN) classification depends significantly on the searched neighborhoods of test samples, namely, the neighborhood sizeKand the used distance metric. For these two issues, many methods either to acquire the adaptiveKor to learn a variant metric have been proposed and yielded appropriate performances. However, most of the existing methods ignore the fact that these two factors can be jointly learned. In this paper, we propose a Joint Evidential K-NN algorithm (JEKNN), which learns the adaptiveKof each sample and distance metric jointly based on the feedback of error function. To break the computational bottleneck of handling large datasets, a distributed version of JEKNN (JEKNN$_{\mathrm{{dis}}}$) is implemented under Apache Spark, i.e., an optimization algorithm based on distributed gradient descent and data parallelism is proposed to accelerate the training stage. Ablation and comparison experiments on small-scale datasets shows the performance improvement from the joint learning and the state-of-the-art accuracy of JEKNN, respectively. Compared to other KNN-based methods designed for Big Data, experimental results on big datasets demonstrate that JEKNN$_{\mathrm{{dis}}}$achieves better scaling efficiency without significant loss of accuracy. Besides, the generalization error bound of the proposed algorithm is also analyzed theoretically.
Chaoyu Gong, James Demmel, Yang You 0001
IEEE Trans. Knowl. Data Eng.1
2023 Adaptive evidential K-NN classification: Integrating neighborhood search and feature weighting
Chaoyu Gong, Zhi-gang Su, Yang You 0001
Inf. Sci.1
2023 A Sparse Reconstructive Evidential K-Nearest Neighbor Classifier for High-Dimensional Data
abstract
The EvidentialK-Nearest Neighbor (EK-NN) classification rule provides a global treatment of uncertainty and imprecision in class labels, and has been widely used in pattern recognition. Nevertheless, EK-NN still suffers from the fixed presupposition of hyper-parameterKwithout prior knowledge, due to the different spatial distribution of neighbors of each pattern in Euclidean space. More concretely, neighbors of some patterns may provide confusing information and then derive wrong classification results. To address this issue, we propose a sparse reconstructive evidentialK-NN (SEK-NN) classifier, appropriately determining an individualKfor each pattern and mapping the correlations between patterns from Euclidean space to a sparse reconstructed space. To match with this sparse reconstructed space, SEK-NN supersedes the Euclidean distance by correlation coefficients to measure the dissimilarities between patterns. When handling high-dimensional data, a parallel version of SEK-NN is implemented under the Apache Spark to speed up the parameter estimation. We respectively test SEK-NN and parallel SEK-NN over 19 middle dimensional datasets, 1 middle volume and 4 high-dimensional datasets that are up to 100 thousand of dimensions. Experimental results show that SEK-NN has great prediction performance and parallel SEK-NN is able to appropriately tackle high-dimensional datasets.
Chaoyu Gong, Zhi-gang Su, Pei-hong Wang, Yang You 0001
IEEE Trans. Knowl. Data Eng.1
2022 Self-reconstructive evidential clustering for high-dimensional data
abstract
Although many algorithms have been presented to tackle the curse of dimensionality in high-dimensional clustering, most of these algorithms require prior knowledge of the number of clusters. Besides, these existing algorithms create only a hard or fuzzy partition for high-dimensional objects, which are often located in highly overlapping areas. The adoption of hard/fuzzy partition ignores the ambiguity in the assignment of objects and may lead to performance degradation. To address these issues, we propose a novel self-reconstructive evidential clustering (SREC) algorithm. After learning the correlations between objects from a self-reconstruction process, SREC provides a human-readable chart. Through this chart, users can select several objects existing in the dataset as the cluster centers, instead of just detecting the number of clusters. Under the framework of evidence theory, SREC derives a more flexible credal partition that improves the fault tolerance of clustering. Ablation study demonstrates the benefits of the self-reconstruction and evidence theory. Comparison experiments on real-world datasets show that SREC consumes competitive running time and performs better than other state-of-the-art algorithms. We also apply SREC in a real-world application scenario to illustrate the rationality of selecting cluster centers by human intervention.
Chaoyu Gong, Di Fu, Yong Liu 0020, Pei-hong Wang, Yang You 0001
ICDE1
2022 Joint Evidential $K$-Nearest Neighbor Classification
abstract
The performance of$K$-nearest neighbor (K-NN) classification depends significantly on the searched neighborhoods of test samples, namely, the neighborhood size$K$and the used distance metric. For the two issues, many methods either to acquire the adaptive$K$or to learn a variant metric have been presented and yielded appropriate performance. However, most of the existing methods ignore the fact that these two factors can be jointly learned. Besides, nearly all the metric learning methods aim to shrink intra-class distance while expanding inter-class distance. In this way, embedding the learned metric directly into the K-NN does not efficiently improve its accuracy. To address these issues, we propose a joint K-NN algorithm with the help of evidence theory, optimizing the joint learning of adaptive$K$and distance matrix based on the feedback from error function. Ablation study demonstrates the performance improvement from the joint learning, and comparison experiments on real-world datasets show that our approach consumes competitive running time and achieves better performance than other state-of-the-art algorithms.
Chaoyu Gong, Yong Liu 0020, Pei-hong Wang, Yang You 0001
ICDE1
2022 Distributed evidential clustering toward time series with big data issue
Chaoyu Gong, Zhi-gang Su, Pei-hong Wang, Yang You 0001
Expert Syst. Appl.1
2022 Clustering based on adaptive local density with evidential assigning strategy
abstract
A new clustering algorithm, based on Adaptive Local Density (ALD) and Evidential K-Nearest Neighbors (EKNN), is proposed here. In density peaks clustering, many other density metrics fail to detect cluster centers on multi-density datasets, however the ALD deals with the tasks very well since it can better utilize the local information. To assign the remaining points after detecting the cluster centers, an assigning strategy in the framework of evidential theory, named EKNN, is created. The advantage of EKNN is twofold. Firstly, by fusing the information of K-Nearest Neighbors, it can reduce the risk of a phenomenon named domino effect: the drawback of one classical clustering, i.e., clustering by fast search and find of density peaks (always named as DPC). Secondly, it can detect border and noise points simultaneously since a credal partition is derived which can mine ambiguity and uncertainty of data structure. Simulations on both synthetic and real-world datasets demonstrate the outstanding performance of ALD-EKNN compared with DPC and some of its successors.
Chaoyu Gong, Pei-hong Wang
Intell. Data Anal.2
2021 Evidential instance selection for K-nearest neighbor classification of big data
Chaoyu Gong, Zhi-gang Su, Pei-hong Wang, Yang You 0001
Int. J. Approx. Reason.1
2021 An evidential clustering algorithm by finding belief-peaks and disjoint neighborhoods
Chaoyu Gong, Zhi-gang Su, Pei-hong Wang
Pattern Recognit.1
2020 Cumulative belief peaks evidential K-nearest neighbor clustering
Chaoyu Gong, Zhi-gang Su, Pei-hong Wang
Knowl. Based Syst.1
2020 An interactive nonparametric evidential regression algorithm with instance selection
Chaoyu Gong, Pei-hong Wang, Zhi-gang Su
Soft Comput.1