Chunzhi Gu

dblp:238/0488 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
16since 2021 · last 2026
0000-0001-7280-337XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Community-level competitive influence payoff maximization
Jun Yu 0012, Chunzhi Gu, Takuya Akashi, Chao Zhang 0030
Knowl. Inf. Syst.3
2026 Few-shot human action anomaly detection via a unified contrastive learning framework
abstract
Human Action Anomaly Detection (HAAD) aims to identify anomalous actions using only normal-action data during training. Most prior work follows a one-model-per-category paradigm, requiring separate training for each action category and numerous normal samples. These requirements hinder scalability and limit applicability in real-world settings, where data are often scarce or novel categories frequently appear. To address these limitations, we propose a unified HAAD framework that supports few-shot settings. Our method learns a category-agnostic representation via contrastive learning and detects anomalies by comparing a test sample with a small support set of normal examples. To improve inter-category generalization and intra-category robustness, we introduce diffusion-based generative motion augmentation to synthesize diverse, realistic training samples. To the best of our knowledge, this is the first work to tailor diffusion-driven motion augmentation specifically to contrastive learning for action anomaly detection. Our few-shot design is particularly useful for monitoring applications, where normality often varies across individuals and contexts. Experiments on HumanAct12 demonstrate state-of-the-art performance on both seen and unseen categories, while improving training efficiency and scalability for few-shot HAAD.
Koichiro Kamide, Shunsuke Sakai, Shun Maeda, Chunzhi Gu, Chao Zhang 0030
Knowl. Based Syst.4
2026 Towards consistent sketch-guided local 3D shape editing
abstract
• We developed a diffusion-based generative framework called SEN, designed for sketch-guided local shape editing. This framework ensures consistency in geometric coherence and preservation. • We introduced an attention fusion module that effectively connects the cross-modal conditions to produce the final shape editing results. • We developed a sketch generator to directly generate sketch images from point clouds to facilitate our SEN training. Modeling complex 3D shapes from 2D sketch images has gained significant attention due to recent advancements in deep generative models. While many efforts have focused on producing geometries that faithfully reflect sketches, there has been limited exploration of sketch-guided local editing. Existing shape generation techniques often modify not only the intended edited areas but also the unedited parts, compromising geometric coherence and fairness among different parts of the shape. This paper introduces a shape editing network (SEN) under the diffusion-based generative framework to specifically enhance consistency in sketch-guided local shape editing. Our core insight is capturing partial geometric structure correlations to achieve consistent editing results. To accomplish this, the SEN is designed to conditionally learn local edits of a partial shape while considering the geometries of the unedited parts and the sketch image. We also devise an attention fusion module that progressively blends and aligns cross-modal conditioning streams between the latent variables of the sketch and point cloud representations. Additionally, by incorporating a guided sampling strategy for diffusion modeling, our SEN allows for more flexible control, even with complex or limited conditioning patterns. Our method significantly improves the geometric coherence of the synthesized parts while preserving the shapes of unedited areas. We demonstrate the advantages of our approach qualitatively and quantitatively across various shape categories, outperforming state-of-the-art methodologies.
Tomohiro Aizawa, Chunzhi Gu, Shigeru Kuriyama
Pattern Recognit.2
2026 Consensus-aware sparse subspace clustering via coreset selection
abstract
Subspace clustering aims to recover low-dimensional subspaces from high-dimensional data and assign each data point to its corresponding subspace. State-of-the-art subspace clustering methods rely on self-expressive models to represent each data point as a linear combination of other data points. However, these methods typically suffer from scalability challenges when dealing with large-scale datasets. In particular, the limited capacity to capture similarities among data points within the same subspace leads to poor connectivity of the affinity matrix, often resulting in over-segmentation. In this paper, we propose a scalable self-expressive model based on consensus learning with selectively sampled subsets to address the connectivity issue. Our core insight is that an ideal alignment between global and local solutions across multiple small subsets plays a key role in promoting dense connections in the affinity matrix for large-scale datasets. To this end, our model is designed to flexibly fuse sparse local solutions obtained from small subsets in a consensus-aware manner to derive a global solution that captures richer pairwise relationships within each subspace. Moreover, we introduce a selective subsampling strategy to generate subsets that effectively approximate the spatial support of the original dataset. This contributes to the reliability of solved local solutions, which further enables robust clustering even in imbalanced datasets. Extensive experiments on synthetic and five real-world datasets show that our method achieves state-of-the-art performance, in terms of clustering accuracy and connectivity.
Katsuya Hotta, Chunzhi Gu, Chao Zhang 0030
Pattern Recognit.2
2026 Frequency-guided multi-level human action anomaly detection with normalizing flows
abstract
We introduce the task of human action anomaly detection (HAAD), which aims to identify anomalous motions in an unsupervised manner given only the pre-determined normal category of training action samples. Compared to prior human-related anomaly detection tasks which primarily focus on unusual events from videos, HAAD involves the learning of specific action labels to recognize semantically anomalous human behaviors. To address this task, we propose a normalizing flow (NF)-based detection framework where the sample likelihood is effectively leveraged to indicate anomalies. As action anomalies often occur in some specific body parts, in addition to the full-body action feature learning, we incorporate extra encoding streams into our framework for finer modeling of body subsets. Our framework is thus multi-level to jointly discover global and local motion anomalies. Furthermore, to show awareness of the potentially jittery data during recording, we resort to discrete cosine transformation by converting the action samples from the temporal to the frequency domain to mitigate the issue of data instability. Extensive experimental results on two human action datasets demonstrate that our method outperforms the baselines formed by adapting state-of-the-art human activity AD approaches to our task of HAAD.
Shun Maeda, Chunzhi Gu, Jun Yu 0012, Shogo Tokai, Shangce Gao, Chao Zhang 0030
Pattern Recognit.2
2025 Incremental pseudo-labeling for black-box unsupervised domain adaptation
abstract
Black-Box unsupervised domain adaptation (BBUDA) learns knowledge only with the prediction of target data from the source model without access to the source data and source model, which attempts to alleviate concerns about the privacy and security of data. However, incorrect pseudo-labels are prevalent in the prediction generated by the source model due to the cross-domain discrepancy, which may substantially degrade the performance of the target model. To address this problem, we propose a novel approach that incrementally selects high-confidence pseudo-labels to improve the generalization ability of the target model. Specifically, we first generate pseudo-labels using a source model and train a crude target model by a vanilla BBUDA method. Second, we iteratively select high-confidence data from the low-confidence data pool by thresholding the softmax probabilities, prototype labels, and intra-class similarity. Then, we iteratively train a stronger target network based on the crude target model to correct the wrongly labeled samples to improve the accuracy of the pseudo-label. Experimental results demonstrate that the proposed method achieves state-of-the-art black-box unsupervised domain adaptation performance on three benchmark datasets.
Yawen Zou, Chunzhi Gu, Jun Yu 0012, Shangce Gao, Chao Zhang 0030
J. Vis. Commun. Image Represent.2
2025 Learning to Discriminate While Contrasting: Combating False Negative Pairs With Coupled Contrastive Learning for Incomplete Multi-View Clustering
abstract
The task of incomplete multi-view clustering (IMvC) aims to partition multi-view data with a lack of completeness into different clusters. The incompleteness can be typically categorized into the case of instance-missing and view-unaligned MvC. However, prior methods either consider each of them or struggle to pursue consistent latent representations among views. In this paper, we propose two forms of contrastive learning paradigms to jointly handle both cases for IMvC. Specifically, we design an instance-oriented contrastive (IOC) learning strategy to achieve intra-class consistency. As negative samples within different datasets can exhibit diverse distributions, we formulate a parameterized boundary for IOC learning to flexibly deal with such differing data modes. To preserve inter-view consistency, we further devise category-oriented contrastive (COC) learning such that data from different views can be seamlessly integrated into a combined semantic space. We also recover the missing instances with the learned latent representations in a reconstructing manner for realigning the incomplete multi-view data to facilitate clustering. Our approach unifies the solution to both incomplete cases into one formulation. To demonstrate the effectiveness of our model, we conduct four types of MvC tasks on six benchmark multi-view datasets and compare our method against state-of the-art IMvC methods. Extensive experiments show that our method achieves state-of-the-art performance, quantitatively and qualitatively.
Katsuya Hotta, Chunzhi Gu, Ao Li 0002, Jun Yu 0012, Chao Zhang 0030
IEEE Trans. Knowl. Data Eng.3
2025 Diverse Code Query Learning for Speech-Driven Facial Animation
abstract
Speech-driven facial animation aims to synthesize lip-synchronized 3D talking faces following the given speech signal. Prior methods to this task mostly focus on pursuing realism with deterministic systems, yet characterizing the potentially stochastic nature of facial motions has been to date rarely studied. While generative modeling approaches can easily handle the one-to-many mapping by repeatedly drawing samples, ensuring a diverse mode coverage of plausible facial motions on small-scale datasets remains challenging and less explored. In this article, we propose predicting multiple samples conditioned on the same audio signal and then explicitly encouraging sample diversity to address diverse facial animation synthesis. Our core insight is to guide our model to explore the expressive facial latent space with a diversity-promoting loss such that the desired latent codes for diversification can be ideally identified. To this end, building upon the rich facial prior learned with vector-quantized variational auto-encoding mechanism, our model temporally queries multiple stochastic codes which can be flexibly decoded into a diverse yet plausible set of speech-faithful facial motions. To further allow for control over different facial parts during generation, the proposed model is designed to predict different facial portions of interest in a sequential manner, and compose them to eventually form full-face motions. Our paradigm realizes both diverse and controllable facial animation synthesis in a unified formulation. We experimentally demonstrate that our method yields state-of-the-art performance both quantitatively and qualitatively, especially regarding sample diversity.
Chunzhi Gu, Shigeru Kuriyama, Katsuya Hotta
IEEE Trans. Vis. Comput. Graph.1
2024 Handling Class Imbalance in Black-Box Unsupervised Domain Adaptation with Synthetic Minority Over-Sampling
abstract
Black-box unsupervised domain adaptation (BBUDA) is a challenging task that transfers knowledge from the source domain to the target domain without access to the source data and source model, thus alleviating public concerns about data security. However, BBUDA requires the source model to function as a black-box predictor for the target data, and the pseudo-labels often exhibit class imbalance, which degrades the performance. To tackle this problem, we propose employing the synthetic minority oversampling technique (SMOTE) and adaptive sampling to rebalance data. Given that predictions often contain errors, we first select reliable high-confidence data before using SMOTE to generate synthetic samples for the minority class. Second, we incrementally select high-confidence data from the remaining low-confidence data with an adaptive sampling rate for each class, in which the minority class (with the fewest samples) is assigned a higher sampling rate and the majority class (with the most samples) is assigned a lower sampling rate. The experimental results demonstrate that our method can mitigate the class imbalance and further improve the performance of the target model.
Yawen Zou, Chunzhi Gu, Guang Li 0008, Jun Yu 0012, Chao Zhang 0030
VCIP2
2024 A multi-in and multi-out dendritic neuron model and its optimization
Jun Yu 0012, Chunzhi Gu, Shangce Gao, Chao Zhang 0030
Knowl. Based Syst.3
2024 Learning disentangled representations for controllable human motion prediction
Chunzhi Gu, Jun Yu 0012, Chao Zhang 0030
Pattern Recognit.1
2024 Orientation-aware leg movement learning for action-driven human motion prediction
Chunzhi Gu, Chao Zhang 0030, Shigeru Kuriyama
Pattern Recognit.1
2023 Teacher-student network for 3D point cloud anomaly detection with few normal samples
Jianjian Qin, Chunzhi Gu, Jun Yu 0012, Chao Zhang 0030
Expert Syst. Appl.2
2022 Learning to predict diverse human motions from a single image via mixture density networks
Chunzhi Gu, Yan Zhao 0038, Chao Zhang 0030
Knowl. Based Syst.1
2022 Example-based color transfer with Gaussian mixture modeling
Chunzhi Gu, Xuequan Lu, Chao Zhang 0030
Pattern Recognit.1
2021 Blur Removal Via Blurred-Noisy Image Pair
abstract
Complex blur such as the mixup of space-variant and space-invariant blur, which is hard to model mathematically, widely exists in real images. In this article, we propose a novel image deblurring method that does not need to estimate blur kernels. We utilize a pair of images that can be easily acquired in low-light situations: (1) a blurred image taken with low shutter speed and low ISO noise; and (2) a noisy image captured with high shutter speed and high ISO noise. Slicing the blurred image into patches, we extend the Gaussian mixture model (GMM) to model the underlying intensity distribution of each patch using the corresponding patches in the noisy image. We compute patch correspondences by analyzing the optical flow between the two images. The Expectation Maximization (EM) algorithm is utilized to estimate the parameters of GMM. To preserve sharp features, we add an additional bilateral term to the objective function in the M-step. We eventually add a detail layer to the deblurred image for refinement. Extensive experiments on both synthetic and real-world data demonstrate that our method outperforms state-of-the-art techniques, in terms of robustness, visual quality, and quantitative metrics.
Chunzhi Gu, Xuequan Lu, Ying He 0001, Chao Zhang 0030
IEEE Trans. Image Process.1