Zhenyu Kuang

dblp:269/7172 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Image and video processing · 100%
Artificial intelligence
1 paper
Generative modeling · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Image and video processing › image restoration
adverse weather image restoration
1.012026
AWM-Fuse: Multi-Modality Image Fusion for Adverse Weather via Global and Local Text Perception · IEEE Trans. Image Process. 2026
Image and video processing › image fusion
multi-modal image fusion
1.012026
AWM-Fuse: Multi-Modality Image Fusion for Adverse Weather via Global and Local Text Perception · IEEE Trans. Image Process. 2026
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.312026
AWM-Fuse: Multi-Modality Image Fusion for Adverse Weather via Global and Local Text Perception · IEEE Trans. Image Process. 2026

Methods — techniques the papers use, named apart from their topics

text perception · 2.0ChatGPT · 2.0BLIP captioning · 2.0
YearPublicationVenuePosition
2026 Balancer: Temporal knowledge graph embedding for novel events reasoning with contrastive learning
Zhenyu Kuang, Bo Qu, Xiang Li 0010, Cong Li 0009
Knowl. Based Syst.1
2026 AWM-Fuse: Multi-Modality Image Fusion for Adverse Weather via Global and Local Text Perception
abstract
Multi-modality image fusion (MMIF) in adverse weather aims to address the loss of visual information caused by weather-related degradations, providing clearer scene representations. Although a few studies have attempted to incorporate textual information to improve semantic perception, they often lack effective categorization and thorough analysis of textual content. To address these limitations, we propose AWM-Fuse, a unified fusion framework that handles diverse weather degradations via global and local text perception with shared parameters. In particular, a global text perception module leverages BLIP-generated captions to extract overall scene features and identify primary degradation types, thus promoting generalization across various adverse weather conditions. Complementing this, the local module employs detailed scene descriptions produced by ChatGPT to concentrate on specific degradation effects through concrete textual cues, enabling the recovery of subtle details. Furthermore, textual descriptions are used to constrain the generation of fused images, effectively steering the network learning process toward better alignment with semantic labels, thereby promoting the learning of more meaningful visual features. To facilitate text-guided fusion under adverse weather, we construct AWMM-Text, a large-scale benchmark providing paired global and local annotations for multi-modality image pairs. Extensive experiments demonstrate that AWM-Fuse consistently outperforms state-of-the-art methods under complex weather conditions and on multiple downstream tasks. Our code is available at https://github.com/Feecuin/AWM-Fuse.
Xilai Li, Huichun Liu, Xiaosong Li 0004, Tao Ye 0002, Zhenyu Kuang, Huafeng Li 0001
IEEE Trans. Image Process.5
2025 Generalizable Prompts Guided by Image-Redundant Separation for Vehicle Reidentification
abstract
Vehicle reidentification (reID) is a critical computer vision task with applications in video surveillance and autonomous vehicles. While significant progress has been made in recent years, domain generalization (DG) in reID remains a challenging and valuable research direction. Learning discriminative features that capture the intrinsic characteristics of vehicles, rather than domain-specific details, is paramount in addressing the domain shift problem, which encompasses disparities in data distribution, feature distribution, and label distribution. Recently, contrastive language image pretraining (CLIP) has attracted widespread attention because of its capacity to generalize knowledge across different domains or contexts. When fine-tuned for DG tasks, it can leverage this broad knowledge to perform well in domains or on tasks it has not specifically seen during training. The foremost work in this context is CLIP-reID, showcasing outstanding experimental performance on vehicle datasets through the integration of learnable prompts. However, the process of acquiring learnable prompts inevitably incorporates noisy text descriptions, such as background and camera style information, resulting in its limitations in DG tasks. To address this distinctive issue, we propose a CLIP-based Image-Redundant Separation (CIRS) framework to remove redundant domain-specific information and then implement visual-text alignment of CLIP. Specifically, we employ a classic variational autoencoder for image reconstruction, which can encourage the images generated by the vector quantized-variational autoencoder (VQ-VAE) network to contain features unrelated to vehicle IDs. Under the precise guidance of the image-redundant separation framework, a set of generalizable and learnable prompts for each vehicle can be effectively generated for reID. Extensive experimental results indicate that our method has achieved remarkable performance on several public datasets.
Zhenyu Kuang, Lidong Cheng, Yue Huang 0001, Xinghao Ding
IEEE Internet Things J.1
2025 PORSCHE: Progressive Optimization and Robust Spatial Convolution for Hybrid Enhancement in Visible-Infrared Vehicle Re-Identification
abstract
Visible-infrared vehicle re-identification has become crucial for stable 24-h surveillance of Visual Internet of Things (VIoT). It aims to leverage the shared information between different modalities to retrieve specific objects. Previous works primarily focus on projecting images from two modalities into a common embedding space to measure their similarity scores. However, the inherent distribution discrepancies between different modalities often lead models to rely on spurious features that are unrelated to vehicle identity, making effective feature fusion challenging. To address this unique problem, we propose the Progressive Optimization and Robust Spatial Convolution for Hybrid Enhancement (PORSCHE) model, which reduces the negative effects of spurious correlations and biases toward training pairs. Specifically, we introduce the Patch-wise Matching (PAM) module, which performs initial coarse-grained alignment between different modalities. Building upon this foundation, we develop the Point-wise Matching (POM) module to achieve fine-grained discriminative alignment through precise point-level feature matching, thereby enhancing identity-specific representation. To optimize these complementary PAM and POM components effectively, we implement a progressive training strategy that hierarchically refines feature representations from local patches to global structures, ensuring stable learning of modality-invariant characteristics. This coarse-to-fine architecture enables gradual fusion and alignment across modalities at both patch and point levels, effectively capturing the essential discriminative features required for robust cross-modality retrieval. Extensive experimental results on MSVR310, WMVeID863, and RGBN300 benchmarks demonstrate the effectiveness of our proposed method. The code will be available at https://github.com/HowardLiu28/PORSCHE.
Yinhao Liu, Zhenyu Kuang, Yige Ma, Xinghao Ding, Yue Huang 0001, Congbo Cai, Xiaosong Li 0004
IEEE Internet Things J.2
2025 PRADA: Prompt-guided Representation Alignment and Dynamic Adaption for time series forecasting
Yinhao Liu, Zhenyu Kuang, Xinghao Ding
Knowl. Based Syst.2
2024 AIVR-Net: Attribute-based invariant visual representation learning for vehicle re-identification
Zhenyu Kuang, Lidong Cheng, Yinhao Liu, Xinghao Ding, Yue Huang 0001
Knowl. Based Syst.2
2023 Radar HRRP Unseen Class Recognition Based on the Joint Dictionary Learning
abstract
Existing task settings and methods for radar high resolution range profile (HRRP) recognition are limited in addressing open challenges. To avoid labor-intensive data collection and model retraining, we formulate a new task called HRRP unseen class recognition, where the testing classes are unknown during training. To perform this task, we utilize metric learning to explore the potential information of unseen categories. Due to the target-aspect sensitivity problem of HRRP, feature extraction is a key step for recognition. Therefore, we take into account the time, frequency and aspect-angle characteristics of targets. Then a joint dictionary learning method is proposed to align different modalities in the latent common space to capture the intrinsic invariant representations of the unseen class targets. A variety of experiments demonstrate the effectiveness of our method.
Chuchu He, Zhenyu Kuang, Yijin Zhong, Xinghao Ding, Yue Huang 0001
ICIP2
2023 Boosting Generalization Performance in Person Re-identification
Lidong Cheng, Zhenyu Kuang, Xinghao Ding, Yue Huang 0001
PRCV (10)2
2023 NSGAIII based on utopian point improvements and its application in wastewater treatment process
Zhenyu Kuang, Zhongda Tian 0001, Shujiang Li
Expert Syst. Appl.1
2023 Joint Image and Feature Levels Disentanglement for Generalizable Vehicle Re-identification
abstract
Domain generalization (DG), which doesn’t require any data from target domains during training, is more challenging but practical than unsupervised domain adaptation (UDA). Since different vehicles of the same type have a similar appearance, neural networks always rely on a small amount of useful information to distinguish them, meaning that is more significant to remove ID-unrelated information for vehicle re-Identification (re-ID). Therefore, it is the key to eliminating the interference of a large amount of redundant information for the generalizable vehicle re-ID method. To address this unique challenge, we propose a novel disentanglement learning method that encourages variational autoencoder (VAE) network to reduce ID-unrelated features of vehicles by minimizing image reconstruction errors and providing sufficient representation to vehicle labels. To capture the intrinsic characteristics associated with the DG task, our core idea is to build the identity information streaming framework to separate ID-related and ID-unrelated information at the image and feature levels. In contrast with the general decoupling methods, our method leverages the decoupling of joint image and feature levels to extract more generalizable features. Furthermore, we present a brand-new vehicle dataset of truck types named “Optimus Prime (Opri)”, which includes multiple images of each truck captured by cameras at different high-speed toll gates. Experimental results on public datasets demonstrate that our method can achieve promising results and outperform several state-of-the-art approaches. Our codes and models are available at JIFD.
Zhenyu Kuang, Chuchu He, Yue Huang 0001, Xinghao Ding, Huafeng Li 0001
IEEE Trans. Intell. Transp. Syst.1
2022 Dual Domain Multi-Task Model for Vehicle Re-Identification
abstract
Vehicle re-identification (re-id) is an essential task in the field of intelligent transportation systems (ITS). The main goal of re-id is to find the same vehicle in different scenarios, which can is still a challenging task in both ITS and computer vision (CV). The existing vehicle re-identification methods simply combine the coarse-grained and the fine-grained attributes together with multi-task training. However, such combination may still have limited performance in vehicles with trivial appearance differences, or with rare models and colors. To solve this problem, we propose a simple yet effective framework, called dual domain multi-task model (DDM), that divides the vehicle images into two domains based on the frequency. And then two parallel branches are proposed to recover the two domains. Furthermore, a multi-task method is proposed, which combines the classification loss in color and model together with triplet loss for fine-grained distance measurement. Besides, a progressive strategy is used in the training process. Two public datasets, PKU VehicleID and VeRi are used to validate the proposed DDM. The experimental results demonstrate that the proposed approach outperforms the existing methods on both datasets.
Yue Huang 0001, Borong Liang, Weiping Xie, Yinghao Liao, Zhenyu Kuang, Yihong Zhuang, Xinghao Ding
IEEE Trans. Intell. Transp. Syst.5
2020 Structure alignment of attributes and visual features for cross-dataset person re-identification
Huafeng Li 0001, Zhenyu Kuang, Zhengtao Yu 0001, Jiebo Luo 0001
Pattern Recognit.2