Zheng Wu 0004

dblp:20/144-4 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
16since 2021 · last 2026
0000-0001-5621-9878ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DBL: Dual-Level balanced learning for long-Tailed classification
Zheng Wu 0004, Kehua Guo, Bin Hu 0021, Xiangyuan Zhu, Rui Ding 0017
Pattern Recognit.1
2026 ReE3D: Boosting Novel View Synthesis for Monocular Images Using Residual Encoders
abstract
In recent years, novel view synthesis from a monocular image has become a research hot-spot that attracts significant attention. Some recent work identifies latent vectors for high-quality view generation via iterative optimisation, which is a time-consuming process. In contrast, some others utilise an encoder learning a mapping function to approximately estimate optimal latent codes, which significantly reduces its processing time but sacrifices reconstruction quality. Consequently, how to balance synthesis quality and its generation efficiency still remains challenging. In this paper, we propose a residual-based encoder to incorporate with a 3D Generative Adversarial Networks (GAN), named ReE3D, for novel view synthesis. It applies an iterative prediction of latent codes to ensure much higher quality of novel view synthesis with an insignificant increase of processing time when compared to existing encoder-based 3D GAN inversion methods. Additionally, we enforce a novel geometric loss constraint on the encoder to predict view-invariant latent codes, thus effectively mitigating the trade-off between geometric and texture quality in 3D GAN inversion. Extensive experimental results demonstrate that our extended encoder-based method has achieved best trade-off performance in terms of novel view synthesis quality and its execution time. Our method has gained comparable synthesis quality with exponentially decreased processing time when compared to iterative optimisation methods, while improved synthesis performance of encoder-based methods significantly.
Kehua Guo, Tianyu Chen 0004, Bin Hu 0021, Zheng Wu 0004, Shaojun Guo, Hui Fang 0003
IEEE Trans. Multim.5
2026 Multi-feature aggregation attention for efficient image super-resolution
Xiangyuan Zhu, Xuchong Liu, Zheng Wu 0004
Vis. Comput.4
2025 Deep Learning for Multiple Sclerosis on AI-Computing Networks: A Systematic Review
abstract
Recent advances in AI-computing networks (ACN) provide a timely backdrop for assessing deep-learning (DL) research in healthcare. This systematic review synthesizes DL applications in multiple sclerosis (MS) and evaluates their readiness for ACN-enabled deployment. A Web of Science Core Collection search (2014-2024) retrieved 438 records; 264 met stringent inclusion criteria. Bibliometric and knowledge-mapping analyses reveal steady growth in publications and citations, with MRIbased UNet variants dominating lesion-segmentation and diseaseclassification tasks. Emerging themes include gait-sensor analytics, longitudinal progression modelling, quantitative susceptibility mapping, and the growing use of transfer and federated learning to overcome data scarcity and privacy barriers. These resourceaware strategies signal a shift toward distributed training and inference paradigms that align naturally with ACN architectures. Nevertheless, few studies report multi-institutional experiments or network-level performance metrics, underscoring the need for tighter integration between DL methods and ACN infrastructure. We highlight research gaps-particularly in cross-site model orchestration and low-latency, on-device inference-that AIcomputing networks are well positioned to address, enabling scalable and interoperable DL services for MS diagnosis and prognosis.
Zheng Wu 0004, Xiangyuan Zhu, Rui Ding 0017, Kehua Guo
HPCC1
2025 Lightweight image super-resolution with tokenized dynamic embedding network
Xiangyuan Zhu, Xuchong Liu, Zheng Wu 0004
Knowl. Based Syst.3
2025 Backdoor Defense in Transportation Cyber-Physical Systems Using Frequency Domain Hybrid Distillation
abstract
In the context of transportation cyber-physical systems (T-CPS), backdoor attacks leveraging traffic images have emerged as a significant security threat. As T-CPS increasingly relies on visual information, such as real-time images captured by traffic cameras, for tasks like traffic sign recognition and autonomous driving, the risk of image-based backdoor attacks has grown substantially. Although various detection-based defense techniques have shown some success in identifying backdoored models, they often fail to fully eliminate backdoor effects, leaving residual security risks. To address this challenge, we propose a Frequency-Domain Hybrid Distillation (FDHD) method for backdoor defense, which effectively weakens the association between backdoor triggers and target labels by combining distillation mechanisms in both the frequency and pixel domains. Furthermore, we design a loss function that integrates feature reconstruction with adaptive alignment, enhancing the student network’s ability to mimic the teacher network and thereby bolstering the backdoor defense capability. Extensive experiments conducted by FDHD on multiple benchmark datasets against the five latest attacks demonstrate that our proposed defense method effectively reduces backdoor threats while maintaining high accuracy in predicting clean samples. This approach will protect against image-based backdoor attacks in T-CPS and lay the foundation for enhancing future traffic safety.
Bin Hu 0021, Kehua Guo, Zheng Wu 0004, Xianhong Wen, Xiaokang Zhou
IEEE Trans. Intell. Transp. Syst.3
2025 GAN Prior-Enhanced Novel View Synthesis From Monocular Degraded Images
abstract
With the escalating demand for three-dimensional visual applications such as gaming, virtual reality, and autonomous driving, novel view synthesis has become a critical area of research. Current methods mainly depend on multiple views of the same subject to achieve satisfactory results, but there is often a significant lack of available data. Typically, only a single degraded image is available for reconstruction, which may be affected by occlusion, low resolution, or absence of color information. To overcome this limitation, we propose a two-stage feature matching approach designed specifically for single degraded images, leading to the synthesis of high-quality novel perspective images. This method involves the sequential use of an encoder for feature extraction followed by the fine-tuning of a generator for feature matching. Additionally, the integration of an information filtering module proposed by us during the GAN inversion process helps eliminate misleading information present in degraded images, thereby correcting the inversion direction. Extensive experimental results show that our method outperforms existing state-of-the-art single-view novel view synthesis techniques in handling challenges like occluded, grayscale, and low-resolution images. Moreover, the efficacy of our method remains unparalleled even when aforementioned method integrated with image restoration algorithms.
Kehua Guo, Zheng Wu 0004, Xianhong Wen, Shaojun Guo, Tianyu Chen 0004
IEEE Trans. Multim.2
2024 Modal adaptive super-resolution for medical images via continual learning
Zheng Wu 0004, Feihong Zhu, Kehua Guo, Chao Liu 0058, Hui Fang 0003
Signal Process.1
2024 GRTR: Gradient Rebalanced Traffic Sign Recognition for Autonomous Vehicles
abstract
Traffic sign recognition is a crucial aspect of autonomous vehicle research, and deep learning techniques have significantly contributed to its progress. Nevertheless, the distribution of traffic sign information in natural complex road conditions is long-tailed, and traffic sign identification in complex road conditions has become a significant barrier to autonomous vehicle applications. The imbalanced distribution of information on the dataset migrates to the feature space during training, resulting in imbalanced classifier prediction. In this paper, we propose the gradient rebalanced traffic sign recognition (GRTR) method to address this problem for the first time. GRTR first evaluates the prediction and classification bias of the classifier using the fitted deviation between the model’s output probability and the ground-truth distributions. Then, GRTR dynamically adjusts the correction and compensation factors following the classifier’s prediction and classification biases. GRTR rebalances the positive and negative sample gradients for each category based on the synergistic effect of the correction and compensation factors to prevent the transfer of distribution imbalance and to significantly enhance the performance of the traffic sign classifier under difficult road conditions. Experimental results demonstrate that our GRTR achieves state-of-the-art performance on long-tailed traffic sign and multilabel datasets.Note to Practitioners—Most traffic sign recognition algorithms are still designed based on the assumption of a balanced distribution of traffic signs in the dataset. Real-world autonomous vehicles require traffic sign recognition on datasets with severely imbalanced distributions. This paper proposes a general approach to solving the long-tailed traffic sign recognition problem.
Kehua Guo, Zheng Wu 0004, Weizheng Wang 0001, Xiaokang Zhou, G. Thippa Reddy, Chao Liu 0058
IEEE Trans Autom. Sci. Eng.2
2024 Progressive Diversity Generation for Single Domain Generalization
abstract
Single domain generalization (single-DG) is a realistic yet challenging domain generalization scenario where a model trained on a single domain generalization scenario where a model trained on a single domain generalizes well to multiple unseen domains. Unlike typical single-DG methods that are essentially supervised data augmentation and focus mainly on the novelty of images, we propose a simple adversarial augmentation method, termed Progressive Diversity Generation (PDG), to synthesize novel and diverse images in a fully unsupervised manner. Specifically, PDG minimizes the uncertainty coefficient to ensure that synthesized images are novel. By modeling conditional probabilities with an auxiliary network, we transfer the adversarial process from semantics to images, thus eliminating dependency on labels. To enhance diversity, we propose the$f$-diversity, a collection of correlation or similarity measures, to allow our model to generate potential images from diverse perspectives. The proposed architecture combines a multi-attribute generator with a progressive generation framework to improve model performance. PDG is the unsupervised and easy-to-implement method that solves single-DG with only synthesized (source) images. Extensive experiments on multiple single-DG benchmarks show that PDG achieves remarkable results and outperforms existing supervised and unsupervised methods by a large margin in single domain generalization. Source code and data are available:https://github.com/Ruiding1/PDG.
Rui Ding 0017, Kehua Guo, Xiangyuan Zhu, Zheng Wu 0004, Hui Fang 0003
IEEE Trans. Multim.4
2023 Gradient-Based Graph Attention for Scene Text Image Super-resolution
abstract
Scene text image super-resolution (STISR) in the wild has been shown to be beneficial to support improved vision-based text recognition from low-resolution imagery. An intuitive way to enhance STISR performance is to explore the well-structured and repetitive layout characteristics of text and exploit these as prior knowledge to guide model convergence. In this paper, we propose a novel gradient-based graph attention method to embed patch-wise text layout contexts into image feature representations for high-resolution text image reconstruction in an implicit and elegant manner. We introduce a non-local group-wise attention module to extract text features which are then enhanced by a cascaded channel attention module and a novel gradient-based graph attention module in order to obtain more effective representations by exploring correlations of regional and local patch-wise text layout properties. Extensive experiments on the benchmark TextZoom dataset convincingly demonstrate that our method supports excellent text recognition and outperforms the current state-of-the-art in STISR. The source code is available at https://github.com/xyzhu1/TSAN.
Xiangyuan Zhu, Kehua Guo, Hui Fang 0003, Rui Ding 0017, Zheng Wu 0004, Gerald Schaefer
AAAI5
2023 Single Domain Generalization via Unsupervised Diversity Probe
abstract
Single domain generalization (SDG) is a realistic yet challenging domain generalization scenario that aims to generalize a model trained on a single domain to multiple unseen domains. Typical SDG methods are essentially supervised data augmentation strategies, which tend to enhance the novelty rather than the diversity of augmented samples. Insufficient diversity may jeopardize the model generalization ability. In this paper, we propose a novel adversarial method, termed Unsupervised Diversity Probe (UDP), to synthesize novel and diverse samples in fully unsupervised settings. More specifically, to ensure that samples are novel, we study SDG from an information-theoretic perspective that minimizes the uncertainty coefficients between synthesized and source samples. Considering that the variation in a single source domain is limited, we introduce a regularization imposed on the auxiliary module that synthesizes variable samples, incorporated with uncertainty coefficients in an adversarial manner to complement the diversity. Subsequently, an available region is utilized to guarantee the samples' safety. For the network architecture, we design a simple probe module that can synthesize samples in several different aspects. UDP is an unsupervised and easy-to-implement method that solves SDG using only synthetic (source) samples, thus reducing the dependence on task models. Extensive experiments on three benchmark datasets show that UDP achieves remarkable results and outperforms existing supervised and unsupervised methods by a large margin in single domain generalization.
Kehua Guo, Rui Ding 0017, Tian Qiu 0002, Xiangyuan Zhu, Zheng Wu 0004, Hui Fang 0003
ACM Multimedia5
2023 Stereoscopic image super-resolution with interactive memory learning
Xiangyuan Zhu, Kehua Guo, Tian Qiu 0002, Hui Fang 0003, Zheng Wu 0004, Xuyang Tan, Chao Liu 0058
Expert Syst. Appl.5
2023 NC2E: boosting few-shot learning with novel class center estimation
Zheng Wu 0004, Changchun Shen, Kehua Guo
Neural Comput. Appl.1
2022 ComGAN: Unsupervised Disentanglement and Segmentation via Image Composition
abstract
We propose ComGAN, a simple unsupervised generative model, which simultaneously generates realistic images and high semantic masks under an adversarial loss and a binary regularization. In this paper, we first investigate two kinds of trivial solutions in the compositional generation process, and demonstrate their source is vanishing gradients on the mask. Then, we solve trivial solutions from the perspective of architecture. Furthermore, we redesign two fully unsupervised modules based on ComGAN (DS-ComGAN), where the disentanglement module associates the foreground, background and mask with three independent variables, and the segmentation module learns object segmentation. Experimental results show that (i) ComGAN's network architecture effectively avoids trivial solutions without any supervised information and regularization; (ii) DS-ComGAN achieves remarkable results and outperforms existing semi-supervised and weakly supervised methods by a large margin in both the image disentanglement and unsupervised segmentation tasks. It implies that the redesign of ComGAN is a possible direction for future unsupervised work.
Rui Ding 0017, Kehua Guo, Xiangyuan Zhu, Zheng Wu 0004
NeurIPS4
2021 BSSPD: A Blockchain-Based Security Sharing Scheme for Personal Data with Fine-Grained Access Control
abstract
Privacy protection and open sharing are the core of data governance in the AI‐driven era. A common data‐sharing management platform is indispensable in the existing data‐sharing solutions, and users upload their data to the cloud server for storage and dissemination. However, from the moment users upload the data to the server, they will lose absolute ownership of their data, and security and privacy will become a critical issue. Although data encryption and access control are considered up‐and‐coming technologies in protecting personal data security on the cloud server, they alleviate this problem to a certain extent. However, it still depends too much on a third‐party organization’s credibility, the Cloud Service Provider (CSP). In this paper, we combined blockchain, ciphertext‐policy attribute‐based encryption (CP‐ABE), and InterPlanetary File System (IPFS) to address this problem to propose a blockchain‐based security sharing scheme for personal data named BSSPD. In this user‐centric scheme, the data owner encrypts the sharing data and stores it on IPFS, which maximizes the scheme’s decentralization. The address and the decryption key of the shared data will be encrypted with CP‐ABE according to the specific access policy, and the data owner uses blockchain to publish his data‐related information and distribute keys for data users. Only the data user whose attributes meet the access policy can download and decrypt the data. The data owner has fine‐grained access control over his data, and BSSPD supports an attribute‐level revocation of a specific data user without affecting others. To further protect the data user’s privacy, the ciphertext keyword search is used when retrieving data. We analyzed the security of the BBSPD and simulated our scheme on the EOS blockchain, which proved that our scheme is feasible. Meanwhile, we provided a thorough analysis of the storage and computing overhead, which proved that BSSPD has a good performance.
Hongmin Gao 0002, Zhaofeng Ma, Shoushan Luo, Zheng Wu 0004
Wirel. Commun. Mob. Comput.5