Huanqiang Zeng

dblp:25/8798 · DBLP profile ↗
← Back
151ranked-venue papers
13as first author
103since 2021 · last 2027
0000-0002-2802-7745ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 97 · 12 first-author · 57 since 2021Artificial intelligence and machine learning · 43 · 1 first-author · 40 since 2021Computer networks · 10 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Systems, architecture and hardware · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2027 Cross-modal attention fusion guided semantic unification for image-text matching
Jinhui Hou, Huanqiang Zeng, C. L. Philip Chen
Expert Syst. Appl.4
2026 Enhancing Few-Shot marble slab surface defect detection: A diffusion framework with knowledge distillation and semantic guidance
Longtao Chen, Jinjie Zheng, Fenglei Xu, Fa Zhu, Ajith Abraham, Huanqiang Zeng
Eng. Appl. Artif. Intell.6
2026 Industrial visual defect detection oriented context-modulated and cross-layer interaction network for image super-resolution
Xiancheng Zhu, Detian Huang, Jianqing Zhu, Huanqiang Zeng
Expert Syst. Appl.5
2026 Text-to-video person re-identification benchmark: Dataset and dual-modal contextual alignment
Jiajun Su, Simin Zhan, Pudu Liu, Jianqing Zhu, Huanqiang Zeng
Neurocomputing5
2026 Gradient degradation-aware rate control for VVC using Nash equilibrium
Chenpeng Lu, Huanqiang Zeng, Chao Jiao, Jing Chen 0001, Huijie Zheng
J. Vis. Commun. Image Represent.2
2026 Accelerating inter-frame prediction in Versatile Video Coding via deep learning-based mode selection
Jing Chen 0001, Huanqiang Zeng, Wenjie Xiang, Yuting Zuo
J. Vis. Commun. Image Represent.3
2026 Lightweight image super-resolution network with adaptive token selection and feature enhancement
Detian Huang, Mingxin Lin, Xinwei Gan, Luanyuan Dai, Huanqiang Zeng
Knowl. Based Syst.5
2026 UDD: Unsupervised denoising diffusion for noisy multi-focus image fusion
Pudu Liu, Jiajun Su, Simin Zhan, Jianqing Zhu, Huanqiang Zeng
Neural Networks6
2026 Focusing on pedestrians like human for clothes changing person re-identification
abstract
• Based on the ensemble coding hypothesis in cognitive neuroscience, we achieve data augmentation by simulating human focus capability. • Our method is the first local detail learning data augmentation designed specifically for clothes changing person re-identification. • Our method achieves SOTA on three public datasets and outperforms various traditional data augmentation methods. Current approaches focus mainly on the design of networks to learn key identity features from local body components for clothes-changing person re-identification (CC-ReID). In this paper, we propose a humanoid focus-inspired image augmentation (HFIA) method, which is intuitive image processing rather than a sophisticated network architecture designed to enhance local nuances of pedestrian images. Based on pedestrian silhouettes, we roughly divide a pedestrian image into five body components, that is, head-shoulder, upper left torso, upper right torso, lower left torso, and lower right torso. The HFIA has two key designs to deal with these components: the central emphasis strategy (CES) and the component continuity processing (CCP). For each component, leveraging the natural tendency of human visual attention towards central regions, the CES constructs an enlargement grid, where the closer the center, the greater the enlargement. To maintain the continuity of assembly, the CCP performs an overall alignment of component centers, that is, all components share the same normalized vertical coordinate and the left and right torsos have mirrored horizontal coordinates. Furthermore, the CCP implements a smoothing post-processing to uniformly erase the discontinuity between the head-shoulder, upper left torso, and upper right torso. Experiments show the state-of-the-art performance of HFIA.
Wenjie Pan, Jianqing Zhu, Xiaolin Cui, Huanqiang Zeng, Yibing Zhan
Neural Networks4
2026 Hybrid-stage association with dynamicity adaptation and enhanced cues for multi-object tracking and segmentation
Longtao Chen, Guoxing Liao, Yifan Shi 0001, Jing Lou, Fenglei Xu, Huanqiang Zeng
Pattern Recognit.6
2026 NGSCD: Nighttime glow suppression for visibility enhancement via structural cascade decomposition model
Jingrong Guo, Huanqiang Zeng
Pattern Recognit.4
2026 Progressive local self-attention for content-aligned super-resolution
Detian Huang, Xiancheng Zhu, Fei Shen 0004, Taotao Lai, Huanqiang Zeng, Junhui Hou
Pattern Recognit.6
2026 Temporal-invariant video contrastive learning: A novel perspective from brain knowledge
Wei Lin 0021, Changxing Jing, Xinghao Ding, Huanqiang Zeng
Pattern Recognit.4
2026 Exploring non-local spatial-angular correlations with a hybrid Mamba-Transformer framework for light field super-resolution
Haosong Liu, Xiancheng Zhu, Huanqiang Zeng, Jianqing Zhu, Jiuwen Cao, Junhui Hou
Pattern Recognit.3
2026 Point cloud compression using graph neural networks
Xiangjie Zhang, Huanqiang Zeng, Xin-Rong Gong, Junhui Hou, Jianqing Zhu
Pattern Recognit.2
2026 Structure-Aware Alignment for Day-Night Cross-Domain Vehicle Re-Identification
abstract
Day-night cross-domain vehicle re-identification (DN-ReID) is fundamentally challenged by drastic illumination changes that create substantial domain gaps and hinder consistent feature representation. Most existing methods focus on aligning distribution statistics but often overlook essential structural relationships, such as cross-domain clustering and geometric topology. To address this, we propose a structure-aware alignment (SAA) method that, for the first time, formulates centered kernel alignment as a trainable loss for structural alignment between domains. This approach explicitly aligns cross-domain relational structures, thereby supporting the transfer of intra-class compactness and inter-class separability from the source to the target domain for robust cross-domain matching. Extensive experiments on DN-348 and DN-Wild demonstrate that our approach consistently outperforms state-of-the-art methods, achieving a 2.7% mAP improvement on DN-348.
Jingyi Zhuang, Baihui Sa, Jinjie Zheng, Liu Liu 0014, Jianqing Zhu, Huanqiang Zeng
IEEE Signal Process. Lett.6
2026 Dual-Domain Deep Learning-Assisted NOMA-CSK Systems for Secure and Efficient Vehicular Communications
abstract
Ensuring secure and efficient multi-user (MU) transmission is critical for vehicular communication systems. Chaos-based modulation schemes have garnered considerable interest owing to their inherent advantages in physical layer security. However, most existing MU chaotic communication systems, particularly those based on non-coherent detection, suffer from low spectral efficiency due to reference signal overhead and limited user connectivity under orthogonal multiple access (OMA). Although non-orthogonal schemes such as sparse code multiple access (SCMA)-based differential chaos shift keying (DCSK) have been explored, they incur high computational complexity and exhibit inflexible scalability owing to fixed codebook designs. This paper proposes a deep learning-assisted power-domain non-orthogonal multiple access chaos shift keying (DL-NOMA-CSK) system for vehicular communications. A deep neural network (DNN)-based demodulator is designed to learn the intrinsic characteristics of chaotic signals during offline training, thereby eliminating the need for chaotic synchronization or reference signal transmission. The demodulator employs a dual-domain feature extraction architecture that jointly processes time-domain and frequency-domain information of the received chaotic signals, enhancing feature learning under dynamic channel conditions. The DNN is integrated into a successive interference cancellation (SIC) framework to mitigate error propagation. Theoretical analysis and extensive simulations demonstrate that the proposed system achieves superior performance in terms of spectral efficiency (SE), energy efficiency (EE), bit error rate (BER), security, and robustness compared to conventional MU-DCSK and existing deep learning-aided schemes. These advantages confirm the practical viability of the proposed system for secure vehicular communications.
Jundong Chen 0004, Huanqiang Zeng, Guofa Cai, Georges Kaddoum
IEEE Trans. Commun.3
2026 Collaborative Refinement Guidance for High-Fidelity Defect Image Synthesis
abstract
Automated Visual Inspection is a cornerstone of modern manufacturing, yet the development of robust deep learning models is frequently impeded by the scarcity and imbalance of training data. This challenge is particularly acute for industrial defects, which often manifest as subtle anomalies intrinsically linked to complex, structured backgrounds. To address this challenge, CRG-DefectDiffuser is proposed as a collaborative generative framework for high-fidelity defect synthesis. At its core lies the Collaborative Refinement Guidance (CRG) mechanism, which orchestrates two specialized diffusion models: one trained on abundant defect-free images to master background context, and another trained on scarce defect patches to encode fine-grained defect semantics. The CRG mechanism steers the synthesis by dynamically generating a guidance map, which is refined through a four-stage process to ensure that morphologically accurate defects are seamlessly integrated into the appropriate background context. Augmenting training data with our method boosts the defect detection mAP@50-95 from a baseline of 0.496 to 0.557, corresponding to a 12.3 percentage point relative improvement. The framework also demonstrates superior scalability, with performance gains continuing up to the three-fold data augmentation evaluated in our experiments, a point where competing methods often falter. These results establish CRG-DefectDiffuser as an effective and practical solution to data scarcity in industrial visual inspection, with strong potential for generalization across diverse manufacturing scenarios. The source code is publicly available at https://github.com/weihang-luo/ CRG-DefectDiffuser.
Weihang Luo, Xingen Gao, Jianxiongwen Huang, Hongyi Zhang 0003, Juqiang Lin, Huanqiang Zeng
IEEE Trans. Circuits Syst. Video Technol.8
2026 Selection, Aggregation, and Enhancement: Trajectory Consistent Diffusion Model for Image Super-Resolution
abstract
Diffusion models have shown strong promise for image super-resolution (ISR). However, current approaches often underuse pretrained diffusion backbones and lack constraints on the sampling trajectory, which degrades structural consistency and fine details. For that, we introduce the trajectory consistent diffusion model (TCDM) for super-resolution, which jointly optimizes the sampling process through lightweight components and inference-time strategies while keeping the diffusion backbone frozen, yielding high-fidelity, detail-rich reconstructions. First, we propose a dynamic semantic selection (DSS) mechanism that records early intermediates, matches them to upsampled low-resolution features, and reconditions sampling with the best match to reduce the mismatch between conditioning and noise scale. Next, we design a cross-step aggregation guidance (CAG) strategy that aggregates features from the current state with the selected intermediate to enforce trajectory-level consistency in noise prediction. Finally, we present a plug-and-play frequency enhancement adapter (FE-Adapter) that injects different frequency-domain cues into the encoder during training, strengthening high-frequency perception while preserving global structures. Extensive experiments on multiple ISR benchmarks show that TCDM achieves strong structural fidelity and competitive no-reference perceptual quality, offering a favorable fidelity-perception trade-off.
Detian Huang, Yaohui Guo, Luanyuan Dai, Fei Shen 0004, Huanqiang Zeng
IEEE Trans. Image Process.6
2026 Exploring Hierarchical Cross-Modal Correlation Consistency for Partial Mismatching
abstract
Cross-modal retrieval facilitates more flexible information access and improves semantic understanding across different modalities. However, traditional cross-modal retrieval models rely on well-aligned datasets, which are often labor-intensive and costly to obtain. In real-world applications, data inevitably includes mismatched pairs, and these semantically inconsistent pairs can significantly degrade retrieval performance. Previous approaches have assumed ideal loss value distributions to optimize models for accurate semantic matching through soft-label estimation. However, the absence of hierarchical semantic correlation learning limits the effectiveness of these models in scenarios involving partial mismatches. To address these challenges, we propose Exploring Hierarchical Cross-Modal Correlation Consistency (EH3C) for cross-modal retrieval under partially mismatched conditions. Specifically, our approach first leverages neighborhood correlation distributions among samples to optimize cross-modal alignment, without assuming ideal distributions. This allows for the measurement of soft matching degrees between cross-modal data pairs and facilitates the effective learning of their positive correlations. Next, we enhance inter-class separability through intra-modal correlation learning by exploiting negative correlations between reliable negative sample pairs, thus enabling a more comprehensive exploration of cross-modal correlations. Finally, to assess the effectiveness and robustness of our approach, we conducted extensive experiments on three benchmark datasets. The results demonstrate that the proposed EH3C significantly improves cross-modal retrieval performance in scenarios involving partial mismatches.
Zhiwen Yu 0002, Jun Yu 0002, Huanqiang Zeng, Zhuoyao Wang 0001, C. L. Philip Chen
IEEE Trans. Image Process.4
2026 Adaptive Compressed-Based Privacy-Preserving Large Language Model for Sensitive Healthcare
abstract
The emergence of large language models (LLMs) has been a key enabler of technological innovation in healthcare. People can conveniently obtain a more accurate medical consultation service by utilizing LLMs' powerful knowledge inference capability. However, existing LLMs require users to upload explicit requests during remote healthcare consultations, which involves the risk of exposing personal privacy. Furthermore, the reliability of the response content generated by LLMs is not guaranteed. To tackle the above challenges, this paper proposes a novel privacy-preserving LLM for user-activated health, called Adaptive Compressed-based Privacy-preserving LLM (ACP2LLM). Specifically, an adaptive token compression method based on information entropy is carefully designed to ensure that ACP2LLM can preserve user-sensitive information when invoking the medical consultation of LLMs deployed on the cloud platform. Moreover, a multi-doctor one-chief physician mechanism is proposed to rationally split and collaboratively infer the patients' requests to achieve the privacy-utility trade-off. Notably, the proposed ACP2LLM also provides highly competitive performance in various token compression rates. Extensive experiments on multiple Medical Question and Answers datasets demonstrate that the proposed ACP2LLM has strong privacy protection capabilities and high answer precision, outperforming current state-of-the-art LLM methods.
Xin-Rong Gong, Jiaran Gao, Yifan Shi 0001, Huanqiang Zeng, Kaixiang Yang 0001
IEEE J. Biomed. Health Informatics6
2026 Self-Paced Attribute Prototype Contrastive Learning for Cross-Modal Materials Perception
abstract
The contactless interaction approach enables robots to identify object attributes via sensors and advanced perceptual technologies, eliminating the need for physical contact. This approach enhances interaction flexibility and safety while offering a richer, more intuitive human-robot experience. Current research focuses on cross-modal retrieval through non-contact feature extraction, enabling material attribute estimation and advancing object perception. However, existing cross-modal material retrieval methods overlook learning common attributes shared within the same material type. To address this, we propose Materials Perception with Self-paced Prototype contrastive Learning (MPSPL). First, attribute prototype contrastive learning extracts shared characteristics within the same material type, aggregates similar samples, and enhances the perception of deep model material attributes. Second, a self-paced learning strategy dynamically adjusts the contrastive loss temperature coefficient, guiding the model to discern discriminative features across materials and progressively recognize inherent attributes. Finally, extensive experiments validate our method under both known and unknown material category conditions. Experiments demonstrate the superiority of our proposed method over state-of-the-art approaches.
Zhiwen Yu 0002, Kaixiang Yang 0001, Huanqiang Zeng, Jun Jiang 0003, C. L. Philip Chen
IEEE Trans. Multim.4
2026 Image Super-Resolution Using Hierarchical Cross-Scale Self-Similarity
abstract
Previous studies have revealed that extending the spatial range of informative pixels offers positive performance gains for image Super-Resolution (SR). To activate more informative pixels, considerable efforts have been devoted to exploring various variants of non-local attention mechanisms for capturing image self-similarity. However, even the state-of-the-art non-local attention mechanisms ignore an inherent property of images, namely hierarchical cross-scale self-similarity. In this paper, we propose the first Hierarchical Cross-Scale Attention (HCSA). Specifically, we first extend the search space to multiple feature maps from a single feature map, and then model cross-scale feature correspondences among different layers. This allows HCSA to activate more informative pixels for image SR by adaptively rescaling and aggregating input pixels and large-scale patches within different feature maps. To ensure accurate cross scale feature matching, we propose to replace plain down sampling operations (e.g., interpolation, pooling) with Haar Wavelet Transform (HWT) encoding, which transfers spatial information of feature maps into the channel dimension, effectively avoiding important information loss. Considering that softmax normal ization in the standard non-local attention often leads to homogeneous feature aggregation due to the amplification of small similarity weights, we propose a simple yet effective Adaptive Selection (AS) operator. This operator generates a learnable sparse mask to remove redundant features, enabling HCSA to perform discriminative feature aggregation. As a generic building block, the proposed HCSA can be flexibly integrated into existing CNN- or Transformer-based SR models, significantly strengthening cross-layer information interaction and cross-scale feature representation. Quantitative and qualitative results demonstrate that our HCSA facilitates existing SR models to achieve superior accuracy and visual quality.
Xiancheng Zhu, Detian Huang, Taiheng Zeng, Xiaoqian Huang, Zhenzhen Hu 0004, Huanqiang Zeng
IEEE Trans. Multim.6
2026 Uncertainty-driven Progressive Single Image De-raining
abstract
Over the past years, progressive methods have demonstrated promising performance in single image de-raining task. Nonetheless, current methods still struggle to precisely remove rain and preserve more image details during the progressive de-raining process, resulting in undesirable local artifacts or image detail loss. To tackle these limitations, a novel progressive approach, called Uncertainty-driven Progressive Single Image De-raining (UPSID), is proposed. Firstly, a powerful internal-and-external dense sub-network is devised, which effectively integrates three proper and flexible components, including dense connection, long short-term memory, and channel attention. Subsequently, the sub-network is further unfolded into multiple recurrent stages to form a progressive de-raining network. Finally, the overall progressive de-raining network is trained with an adaptive weighted loss to focus more on challenging pixels that characterize rain or texture/edge regions. Extensive quantitative and qualitative experiments confirm that the proposed UPSID outperforms multiple state-of-the-art algorithms, including single-stage, progressive, and uncertainty-driven single image de-raining methods. Additionally, this article also demonstrates the superiority of UPSID for other similar image restoration tasks such as single image de-snowing. The code will be publicly available at https://github.com/Lcai-QZ/UPSID .
Jianqing Zhu, Huanqiang Zeng, Tao Zhu 0002, Jing Chen 0001, Wenkang Su 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2025 Fair Training with Zero Inputs
abstract
There are two manifestations of classification fairness. One is the preference for head classes with more instances due to the long-tail (LT) distribution of training data. The other is the clever Hans (CH) effect, where non-discriminative features are mistakenly used for classification. In this paper, we find that using category-agnostic zero-valued data can simultaneously reveal both types of unfairness. Based on this, we propose a zero uniformity training (ZUT) framework to optimize classification fairness. The ZUT framework inputs category-agnostic zero-valued data into the model in parallel and uses zero uniformity loss (ZUL) to optimize classification fairness. The ZUL loss mitigates bias towards specific classes by unifying the classification features corresponding to zero-valued data. The ZUT framework is compatible with various classification-based tasks. Experiments show that the ZUT framework can improve the performance of multiple state-of-the-art methods in image classification, person re-identification, and semantic segmentation.
Wenjie Pan, Jianqing Zhu, Huanqiang Zeng
AAAI3
2025 MSD: Mask-Guided and Semantic-Guided Diffusion-Based Framework for Stone Surface Defect Detection
Longtao Chen, Jinjie Zheng, Fenglei Xu, Jing Lou, Huanqiang Zeng
CVM (1)5
2025 Bridging Clustering and Training Gap in Unsupervised Vessel Re-identification
abstract
Unsupervised vessel re-identification is challenged by a gap between clustering and training stages, where clustering relies on features from original images while training utilizes features from enhanced images. To address this, we propose a novel consistent contrastive learning (CCL) approach that innovatively applies earth mover’s distance (EMD) as a consistency loss function. EMD effectively aligning both global and local structures of original and enhanced features, thus bridging the gap between clustering and training stages. Unlike traditional metrics such as Euclidean distance (ED) and KL divergence, which primarily focus on pointwise or global similarity, EMD considers both overall distribution and local feature importance, enabling more robust feature alignment. Moreover, compared to the structural similarity index metric (SSIM), EMD dynamically adjusts local feature contributions using weight information, ensuring precise matching even in complex scenarios. Thanks to these advantages, our method demonstrates a 4.0% improvement in mean average precision over the original VesselReID benchmark, highlighting its effectiveness in addressing the challenges of unsupervised vessel re-identification.
Jianqing Zhu, Linhan Huang, Huanqiang Zeng
IJCNN5
2025 Dynamicity Adaptation for Multi-object Tracking and Segmentation: Toward Improved Association Correction
abstract
Dynamicity is a critical and highly challenging aspect in Multi-Object Tracking and Segmentation (MOTS), significantly impeding the effective integration of diverse association cues. High dynamicity, such as severe occlusion or deformation, can distort appearance cues, leading to inaccurate inter-object relationships and misleading results. Conversely, in low dynamicity states, spatiotemporal consistency of appearance cues aids in recovering object states. To address this issue, we propose a straightforward, effective, and versatile Dynamicity Adaptation for Multi-object Tracking and Segmentation, named DA-Track. First, we leverage the sensitivity of appearance cues to dynamicity through pre-association, capturing dynamic behavior in objects. Second, Dynamicity Adaptation incorporates Dynamicity Selection to identify reliable appearance cues based on pre-association results and Occlusion Dynamicity Fusing to adaptively integrate appearance and motion cues by analyzing historical mask variations. Experiments on MOTS20 and KITTI MOTS datasets demonstrate DA-Track’s robust and reliable performance across diverse scenarios.
Longtao Chen, Guoxing Liao, Jing Lou, Fenglei Xu, Bingwen Hu, Lineng Chen, Huanqiang Zeng
IROS7
2025 EPDiff: Enhancing Prior-guided Diffusion model for Real-world Image Super-Resolution
abstract
Diffusion Models (DMs) have achieved promising success in Real-world Image Super-Resolution (Real-ISR), where they reconstruct High-Resolution (HR) images from available Low-Resolution (LR) counterparts with unknown degradation by leveraging pre-trained Text-to-Image (T2I) diffusion models. However, due to the randomness nature of DMs and the severe degradation commonly presented in LR images, most DMs-based Real-ISR methods neglect the structure-level and semantic information, which results in reconstructed HR images suffering not only from important edge missing, but also from undesired regional information confusion. To tackle these challenges, we propose an Enhancing Prior-guided Diffusion model (EPDiff) for Real-ISR, which leverages high-frequency priors and semantic guidance to generate reconstructed images with realistic details. Firstly, we design a Guide Adapter (GA) module that extracts latent texture and edge features from LR images to provide high-frequency priors. Subsequently, we introduce a Semantic Prompt Extractor (SPE) that generates high-quality semantic prompts to enhance image understanding. Additionally, we build a Feature Rectify ControlNet (FRControlNet) to refine feature modulation, enabling realistic detail generation. Extensive experiments demonstrate that the proposed EPDiff outperforms state-of-the-art methods on both synthetic and real-world datasets.
Detian Huang, Miaohua Ruan, Yaohui Guo, Zhenzhen Hu 0004, Huanqiang Zeng
Comput. Vis. Image Underst.5
2025 Enhanced Cross Layer Refinement Network for robust lane detection across diverse lighting and road conditions
Weilong Dai, Huanqiang Zeng
Eng. Appl. Artif. Intell.5
2025 CMASR: Lightweight image super-resolution with cluster and match attention
Detian Huang, Mingxin Lin, Huanqiang Zeng
Image Vis. Comput.4
2025 Image re-identification: Where self-supervision meets vision-language learning
Yuying Liang, Huakun Huang, Huanqiang Zeng
Image Vis. Comput.5
2025 Local flow propagation and global multi-scale dilated Transformer for video inpainting
Yuting Zuo, Jing Chen 0001, Kaixing Wang, Huanqiang Zeng
J. Vis. Commun. Image Represent.5
2025 A federated vehicle re-identification benchmark and performance optimization using dual-phase contrastive dynamic aggregation
Linhan Huang, Jianqing Zhu, Huanqiang Zeng
Neural Networks5
2025 Memory-augmented shuffled meta learning for visible-infrared person re-identification
Hanxiao Wu, Jianqing Zhu, Liu Liu 0014, Huanqiang Zeng
Neural Networks6
2025 Learning Efficient and Effective Trajectories for Differential Equation-Based Image Restoration
abstract
The differential equation-based image restoration approach aims to establish learnable trajectories connecting high-quality images to a tractable distribution, e.g., low-quality images or a Gaussian distribution. In this paper, we reformulate the trajectory optimization of this kind of method, focusing on enhancing both reconstruction quality and efficiency. Initially, we navigate effective restoration paths through a reinforcement learning process, gradually steering potential trajectories toward the most precise options. Additionally, to mitigate the considerable computational burden associated with iterative sampling, we propose cost-aware trajectory distillation to streamline complex paths into several manageable steps with adaptable sizes. Moreover, we fine-tune a foundational diffusion model (FLUX) with 12B parameters by using our algorithms, producing a unified framework for handling 7 kinds of image restoration tasks. Extensive experiments showcase the significant superiority of the proposed method, achieving a maximum PSNR improvement of 2.1 dB over state-of-the-art methods, while also greatly enhancing visual perceptual quality.
Jinhui Hou, Hui Liu 0032, Huanqiang Zeng, Junhui Hou
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 HIA-Net: Hierarchical Interactive Alignment Network for Multimodal Few-Shot Emotion Recognition
abstract
Physiological multimodal emotion recognition (PMER) has become a key research direction for advancing human-computer interaction and affective computing. However, current PMER methods are affected by significant individual differences and the limited number of samples, making it challenging to capture complex emotional states comprehensively. To address aforementioned issues, this letter proposes a novel multimodal Few-Shot emotion recognition model, called Hierarchical Interactive Alignment Network (HIA-Net). Specifically, the Hierarchical Adaptive Interactive Attention (HAIA) module of HIA-Net is proposed to capture multidimensional emotional features and aggregate the cross-modal information effectively. Additionally, a cross-domain optimization strategy based on the maximum mean discrepancy is proposed to enhance the HIA-Net's adaptability across varying data distributions. Experimental results show that HIA-Net achieves state-of-the-art performance under Few-Shot experimental paradigms on the SEED and SEED-FRA datasets. Our codes are available athttps://github.com/As1yk/HIA-Net.
Yuankang Fu, Kaixiang Yang 0001, Xin-Rong Gong, Huanqiang Zeng
IEEE Signal Process. Lett.5
2025 Multi-Modal Prior-Guided Diffusion Model for Blind Image Super-Resolution
abstract
Recently, diffusion models have achieved remarkable success in blind image super-resolution. However, most existing methods rely solely on uni-modal degraded low-resolution images to guide diffusion models for restoring high-fidelity images, resulting in inferior realism. In this letter, we propose a Multi-modal Prior-Guided diffusion model for blind image Super-Resolution (MPGSR), which fine-tunes Stable Diffusion (SD) by utilizing the superior visual-and-textual guidance for restoring realistic high-resolution images. Specifically, our MPGSR involves two stages, i.e., multi-modal guidance extraction and adaptive guidance injection. For the former, we propose a composited transformer and further incorporate it with GPT-CLIP to extract the representative visual-and-textual guidance. For the latter, we design a feature calibration ControlNet to inject the visual guidance and employ the cross-attention layer provided by the frozen SD to inject the textual guidance, thus effectively activating the powerful text-to-image generation potential. Extensive experiments show that our MPGSR outperforms state-of-the-art methods in restoration quality and convergence time.
Detian Huang, Jiaxun Song, Xiaoqian Huang, Zhenzhen Hu 0004, Huanqiang Zeng
IEEE Signal Process. Lett.5
2025 High-Capacity Reversible Data Hiding in Encrypted Images With Edge-Directed Prediction and Adaptive Entropy Coding
abstract
As cloud services rapidly evolve and the demand for privacy protection grows, reversible data hiding in encrypted images (RDHEI) has gained significant attention. To enhance the data embedding capacity of RDHEI, this paper proposes an adaptive prediction-error entropy encoding framework that dynamically allocates edge-directed prediction (EDP) errors to either separate or grouped entropy coding modes, thereby optimizing net payload capacity. The image owner first predicts the pixel values of the cover image using EDP algorithm and calculates the prediction errors. After computing the prediction errors, the optimal thresholds are adaptively determined by minimizing the expected codeword length through entropy coding theory. Using these optimized thresholds, the prediction errors are classified into separate or grouped encoding categories, and losslessly compressed via arithmetic encoding. Through the processes of image encryption and self-embedding, an encrypted image with embedding room is generated and subsequently uploaded to the cloud server. The data hider can easily locate the data embedding room in the encrypted domain of the image and embed encrypted additional data to obtain the marked encrypted image. The authorized recipient can correctly extract the embedded data, restore the original plaintext image without any loss, or do both. The experimental results demonstrate the effectiveness of the proposed approach, surpassing many state-of-the-art RDHEI techniques.
Yingqiang Qiu, Kaimeng Chen, Xiaodan Lin, Guogang Li, Huanqiang Zeng
IEEE Signal Process. Lett.5
2025 Hierarchical Feature Fusion CNN: Fast Intra Prediction Mode Decision for VVC Screen Content Coding
abstract
Versatile Video Coding (VVC) inherits Screen Content Coding (SCC) tools such as Intra Block Copy (IBC) and Palette mode (PLT) from High Efficiency Video Coding Screen Content Coding (HEVC-SCC), which is known as VVC-SCC. VVC-SCC can effectively improve the efficiency of screen content encoding, but it can also lead to higher encoding complexity. In order to reduce the encoding complexity of VVC-SCC, we design a Hierarchical Feature Fusion Convolutional Neural Network (HFF-CNN) for predicting the current CU best intra prediction mode. The encoder determines the current CU intra prediction mode based on the network out put best prediction mode, angle intra prediction mode indexes, and adjacent CU mode probability, skipping unnecessary rate distortion cost calculations and speeding up the encoding process. Experimental results show that the proposed model reduces the intra frame encoding time of VCC-SCC by 36.6% while increasing the average BDBR by 0.44%. Compared to state-of-the-art algorithms, it exhibits a better balance between the rate distortion performance and the encoding complexity.
Jiaxin Zeng, Jing Chen 0001, Huanqiang Zeng
IEEE Signal Process. Lett.3
2025 A Novel Deep Learning-Based Receiver for Non-Coherent Chaotic Communication Systems With Temporal Dependencies and Spectral Properties of Chaotic Signals
abstract
Chaotic signals are spread-spectrum nonlinear signals with initial value sensitivity that can potentially provide safe and anti-jamming digital communication systems. However, to reduce receiver complexity and achieve smooth demodulation, traditional chaos-based wireless communication systems require the transmission of additional chaotic reference signals, which reduces the spectral efficiency and degrades the security characteristics of the chaotic signals. Recent work has explored deep learning (DL)-aided transceivers to address these issues. However, these techniques have not yet fully exploited the inherent characteristics of chaotic signals, e.g., spectral properties. To maximize the potential of chaotic signals in wireless communication, this paper proposes a power spectral density-based deep learning chaos shift keying (PSD-DLCSK) receiver. The proposed PSD-DLCSK scheme effectively utilizes the spectral properties of chaotic signals, where the PSD of received signals serves as input to a deep neural network (DNN) for symbol detection. The proposed DNN architecture combines long short-term memory networks with self-attention mechanisms to effectively capture the temporal dependencies and spectral properties of chaotic signals, enabling reliable intelligent demodulation without chaotic reference signals. Extensive simulations demonstrate that PSD-DLCSK achieves superior bit error rate (BER) performance compared to traditional receivers as well as other DL-aided schemes over additive white Gaussian noise and multipath Rayleigh fading channels.
Jundong Chen 0004, Huanqiang Zeng, Guofa Cai, Haoyu Zhou, Ziqin Shen, Georges Kaddoum
IEEE Trans. Commun.3
2025 Harmonizing Metric Discrepancy for Cross-Modal Object Re-Identification
abstract
Visible and infrared cross-modal re-identification tasks often encounter significant modal discrepancies, which undermine the effectiveness of feature extraction and compromise the reliability of similarity metrics. These discrepancies pose a substantial challenge for accurately matching data across different modalities. To address these issues, we propose a novel approach centered on the maximum mean metric discrepancy (MMMD). We leverage kernel-based statistical techniques to effectively capture and quantify the disparities in cross-modal metrics, providing a robust framework for aligning metrics from different modalities. Building upon the foundation of MMMD, we develop the metric discrepancy harmonization (MDH) method. This method integrates a temperature-controlled optimization technique designed to enhance metric alignment across various modal configurations, ensuring more consistent and reliable performance. By focusing on metric alignment, our approach enhances the accuracy of cross-modal re-identification tasks. Comprehensive evaluations on the LLCM, RGBN300, and SYSU-MM01 datasets demonstrate that our approach achieves state-of-the-art performance.
Linhan Huang, Liu Liu 0014, Jianqing Zhu, Huanqiang Zeng
IEEE Trans. Circuits Syst. Video Technol.5
2025 A No-Reference Quality Assessment Model for Screen Content Videos via Hierarchical Spatiotemporal Perception
abstract
In this paper, a novel deep learning-based no-reference video quality assessment (NR-VQA) model for screen content videos (SCVs) is proposed, called the hierarchical spatiotemporal perceptual quality model (HSPQ). Firstly, the human visual system (HVS) perceives SCVs hierarchically, with varying sensitivity and attention to diverse attribute regions. Secondly, the visual redundancies are copious in the spatiotemporal domain of SCVs, degrading video quality to some extent. Based on these characteristics, the SCVs are decomposed into three hierarchical levels (i.e., patch level, frame level, and video level), which contain quality-related spatiotemporal information. Specifically, the visual saliency is first utilized for more salient textual and pictorial patches selection, and then, a dual-channel convolutional neural network integrating spatial-gate feature enhancement module (SGFEM) is designed to evaluate the quality of patches based on their attributes at the patch level separately. With spatial correlation, an adaptive blur-focused visual mechanism-based weighting strategy (BFWS) is proposed for converting quality scores from patch level to frame level. Finally, the video-level quality score, which reflects the temporal perceptual quality degradation, is combined to provide a comprehensive evaluation of distorted SCV quality. Experiments conducted on the Screen Content Video Database (SCVD) and Compressed Screen Content Video Quality (CSCVQ) databases demonstrate that our proposed HSPQ model aligns better with the visual perception of SCVs by the HVS. Moreover, it exhibits strong robustness compared to multiple classic and state-of-the-art image/video quality assessment models.
Huanqiang Zeng, Jing Chen 0001, Yifan Shi 0001, Junhui Hou
IEEE Trans. Circuits Syst. Video Technol.2
2025 Torch-Advent-Civilization-Evolution: Accelerating Diffusion Model for Image Restoration
abstract
Recently, diffusion models as a hot paradigm have shown considerable superiority in image restoration with an unsupervised manner. However, they require iterative refinement from isotropic Gaussian through thousands of steps to produce a sample with exceptional quality. Most existing methods are devoted to designing fast solvers for the reverse stochastic differential equation (SDE) to accelerate sampling, while neglecting the potential of forward SDE. To better stimulate this potential, we propose the Torch-Advent-Civilization-Evolution (TACE), a novel diffusion model-based zero-shot framework for image restoration. Specifically, we propose the “Torch”, a latent vector that explicitly contains content information from the measurement image. By utilizing the Torch instead of isotropic Gaussian as initialization, our TACE significantly accelerates image restoration with better consistency and realness. To acquire the Torch from the forward process, we propose Prometheus SDEs, a cluster of equivalent SDEs. Furthermore, we construct a Conditional Guidance Projection (CGP) for the reverse SDE to strengthen the consistency of restored images. Finally, we design a Civilization Shuttle Strategy (CSS) for the generation process to enhance the realness of restored images. Extensive experiments validate that our TACE achieves state-of-the-art performance with fewer sampling steps in various typical tasks, such as super-resolution, deblurring, and colorization.
Jiaxun Song, Detian Huang, Xiaoqian Huang, Miaohua Ruan, Huanqiang Zeng
IEEE Trans. Circuits Syst. Video Technol.5
2025 Enhanced Cross-Modal Hashing via Hybrid Distillation and Structural Refinement
abstract
Since cross-modal hashing requires minimal storage and computation, it is becoming increasingly popular with the exponential growth of multimedia content on the internet. However, the lack of accurate supervisory data has curtailed the effectiveness of unsupervised hashing techniques. Conversely, supervised hashing strategies necessitate considerable human and financial resources for data annotation. To address this limitation, we propose a novel semi-supervised cross-modal hashing method called Enhanced Cross-Modal Hashing via Hybrid Distillation and Structural Refinement (HDSR). Specifically, we first learn the features of inter-modal and inter-instance similarity relationships through pointwise semantic alignment and listwise similarity partial order learning, respectively, to extract refined structural representations from partially labeled data. Secondly, by fusing inter-modal similarity to construct higher-order affinity matrices, we precisely delineate the semantic correlation information across cross-modal data, facilitating stable self-supervised training of unlabeled data through the application of momentum fusion strategies. Finally, the refined structural representation of labeled data is transferred into unlabeled branches through hybrid distillation, enhancing the performance of cross-modal hash learning by generating compact and accurate hash codes. The proposed HDSR is compared with several state-of-the-art deep cross-modal hashing methods on three widely used benchmark databases, and the experimental results verify its efficiency and superiority.
Zhiwen Yu 0002, Kaixiang Yang 0001, Jun Yu 0002, Huanqiang Zeng, C. L. Philip Chen
IEEE Trans. Image Process.5
2025 On the Behavior of Contrastive Regularization in Improving Chinese Text Recognizer
abstract
The dense representation space in Chinese scene text recognition (STR) makes discriminating between categories highly challenging, because of the large candidate category set. Mainstream STR methods have achieved remarkable advancements by leveraging linguistic knowledge to implicitly address this challenge. In this paper, inspired by the correlation between recognizer performance and the distributional properties of character representations, as well as the inherent consistency between this correlation and supervised contrastive learning (SupCon), we thoroughly investigate how to integrate SupCon with an STR model to alleviate this challenge, and elucidate some dynamic behaviors underlying the performance improvements. Specifically, we analyze the SupCon-STR models instantiated with different projectors and evaluate their distributional properties through metrics, including intra-class compactness, inter-class separability, and feature redundancy, while assessing performances that involve in-domain accuracy and cross-domain recognition generalization. The main results reveal how the temperature$\tau$and projectors affect the representation distribution, and highlight that suitable intra-class compactness and sufficient inter-class separability are key factors for delivering competitive performances in both in-domain and cross-domain STR scenarios. Moreover, these results also provide valuable insights into the design of SupCon-STR architectures for diverse resource constraints. Taking existing Chinese STR models as baselines, and combining SupCon-STR with them, the average improvements in cross-domain recognition performance are over 5% across 7 testing datasets. A new state-of-the-art accuracy of 77.19% on the ChineseScenebenchmark is also established.
Tianlei Wang, Huanqiang Zeng, Jiuwen Cao
IEEE Trans. Multim.3
2025 Ensemble Prototype Networks for Unsupervised Cross-Modal Hashing With Cross-Task Consistency
abstract
In the swiftly advancing realm of information retrieval, unsupervised cross-modal hashing has emerged as a focal point of research, taking advantage of the inherent advantages of the multifaceted and dynamism inherent in multimedia data. Existing unsupervised cross-modal hashing methods rely mainly on initial pre-trained correlations among cross-modal features, and the inaccurate neighborhood correlations impacts the presentation of common semantics throughout the optimization. To address the aforementioned issues, we proposeEnsemblePrototypeNetworks (EPNet), which delineates class attributes of cross-modal instances through an ensemble clustering methodology. EPNet seeks to extract correlation information between instances by leveraging local correlation aggregation and ensemble clustering from multiple perspectives, aiming to reduce initialization effects and enhance cross-modal representations. Specifically, the local correlation aggregation is first proposed within a batch of semantic affinity relationships to generate a precise and compact hash code among cross-modal instances. Secondly, the ensemble prototype module is employed to discern the class attributes of deep features, thereby aiding the model in extracting more universally applicable feature representations. Thirdly, an early attempt to constrict the representational congruity of local semantic affinity relationships and deep feature ensemble prototype correlations using cross-task consistency loss aims to enhance the representation of cross-modal common semantic features. Finally, EPNet outperforms several state-of-the-art cross-modal retrieval methods on three real-world image-text datasets in extensive experiments.
Huanqiang Zeng, Yifan Shi 0001, Jianqing Zhu, Kaixiang Yang 0001, Zhiwen Yu 0002
IEEE Trans. Multim.2
2025 Unsupervised 3D Point Cloud Completion via Multi-View Adversarial Learning
abstract
In real-world scenarios, scanned point clouds are often incomplete due to occlusion issues. The tasks of self-supervised and weakly-supervised point cloud completion involve reconstructing missing regions of these incomplete objects without the supervision of complete ground truth. Current methods either rely on multiple views of partial observations for supervision or overlook the intrinsic geometric similarity that can be identified and utilized from the given partial point clouds. In this paper, we propose MAL-UPC, a framework that effectively leverages both region-level and category-specific geometric similarities to complete missing structures. Our MAL-UPC does not require any 3D complete supervision and only necessitates single-view partial observations in the training set. Specifically, we first introduce a Pattern Retrieval Network to retrieve similar position and curvature patterns between the partial input and the predicted shape, then leverage these similarities to densify and refine the reconstructed results. Additionally, we render the reconstructed complete shape into multi-view depth maps and design an adversarial learning module to learn the geometry of the target shape from category-specific single-view depth images of the partial point clouds in the training set. To achieve anisotropic rendering, we design a density-aware radius estimation algorithm to improve the quality of the rendered images. Our MAL-UPC outperforms current state-of-the-art self-supervised methods and even some unpaired approaches.
Lintai Wu, Xianjing Cheng, Yong Xu 0001, Huanqiang Zeng, Junhui Hou
IEEE Trans. Vis. Comput. Graph.4
2024 Adaptive Data Association for Enhanced Multi-object Tracking and Segmentation with Pre-matching and Selective Association
Longtao Chen, Guoxing Liao, Gaofeng Zhu, Huanqiang Zeng
ICPR (22)4
2024 Audio-Visual Cross-Modal Generation with Multimodal Variational Generative Model
abstract
Audio and Visual are two important visual modalities in video content understanding. However, the absence of one modality may be observed in practical applications due to the real environmental factors, which leads to the information loss. Therefore, audio and visual fusion is focused on using the shared and complementary information between modalities to recover the missing modalities from the available data modalities. In this paper, an Adversarial Hierarchical Variational Auto-Encoder (Adv-HVAE) model is proposed to solve this problem of modality data loss. A multimodal representation is first learned using a hierarchical Variational Autoencoder (VAE) model that enables the generation of missing modal data under any subset of available modalities. Also to obtain a more robust multimodal representation, a feature generation network is utilized to approximate the latent distribution of missing modalities. Finally, the adversarial training network is shown to be effective in improving the data quality generated through the Adv-HVAE framework. Experimental results demonstrate that Adv-HVAE achieves best generation results on two benchmark datasets, avMNIST and Sub-URMP.
Zhubin Xu, Tianlei Wang, Dinghan Hu, Huanqiang Zeng, Jiuwen Cao
ISCAS5
2024 PrefPaint: Aligning Image Inpainting Diffusion Model with Human Preference
abstract
In this paper, we make the first attempt to align diffusion models for image inpainting with human aesthetic standards via a reinforcement learning framework, significantly improving the quality and visual appeal of inpainted images. Specifically, instead of directly measuring the divergence with paired images, we train a reward model with the dataset we construct, consisting of nearly 51,000 images annotated with human preferences. Then, we adopt a reinforcement learning process to fine-tune the distribution of a pre-trained diffusion model for image inpainting in the direction of higher reward. Moreover, we theoretically deduce the upper bound on the error of the reward model, which illustrates the potential confidence of reward estimation throughout the reinforcement alignment process, thereby facilitating accurate regularization. Extensive experiments on inpainting comparison and downstream tasks, such as image extension and 3D reconstruction, demonstrate the effectiveness of our approach, showing significant improvements in the alignment of inpainted images with human preference compared with state-of-the-art methods. This research not only advances the field of image inpainting but also provides a framework for incorporating human preference into the iterative refinement of generative models based on modeling reward accuracy, with broad implications for the design of visually driven AI applications. Our code and dataset are publicly available at \url{https://prefpaint.github.io}.
Kendong Liu, Chuanhao Li 0002, Hui Liu 0032, Huanqiang Zeng, Junhui Hou
NeurIPS5
2024 Mask-Guided Clothes-Irrelevant and Background-Irrelevant Network with Knowledge Propagation for Cloth-Changing Person Re-identification
Gaofeng Zhu, Longtao Chen, Guoxing Liao, Huanqiang Zeng
PRCV (12)5
2024 Global Instance Relation Distillation for convolutional neural network compression
Haolin Hu, Huanqiang Zeng, Yifan Shi 0001, Jianqing Zhu, Jing Chen 0001
Neural Comput. Appl.2
2024 Unsupervised video-based action recognition using two-stream generative adversarial network
Wei Lin 0021, Huanqiang Zeng, Jianqing Zhu, Chih-Hsien Hsia, Junhui Hou, Kai-Kuang Ma
Neural Comput. Appl.2
2024 Cross-modal group-relation optimization for visible-infrared person re-identification
Jianqing Zhu, Hanxiao Wu, Yuqing Fu, Huanqiang Zeng, Liu Liu 0014, Zhen Lei 0001
Neural Networks6
2024 Deep Diversity-Enhanced Feature Representation of Hyperspectral Images
abstract
In this paper, we study the problem of efficiently and effectively embedding the high-dimensional spatio-spectral information of hyperspectral (HS) images, guided by feature diversity. Specifically, based on the theoretical formulation that feature diversity is correlated with the rank of the unfolded kernel matrix, we rectify 3D convolution by modifying its topology to enhance the rank upper-bound. This modification yields a rank-enhanced spatial-spectral symmetrical convolution set (ReS$^{3}$-ConvSet), which not only learns diverse and powerful feature representations but also saves network parameters. Additionally, we also propose a novel diversity-aware regularization (DA-Reg) term that directly acts on the feature maps to maximize independence among elements. To demonstrate the superiority of the proposed ReS$^{3}$-ConvSet and DA-Reg, we apply them to various HS image processing and analysis tasks, including denoising, spatial super-resolution, and classification. Extensive experiments show that the proposed approaches outperform state-of-the-art methods both quantitatively and qualitatively to a significant extent. The code is publicly available athttps://github.com/jinnh/ReSSS-ConvSet.
Jinhui Hou, Junhui Hou, Hui Liu 0032, Huanqiang Zeng, Deyu Meng
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 Pairwise difference relational distillation for object re-identification
Hanxiao Wu, Yihong Lin, Jianqing Zhu, Huanqiang Zeng
Pattern Recognit.5
2024 Distillation embedded absorbable pruning for fast object re-identification
Hanxiao Wu, Jianqing Zhu, Huanqiang Zeng
Pattern Recognit.4
2024 Modality-Consistent Attention for Visible-Infrared Vehicle Re-Identification
abstract
Visible-infrared vehicle re-identification (VIVR) seeks to match vehicle images of the same identity taken by cameras of different modalities. The noticeable disparity between visible and infrared modalities leads to attention deviations, causing deep models to incorrectly focus on different local regions of vehicles in visible and infrared images. We observed that the spatial distributions of distinguishing local regions, such as logos, front windows, and wheels, exhibit similarity in average images obtained from both visible and infrared images. Based on this, we propose a modality-consistent attention (MCA) approach for VIVR. Unlike image-level attention, our MCA is identity-level attention that holistically emphasizes the distinguishing regions of a vehicle identity across multiple images captured from various viewpoints. Furthermore, we constrain the differences between the identity-level spatial attention masks resulting from visible and infrared modalities. This approach helps deep networks focus consistently on learning the distinguishing local characteristics of vehicles across different modalities and viewpoints. Our experiments on RGBN300 and MSVR310 datasets demonstrate that our approach achieves state-of-the-art performance.
Jiajun Su, Jianqing Zhu, Liu Liu 0014, Huanqiang Zeng
IEEE Signal Process. Lett.5
2024 A Benchmark for Vehicle Re-Identification in Mixed Visible and Infrared Domains
abstract
We propose a new benchmark for vehicle re-identification in mixed visible and infrared domains. Unlike cross-modal vehicle re-identification, we focus on a more realistic scenario, namely, mixed-modal vehicle re-identification, in which both probe and gallery sets contain visible and infrared images. We provide auto-cropped visible and infrared images, simulating data from actual surveillance systems. We design a mixed triplet loss function for model training. We report both mixed-modal and cross-modal retrieval performance. Our mixed method performs well in cross-modal retrieval, e.g., the Rank-1 identification rate is 79.19% in the visible-to-infrared retrieval mode.
Simin Zhan, Jianqing Zhu, Huanqiang Zeng
IEEE Signal Process. Lett.5
2024 Collaborative Knowledge Distillation
abstract
Existing research on knowledge distillation has primarily concentrated on the task of facilitating student networks in acquiring the complete knowledge imparted by teacher networks. However, recent studies have shown that good networks are not suitable for acting as teachers, and there is a positive correlation between distillation performance and teacher prediction uncertainty. To address this finding, this paper thoroughly analyzes in depth the reasons why the teacher network affects the distillation performance, gives full play to the participation of the student network in the process of knowledge distillation, and assists the teacher network in distilling the knowledge that is suitable for their learning. In light of this premise, a novel approach known as Collaborative Knowledge Distillation (CKD) is introduced, which is founded upon the concept of "Tailoring the Teaching to the Individual". Compared with Baseline, this paper’s method improves students’ accuracy by an average of 3.42% in CIFAR-100 experiments, and by an average of 1.71% compared with the classical Knowledge Distillation (KD) method. The ImageNet experiments conducted revealed a significant improvement of 2.04% in the Top-1 accuracy of the students.
Junhuang Wang, Jianqing Zhu, Huanqiang Zeng
IEEE Trans. Circuits Syst. Video Technol.5
2024 Expanding and Refining Hybrid Compressors for Efficient Object Re-Identification
abstract
Recent object re-identification (Re-ID) methods gain high efficiency via lightweight student models trained by knowledge distillation (KD). However, the huge architectural difference between lightweight students and heavy teachers causes students to have difficulties in receiving and understanding teachers' knowledge, thus losing certain accuracy. To this end, we propose a refiner-expander-refiner (RER) structure to enlarge a student's representational capacity and prune the student's complexity. The expander is a multi-branch convolutional layer to expand the student's representational capacity to understand a teacher's knowledge comprehensively, which does not require any feature-dimensional adapter to avoid knowledge distortions. The two refiners are 1×1 convolutional layers to prune the input and output channels of the expander. In addition, in order to alleviate the competition accuracy-related and pruning-related gradients, we design a common consensus gradient resetting (CCGR) method, which discards unimportant channels according to the intersection of each sample's unimportant channel judgment. Finally, the trained RER can be simplified into a slim convolutional layer via re-parameterization to speed up inference. As a result, we propose an expanding and refining hybrid compressing (ERHC) method. Extensive experiments show that our ERHC has superior inference speed and accuracy, e.g., on the VeRi-776 dataset, given the ResNet101 as a teacher, ERHC saves 75.33% model parameters (MP) and 74.29% floating-point of operations (FLOPs) without sacrificing accuracy.
Hanxiao Wu, Jianqing Zhu, Huanqiang Zeng, Jing Zhang 0037
IEEE Trans. Image Process.4
2024 Width-Adaptive CNN: Fast CU Partition Prediction for VVC Screen Content Coding
abstract
Screen content coding (SCC) in Versatile Video Coding (VVC) improves the coding efficiency of screen content videos (SCVs) significantly but results in high computational complexity due to the quad-tree plus multi-type tree (QTMT) structure of the coding unit (CU) partitioning. Therefore, we make the first attempt to reduce the encoding complexity from the perspective of CU partitioning for SCC in VVC. To this end, a fast CU partition prediction method is technically developed for VVC-SCC. First, to solve the problem of lacking sufficient SCC training data, SCVs are collected to establish a database containing CUs of various sizes and corresponding partition labels. Second, to determine the partition decision in advance, a novel WA-CNN model is proposed, which is capable of predicting two large CUs for VVC-SCC by adjusting the feature channels based on the size of input CU blocks. Finally, considering the imbalanced proportion of diverse partition decisions, a loss function with the weight that equalizes the contribution of imbalanced data is formulated to train the proposed WA-CNN model. Experimental results show that the proposed model reduces the SCC intra-encoding time by 35.65%${\sim }$38.31% with an average of 1.84%${\sim }$2.42% BDBR increase.
Chao Jiao, Huanqiang Zeng, Jing Chen 0001, Chih-Hsien Hsia, Tianlei Wang, Kai-Kuang Ma
IEEE Trans. Multim.2
2024 Consensus Clustering With Co-Association Matrix Optimization
abstract
Consensus clustering can derive a more promising and robust clustering result by integrating multiple partitions strategically. However, there are several limitations in the existing approaches: 1) most of the methods compute the ensemble-information matrix heuristically and lack of sufficient optimization; 2) the information from the original dataset is rarely considered; and 3) the noise in both label space and feature space is ignored. To address these issues, we proposed a novel consensus clustering method with co-association matrix optimization (CC-CMO), which aims at improving the co-association matrix by taking abundant information from both label space and feature space into consideration. In label space, CC-CMO derives a weighted partition matrix capturing the intercluster correlation and further designs a least squares regression (LSR) model to explore the global structure of data. In feature space, CC-CMO minimizes the reconstruction error with doubly stochastic normalization in the projective subspace to eliminate noise features as well as learn the local affinity of data. To improve the co-association matrix by jointly considering the subspace representation, global structure, and local affinity of data, we explicitly propose a unified optimization framework and design an alternating optimization algorithm for the optimal co-association matrix. Extensive experiments on a variety of real-world datasets demonstrate the superior performance of CC-CMO to the state-of-the-art consensus clustering approaches.
Yifan Shi 0001, Zhiwen Yu 0002, C. L. Philip Chen, Huanqiang Zeng
IEEE Trans. Neural Networks Learn. Syst.4
2023 Global Structure-Aware Diffusion Process for Low-light Image Enhancement
abstract
This paper studies a diffusion-based framework to address the low-light image enhancement problem. To harness the capabilities of diffusion models, we delve into this intricate process and advocate for the regularization of its inherent ODE-trajectory. To be specific, inspired by the recent research that low curvature ODE-trajectory results in a stable and effective diffusion process, we formulate a curvature regularization term anchored in the intrinsic non-local structures of image data, i.e., global structure-aware regularization, which gradually facilitates the preservation of complicated details and the augmentation of contrast during the diffusion process. This incorporation mitigates the adverse effects of noise and artifacts resulting from the diffusion process, leading to a more precise and flexible enhancement. To additionally promote learning in challenging regions, we introduce an uncertainty-guided regularization technique, which wisely relaxes constraints on the most extreme regions of the image. Experimental evaluations reveal that the proposed diffusion-based framework, complemented by rank-informed regularization, attains distinguished performance in low-light enhancement. The outcomes indicate substantial advancements in image quality, noise suppression, and contrast amplification in comparison with state-of-the-art methods. We believe this innovative approach will stimulate further exploration and advancement in low-light image processing, with potential implications for other applications of diffusion models. The code is publicly available at https://github.com/jinnh/GSAD.
Jinhui Hou, Junhui Hou, Hui Liu 0032, Huanqiang Zeng, Hui Yuan 0001
NeurIPS5
2023 Attribute-Image Person Re-identification via Modal-Consistent Metric Learning
Jianqing Zhu, Liu Liu 0014, Yibing Zhan, Xiaobin Zhu 0001, Huanqiang Zeng, Dacheng Tao
Int. J. Comput. Vis.5
2023 Visible-infrared person re-identification using high utilization mismatch amending triplet loss
Jianqing Zhu, Hanxiao Wu, Huanqiang Zeng, Xiaobin Zhu 0001, Jingchang Huang, Canhui Cai
Image Vis. Comput.4
2023 Screen content video quality assessment based on spatiotemporal sparse feature
Huanqiang Zeng, Hailiang Huang 0002, Shan Cheng, Junhui Hou
J. Vis. Commun. Image Represent.2
2023 Content-Aware Warping for View Synthesis
abstract
Existing image-based rendering methods usually adopt depth-based image warping operation to synthesize novel views. In this paper, we reason the essential limitations of the traditional warping operation to be the limited neighborhood and only distance-based interpolation weights. To this end, we propose content-aware warping, which adaptively learns the interpolation weights for pixels of a relatively large neighborhood from their contextual information via a lightweight neural network. Based on this learnable warping module, we propose a new end-to-end learning-based framework for novel view synthesis from a set of input source views, in which two additional modules, namely confidence-based blending and feature-assistant spatial refinement, are naturally proposed to handle the occlusion issue and capture the spatial correlation among pixels of the synthesized view, respectively. Besides, we also propose a weight-smoothness loss term to regularize the network. Experimental results on light field datasets with wide baselines and multi-view datasets show that the proposed method significantly outperforms state-of-the-art methods both quantitatively and visually. The source code is publicly available at https://github.com/MantangGuo/CW4VS.
Mantang Guo, Junhui Hou, Jing Jin 0006, Hui Liu 0032, Huanqiang Zeng, Jiwen Lu
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Multiscale Attentive Image De-Raining Networks via Neural Architecture Search
abstract
Multi-scale architectures and attention modules have shown effectiveness in many deep learning-based image de-raining methods. However, manually designing and integrating these two components into a neural network requires a bulk of labor and extensive expertise. In this article, a high-performance multi-scale attentive neural architecture search (MANAS) framework is technically developed for image de-raining. The proposed method formulates a new multi-scale attention search space with multiple flexible modules that are favorite to the image de-raining task. Under the search space, multi-scale attentive cells are built, which are further used to construct a powerful image de-raining network. The internal multi-scale attentive architecture of the de-raining network is searched automatically through a gradient-based search algorithm, which avoids the daunting procedure of the manual design to some extent. Moreover, in order to obtain a robust image de-raining model, a practical and effective multi-to- one training strategy is also presented to allow the de-raining network to get sufficient background information from multiple rainy images with the same background scene, and meanwhile, multiple loss functions including external loss, internal loss, architecture regularization loss, and model complexity loss are jointly optimized to achieve robust de-raining performance and controllable model complexity. Extensive experimental results on both synthetic and realistic rainy images, as well as the down-stream vision applications (i.e., objection detection and segmentation) consistently demonstrate the superiority of our proposed method. The code is publicly available athttps://github.com/lcai-gz/MANAS.
Yuli Fu 0001, Wanliang Huo, Youjun Xiang, Tao Zhu 0002, Huanqiang Zeng, Delu Zeng
IEEE Trans. Circuits Syst. Video Technol.7
2023 CLSR: Cross-Layer Interaction Pyramid Super-Resolution Network
abstract
Convolutional Neural Network (CNN) achieves impressive success in image super-resolution (SR), where global context interaction is critical for reconstructing reliable edge and texture details. However, most CNN-based SR models focus on modeling the global contextual information within a single feature map by using attention mechanisms, and ignore the dependencies among hierarchical features, resulting in blurred or even distorted detail restoration, especially for SR tasks with large scaling factors (i.e.,$\times 4$,$\times 8$). To tackle the above issue, we propose a Cross-Layer interaction pyramid Super-Resolution (CLSR) network that reconstructs the desired SR images progressively in a coarse-to-fine fashion. Specifically, we propose a novel Cross-Layer Non-Local attention (CLNL) for accurate detail restoration. Through explicitly modeling the long-range feature-wise similarities within and between layers, the proposed CLNL is able to discriminatively explore complementary patches from hierarchical features to reconstruct the target LR patches. Then, to further strengthen the information interaction among hierarchical features at different scales, we propose a novel Gradient Consistency-Aware learning framework (GCA) by constructing a closed loop (LR$\rightarrow $HR$\rightarrow $LR) on the gradient space. The proposed GCA is able to effectively capture the interdependence between LR and HR gradient maps to guide our CLSR for reliable detail restoration. Extensive experiments validate that our CLSR outperforms the state-of-the-art methods in terms of both reconstruction accuracy and visual quality.
Detian Huang, Xiancheng Zhu, Huanqiang Zeng
IEEE Trans. Circuits Syst. Video Technol.4
2023 Unsupervised Video-Based Action Recognition With Imagining Motion and Perceiving Appearance
abstract
Video-based action recognition is a challenging task, which demands carefully considering the temporal property of videos in addition to the appearance attributes. Particularly, the temporal domain of raw videos usually contains significantly more redundant or irrelevant information than still images. For that, this paper proposes an unsupervised video-based action recognition approach with imagining motion and perceiving appearance, called IMPA, by comprehensively learning the spatio-temporal characteristics inherited in videos, with a particular emphasis on the moving object for action recognition. Specifically, a self-supervised Motion Extracting Block (MEB) is designed to extract the principal motion features by focusing on the large movement of the moving object, based on the observation that humans can infer complete motion trajectories from partial moving objects. To further take the indispensable appearance attribute in videos into account, an unsupervised Appearance Learning Block (ALB) is developed to perceive the static appearance, thus in combination with the MEB to recognize actions. Extensive validation experiments and ablation studies on multiple datasets demonstrate that our proposed IMPA approach obtains superior performance and surpasses other classical and state-of-the-art unsupervised action recognition methods.
Wei Lin 0021, Yihong Zhuang, Xinghao Ding, Xiaotong Tu, Yue Huang 0001, Huanqiang Zeng
IEEE Trans. Circuits Syst. Video Technol.7
2023 Stacked One-Class Broad Learning System for Intrusion Detection in Industry 4.0
abstract
With the vigorous development of Industry 4.0, industrial Big Data has turned into the core element of the Industrial Internet of Things. As one of the most fundamental and indispensable components in industrial cyber-physical systems (CPS), intelligent anomaly detection is still an essential and challenging issue. However, with the development of the network, there may exist unknown types of attacks, which are difficult to collect. Facing one-class industrial intrusion detection scenario that the collected training data only includes normal state, the one-class broad learning system (OCBLS) and the stacked OCBLS (ST-OCBLS) algorithms are developed. Benefiting from the characteristics of BLS, our proposed approaches retain the advantage of efficient training process. Moreover, the high-level hidden features of the network traffic data can be learned through the progressive encoding and decoding mechanism in ST-OCBLS. Extensive comparative experiments on several real-world intrusion detection tasks are carried out to demonstrate that our proposed methods have competitive performance and high efficiency in the face of complex network data and diversified types of intrusions. Overall, this article provides a new alternative solution for network intrusion detection in Industry 4.0.
Kaixiang Yang 0001, Yifan Shi 0001, Zhiwen Yu 0002, Qinmin Yang, Arun Kumar Sangaiah, Huanqiang Zeng
IEEE Trans. Ind. Informatics6
2023 Self-Supervised Video-Based Action Recognition With Disturbances
abstract
Self-supervised video-based action recognition is a challenging task, which needs to extract the principal information characterizing the action from content-diversified videos over large unlabeled datasets. However, most existing methods choose to exploit the natural spatio-temporal properties of video to obtain effective action representations from a visual perspective, while ignoring the exploration of the semantic that is closer to human cognition. For that, a self-supervised Video-based Action Recognition method with Disturbances called VARD, which extracts the principal information of the action in terms of the visual and semantic, is proposed. Specifically, according to cognitive neuroscience research, the recognition ability of humans is activated by visual and semantic attributes. An intuitive impression is that minor changes of the actor or scene in video do not affect one person's recognition of the action. On the other hand, different humans always make consistent opinions when they recognize the same action video. In other words, for an action video, the necessary information that remains constant despite the disturbances in the visual video or the semantic encoding process is sufficient to represent the action. Therefore, to learn such information, we construct a positive clip/embedding for each action video. Compared to the original video clip/embedding, the positive clip/embedding is disturbed visually/semantically by Video Disturbance and Embedding Disturbance. Our objective is to pull the positive closer to the original clip/embedding in the latent space. In this way, the network is driven to focus on the principal information of the action while the impact of sophisticated details and inconsequential variations is weakened. It is worthwhile to mention that the proposed VARD does not require optical flow, negative samples, and pretext tasks. Extensive experiments conducted on the UCF101 and HMDB51 datasets demonstrate that the proposed VARD effectively improves the strong baseline and outperforms multiple classical and advanced self-supervised action recognition methods.
Wei Lin 0021, Xinghao Ding, Yue Huang 0001, Huanqiang Zeng
IEEE Trans. Image Process.4
2023 DeflickerCycleGAN: Learning to Detect and Remove Flickers in a Single Image
abstract
Eliminating the flickers in digital images captured by rolling shutter cameras is a fundamental and important task in computer vision applications. The flickering effect in a single image stems from the mechanism of asynchronous exposure of rolling shutters employed by cameras equipped with CMOS sensors. In an artificial lighting environment, the light intensity captured at different time intervals varies due to the fluctuation of the power grid, ultimately resulting in the flickering artifact in the image. Up to date, there are few studies related to single image deflickering. Further, it is even more challenging to remove flickers without a priori information, e.g., camera parameters or paired images. To address these challenges, we propose an unsupervised framework termed DeflickerCycleGAN, which is trained on unpaired images for end-to-end single image deflickering. Besides the cycle-consistency loss to maintain the similarity of image contents, we meticulously design another two novel loss functions, i.e., gradient loss and flicker loss, to reduce the risk of edge blurring and color distortion. Moreover, we provide a strategy to determine whether an image contains flickers or not without extra training, which leverages an ensemble methodology based on the output of two previously trained markovian discriminators. Extensive experiments on both synthetic and real datasets show that our proposed DeflickerCycleGAN not only achieves excellent performance on flicker removal in a single image but also shows high accuracy and competitive generalization ability on flicker detection, compared to that of a well-trained classifier based on ResNet50.
Xiaodan Lin, Yangfu Li, Jianqing Zhu, Huanqiang Zeng
IEEE Trans. Image Process.4
2023 GiT: Graph Interactive Transformer for Vehicle Re-Identification
abstract
Transformers are more and more popular in computer vision, which treat an image as a sequence of patches and learn robust global features from the sequence. However, pure transformers are not entirely suitable for vehicle re-identification because vehicle re-identification requires both robust global features and discriminative local features. For that, a graph interactive transformer (GiT) is proposed in this paper. In the macro view, a list of GiT blocks are stacked to build a vehicle re-identification model, in where graphs are to extract discriminative local features within patches and transformers are to extract robust global features among patches. In the micro view, graphs and transformers are in an interactive status, bringing effective cooperation between local and global features. Specifically, one current graph is embedded after the former level's graph and transformer, while the current transform is embedded after the current graph and the former level's transformer. In addition to the interaction between graphs and transforms, the graph is a newly-designed local correction graph, which learns discriminative local features within a patch by exploring nodes' relationships. Extensive experiments on three large-scale vehicle re-identification datasets demonstrate that our GiT method is superior to state-of-the-art vehicle re-identification approaches.
Fei Shen 0004, Jianqing Zhu, Xiaobin Zhu 0001, Huanqiang Zeng
IEEE Trans. Image Process.5
2023 Adaptive Ensemble Clustering With Boosting BLS-Based Autoencoder
abstract
Ensemble clustering has an advantage in producing a more promising and robust clustering result by combining multiple partitions strategically. The quality of both base partitions and co-association matrix plays an essential role in improving the consensus partition. However, the current ensemble clustering methods have several limitations: 1) The noise in high-dimensional feature space is ignored; 2) The independent base partition generation process does not pay attention to ambiguous samples; 3) The co-association matrix and the weights of base partitions commonly lack of theoretical optimization. In order to address these issues, we propose an adaptive ensemble clustering framework with boosting BLS-based autoencoder (BoostAEC). In the generation step, a boosting BLS-based autoencoder (BoostBLSAE) is designed to generate base partitions sequentially, which learns compressed feature subspaces for ambiguous samples and adaptively evaluates the corresponding weights of reliability. In the integration step, we construct a fuzzy membership function to capture the inter-cluster correlation and explicitly propose a consensus objective function to optimize the unified co-association matrix by considering the weighted base partitions. Extensive experiments on various real-world datasets demonstrate the superior performance of BoostAEC to the state-of-the-art ensemble clustering methods.
Yifan Shi 0001, Kaixiang Yang 0001, Zhiwen Yu 0002, C. L. Philip Chen, Huanqiang Zeng
IEEE Trans. Knowl. Data Eng.5
2023 3D-Gradient Guided Rate Control Model for Screen Content Video Coding
abstract
Compared with natural videos,screen content videos(SCVs) have particular features, such as fruitful sharper edges, lots of computer-generated graphics and texts, a large amount of flat areas. New tools are adopted toHEVC extensions on Screen Content Coding(HEVC-SCC), the traditional video rate control methods for natural videos are not effective for SCVs. For that, a3D-gradient guided rate control modelfor SCV coding, named 3DG-RC, is proposed to allocate bitrate more efficiently serving for SCVs. By considering the particular spatial-temporal characteristics of SCVs, the spatial and temporal feature extraction scheme is developed by using 3D-gradient filter and performed on the SCV to extract the spatial and temporal features simultaneously for guiding the bit allocation. The spatial-temporal feature similarity between three original reference SCV frames and their reconstructed ones is used to estimate the encoding parameters of the current block and frame. Experimental results demonstrate that compared with the classical and state-of-the-art rate control methods for HEVC-SCC, the proposed 3DG-RC algorithm achieves significant bitrate mismatch reduction and coding efficiency improvement for HEVC-SCC. In specific, the proposed 3DG-RC model outperforms the rate control model in SCM-8.8 with over 41.33% and 37.95% BD-BR savings on average, forlow delay B(LDB) andrandom access(RA) coding structure, respectively.
Jing Chen 0001, Huanqiang Zeng, Chih-Hsien Hsia, Tianlei Wang, Kai-Kuang Ma
IEEE Trans. Multim.3
2023 Deep Cross-Modal Hashing Based on Semantic Consistent Ranking
abstract
The amount of multi-modal data available on the Internet is enormous. Cross-modal hash retrieval maps heterogeneous cross-modal data into a single Hamming space to offer fast and flexible retrieval services. However, existing cross-modal methods mainly rely on the feature-level similarity between multi-modal data and ignore the relationship between relative rankings and label-level fine-grained similarity of neighboring instances. To overcome these issues, we propose a novelDeepCross-modalHashing based onSemanticConsistentRanking (DCH-SCR) that comprehensively investigates the intra-modal semantic similarity relationship. Firstly, to the best of our knowledge, it is an early attempt to preserve semantic similarity for cross-modal hashing retrieval by combining label-level and feature-level information. Secondly, the inherent gap between modalities is narrowed by developing a ranking alignment loss function. Thirdly, the compact and efficient hash codes are optimized based on the common semantic space. Finally, we use the gradient to specify the optimization direction and introduce the Normalized Discounted Cumulative Gain (NDCG) to achieve varying optimization strengths for data pairs with different similarities. Extensive experiments on three real-world image-text retrieval datasets demonstrate the superiority of DCH-SCR over several state-of-the-art cross-modal retrieval methods.
Huanqiang Zeng, Yifan Shi 0001, Jianqing Zhu, Chih-Hsien Hsia, Kai-Kuang Ma
IEEE Trans. Multim.2
2022 Deep Rank Cross-Modal Hashing with Semantic Consistent for Image-Text Retrieval
abstract
Cross-modal hashing retrieval approaches maps heterogeneous multi-modal data into a common hamming space to achieve efficient and flexible retrieval performance. However, existing cross-modal methods mainly exploit feature-level similarity between multi-modal data, the label-level similarity and relative ranking relationship between adjacent instances have been ignored. To address these problems, we propose a novel Deep Rank Cross-modal Hashing(DRCH) method that fully explores the intra-modal semantic similarity relationship. Firstly, DRCH preserves semantic similarity by combining both label-level and feature-level information. Secondly, the inherent gap between modalities are narrowed by proposing a ranking alignment loss function. Finally, the compact and efficient hash codes are optimized from the common semantic space. Extensive experiments on two real-world image-text retrieval datasets demonstrate the superiority of DRCH compared with several state-of-the-art(SOTA) methods.
Huanqiang Zeng, Yifan Shi 0001, Jianqing Zhu, Kai-Kuang Ma
ICASSP2
2022 A Cony-Attention Network for Detecting the Presence of ENF Signal in Short-Duration Audio
abstract
Detecting the presence of the electric network frequency (ENF) signal in audio recordings is a prerequisite of applying the ENF criterion that plays an essential role in numerous forensic applications. However, existing detection methods are powerless to handle short-duration audio recordings that have attracted considerable attention due to the popularity of voice messaging apps. This paper proposes a novel deep learning-based approach for ENF detection in short audio recordings, reducing the minimum operating range of audio duration to 1/10 of the state-of-the-art methods. Meanwhile, a convolutional attention network termed Conv-AttNet is proposed to improve the detection performance of convolutional neural networks (CNN) through the attention mechanism. Experiments on both synthetic and real-world audio recordings reveal that Conv-AttNet is able to detect the ENF signal buried in only 2 seconds of audio recordings, surpassing both matched filtering and typical CNN like ResNet50. In addition, the detection accuracy can be further increased by utilizing audio recordings of longer duration.
Yangfu Li, Xiaodan Lin, Yingqiang Qiu, Huanqiang Zeng
MMSP4
2022 A sample-proxy dual triplet loss function for object re-identification
abstract
Abstract Object re‐identification, such as vehicle re‐identification or pedestrian re‐identification, plays a significant role in intelligent video surveillance systems for public security. Due to viewpoint variations and appearance changes, both pedestrians and vehicles usually have complex intra‐class variations. However, most existing object re‐identification methods often use a sample‐level triplet loss function cooperating with a single‐proxy softmax loss function, which could not handle complex intra‐class variations well. In this paper, a sample‐proxy dual triplet (SPDT) loss function is proposed, which works with a multi‐proxy softmax (MPS) loss function. The MPS loss function is in charge of learning multiple proxies to represent a class. The SPDT loss function is responsible for enlarging inter‐class distances as well as shrinking intra‐class distances on both sample and proxy levels. Therefore, the method not only handles multi‐proxy intra‐class variations but also fully learns discrimination on samples and proxies. Experiments on two large datasets, that is, VeRi776 and DukeMTMC‐reID, demonstrate that the method is superior to state‐of‐the‐art object re‐identification approaches.
Hanxiao Wu, Fei Shen 0004, Jianqing Zhu, Huanqiang Zeng, Xiaobin Zhu 0001, Zhen Lei 0001
IET Image Process.4
2022 An Efficient Multiresolution Network for Vehicle Reidentification
abstract
In general, vehicle images have varying resolutions due to vehicles’ movements and different camera settings. However, most existing vehicle reidentification models are single-resolution deep networks trained with preuniformly resizing vehicle images, which underestimate adverse effects of varying resolutions and lead to unsatisfactory performance. A straightforward solution for dealing with varying resolutions is to train multiple vehicle reidentification models. Each model is independently trained with images of a specific resolution. However, this straightforward solution requires significant overhead and ignores intrinsic associations among different resolution images. For that, an efficient multiresolution network (EMRN) is proposed for vehicle reidentification in this article. First, EMRN embeds a newly designed multiresolution feature dimension uniform module (MR-FDUM) behind a traditional backbone network (i.e., ResNet-50). As a result, the whole model can extract fixed dimensional features from different resolution images so that it can be trained with one loss function of fixed dimensional parameters rather than training multiple models. Second, a multiresolution image randomly feeding strategy is designed to train EMRN, making each minibatch data of a random resolution during the training process. Consequently, EMRN can implicitly learn collaborative multiresolution features via only a unitary deep network. The experiments on three large-scale data sets, i.e., VeRi776, VehicleID, and VRIC, demonstrate that EMRN is superior to state-of-the-art vehicle reidentification methods.
Fei Shen 0004, Jianqing Zhu, Xiaobin Zhu 0001, Jingchang Huang, Huanqiang Zeng, Zhen Lei 0001, Canhui Cai
IEEE Internet Things J.5
2022 Clustering-Guided Pairwise Metric Triplet Loss for Person Reidentification
abstract
Most of the loss functions proposed for person reidentification (Re-ID) are expected to be easy to deploy, efficiently improve network performance, and will not introduce redundant parameters. This study proposes a no-parameter and generic clustering-guided pairwise metric triplet (CPM-Triplet) loss based on the hard sample mining triplet loss for the metric learning loss. CPM-Triplet loss deploys two metrics: 1) the Euclidean metric and 2) the cosine metric, to complementarily improve the metric learning of the model. Paralleled to the Euclidean metric, the cosine metric quantifies the sample similarity in a different way to the Euclidean metric, which takes a different perspective to explore the distribution of samples. But the pairwise metric mainly improves the precision between dissimilar samples of the same label and could not solve the problem of excessive outliers. Therefore, the clustering-guided correction term was deployed to apply to all samples with the same label to mine the similarity in the samples, while weakening the influence of outliers in CPM-Triplet loss. Experiments conducted on four benchmark data sets show that the combination of the CPM-Triplet loss and the widely used Bag-of-Tricks baseline generally outperforms the baseline and numerous state-of-the-art methods studied in this article. The source code would be available athttps://github.com/weiyu-zeng/CPM-Triplet-loss.
Weiyu Zeng, Tianlei Wang, Jiuwen Cao, Huanqiang Zeng
IEEE Internet Things J.5
2022 Proximal-Gen for fast compressed sensing recovery
Yuli Fu 0001, Tao Zhu 0002, Youjun Xiang, Huanqiang Zeng
J. Vis. Commun. Image Represent.5
2022 Spatial-frequency HEVC multiple description video coding with adaptive perceptual redundancy allocation
Feifeng Wang, Jing Chen 0001, Huanqiang Zeng, Canhui Cai
J. Vis. Commun. Image Represent.3
2022 Deep Coarse-to-Fine Dense Light Field Reconstruction With Flexible Sampling and Geometry-Aware Fusion
abstract
A densely-sampled light field (LF) is highly desirable in various applications, such as 3-D reconstruction, post-capture refocusing and virtual reality. However, it is costly to acquire such data. Although many computational methods have been proposed to reconstruct a densely-sampled LF from a sparsely-sampled one, they still suffer from either low reconstruction quality, low computational efficiency, or the restriction on the regularity of the sampling pattern. To this end, we propose a novel learning-based method, which accepts sparsely-sampled LFs with irregular structures, and produces densely-sampled LFs with arbitrary angular resolution accurately and efficiently. We also propose a simple yet effective method for optimizing the sampling pattern. Our proposed method, an end-to-end trainable network, reconstructs a densely-sampled LF in a coarse-to-fine manner. Specifically, the coarse sub-aperture image (SAI) synthesis module first explores the scene geometry from an unstructured sparsely-sampled LF and leverages it to independently synthesize novel SAIs, in which a confidence-based blending strategy is proposed to fuse the information from different input SAIs, giving an intermediate densely-sampled LF. Then, the efficient LF refinement module learns the angular relationship within the intermediate result to recover the LF parallax structure. Comprehensive experimental evaluations demonstrate the superiority of our method on both real-world and synthetic LF images when compared with state-of-the-art methods. In addition, we illustrate the benefits and advantages of the proposed approach when applied in various LF-based applications, including image-based rendering and depth estimation enhancement. The code is available at https://github.com/jingjin25/LFASR-FS-GAF.
Jing Jin 0006, Junhui Hou, Jie Chen 0026, Huanqiang Zeng, Sam Kwong, Jingyi Yu 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Point Cloud Quality Assessment via 3D Edge Similarity Measurement
abstract
In this letter, a new full-reference metric is presented to assess the perceptual quality of the point clouds (PCs). The human visual system (HVS) always shows a high sensitivity to the three-dimensional (3D) edge features inherent in the PCs. With this motivation, the three-dimensional edge similarity-based model (TDESM) is proposed, which makes the first attempt to apply 3D Difference of Gaussian (3D-DOG) on point cloud quality assessment (PCQA). Specifically, the 3D edge features are captured by convolving the dual-scale 3D-DOG filters with both reference and distorted PCs. The quality scores of distorted PCs are generated by combining the 3D edge similarity measured from different scales. The experiments are conducted on four publicly available PCQA datasets, i.e., Torlig2018, M-PCCD, ICIP2020, and SJTU-PCQA. Compared with multiple state-of-the-art PCQA metrics, our proposed approach is able to be higher consistent with the subjective perception on the PCs.
Zian Lu, Hailiang Huang 0002, Huanqiang Zeng, Junhui Hou, Kai-Kuang Ma
IEEE Signal Process. Lett.3
2022 Joint Depth and Density Guided Single Image De-Raining
abstract
Single image de-raining is an important and highly challenging problem. To address this problem, some depth or density guided single-image de-raining methods have been developed with encouraging performance. However, these methods individually use the depth or the density to guide the network to conduct image de-raining. In this paper, a noveljoint depth and density guided de-raining(JDDGD) method is technically developed. The JDDGD starts with adepth-density inference network(DDINet) to extract the depth and density information from an input rainy image, followed by adepth-density-basedconditional generative adversarial network (DD-CGAN) to exploit the depth and density information provided by DDINet to achieve adaptive rain streak and fog removal. To prevent the spatially-varying local artifacts, an effectiveglobal-local discriminatorsstructure is introduced in the proposed DD-CGAN to globally and locally inspect the generated images. In addition, multiple loss functions includingmulti-scale pixel loss,multi-scale perceptual loss, andglobal-local generative adversarial lossare also jointly used to train our model to achieve the best performance. Both quantitative and qualitative results show that the proposed JDDGD method achieves superior performance than previousnon-guided,density-guided, anddepth-guided de-rainingmethods.
Yuli Fu 0001, Tao Zhu 0002, Youjun Xiang, Huanqiang Zeng
IEEE Trans. Circuits Syst. Video Technol.6
2022 A Hybrid Compression Framework for Color Attributes of Static 3D Point Clouds
abstract
The emergence of 3D point clouds (3DPCs) is promoting the rapid development of immersive communication, autonomous driving, and so on. Due to the huge data volume, the compression of 3DPCs is becoming more and more attractive. We propose a novel and efficient color attribute compression method for static 3DPCs. First, a 3DPC is partitioned into several sub-point clouds by color distribution analysis. Each sub-point cloud is then decomposed into a lot of 3D blocks by an improved k-d tree-based decomposition algorithm. Afterwards, a novel virtual adaptive sampling-based sparse representation strategy is proposed for each 3D block to remove the redundancy among points, in which the bases of the graph transform (GT) and the discrete cosine transform (DCT) are used as candidates of the complete dictionary. Experimental results over 10 common 3DPCs demonstrate that the proposed method can achieve superior or comparable coding performance when compared with the current state-of-the-art methods.
Hao Liu 0044, Hui Yuan 0001, Qi Liu 0029, Junhui Hou, Huanqiang Zeng, Sam Kwong
IEEE Trans. Circuits Syst. Video Technol.5
2022 High-Capacity Framework for Reversible Data Hiding in Encrypted Image Using Pixel Prediction and Entropy Encoding
abstract
While the existing reserving room before encryption (RRBE) based reversible data hiding in encrypted image (RDHEI) schemes can achieve decent embedding capacity, the capacity of the existing vacating room by encryption (VRBE) based schemes is relatively low. To address this issue, this paper proposes a generalized framework for high-capacity RDHEI for both the RRBE and VRBE cases. First, an efficient embedding room generation algorithm (ERGA) is designed to produce large embedding room using pixel prediction and entropy encoding. Then, we propose two RDHEI schemes, one for RRBE, another for VRBE. In the RRBE scenario, the image owner generates the embedding room with ERGA and encrypts the preprocessed image using stream cipher with two encryption keys. Then, the data hider locates the embedding room and embeds the additional encrypted data. In the VRBE scenario, the cover image is encrypted by an improved block modulation and permutation encryption algorithm, where the spatial redundancy in the plain-text image is greatly preserved. Then, the data hider applies ERGA on the encrypted image to generate the embedding room and conducts data embedding. For both schemes, receivers with different authentication keys can conduct either error-free data extraction or error-free image recovery. The experimental results show that the two proposed schemes outperform many state-of-the-art RDHEI schemes. Besides, they can ensure high security level, where the original image can be hardly discovered from the encrypted version before or after data hiding by unauthorized users.
Yingqiang Qiu, Qichao Ying, Yuyan Yang, Huanqiang Zeng, Sheng Li 0006, Zhenxing Qian
IEEE Trans. Circuits Syst. Video Technol.4
2022 Deep Posterior Distribution-Based Embedding for Hyperspectral Image Super-Resolution
abstract
In this paper, we investigate the problem of hyperspectral (HS) image spatial super-resolution via deep learning. Particularly, we focus on how to embed the high-dimensional spatial-spectral information of HS images efficiently and effectively. Specifically, in contrast to existing methods adopting empirically-designed network modules, we formulate HS embedding as an approximation of the posterior distribution of a set of carefully-defined HS embedding events, including layer-wise spatial-spectral feature extraction and network-level feature aggregation. Then, we incorporate the proposed feature embedding scheme into a source-consistent super-resolution framework that is physically-interpretable, producing PDE-Net, in which high-resolution (HR) HS images are iteratively refined from the residuals between input low-resolution (LR) HS images and pseudo-LR-HS images degenerated from reconstructed HR-HS images via probability-inspired HS embedding. Extensive experiments over three common benchmark datasets demonstrate that PDE-Net achieves superior performance over state-of-the-art methods. Besides, the probabilistic characteristic of this kind of networks can provide the epistemic uncertainty of the network outputs, which may bring additional benefits when used for other HS image-based applications. The code will be publicly available at https://github.com/jinnh/PDE-Net.
Jinhui Hou, Junhui Hou, Huanqiang Zeng, Jinjian Wu, Jiantao Zhou 0001
IEEE Trans. Image Process.4
2022 A Spatial and Geometry Feature-Based Quality Assessment Model for the Light Field Images
abstract
This paper proposes a new full-reference image quality assessment (IQA) model for performing perceptual quality evaluation on light field (LF) images, called the spatial and geometry feature-based model (SGFM). Considering that the LF image describe both spatial and geometry information of the scene, the spatial features are extracted over the sub-aperture images (SAIs) by using contourlet transform and then exploited to reflect the spatial quality degradation of the LF images, while the geometry features are extracted across the adjacent SAIs based on 3D-Gabor filter and then explored to describe the viewing consistency loss of the LF images. These schemes are motivated and designed based on the fact that the human eyes are more interested in the scale, direction, contour from the spatial perspective and viewing angle variations from the geometry perspective. These operations are applied to the reference and distorted LF images independently. The degree of similarity can be computed based on the above-measured quantities for jointly arriving at the final IQA score of the distorted LF image. Experimental results on three commonly-used LF IQA datasets show that the proposed SGFM is more in line with the quality assessment of the LF images perceived by the human visual system (HVS), compared with multiple classical and state-of-the-art IQA models.
Hailiang Huang 0002, Huanqiang Zeng, Junhui Hou, Jing Chen 0001, Jianqing Zhu, Kai-Kuang Ma
IEEE Trans. Image Process.2
2022 Screen Content Video Quality Assessment Model Using Hybrid Spatiotemporal Features
abstract
In this paper, a full-reference video quality assessment (VQA) model is designed for the perceptual quality assessment of the screen content videos (SCVs), called the hybrid spatiotemporal feature-based model (HSFM). The SCVs are of hybrid structure including screen and natural scenes, which are perceived by the human visual system (HVS) with different visual effects. With this consideration, the three dimensional Laplacian of Gaussian (3D-LOG) filter and three dimensional Natural Scene Statistics (3D-NSS) are exploited to extract the screen and natural spatiotemporal features, based on the reference and distorted SCV sequences separately. The similarities of these extracted features are then computed independently, followed by generating the distorted screen and natural quality scores for screen and natural scenes. After that, an adaptive screen and natural quality fusion scheme through the local video activity is developed to combine them for arriving at the final VQA score of the distorted SCV under evaluation. The experimental results on the Screen Content Video Database (SCVD) and Compressed Screen Content Video Quality (CSCVQ) databases have shown that the proposed HSFM is more in line with the perceptual quality assessment of the SCVs perceived by the HVS, compared with a variety of classic and latest IQA/VQA models.
Huanqiang Zeng, Hailiang Huang 0002, Junhui Hou, Jiuwen Cao, Yongtao Wang, Kai-Kuang Ma
IEEE Trans. Image Process.1
2021 Object Re-identification Using Teacher-Like and Light Students
Hanxiao Wu, Fei Shen 0004, Jianqing Zhu, Huanqiang Zeng
BMVC5
2021 Semantic-embedded Unsupervised Spectral Reconstruction from Single RGB Images in the Wild
abstract
This paper investigates the problem of reconstructing hyperspectral (HS) images from single RGB images captured by commercial cameras, without using paired HS and RGB images during training. To tackle this challenge, we propose a new lightweight and end-to-end learning-based framework. Specifically, on the basis of the intrinsic imaging degradation model of RGB images from HS images, we progressively spread the differences between input RGB images and re-projected RGB images from recovered HS images via effective unsupervised camera spectral response function estimation. To enable the learning without paired ground-truth HS images as supervision, we adopt the adversarial learning manner and boost it with a simple yet effective ℒ1gradient clipping scheme. Besides, we embed the semantic information of input RGB images to locally regularize the unsupervised learning, which is expected to promote pixels with identical semantics to have consistent spectral signatures. In addition to conducting quantitative experiments over two widely-used datasets for HS image reconstruction from synthetic RGB images, we also evaluate our method by applying recovered HS images from real RGB images to HS-based visual tracking. Extensive results show that our method significantly outperforms state-of-the-art unsupervised methods and even exceeds the latest supervised method under some settings. The source code is public available at https://github.com/zbzhzhy/Unsupervised-Spectral-Reconstruction.
Hui Liu 0032, Junhui Hou, Huanqiang Zeng, Qingfu Zhang 0001
ICCV4
2021 Learning Spatial-angular Fusion for Compressive Light Field Imaging in a Cycle-consistent Framework
abstract
This paper investigates the 4-D light field (LF) reconstruction from 2-D measurements captured by the coded aperture camera. To tackle such an ill-posed inverse problem, we propose a cycle-consistent reconstruction network (CR-Net). To be specific, based on the intrinsic linear imaging model of the coded aperture, CR-Net reconstructs an LF through progressively eliminating the residuals between the projected measurements from the reconstructed LF and input measurements. Moreover, to address the crucial issue of extracting representative features from high-dimensional LF data efficiently and effectively, we formulate the problem in a probability space and propose to approximate a posterior distribution of a set of carefully-defined LF processing events, including both layer-wise spatial-angular feature extraction and network-level feature aggregation. Through droppath from a densely-connected template network, we derive an adaptively learned spatial-angular fusion strategy, which is sharply contrasted with existing manners that combine spatial and angular features empirically. Extensive experiments on both simulated measurements and measurements by a real coded aperture camera demonstrate the significant advantage of our method over state-of-the-art ones, i.e., our method improves the reconstruction quality by 4.5 dB.
Xianqiang Lyu, Mantang Guo, Jing Jin 0006, Junhui Hou, Huanqiang Zeng
ACM Multimedia6
2021 A spatial structural similarity triplet loss for auxiliary vehicle re-identification
Jianqing Zhu, Liu Liu 0014, Xiaobin Zhu 0001, Huanqiang Zeng
Sci. China Inf. Sci.4
2021 Cascading Scene and Viewpoint Feature Learning for Pedestrian Gender Recognition
abstract
Pedestrian gender recognition plays an important role in smart city. To effectively improve the pedestrian gender recognition performance, a new method, called cascading scene and viewpoint feature learning (CSVFL), is proposed in this article. The novelty of the proposed CSVFL lies on the joint consideration of two crucial challenges in pedestrian gender recognition, namely, scene and viewpoint variation. For that, the proposed CSVFL starts with the scene transfer (ST) scheme, followed by the viewpoint adaptation (VA) scheme in a cascading manner. Specifically, the ST scheme exploits the key pedestrian segmentation network to extract the key pedestrian masks for the subsequent key pedestrian transfer generative adversarial network, with the goal of encouraging the input pedestrian image to have the similar style to the target scene while preserving the image details of the key pedestrian as much as possible. Afterward, the obtained scene-transferred pedestrian images are fed to train the deep feature learning network with the VA scheme, in which each neuron will be enabled/disabled for different viewpoints depending on whether it has contribution on the corresponding viewpoint. Extensive experiments conducted on the commonly used pedestrian attribute data sets have demonstrated that the proposed CSVFL approach outperforms multiple recently reported pedestrian gender recognition methods.
Huanqiang Zeng, Jianqing Zhu, Jiuwen Cao, Yongtao Wang, Kai-Kuang Ma
IEEE Internet Things J.2
2021 A fast algorithm based on gray level co-occurrence matrix and Gabor feature for HEVC screen content coding
Jing Chen 0001, Jianshan Ou, Huanqiang Zeng, Canhui Cai
J. Vis. Commun. Image Represent.3
2021 A Light Field Image Quality Assessment Model Based on Symmetry and Depth Features
abstract
This paper presents a new full-reference image quality assessment (IQA) method for conducting the perceptual quality evaluation of the light field (LF) images, called the symmetry and depth feature-based model (SDFM). Specifically, the radial symmetry transform is first employed on the luminance components of the reference and distorted LF images to extract their symmetry features for capturing the spatial quality of each view of an LF image. Second, the depth feature extraction scheme is designed to explore the geometry information inherited in an LF image for modeling its LF structural consistency across views. The similarity measurements are subsequently conducted on the comparison of their symmetry and depth features separately, which are further combined to achieve the quality score for the distorted LF image. Note that the proposed SDFM that explores the symmetry and depth features is conformable to the human vision system, which identifies the objects by sensing their structures and geometries. Extensive simulation results on the dense light fields dataset have clearly shown that the proposed SDFM outperforms multiple classical and recently developed IQA algorithms on quality evaluation of the LF images.
Huanqiang Zeng, Junhui Hou, Jing Chen 0001, Jianqing Zhu, Kai-Kuang Ma
IEEE Trans. Circuits Syst. Video Technol.2
2021 Hyperspectral Image Super-Resolution via Deep Progressive Zero-Centric Residual Learning
abstract
This paper explores the problem of hyperspectral image (HSI) super-resolution that merges a low resolution HSI (LR-HSI) and a high resolution multispectral image (HR-MSI). The cross-modality distribution of the spatial and spectral information makes the problem challenging. Inspired by the classic wavelet decomposition-based image fusion, we propose a novel lightweight deep neural network-based framework, namely progressive zero-centric residual network (PZRes-Net), to address this problem efficiently and effectively. Specifically, PZRes-Net learns a high resolution and zero-centric residual image, which contains high-frequency spatial details of the scene across all spectral bands, from both inputs in a progressive fashion along the spectral dimension. And the resulting residual image is then superimposed onto the up-sampled LR-HSI in a mean-value invariant manner, leading to a coarse HR-HSI, which is further refined by exploring the coherence across all spectral bands simultaneously. To learn the residual image efficiently and effectively, we employ spectral-spatial separable convolution with dense connections. In addition, we propose zero-mean normalization implemented on the feature maps of each layer to realize the zero-mean characteristic of the residual image. Extensive experiments over both real and synthetic benchmark datasets demonstrate that our PZRes-Net outperforms state-of-the-art methods to a significant extent in terms of both 4 quantitative metrics and visual quality, e.g., our PZRes-Net improves the PSNR more than 3dB, while saving 2.3× parameters and consuming 15× less FLOPs. The code is publicly available at https://github.com/zbzhzhy/PZRes-Net.
Junhui Hou, Jie Chen 0026, Huanqiang Zeng, Jiantao Zhou 0001
IEEE Trans. Image Process.4
2021 Maximum Correntropy Criterion-Based Hierarchical One-Class Classification
abstract
Due to the effectiveness of anomaly/outlier detection, one-class algorithms have been extensively studied in the past. The representatives include the shallow-structure methods and deep networks, such as the one-class support vector machine (OC-SVM), one-class extreme learning machine (OC-ELM), deep support vector data description (Deep SVDD), and multilayer OC-ELM (ML-OCELM/MK-OCELM). However, existing algorithms are generally built on the minimum mean-square-error (mse) criterion, which is robust to the Gaussian noises but less effective in dealing with large outliers. To alleviate this deficiency, a robust maximum correntropy criterion (MCC)-based OC-ELM (MC-OCELM) is first proposed and then further extended to a hierarchical network to enhance its capability in characterizing complex and large data (named HC-OCELM). The gradient derivation combining with a fixed-point iterative updation scheme is adopted for the output weight optimization. Experiments on many benchmark data sets are conducted for effectiveness validation. Comparisons to many state-of-the-art approaches are provided for the superiority demonstration.
Jiuwen Cao, Haozhen Dai, Bai Ying Lei, Chun Yin, Huanqiang Zeng, Anton Kummert
IEEE Trans. Neural Networks Learn. Syst.5
2020 Fast compressed sensing recovery using generative models and sparse deviations modeling
abstract
This paper develops an algorithm to effectively explore the advantages of both sparse vector recovery methods and generative model-based recovery methods for solving compressed sensing recovery problem. The proposed algorithm mainly consists of two steps. In the first step, a network-based projected gradient descent (NPGD) is introduced to solve a non-convex optimization problem, obtaining a preliminary recovery of the original signal. Then with the obtained preliminary recovery, a l1norm regularized optimization problem is solved by optimizing for sparse deviation vectors. Experimental results on two bench-mark datasets for image compressed sensing clearly demonstrate that the proposed recovery algorithm can bring about high computation speed, while decreasing the reconstruction error continuously with increasing the number of measurements.
Yuli Fu 0001, Youjun Xiang, Tao Zhu 0002, Huanqiang Zeng
VCIP6
2020 Learning Matching Behavior Differences for Compressing Vehicle Re-identification Models
abstract
Vehicle re-identification matching vehicles captured by different cameras has great potential in the field of public security. However, recent vehicle re-identification approaches exploit complex networks, causing large computations in their testing phases. In this paper, we propose a matching behavior difference learning (MBDL) method to compress vehicle re-identification models for saving testing computations. In order to represent the matching behavior evolution across two different layers of a deep network, a matching behavior difference (MBD) matrix is designed. Then, our MBDL method minimizes the L1 loss function among MBD matrixes from a small student network and a complex teacher network, ensuring the student network use less computations to simulate the teacher network's matching behaviors. During the testing phase, only the small student network is utilized so that testing computations can be significantly reduced. Experiments on VeRi776 and VehicleID datasets show that MBDL outperforms many state-of-the-art approaches in terms of accuracy and testing time performance.
Jianqing Zhu, Huanqiang Zeng, Canhui Cai, Lixin Zheng
VCIP3
2020 Object Reidentification via Joint Quadruple Decorrelation Directional Deep Networks in Smart Transportation
abstract
Object reidentification with the goal of matching pedestrian or vehicle images captured from different camera viewpoints is of considerable significance to public security. Quadruple directional deep learning features (QD-DLFs) can comprehensively describe object images. However, the correlation among QD-DLFs is an unavoidable problem, since QD-DLFs are learned with quadruple independent directional deep networks (QIDDNs) driven with the same training data, and each network holds the same basic deep feature learning architecture (BDFLA). The correlation among QD-DLFs is harmful to the complementarity of QD-DLFs, restricting the object reidentification performance. For that, we propose joint quadruple decorrelation directional deep networks (JQD3Ns) to reduce the correlation among the learned QD-DLFs. In order to jointly train JQD3Ns, besides the softmax loss functions, a parameter correlation cost function is proposed to indirectly reduce the correlation among QD-DLFs by enlarging the dissimilarity among the parameters of JQD3Ns. Extensive experiments on three publicly available large-scale data sets demonstrate that the proposed JQD3Ns approach is superior to multiple state-of-the-art object reidentification methods.
Jianqing Zhu, Jingchang Huang, Huanqiang Zeng, Xiaoqing Ye, Baoqing Li, Zhen Lei 0001, Lixin Zheng
IEEE Internet Things J.3
2020 Body Symmetry and Part-Locality-Guided Direct Nonparametric Deep Feature Enhancement for Person Reidentification
abstract
In recent years, deep learning (DL) has been successfully and widely applied in the person reidentification (Re-ID). However, the DL-based person Re-ID methods face a bottleneck that the scales of most existing person Re-ID databases are not large enough for training very deep models. To address this problem, a body symmetry and part-locality-guided direct nonparametric deep feature enhancement (DNDFE) method is proposed in this article. Based on the observation that the body symmetry and part locality are two important appearance properties inherited in the upright walking persons, the proposed method designs two nonparametric layers, namely, the body symmetry average pooling and local normalization layers, to construct a DNDFE module to well explore the body symmetry and part locality properties. The proposed DNDFE module could be directly embedded between the traditional deep feature learning module and similarity learning module to enhance the DL features so as to improve the person Re-ID performance. The experimental results have shown that the proposed DNDFE method is superior to multiple state-of-the-art person Re-ID methods in terms of accuracy and efficiency.
Jianqing Zhu, Huanqiang Zeng, Jingchang Huang, Xiaobin Zhu 0001, Zhen Lei 0001, Canhui Cai, Lixin Zheng
IEEE Internet Things J.2
2020 Joint Pyramid Feature Representation Network for Vehicle Re-identification
Xiangwei Lin, Huanqiang Zeng, Jinhui Hou, Jiuwen Cao, Jianqing Zhu, Jing Chen 0001
Mob. Networks Appl.2
2020 Reversible data hiding in encrypted images using adaptive reversible integer transformation
Yingqiang Qiu, Zhenxing Qian, Huanqiang Zeng, Xiaodan Lin, Xinpeng Zhang 0001
Signal Process.3
2020 3D Point Cloud Attribute Compression via Graph Prediction
abstract
3D point clouds associated with attributes are considered as a promising data representation for immersive communication. The large amount of data, however, poses great challenges to the subsequent transmission and storage processes. In this letter, we propose a new compression scheme for the color attribute of static voxelized 3D point clouds. Specifically, we first partition the colors of a 3D point cloud into clusters by applying k-d tree to the geometry information, which are then successively encoded. To eliminate the redundancy, we propose a novel prediction module, namely graph prediction, in which a small number of representative points selected from previously encoded clusters are used to predict the points to be encoded by exploring the underlying graph structure constructed from the geometry information. Furthermore, the prediction residuals are transformed with the graph transform, and the resulting transform coefficients are finally uniformly quantified and entropy encoded. Experimental results show that the proposed compression scheme is able to achieve better rate-distortion performance at a lower computational cost when compared with state-of-the-art methods.
Shuai Gu, Junhui Hou, Huanqiang Zeng, Hui Yuan 0001
IEEE Signal Process. Lett.3
2020 Screen Content Video Quality Assessment: Subjective and Objective Study
abstract
In this paper, we make the first attempt to study the subjective and objective quality assessment for the screen content videos (SCVs). For that, we construct the first large-scale video quality assessment (VQA) database specifically for the SCVs, called the screen content video database (SCVD). This SCVD provides 16 reference SCVs, 800 distorted SCVs, and their corresponding subjective scores, and it is made publicly available for research usage. The distorted SCVs are generated from each reference SCV with 10 distortion types and 5 degradation levels for each distortion type. Each distorted SCV is rated by at least 32 subjects in the subjective test. Furthermore, we propose the first full-reference VQA model for the SCVs, called the spatiotemporal Gabor feature tensor-based model (SGFTM), to objectively evaluate the perceptual quality of the distorted SCVs. This is motivated by the observation that 3D-Gabor filter can well stimulate the visual functions of the human visual system (HVS) on perceiving videos, being more sensitive to the edge and motion information that are often-encountered in the SCVs. Specifically, the proposed SGFTM exploits 3D-Gabor filter to individually extract the spatiotemporal Gabor feature tensors from the reference and distorted SCVs, followed by measuring their similarities and later combining them together through the developed spatiotemporal feature tensor pooling strategy to obtain the final SGFTM score. Experimental results on SCVD have shown that the proposed SGFTM yields a high consistency on the subjective perception of SCV quality and consistently outperforms multiple classical and state-of-the-art image/video quality assessment models.
Shan Cheng, Huanqiang Zeng, Jing Chen 0001, Junhui Hou, Jianqing Zhu, Kai-Kuang Ma
IEEE Trans. Image Process.2
2020 3D Point Cloud Attribute Compression Using Geometry-Guided Sparse Representation
abstract
3D point clouds associated with attributes are considered as a promising paradigm for immersive communication. However, the corresponding compression schemes for this media are still in the infant stage. Moreover, in contrast to conventional image/video compression, it is a more challenging task to compress 3D point cloud data, arising from the irregular structure. In this paper, we propose a novel and effective compression scheme for the attributes of voxelized 3D point clouds. In the first stage, an input voxelized 3D point cloud is divided into blocks of equal size. Then, to deal with the irregular structure of 3D point clouds, a geometry-guided sparse representation (GSR) is proposed to eliminate the redundancy within each block, which is formulated as an ℓ0-norm regularized optimization problem. Also, an inter-block prediction scheme is applied to remove the redundancy between blocks. Finally, by quantitatively analyzing the characteristics of the resulting transform coefficients by GSR, an effective entropy coding strategy that is tailored to our GSR is developed to generate the bitstream. Experimental results over various benchmark datasets show that the proposed compression scheme is able to achieve better rate-distortion performance and visual quality, compared with state-of-the-art methods.
Shuai Gu, Junhui Hou, Huanqiang Zeng, Hui Yuan 0001, Kai-Kuang Ma
IEEE Trans. Image Process.3
2020 Color Image Demosaicing Using Progressive Collaborative Representation
abstract
In this paper, a progressive collaborative representation (PCR) framework is proposed that is able to incorporate any existing color image demosaicing method for further boosting its demosaicing performance. Our PCR consists of two phases: (i) offline training and (ii) online refinement. In phase (i), multiple training-and-refining stages will be performed. In each stage, a new dictionary will be established through the learning of a large number of feature-patch pairs, extracted from the demosaicked images of the current stage and their corresponding original full-color images. After training, a projection matrix will be generated and exploited to refine the current demosaicked image. The updated image with improved image quality will be used as the input for the next training-and-refining stage and performed the same processing likewise. At the end of phase (i), all the projection matrices generated as above-mentioned will be exploited in phase (ii) to conduct online demosaicked image refinement of the test image. Extensive simulations conducted on two commonly-used test datasets (i.e., the IMAX and Kodak) for evaluating the demosaicing algorithms have clearly demonstrated that our proposed PCR framework is able to constantly boost the performance of any image demosaicing method we experimented, in terms of the objective and subjective performance evaluations.
Zhangkai Ni, Kai-Kuang Ma, Huanqiang Zeng, Baojiang Zhong
IEEE Trans. Image Process.3
2020 Light Field Image Quality Assessment via the Light Field Coherence
abstract
In this paper, a novel full-referenceimage quality assessment(IQA) method for evaluating the quality of the distortedlight field(LF) image against its reference LF image is proposed, called thelog-Gabor feature-basedlight field coherence (LGF-LFC). Based on the fact that to compare two LF images, it essentially boils down to measure howcoherentof these two LF images, we attempt to measure the degree of their LFcoherence(LFC). To pursue this goal, the salient features from the reference and distorted LF images under comparison need to be extracted. By considering that the Gabor feature has the ability to well characterize thehuman visual system(HVS) perception, and the special characteristics of the LF images, themulti-scale andsingle-scale Gabor feature extraction schemes are developed to extract the multi-scale log-Gabor features from thesub-aperture images(SAIs) and the single-scale log-Gabor feature from theepi-polar images(EPIs), respectively. Note that the former can reflect the image details (via the SAIs), while the latter indicates the viewing consistency (via the EPI’s depth information). The similarity measurements are subsequently conducted on the comparison of their SAIs and that of their EPIs separately, followed by combining them together for arriving at the final score. Extensive simulation results have clearly demonstrated that the proposed LGF-LFC is more consistent with the perception of the HVS on the quality evaluation of the LF images than multiple classical and state-of-the-art IQA methods.
Huanqiang Zeng, Junhui Hou, Jing Chen 0001, Kai-Kuang Ma
IEEE Trans. Image Process.2
2020 Vehicle Re-Identification Using Quadruple Directional Deep Learning Features
abstract
In order to resist the adverse effect of viewpoint variations, we design quadruple directional deep learning networks to extract quadruple directional deep learning features (QD-DLF) of vehicle images for improving vehicle re-identification performance. The quadruple directional deep learning networks are of similar overall architecture, including the same basic deep learning architecture but different directional feature pooling layers. Specifically, the same basic deep learning architecture that is a shortly and densely connected convolutional neural network is utilized to extract the basic feature maps of an input square vehicle image in the first stage. Then, the quadruple directional deep learning networks utilize different directional pooling layers, i.e., horizontal average pooling layer, vertical average pooling layer, diagonal average pooling layer, and anti-diagonal average pooling layer, to compress the basic feature maps into horizontal, vertical, diagonal, and anti-diagonal directional feature maps, respectively. Finally, these directional feature maps are spatially normalized and concatenated together as a quadruple directional deep learning feature for vehicle re-identification. The extensive experiments on both VeRi and VehicleID databases show that the proposed QD-DLF approach outperforms multiple state-of-the-art vehicle re-identification methods.
Jianqing Zhu, Huanqiang Zeng, Jingchang Huang, Shengcai Liao, Zhen Lei 0001, Canhui Cai, Lixin Zheng
IEEE Trans. Intell. Transp. Syst.2
2019 A Log-Gabor Feature-Based Quality Assessment Model for Screen Content Images
abstract
In this paper, an image quality assessment (IQA) model for conducting objective evaluations of screen content images (S-CIs) is proposed, called the log-Gabor feature-based model (LGFM). From the standpoint of signal representation, the log-Gabor filters outperform the classical Gabor filters since the outputs of the log-Gabor filters are more consistent with the perception of visual cortex in human visual system (HVS). Furthermore, the following two remarkable characteristics of the log-Gabor filters are highly beneficial to develop a more accurate IQA model; i.e., (i) zero response at the DC, and (ii) stronger response at high frequencies. In our proposed L-GFM, the log-Gabor filters are used to extract features from the luminance of the reference SCIs and that of the distorted SCIs for measuring their degree of similarity. Together with the measurements from the other two chrominance components, the final LGFM score will be arrived at the output of the pooling stage. Extensive simulation results have shown that our proposed LGFM is highly consistent with the human perception, compared to other state-of-the-art IQA models.
Kai-Kuang Ma, Huanqiang Zeng
ICIP3
2019 Joint horizontal and vertical deep learning feature for vehicle re-identification
Jianqing Zhu, Huanqiang Zeng, Yongzhao Du, Lixin Zheng, Canhui Cai
Sci. China Inf. Sci.2
2019 Multi-label learning with multi-label smoothing regularization for vehicle re-identification
Jinhui Hou, Huanqiang Zeng, Jianqing Zhu, Jing Chen 0001, Kai-Kuang Ma
Neurocomputing2
2019 UHD Video Coding: A Light-Weight Learning-Based Fast Super-Block Approach
abstract
The ultra high-definition (UHD) video format, which has recently become popular, aims to provide high spatial resolution, high temporal frame rate, high sample bit-depth, and wide pixel color gamut. Despite the continued development of global network capacities, it inevitably causes the increased bandwidth cost of catering to the requirement of delivering UHD video services. To address such challenges, this paper presents an improved super coding unit (SCU) method for UHD video coding in High Efficiency Video Coding (HEVC). Initially, the medium coding unit (MCU) is proposed to avoid unnecessary brute-force coding unit (CU) partitions of SCU. Furthermore, the SCU is proposed to be encoded by Direct-MCU and SCU-to-MCU modes: the Direct-MCU mode is intended to better adapt to the texture-rich region, which guarantees the compression efficiency by avoiding extra-size CU partition; the SCU-to-MCU mode is designed for the homogeneous region of UHD content, which saves the encoding time by skipping fine-grained CU partition search. Moreover, a learning-based fast SCU decision approach is proposed to speed up the determination process of Direct-MCU and SCU-to-MCU, where three representative handcrafted features are extracted. Experimental results show that our method achieves an affordable complexity and excellent coding efficiency (up to 7.30% Bjøntegaard Delta rate savings) in UHD video coding compared to recent HEVC reference software.
Miaohui Wang, Wuyuan Xie, Xiandong Meng, Huanqiang Zeng, King Ngi Ngan
IEEE Trans. Circuits Syst. Video Technol.4
2019 Statistical Early Termination and Early Skip Models for Fast Mode Decision in HEVC INTRA Coding
abstract
In this article, statistical Early Termination (ET) and Early Skip (ES) models are proposed for fast Coding Unit (CU) and prediction mode decision in HEVC INTRA coding, in which three categories of ET and ES sub-algorithms are included. First, the CU ranges of the current CU are recursively predicted based on the texture and CU depth of the spatial neighboring CUs. Second, the statistical model based ET and ES schemes are proposed and applied to optimize the CU and INTRA prediction mode decision, in which the coding complexities over different decision layers are jointly minimized subject to acceptable rate-distortion degradation. Third, the mode correlations among the INTRA prediction modes are exploited to early terminate the full rate-distortion optimization in each CU decision layer. Extensive experiments are performed to evaluate the coding performance of each sub-algorithm and the overall algorithm. Experimental results reveal that the overall proposed algorithm can achieve 45.47% to 74.77%, and 58.09% on average complexity reduction, while the overall Bjøntegaard delta bit rate increase and Bjøntegaard delta peak signal-to-noise ratio degradation are 2.29% and −0.11 dB, respectively.
Yun Zhang 0002, Na Li 0015, Sam Kwong, Gangyi Jiang, Huanqiang Zeng
ACM Trans. Multim. Comput. Commun. Appl.5
2018 A Shortly and Densely Connected Convolutional Neural Network for Vehicle Re-identification
abstract
In this paper, we propose a shortly and densely connected convolutional neural network (SDC-CNN) for vehicle re-identification. The proposed SDC-CNN mainly consists of short and dense units (SDUs), necessary pooling and normalization layers. The main contribution lies at the design of short and dense connection mechanism, which would effectively improve the feature learning ability. Specifically, in the proposed short and dense connection mechanism, each SDU contains a short list of densely connected convolutional layers and each convolutional layer is of the same appropriate channels. Consequently, the number of connections and the input channel of each convolutional layer are limited in each SDU, and the architecture of SDC-CNN is simple. Extensive experiments on both VeRi and VehicleID datasets show that the proposed SDC-CNN is obviously superior to multiple state-of-the-art vehicle re-identification methods.
Jianqing Zhu, Huanqiang Zeng, Zhen Lei 0001, Shengcai Liao, Lixin Zheng, Canhui Cai
ICPR2
2018 A multi-order derivative feature-based quality assessment model for light field image
Huanqiang Zeng, Jing Chen 0001, Jianqing Zhu, Kai-Kuang Ma
J. Vis. Commun. Image Represent.2
2018 A multi-scale contrast-based image quality assessment model for multi-exposure image fusion
Huanqiang Zeng, Jing Chen 0001, Jianqing Zhu, Junhui Hou
Signal Process.3
2018 Screen Content Image Quality Assessment Using Multi-Scale Difference of Gaussian
abstract
In this paper, a novel image quality assessment (IQA) model for the screen content images (SCIs) is proposed by using multi-scale difference of Gaussian (MDOG). Motivated by the observation that the human visual system (HVS) is sensitive to the edges while the image details can be better explored in different scales, the proposed model exploits MDOG to effectively characterize the edge information of the reference and distorted SCIs at two different scales, respectively. Then, the degree of edge similarity is measured in terms of the smaller-scale edge map. Finally, the edge strength computed based on the larger-scale edge map is used as the weighting factor to generate the final SCI quality score. Experimental results have shown that the proposed IQA model for the SCIs produces high consistency with human perception of the SCI quality and outperforms the state-of-the-art quality models.
Ying Fu 0004, Huanqiang Zeng, Lin Ma 0002, Zhangkai Ni, Jianqing Zhu, Kai-Kuang Ma
IEEE Trans. Circuits Syst. Video Technol.2
2018 Deep Hybrid Similarity Learning for Person Re-Identification
abstract
Person re-identification (Re-ID) aims to match person images captured from two non-overlapping cameras. In this paper, a deep hybrid similarity learning (DHSL) method for person Re-ID based on a convolution neural network (CNN) is proposed. In our approach, a light CNN learning feature pair for the input image pair is simultaneously extracted. Then, both the elementwise absolute difference and multiplication of the CNN learning feature pair are calculated. Finally, a hybrid similarity function is designed to measure the similarity between the feature pair, which is realized by learning a group of weight coefficients to project the elementwise absolute difference and multiplication into a similarity score. Consequently, the proposed DHSL method is able to reasonably assign complexities of feature learning and metric learning in a CNN, so that the performance of person Re-ID is improved. Experiments on three challenging person Re-ID databases, QMUL GRID, VIPeR, and CUHK03, illustrate that the proposed DHSL method is superior to multiple state-of-the-art person Re-ID methods.
Jianqing Zhu, Huanqiang Zeng, Shengcai Liao, Zhen Lei 0001, Canhui Cai, Lixin Zheng
IEEE Trans. Circuits Syst. Video Technol.2
2018 A Gabor Feature-Based Quality Assessment Model for the Screen Content Images
abstract
In this paper, an accurate and efficient full-reference image quality assessment (IQA) model using the extracted Gabor features, called Gabor feature-based model (GFM), is proposed for conducting objective evaluation of screen content images (SCIs). It is well-known that the Gabor filters are highly consistent with the response of the human visual system (HVS), and the HVS is highly sensitive to the edge information. Based on these facts, the imaginary part of the Gabor filter that has odd symmetry and yields edge detection is exploited to the luminance of the reference and distorted SCI for extracting their Gabor features, respectively. The local similarities of the extracted Gabor features and two chrominance components, recorded in the LMN color space, are then measured independently. Finally, the Gabor-feature pooling strategy is employed to combine these measurements and generate the final evaluation score. Experimental simulation results obtained from two large SCI databases have shown that the proposed GFM model not only yields a higher consistency with the human perception on the assessment of SCIs but also requires a lower computational complexity, compared with that of classical and state-of-the-art IQA models. The source code for the proposed GFM will be available at http://smartviplab.org/pubilcations/GFM.html.
Zhangkai Ni, Huanqiang Zeng, Lin Ma 0002, Junhui Hou, Jing Chen 0001, Kai-Kuang Ma
IEEE Trans. Image Process.2
2017 Sum-of-gradient based fast intra coding in 3D-HEVC for depth map sequence (SOG-FDIC)
Jing Chen 0001, Huanqiang Zeng, Canhui Cai, Kai-Kuang Ma
J. Vis. Commun. Image Represent.3
2017 ESIM: Edge Similarity for Screen Content Image Quality Assessment
abstract
In this paper, an accurate full-reference image quality assessment (IQA) model developed for assessing screen content images (SCIs), called the edge similarity (ESIM), is proposed. It is inspired by the fact that the human visual system (HVS) is highly sensitive to edges that are often encountered in SCIs; therefore, essential edge features are extracted and exploited for conducting IQA for the SCIs. The key novelty of the proposed ESIM lies in the extraction and use of three salient edge features-i.e., edge contrast, edge width, and edge direction. The first two attributes are simultaneously generated from the input SCI based on a parametric edge model, while the last one is derived directly from the input SCI. The extraction of these three features will be performed for the reference SCI and the distorted SCI, individually. The degree of similarity measured for each above-mentioned edge attribute is then computed independently, followed by combining them together using our proposed edge-width pooling strategy to generate the final ESIM score. To conduct the performance evaluation of our proposed ESIM model, a new and the largest SCI database (denoted as SCID) is established in our work and made to the public for download. Our database contains 1800 distorted SCIs that are generated from 40 reference SCIs. For each SCI, nine distortion types are investigated, and five degradation levels are produced for each distortion type. Extensive simulation results have clearly shown that the proposed ESIM model is more consistent with the perception of the HVS on the evaluation of distorted SCIs than the multiple state-of-the-art IQA methods.
Zhangkai Ni, Lin Ma 0002, Huanqiang Zeng, Jing Chen 0001, Canhui Cai, Kai-Kuang Ma
IEEE Trans. Image Process.3
2016 Robust laplacian matrix learning for smooth graph signals
abstract
We propose a new method for robust learning Laplacian matrices from observed smooth graph signals in the presence of both Gaussian noise and random-valued impulse noise (i.e., outliers). Using the recently developed factor analysis model for representing smooth graph signals in [1], we formulate our learning process as a constrained optimization problem, and adopt the £i-norm for measuring the data fidelity in order to improve robustness. Computational results on three types of synthetic graphs demonstrate that the proposed method outperforms the state-of-the-art methods in terms of commonly used information retrieval metrics, such as F-measure, precision, recall and normalized mutual information. In particular, we observed that F-measure is improved by up to 16%.
Junhui Hou, Lap-Pui Chau, Ying He 0001, Huanqiang Zeng
ICIP4
2016 Screen content image quality assessment using edge model
abstract
Since the human visual system (HVS) is highly sensitive to edges, a novel image quality assessment (IQA) metric for assessing screen content images (SCIs) is proposed in this paper. The turnkey novelty lies in the use of an existing parametric edge model to extract two types of salient attributes - namely, edge contrast and edge width, for the distorted SCI under assessment and its original SCI, respectively. The extracted information is subject to conduct similarity measurements on each attribute, independently. The obtained similarity scores are then combined using our proposed edge-width pooling strategy to generate the final IQA score. Hopefully, this score is consistent with the judgment made by the HVS. Experimental results have shown that the proposed IQA metric produces higher consistency with that of the HVS on the evaluation of the image quality of the distorted SCI than that of other state-of-the-art IQA metrics.
Zhangkai Ni, Lin Ma 0002, Huanqiang Zeng, Canhui Cai, Kai-Kuang Ma
ICIP3
2016 Low complexity depth intra coding in 3D-HEVC based on depth classification
abstract
The latest high efficiency video coding-based three dimensional video coding (3D-HEVC) exploits sophisticated intra prediction scheme to improve the coding performance of the depth video, but incurring heavy computational complexity. To address this problem, a low complexity depth intra coding method is presented for 3D-HEVC based on depth classification. Firstly, a database of depth prediction units (PUs) with three kinds of complexities is collected based on their optimal intra prediction mode. Then, the histogram of oriented gradient (HOG) features are extracted on these established database to train the classifier using support vector machine (SVM). For the current depth PU, the trained classifier is applied to determine its most possible complexity class so as to select the corresponding modes for involving the mode decision process. Experimental results show that the proposed method is able to significantly reduce the computational complexity while keeping almost the same coding performance of depth video and video quality of the synthesized view, compared with the exhaustive mode decision in 3D-HEVC.
Huijie Zheng, Jianqing Zhu, Huanqiang Zeng, Jing Chen 0001, Canhui Cai, Kai-Kuang Ma
VCIP3
2016 Quad binary pattern and its application in mean-shift tracking
Huanqiang Zeng, Jing Chen 0001, Xiaolin Cui, Canhui Cai, Kai-Kuang Ma
Neurocomputing1
2016 Multiple description video coding based on adaptive data reuse
Meng Dong, Huanqiang Zeng, Jing Chen 0001, Canhui Cai, Kai-Kuang Ma
J. Vis. Commun. Image Represent.2
2016 Guest Editorial: Emerging Visual Information Processing Technologies for Multimedia Applications
Huanqiang Zeng
Multim. Tools Appl.1
2016 Perceptual sensitivity-based rate control method for high efficiency video coding
Huanqiang Zeng, Aisheng Yang, King Ngi Ngan, Miaohui Wang
Multim. Tools Appl.1
2016 Gradient Direction for Screen Content Image Quality Assessment
abstract
In this letter, we make the first attempt to explore the usage of the gradient direction to conduct the perceptual quality assessment of the screen content images (SCIs). Specifically, the proposed approach first extracts the gradient direction based on the local information of the image gradient magnitude, which not only preserves gradient direction consistency in local regions, but also demonstrates sensitivities to the distortions introduced to the SCI. A deviation-based pooling strategy is subsequently utilized to generate the corresponding image quality index. Moreover, we investigate and demonstrate the complementary behaviors of the gradient direction and magnitude for SCI quality assessment. By jointly considering them together, our proposed SCI quality metric outperforms the state-of-the-art quality metrics in terms of correlation with human visual system perception.
Zhangkai Ni, Lin Ma 0002, Huanqiang Zeng, Canhui Cai, Kai-Kuang Ma
IEEE Signal Process. Lett.3
2015 Multiple Description Coding for Multi-view Video
Jing Chen 0001, Canhui Cai, Xiaolan Wang 0006, Huanqiang Zeng, Kai-Kuang Ma
ACIVS4
2015 Improved block level adaptive quantization for high efficiency video coding
abstract
As the concept of block level adaptivity becomes an important feature in recent video CODECs, block level adaptive quantization (BLAQ) is being considered in the High Efficiency Video Coding (HEVC) standard. The BLAQ is based on the assumption that each block should have its own quantization parameter (QP), which can adapt to the local content of video sequences much better, and hence the video encoder with adaptive QP can perform a better perceptual quality. However, in the HEVC reference software, the BLAQ is required to obtain a proper QP for each block by the rate distortion optimization (RDO) scheme and so the computational complexity of the encoder increases significantly. In this paper, an improved BLAQ algorithm is proposed to obtain the adaptive QP for each block. The simulation results show that the proposed method can save more bits as well as require lower computational complexity, compared to the traditional method.
Miaohui Wang, King Ngi Ngan, Hongliang Li 0001, Huanqiang Zeng
ISCAS4
2015 SIFT-flow-based color correction for multi-view video
Huanqiang Zeng, Kai-Kuang Ma, Canhui Cai
Signal Process. Image Commun.1
2014 Fast Multiview Video Coding Using Adaptive Prediction Structure and Hierarchical Mode Decision
abstract
The multiview video coding (MVC) adopts hierarchical B picture prediction structure and offers many prediction modes to effectively remove the spatial, temporal, and inter-view redundancies inherited in multiview video (MVV), but at the price of extremely high computational complexity. To address this problem, a fast MVC method by jointly using adaptive prediction structure (APS) and hierarchical mode decision (HMD) is proposed in this paper. The complexity reduction is achieved by: 1) designing four APSs for different MVV contents based on the fact that the contribution of the inter-view prediction varies from sequence to sequence and 2) developing an HMD scheme based on the observation that the relationship between the rate distortion (RD) cost and size of prediction mode is a unimodal function. In particular, for the current group of picture of the input MVV, the prediction structure is adaptively selected based on its characteristic, which is measured by the ratio of the average RD cost of the base view frames to the sum of the average RD cost of the base view frames and that of anchor frames in nonbase views, and then an HMD scheme is further performed to skip the checking process of those unlikely modes. The experimental results have shown that compared with the exhaustive mode decision in the MVC, the proposed algorithm achieves a reduction of the computational complexity by 83.49% on average, whereas incurring only a 0.086 dB loss in Bjontegaard delta peak signal-to-noise ratio and 2.97% increment on the total Bjontegaard delta bit rate.
Huanqiang Zeng, Xiaolan Wang 0006, Canhui Cai, Jing Chen 0001
IEEE Trans. Circuits Syst. Video Technol.1
2013 A rate distortion optimized transform for motion compensation residual
abstract
In this paper, we propose a rate distortion optimization based content adaptive transform method for motion compensation residuals. The proposed method utilizes pixel rearrangement to dynamically adjust the transform kernels to adapt to the residual content. Comparing with the traditional adaptive transforms, the highlight of this work is that it obtains the transform kernels from the decoded block, and hence it consumes only one overhead bit for each transform unit. Moreover, rate distortion optimization scheme is used to choose the best candidate kernels. Experimental results show that the proposed method achieves an average 0.35 dB gain of PSNR in comparison with the key technical areas (KTA) encoder.
Miaohui Wang, King Ngi Ngan, Huanqiang Zeng
PCS3
2013 Perceptual adaptive Lagrangian multiplier for high efficiency video coding
abstract
In high efficiency video coding (HEVC), Lagrangian rate distortion optimization (RDO) technique is used to optimize the rate distortion (RD) performance. However, the corresponding Lagrangian multiplier does not consider the perceptual characteristic of the input video and thus is not effective for perceptual video coding. To address this problem, an efficient perceptual adaptive Lagrangian multiplier for HEVC is proposed. Based on the human visual system (HVS) observation that the region with less perceptual sensitivity can tolerate more distortion, the Lagrangian multiplier is adaptively adjusted for each coding tree unit (CTU) based on its perceptual sensitivity so that the perceptual quality of the reconstructed video can be improved. The above-mentioned perceptual sensitivity for each CTU is measured according to two perceptual features—spatial energy ratio and temporal motion activity. Experimental results have shown that the proposed method is able to significantly improve the perceptual RD performance, compared with the original HEVC.
Huanqiang Zeng, King Ngi Ngan, Miaohui Wang
PCS1
2013 Efficient early direct mode decision for multi-view video coding
Fengsui Wang, Huanqiang Zeng, Qinghong Shen, Sidan Du
Signal Process. Image Commun.2
2012 Content-adaptive temporal consistency enhancement for depth video
abstract
The video plus depth format, which is composed of the texture video and the depth video, has been widely used for free viewpoint TV. However, the temporal inconsistency is often encountered in the depth video due to the error incurred in the estimation of the depth values. This will inevitably deteriorate the coding efficiency of depth video and the visual quality of synthesized view. To address this problem, a content-adaptive temporal consistency enhancement (CTCE) algorithm for the depth video is proposed in this paper, which consists of two sequential stages: (1) classification of stationary and non-stationary regions based on the texture video, and (2) adaptive temporal consistency filtering on the depth video. The result of the first stage is used to steer the second stage so that the filtering process will be conducted in an adaptive manner. Extensive experimental results have shown that the proposed CTCE algorithm can effectively mitigate the temporal inconsistency in the original depth video and consequently improve the coding efficiency of depth video and the visual quality of synthesized view.
Huanqiang Zeng, Kai-Kuang Ma
ICIP1
2011 Fast Mode Decision for Multiview Video Coding Using Mode Correlation
abstract
Exhaustive mode decision has been exploited in multiview video coding for effectively improving the coding efficiency, but at the expense of yielding much higher computational complexity. In this paper, a fast mode decision algorithm, called the mode correlation-based mode decision (MCMD), is proposed to speed up the encoding process by reducing the number of the modes required to be checked. In our approach, all the prediction modes are first categorized into five motion-activity classes, and only one of them will be chosen to identify the optimal mode in a hierarchical manner, as follows. For each macroblock (MB), the proposed MCMD algorithm always begins with checking whether the rate-distortion cost computed at the SKIP mode (i.e., Class 1) is below an adaptive threshold for providing a possible early termination chance. If this early termination condition is not met, one of the remaining four motion-activity classes will be chosen for further mode checking according to the analysis of the predicted motion vector (PMV) of the current MB. The above-mentioned adaptive threshold and PMV are derived by exploiting the mode correlation between the current MB and a set of adjacent MBs (i.e., region of support) in the current view and its neighboring view. Experimental results have shown that compared with exhaustive mode decision, which is a default approach set in the joint multiview video model (JMVM) reference software, the proposed MCMD algorithm achieves a reduction of the computational complexity by 73.39% on average, while incurring only 0.07 dB loss in peak signal-to-noise ratio (PSNR) and 2.22% increment on the total bit rate.
Huanqiang Zeng, Kai-Kuang Ma, Canhui Cai
IEEE Trans. Circuits Syst. Video Technol.1
2011 Large Disparity Motion Layer Extraction via Topological Clustering
abstract
In this paper, we present a robust and efficient approach to extract motion layers from a pair of images with large disparity motion. First, motion models are established as: 1) initial SIFT matches are obtained and grouped into a set of clusters using our developed topological clustering algorithm; 2) for each cluster with no less than three matches, an affine transformation is estimated with least-square solution as tentative motion model; and 3) the tentative motion models are refined and the invalid models are pruned. Then, with the obtained motion models, a graph cuts based layer assignment algorithm is employed to segment the scene into several motion layers. Experimental results demonstrate that our method can successfully segment scenes containing objects with large interframe motion or even with significant interframe scale and pose changes. Furthermore, compared with the previous method invented by Wills and its modified version, our method is much faster and more robust.
Yongtao Wang, Junbin Gong, Dazhi Zhang, Chenqiang Gao, Jinwen Tian, Huanqiang Zeng
IEEE Trans. Image Process.6
2010 Mode-correlation-based early termination mode decision for multi-view video coding
abstract
Exhaustive mode decision is exploited in multi-view video coding for effectively improving the coding efficiency, but at the expense of yielding higher computational complexity. In this paper, a fast mode decision algorithm, called the mode-correlation-based early termination (MET), is proposed. For each macroblock, the proposed MET algorithm always starts with checking whether the rate-distortion (RD) cost computed at the SKIP mode is below an adaptive threshold for providing a possible early termination chance. This adaptive threshold is calculated by using the mode correlation between the current macroblock and a set of adjacent macroblocks in the current view and its neighboring view. Experimental results have shown that compared with exhaustive mode decision, which is a default approach set in the JMVM reference software, the proposed MET algorithm achieves a reduction of the computational complexity by 65.91% and the total bit rate by 0.98% on average, while incurring only 0.06 dB loss in peak signal-to-noise ratio (PSNR).
Huanqiang Zeng, Kai-Kuang Ma, Canhui Cai
ICIP1
2010 Motion activity-based block size decision for multi-view video coding
abstract
Motion estimation and disparity estimation using variable block sizes have been exploited in multi-view video coding to effectively improve the coding efficiency, but at the expense of yielding higher computational complexity. In this paper, a fast block size decision algorithm, called motion activity-based block size decision (MABSD), is proposed. In our approach, the various motion estimation and disparity estimation block sizes are classified into four classes, and only one of them will be chosen to further identify the optimal block size within that class according to the measured motion activity of the current macroblock. The above-mentioned motion activity can be measured by the maximum city-block distance of a set of motion vectors taken from the adjacent macroblocks in the current view and its neighboring view. Experimental results have shown that compared with exhaustive block size decision, which is a default approach set in the JMVM reference software, the proposed MABSD algorithm achieves a reduction of computational complexity by 42% on average, while incurring only 0.01 dB loss in peak signal-to-noise ratio (PSNR) and 1% increment on the total bit rate.
Huanqiang Zeng, Kai-Kuang Ma, Canhui Cai
PCS1
2010 Hierarchical Intra Mode Decision for H.264/AVC
abstract
The intra mode prediction via exhaustive mode decision exploited in the H.264/advanced video coding effectively improves the coding efficiency, but at the expense of yielding higher computational complexity. In this letter, a fast intra mode decision algorithm, called the hierarchical intra mode decision (HIMD), is proposed to speed up the mode decision process by reducing the number of modes required to be checked for each macroblock. The novelty of the proposed HIMD algorithm lies at the following accounts. 1) An early decision with adaptive thresholding is developed for the mode decision of the luma component. 2) The candidate modes are selected according to their Hadamard distances and prediction directions. 3) Only one of the hierarchical paths will be chosen to compute its least rate-distortion cost. Experimental results have shown that the proposed HIMD algorithm achieves a reduction of 85.75% computational complexity on average, while incurring only 0.164 dB loss in peak signal-to-noise ratio (PSNR) and 2.336% increment on the total bit rate compared with that of exhaustive mode decision, which is a default approach set in the joint model reference software.
Huanqiang Zeng, Kai-Kuang Ma, Canhui Cai
IEEE Trans. Circuits Syst. Video Technol.1
2009 Fast motion estimation for H.264
Canhui Cai, Huanqiang Zeng, Sanjit K. Mitra
Signal Process. Image Commun.2
2009 Fast Mode Decision for H.264/AVC Based on Macroblock Motion Activity
abstract
The intra-mode and inter-mode predictions have been made available in H.264/AVC for effectively improving coding efficiency. However, exhaustively checking for all the prediction modes for identifying the best one (commonly referred to asexhaustivemodedecision) greatly increases computational complexity. In this paper, a fast mode decision algorithm, called themotionactivity-basedmodedecision(MAMD), is proposed to speed up the encoding process by reducing the number of modes required to be checked in a hierarchical manner, and is as follows. For each macroblock, the proposed MAMD algorithm always starts with checking the rate-distortion (RD) cost computed at the SKIP mode for a possible early termination, once the RD cost value is below a predetermined ldquolowrdquo threshold. On the other hand, if the RD cost exceeds another ldquohighrdquo threshold, then this indicates that only the intra modes are worthwhile to be checked. If the computed RD cost falls between the above-mentioned two thresholds, the remaining seven modes, which are classified into three motion activity classes in our work, will be examined, and only one of the three classes will be chosen for further mode checking. The above-mentioned motion activity can be quantitatively measured, which is equal to the maximum city-block length of the motion vector taken from a set of adjacent macroblocks (i.e., region of support, ROS). This measurement is then used to determine the most possible motion-activity class for the current macroblock. Experimental results have shown that, on average, the proposed MAMD algorithm reduces the computational complexity by 62.96%, while incurring only 0.059 dB loss in PSNR (peak signal-to-noise ratio) and 0.19% increment on the total bit rate compared to that of exhaustive mode decision, which is a default approach set in the JM reference software.
Huanqiang Zeng, Canhui Cai, Kai-Kuang Ma
IEEE Trans. Circuits Syst. Video Technol.1