Honggang Chen

dblp:195/2133 · DBLP profile ↗
← Back
66ranked-venue papers
11as first author
56since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 33 · 5 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 5 first-author · 19 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Filter, Correlate, Compress: Training-Free Token Reduction for MLLM Acceleration
abstract
The quadratic complexity of Multimodal Large Language Models (MLLMs) with respect to context length poses significant computational and memory challenges, hindering their real-world deployment. In the paper, we devise a ''filter-correlate-compress'' framework to accelerate the MLLM by systematically optimizing multimodal context length during prefilling. The framework first implements FiCoCo-V, a training-free method operating within the vision encoder. It employs a redundancy-based token discard mechanism that uses a novel integrated metric to accurately filter out redundant visual tokens. To mitigate information loss, the framework introduces a correlation-based information recycling mechanism that allows preserved tokens to selectively recycle information from correlated discarded tokens with a self-preserving compression, thereby preventing the dilution of their own core content. The framework's FiCoCo-L variant further leverages task-aware textual priors to perform token reduction directly within the LLM decoder. Extensive experiments demonstrate that the FiCoCo series effectively accelerates a range of MLLMs, achieves up to 14.7× FLOPs reduction with 93.6% performance retention. Our methods consistently outperform state-of-the-art training-free approaches, showcasing effectiveness and generalizability across model architectures, sizes, and tasks without requiring retraining.
Xuyang Liu 0002, Pengxiang Ding, Honggang Chen, Qingsen Yan, Siteng Huang
AAAI6
2026 Global Compression Commander: Plug-and-Play Inference Acceleration for High-Resolution Large Vision-Language Models
abstract
Large vision-language models (LVLMs) excel at visual understanding but face efficiency challenges due to quadratic complexity when processing long multimodal contexts. While token compression can reduce computational costs, existing approaches are designed for single-view LVLMs and fail to account for the unique multi-view characteristics of high-resolution LVLMs that use dynamic cropping. Current methods treat all tokens uniformly, yet our analysis shows that global thumbnails can naturally guide the compression of local crops by providing holistic context for evaluating informativeness. In this paper, we first analyze the dynamic cropping strategy, revealing both the complementary relationship between thumbnails and crops and the distinct characteristics across different crops. Based on these insights, we propose ''Global Compression Commander'' (GlobalCom2), a novel plug-and-play token compression framework for high-resolution LVLMs. GlobalCom2 uses the thumbnail as a ''commander'' to adaptively guide the compression of local crops, preserving informative details while removing redundancy. Extensive experiments demonstrate that GlobalCom2 maintains over 90% of model performance while compressing 90% of visual tokens, reducing FLOPs to 9.1% and peak memory usage to 60% of the original.
Xuyang Liu 0002, Yingyao Wang, Jiale Yuan, Siteng Huang, Honggang Chen
AAAI9
2026 A lightweight framework for robust object detection in adverse weather based on dual-teacher feature alignment
Hanjun Zheng, Shengjie Ye, Linbo Qing, Honggang Chen
Neurocomputing5
2026 Enhancing arbitrary-scale super-resolution with scale-aware multiscale nonlocal feature extraction and local structure-adaptive upsampling
Honggang Chen, Shuhua Xiong, Xiaohai He
Neurocomputing2
2026 DSRIR: Dynamic spatial refinement learning for progressive all-in-one image restoration
Xiao Liu 0022, Yutong Yang, Zhengyong Wang, Xiaohai He, Honggang Chen, Yi Li 0069, Pingyu Wang
Inf. Process. Manag.6
2026 RelPosGAR: Hierarchical relative position-aware interaction modeling for weakly supervised skeleton-based group activity recognition
Lindong Li, Linbo Qing, Liuyi Tao, Pingyu Wang, Honggang Chen, Owen Noel Newton Fernando, Weisi Lin
Pattern Recognit.5
2026 Blind visual security assessment using a simple parallel dual-stream network
Xiaodong Bi, Xiaohai He, Zeming Zhao, Shuhua Xiong, Honggang Chen, Ray E. Sheriff
Signal Process. Image Commun.5
2026 DVSH-GS: Dense-Voxel and Spherical Harmonics Guided Gaussian Splatting for Sparse View rendering
Fansen Zhou, Honggang Chen
IEEE Signal Process. Lett.3
2026 A Humanoid TacTip Gripper With SSIM-CNN Recognition for Strawberry Harvesting
abstract
Most existing robotic fruit harvesting systems rely on mechanically complex end-effectors that lack delicate tactile dexterity and fail to replicate the sophisticated sensory-motor coordination of human pickers. To enable safe, reliable, and efficient automated harvesting of delicate fruits, this paper presents a novel bio-inspired humanoid TacTip gripper for precision strawberry harvesting. Inspired by human thumbfinger opposition, we design an asymmetric TacTip gripper that integrates a Thumb tactile sensor with a built-in fingernail for stem cutting and a supporting Pillow tactile sensor. We further develop a hybrid SSIM-CNN perception framework that fuses real-time structural similarity index measure (SSIM) from both fingertips with convolutional neural network (CNN) features, enabling precise closed-loop grasp-state detection and gentle force adjustment. In addition, a segmented dynamic system (DS) motion planner decomposes the harvesting task into approach, cut, and place phases, generating reactive, smooth, and biologically plausible trajectories. Experimental results on both laboratory setups and live potted strawberry plants demonstrate reliable full-cycle harvesting with high success rates and minimal fruit damage. The proposed system provides a practical and effective solution for automated delicate fruit harvesting.
Kunlin Guo, Honggang Chen, Jiehao Li, Zhenyu Lu 0001, Chenguang Yang 0001
IEEE Trans Autom. Sci. Eng.2
2026 M2IST: Multi-Modal Interactive Side-Tuning for Efficient Referring Expression Comprehension
abstract
Referring expression comprehension (REC) is a vision-language task to locate a target object in an image based on a language expression. Fully fine-tuning general-purpose pre-trained vision-language foundation models for REC yields impressive performance but becomes increasingly costly. Parameter-efficient transfer learning (PETL) methods have shown strong performance with fewer tunable parameters. However, directly applying PETL to REC faces two challenges: (1) insufficient multi-modal interaction between pre-trained vision-language foundation models, and (2) high GPU memory usage due to gradients passing through the heavy vision-language foundation models. To this end, we present M2IST: Multi-Modal Interactive Side-Tuning with M3ISAs: Mixture of Multi-Modal Interactive Side-Adapters. During fine-tuning, we fix the pre-trained uni-modal encoders and update M3ISAs to enable efficient vision-language alignment for REC. Empirical results reveal that M2IST achieves better performance-efficiency trade-off than full fine-tuning and other PETL methods, requiring only 2.11% tunable parameters, 39.61% GPU memory, and 63.46% training time while maintaining competitive performance. Our code is released at https://github.com/xuyang-liu16/M2IST.
Xuyang Liu 0002, Ting Liu 0018, Siteng Huang, Yi Xin 0003, Yue Hu 0016, Long Qin 0004, Yuanyuan Wu 0001, Honggang Chen
IEEE Trans. Circuits Syst. Video Technol.9
2025 LEARN: Knowledge Adaptation from Large Language Model to Recommendation for Practical Industrial Application
abstract
Contemporary recommendation systems predominantly rely on ID embedding to capture latent associations among users and items. However, this approach overlooks the wealth of semantic information embedded within textual descriptions of items, leading to suboptimal performance and poor generalizations. Leveraging the capability of large language models to comprehend and reason about textual content presents a promising avenue for advancing recommendation systems. To achieve this, we propose an Llm-driven knowlEdge Adaptive RecommeNdation (LEARN) framework that synergizes open-world knowledge with collaborative knowledge. We address computational complexity concerns by utilizing pretrained LLMs as item encoders and freezing LLM parameters to avoid catastrophic forgetting and preserve open-world knowledge. To bridge the gap between the open-world and collaborative domains, we design a twin-tower structure supervised by the recommendation task and tailored for practical industrial application. Through experiments on the real large-scale industrial dataset and online A/B tests, we demonstrate the efficacy of our approach in industry application. We also achieve state-of-the-art performance on six Amazon Review datasets to verify the superiority of our method.
Jian Jia, Yan Li 0043, Honggang Chen, Xuehan Bai, Zhaocheng Liu, Jian Liang 0001, Quan Chen 0006, Han Li 0005, Peng Jiang 0002, Kun Gai
AAAI4
2025 RXNet: cross-modality person re-identification based on a dual-branch network
Weiyang Zhang, Jiong Guo, Qiang Liu 0021, Maoyang Zou, Honggang Chen
Appl. Intell.5
2025 DEMNet: A degradation difference enabled multi-stage network for multiple degradation image restoration
Yutong Yang, Honggang Chen
Knowl. Based Syst.4
2025 DRFormer: A Discriminable and Reliable Feature Transformer for Person Re-Identification
abstract
As person image variations are likely to cause a part misalignment problem, most previous person Re-Identification (ReID) works may adopt local feature partition or additional landmark annotations to acquire aligned person features and boost ReID performance. However, such approaches either only achieve coarse-grained part alignments without considering detailed image variations within each part, or require extra annotated landmarks to train an available pose estimation model. In this work, we propose an effective Discriminable and Reliable Transformer (DRFormer) framework to learn part-aligned person representations with only person identity labels. Specifically, the DRFormer framework consists of Discriminable Feature Transformer (DFT) and Reliable Feature Transformer (RFT) modules, which generate discriminable and reliable high-order features, respectively. For reducing the dimension of high-order features, the DFT module utilizes a Self-Attentive Kronecker Product (SAKP) algorithm to promote the representational capabilities of compressed features via a self-attention strategy. For eliminating the background noise, the RFT module mines the foreground regions to adaptively aggregate foreground features via a Gumbel-Softmax strategy. Moreover, the proposed framework derives from an interpretable motivation and elegantly solves part misalignments without using feature partition or pose estimation. This paper theoretically and experimentally demonstrates the superiority of the proposed DRFormer framework, achieving state-of-the-art performance on various person ReID datasets.
Pingyu Wang, Xingjian Zheng, Linbo Qing, Bonan Li, Zhicheng Zhao 0001, Honggang Chen
IEEE Trans. Inf. Forensics Secur.7
2025 A compressed video quality enhancement algorithm based on CNN and transformer hybrid network
Xiaohai He, Shuhua Xiong, Haibo He, Honggang Chen
J. Supercomput.5
2024 UWB weak signal detection and recovery algorithm based on power spectrum entropy and improved stochastic resonance
abstract
In electromagnetic spectrum monitoring, Ultra-Wideband (UWB) signals, as a kind of low interception rate communication signal, are easy to bring great security risks. UWB signals are easy to be submerged in the background noise and difficult to detect. Therefore, this paper proposes an UWB signal detection and recovery algorithm based on power spectrum entropy and improved stochastic resonance for the above problems. Firstly, the power spectrum entropy algorithm was used to detect the existence of the signal. After confirming the existence of the signal, the adaptive piecewise stochastic resonance system was used to recover the detected signal waveform. On the basis of stochastic resonance processing, the Variational Mode Decomposition (VMD) algorithm combined with wavelet transform is used to further filter the signal noise and improve the recovery degree of the signal waveform. The simulation results show that the proposed algorithm can effectively recovery the UWB signal, and the recovery degree is better than the traditional bistable stochastic resonance model. The measured experiment verifies the effectiveness of the proposed algorithm.
Yanyun Xu, Honggang Chen
HPCC4
2024 VGDIFFZERO: Text-To-Image Diffusion Models Can Be Zero-Shot Visual Grounders
abstract
Large-scale text-to-image diffusion models have shown impressive capabilities for generative tasks by leveraging strong vision-language alignment from pre-training. However, most vision-language discriminative tasks require extensive fine-tuning on carefully-labeled datasets to acquire such alignment, with great cost in time and computing resources. In this work, we explore directly applying a pre-trained generative diffusion model to the challenging discriminative task of visual grounding without any fine-tuning and additional training dataset. Specifically, we propose VGDiffZero, a simple yet effective zero-shot visual grounding framework based on text-to-image diffusion models. We also design a comprehensive region-scoring method considering both global and local contexts of each isolated proposal. Extensive experiments on RefCOCO, RefCOCO+, and RefCOCOg show that VGDiffZero achieves strong performance on zero-shot visual grounding. Our code is available at https://github.com/xuyang-liu16/VGDiffZero.
Xuyang Liu 0002, Siteng Huang, Yachen Kang, Honggang Chen
ICASSP4
2024 DARA: Domain- and Relation-Aware Adapters Make Parameter-Efficient Tuning for Visual Grounding
abstract
Visual grounding (VG) is a challenging task to localize an object in an image based on a textual description. Recent surge in the scale of VG models has substantially improved performance, but also introduced a significant burden on computational costs during fine-tuning. In this paper, we explore applying parameter-efficient transfer learning (PETL) to efficiently transfer the pre-trained vision-language knowledge to VG. Specifically, we propose DARA, a novel PETL method comprising Domain-aware Adapters (DA Adapters) and Relation-aware Adapters (RA Adapters) for VG. DA Adapters first transfer intra-modality representations to be more fine-grained for the VG domain. Then RA Adapters share weights to bridge the relation between two modalities, improving spatial reasoning. Empirical results on widely-used benchmarks demonstrate that DARA achieves the best accuracy while saving numerous updated parameters compared to the full fine-tuning and other PETL methods. Notably, with only 2.13% tunable backbone parameters, DARA improves average accuracy by 0.81% across the three benchmarks compared to the baseline model. Our code is available at https://github.com/liuting20/DARA.
Ting Liu 0018, Xuyang Liu 0002, Siteng Huang, Honggang Chen, Quanjun Yin, Long Qin 0004, Yue Hu 0016
ICME4
2024 GooseFS: Distributed Cache Service to Enhance Cloud Object Storage Performance
abstract
With advantages of scalability and low cost, cloud object storage has become the preferred storage solution for massive data. However, object storage has its own defects, such as poor performance of small I/Os, which restricts object storage from being widely used in some scenarios like Big Data. In this paper, we propose GooseFS, a distributed high-performance cache service built on top of object storage, which accelerates the access to object storage under the storage-compute separation architecture. GooseFS makes great enhancements on small I/Os performance as well as metadata performance for data stored in object storage. GooseFS introduces three key designs: (1) Compute-Side Cache, which reduces the data access latency under random small I/O with multi-granularity cache management and short circuit read mechanism. (2) Storage-Side Cache, which improves the throughput with a high-performance SSD cache pool. (3) Metadata Acceleration, which significantly improves the performance of metadata operations through special metadata organization and lock-free strong consistency cache. Experiment show that compared with native object storage, GooseFS increases the throughput of small random I/O by$5\times-9\times$and improves metadata operation performance by at least$3.9\times$, meeting the performance requirements of various workloads.
Junming Yan, Zhitao Chen, Honggang Chen
NAS7
2024 Semantic and geometric information propagation for oriented object detection in aerial images
Xiaohai He, Honggang Chen, Linbo Qing, Qizhi Teng
Appl. Intell.3
2024 A Multi-Attention Feature Distillation Neural Network for Lightweight Single Image Super-Resolution
abstract
In recent years, remarkable performance improvements have been produced by deep convolutional neural networks (CNN) for single image super-resolution (SISR). Nevertheless, a high proportion of CNN-based SISR models are with quite a few network parameters and high computational complexity for deep or wide architectures. How to more fully utilize deep features to make a balance between model complexity and reconstruction performance is one of the main challenges in this field. To address this problem, on the basis of the well-known information multi-distillation model, a multi-attention feature distillation network termed as MAFDN is developed for lightweight and accurate SISR. Specifically, an effective multi-attention feature distillation block (MAFDB) is designed and used as the basic feature extraction unit in MAFDN. With the help of multi-attention layers including pixel attention, spatial attention, and channel attention, MAFDB uses multiple information distillation branches to learn more discriminative and representative features. Furthermore, MAFDB introduces the depthwise over-parameterized convolutional layer (DO-Conv)-based residual block (OPCRB) to enhance its ability without incurring any parameter and computation increase in the inference stage. The results on commonly used datasets demonstrate that our MAFDN outperforms existing representative lightweight SISR models when taking both reconstruction performance and model complexity into consideration. For example, for × 4 SR on Set5, MAFDN (597K/33.79G) obtains 0.21 dB/0.0037 and 0.10 dB/0.0015 PSNR/SSIM gains over the attention-based SR model AFAN (692K/50.90G) and the feature distillation-based SR model DDistill-SR (675K/32.83G), respectively.
Yongfei Zhang, Xinying Lin, Linbo Qing, Xiaohai He, Yi Li 0069, Honggang Chen
Int. J. Intell. Syst.8
2024 Blind video quality assessment based on Spatio-Temporal Feature Resolver
Xiaodong Bi, Xiaohai He, Shuhua Xiong, Zeming Zhao, Honggang Chen, Ray E. Sheriff
Neurocomputing5
2024 Multi deep invariant feature learning for cross-resolution person re-identification
Weicheng Zhang, Shuhua Xiong, Xiaohai He, Honggang Chen
Inf. Process. Manag.6
2024 A channel-wise contextual module for learned intra video compression
Yanrui Zhan, Shuhua Xiong, Xiaohai He, Honggang Chen
J. Vis. Commun. Image Represent.5
2024 Fast CU partition strategy based on texture and neighboring partition information for Versatile Video Coding Intra Coding
Ruolan Yang, Xiaohai He, Shuhua Xiong, Zeming Zhao, Honggang Chen
Multim. Tools Appl.5
2024 DVC-Net: a new dual-view context-aware network for emotion recognition in the wild
Linbo Qing, Hongqian Wen, Honggang Chen, Rulong Jin, Yongqiang Cheng 0001, Yonghong Peng
Neural Comput. Appl.3
2024 Depth map super-resolution via learned nonlocal model and enhanced local regularization
Xiaohai He, Honggang Chen, Chao Ren 0002
Signal Process.3
2024 HCT: Chinese Medical Machine Reading Comprehension Question-Answering via Hierarchically Collaborative Transformer
abstract
Chinese medical machine reading comprehension question-answering (cMed-MRCQA) is a critical component of the intelligence question-answering task, focusing on the Chinese medical domain question-answering task. Its purpose enable machines to analyze and understand the given text and question and then extract the accurate answer. To enhance cMed-MRCQA performance, it is essential to possess a profound comprehension and analysis of the context, deduce concealed information from the textual content and, subsequently, precisely determine the answer's span. The answer span has predominantly been defined by language items, with sentences employed in most instances. However, it is worth noting that sentences may not be properly split to varying degrees in various languages, making it challenging for the model to predict the answer zone. To alleviate this issue, this paper presents a novel architecture called HCT based on a Hierarchically Collaborative Transformer. Specifically, we presented a hierarchical collaborative method to locate the boundaries of sentence and answer spans separately. First, we designed a hierarchical encoding module to obtain the local semantic features of the corpus; second, we proposed a sentence-level self-attention module and a fused interaction-attention module to get the global information about the text. Finally, the model is trained by combining loss functions. Extensive experiments were conducted on the public dataset CMedMRC and the reconstruction dataset eMedicine to validate the effectiveness of the proposed method. Experimental results showed that the proposed method performed better than the state-of-the-art methods. Using the F1 metric, our model scored 90.4% on the CMedMRC and 73.2% on eMedicine.
Xiaohai He, Luping Liu, Qingmao Fang, Honggang Chen, Yan Liu 0078
IEEE J. Biomed. Health Informatics6
2024 Activating More Information in Arbitrary-Scale Image Super-Resolution
abstract
Single-image super-resolution (SISR) has experienced vigorous growth with the rapid development of deep learning. However, handling arbitrary scales (e.g.,integers, non-integers, or asymmetric) using a single model remains a challenging task. Existing super-resolution (SR) networks commonly employ static convolutions during feature extraction, which cannot effectively perceive changes in scales. Moreover, these continuous-scale upsampling modules only utilize the scale factors, without considering the diversity of local features. To activate more information for better reconstruction, two plug-in and compatible modules for fixed-scale networks are designed to perform arbitrary-scale SR tasks. Firstly, we design a Scale-aware Local Feature Adaptation Module (SLFAM), which adaptively adjusts the attention weights of dynamic filters based on the local features and scales. It enables the network to possess stronger representation capabilities. Then we propose a Local Feature Adaptation Upsampling Module (LFAUM), which combines scales and local features to perform arbitrary-scale reconstruction. It allows the upsampling to adapt to local structures. Besides, deformable convolution is utilized letting more information to be activated in the reconstruction, enabling the network to better adapt to the texture features. Extensive experiments on various benchmark datasets demonstrate that integrating the proposed modules into a fixed-scale SR network enables it to achieve satisfactory results with non-integer or asymmetric scales while maintaining advanced performance with integer scales.
Yaoqian Zhao, Qizhi Teng, Honggang Chen, Shujiang Zhang, Xiaohai He, Yi Li 0069, Ray E. Sheriff
IEEE Trans. Multim.3
2024 DAG-YOLO: A Context-Feature Adaptive fusion Rotating Detection Network in Remote Sensing Images
abstract
Object detection in remote sensing image (RSI) research has seen significant advancements, particularly with the advent of deep learning. However, challenges such as orientation, scale, aspect ratio variations, dense object distribution, and category imbalances remain. To address these challenges, we present DAG-YOLO, a one-stage context-feature adaptive weighted fusion network that incorporates through three innovative parts. First, we integrate 1D Gaussian Angle-coding with YOLOv5 to convert the angle regression task into a classification task, establishing a more robust rotating object detection baseline, GLR-YOLO. Second, we introduce the Dual Branch Context Adaptive Modeling module, which enhances feature extraction capabilities by capturing global context information. Third, we design an adaptive detect head with the Adaptive Global Feature Aggregation and Reweighting (AGFAR) module. AGFAR addresses feature inconsistency among different output layers of the Feature Pyramid Network, retaining useful semantic information and elevating detection accuracy. Extensive experiments on public datasets DOTA-v1.0, DOTA-v1.5, and UCAS-AOD showcase mAP scores of 77.75%, 73.79%, and 90.27%, respectively. Our proposed method has the best performance among the current mainstream SOTA methods, which proves its effectiveness in RSI object detection.
Zhenjiang Guo, Xiaohai He, Linbo Qing, Honggang Chen
ACM Trans. Multim. Comput. Commun. Appl.5
2023 Domain adaptive person re-identification with memory-based circular ranking
Honggang Chen, Nan Guo 0003, Xiaochun Ye, Dongrui Fan
Appl. Intell.1
2023 Nonlocal-guided enhanced interaction spatial-temporal network for compressed video super-resolution
Junxiong Cheng, Shuhua Xiong, Xiaohai He, Chao Ren 0002, Tingrong Zhang, Honggang Chen
Appl. Intell.6
2023 Image classification based on self-distillation
Linbo Qing, Xiaohai He, Honggang Chen, Qiang Liu 0021
Appl. Intell.4
2023 Self-supervised cycle-consistent learning for scale-arbitrary real-world single image super-resolution
Honggang Chen, Xiaohai He, Yuanyuan Wu 0001, Linbo Qing, Ray E. Sheriff
Expert Syst. Appl.1
2023 RestorNet: An efficient network for multiple degradation image restoration
Honggang Chen, Haosong Gou, Zhengyong Wang, Xiaohai He, Linbo Qing, Ray E. Sheriff
Knowl. Based Syst.2
2023 Multi-relation graph convolutional network for Alzheimer's disease diagnosis using structural MRI
Xiaohai He, Linbo Qing, Xiang Chen 0008, Yan Liu 0078, Honggang Chen
Knowl. Based Syst.6
2023 Block-correlation-based intra prediction for VVC
Shuhua Xiong, Xiaohai He, Honggang Chen, Chao Ren 0002
Multim. Tools Appl.4
2023 BDNet: A BERT-based dual-path network for text-to-image cross-modal person re-identification
Qiang Liu 0021, Xiaohai He, Qizhi Teng, Linbo Qing, Honggang Chen
Pattern Recognit.5
2023 Dynamically Optimized Human Eyes-to-Face Generation via Attribute Vocabulary
abstract
Generating face from human eyes, named eyes-to-face generation, is an interesting research topic of face synthesis, which has great potential in the field of public security. One of the main challenges in eyes-to-face generation is the unbalanced information between inputs and outputs, where the outputs are complete facial images while the inputs only contain limited information in the region of eyes. The existing methods generate faces directly from eyes without considering the possibly available facial information (e.g. facial attributes), resulting in inaccurate predictions and high uncertainty in those features less correlated with eyes (e.g. hairstyle, moustache, facial contour). To address this challenge, we propose a two-stage solution (named EA2F-GAN) to dynamically optimize eyes-to-face generation via attribute vocabulary. In addition, a dataset named TEAF is constructed based on the public datasets CelebA and LFW, containing 138,934 triples of eye image, attribute vocabulary, and face image. Sufficient experimental results show that, by incorporating additional facial attributes, our proposed approach can synthesize realistic face with high consistency to the original one, significantly overwhelming state-of-the-art methods.
Xiaodong Luo, Xiaohai He, Xiang Chen 0008, Linbo Qing, Honggang Chen
IEEE Signal Process. Lett.5
2023 Efficient Rate Control in Versatile Video Coding With Adaptive Spatial-Temporal Bit Allocation and Parameter Updating
abstract
Despite the fact that Versatile Video Coding (VVC) has achieved superior coding performance, two major problems remain for the rate control (RC) model in VVC. First, the regions concerned by human eyes are not clear enough in the coded video due to the deviation between the target bit allocation strategy of the coding tree unit (CTU) in RC and the human visual attention mechanism (HVAM). Second, there are significant quality fluctuations in the coded video frames due to the inappropriate updating speed. To address the above problems, we propose an efficient rate control (ERC) model. Specifically, in order to make the coded video more consistent with the attention of human eyes, we extract texture and motion-based spatial-temporal information to guide the bit allocation at the CTU level. Furthermore, based on the quasi-Newton algorithm and bit error, we propose an adaptive parameter updating (APU) method with the proper updating speed to precisely control the bits per frame. The proposed ERC outperforms the default RC model of VVC Test Model (VTM) 9.1 by saving the average Bjøntegaard Delta Rate (BD-Rate) on full-frame video sequences by 3.60% and 4.94% under low delay P (LDP) and random access (RA) configurations respectively, with higher bitrate accuracy. Moreover, the Peak Signal-to-Noise Ratio (PSNR) and actual coded bits per frame in the video coded by the proposed ERC are more stable.
Liqiang He, Xiaohai He, Shuhua Xiong, Zeming Zhao, Honggang Chen
IEEE Trans. Circuits Syst. Video Technol.6
2022 Verifying Privacy-Preserving Financing Orders on a Consortium Blockchain Based on zk-SNARKs
abstract
Due to its efficiency, low overhead, and high scalability, consortium blockchain has been deeply applied in various fields of society. Order financing is one of the scenarios of applying consortium blockchain. Since data on the consortium blockchain is available to the blockchain members, information of a financing order written directly to the blockchain will leak the commercial privacy of the purchaser and supplier. Therefore, the financing order data should be encrypted when published as a transaction on the consortium blockchain. However, the investor needs to verify the financing order data on a consortium blockchain before loaning money to the supplier. It is tricky to efficiently satisfy the verifiability of encrypted financing order data on the consortium blockchain. This work proposes VmppOrder, a verifiable model for privacy-preserving financing orders on a consortium blockchain based on zero-knowledge Succinct Non-interactive ARguments of Knowledge (zk-SNARKs). By the supplier publishing zero-knowledge proofs generated from the financing order, the investor can verify the encrypted financing order published on the consortium blockchain without decrypting it. We elaborate on the specific construction of VmppOrder and analyze the security of the constructed circuit with zero-knowledge proof. We implement a prototype of the model on Hyperledger Fabric based on Libsnark and conduct comprehensive experiments to evaluate its performance. Our experimental results validate the efficiency of the proposed model. Its order proof generation takes about 6.31 seconds, the order verification takes only 2.58 milliseconds, and the transaction processing speed is about 660 transactions per second on a moderately equipped machine.
Xiaoyan Hu 0007, Guang Cheng 0001, Honggang Chen, Zhichao Liang
WCNC6
2022 Dual adaptive alignment and partitioning network for visible and infrared cross-modality person re-identification
Qiang Liu 0021, Qizhi Teng, Honggang Chen, Bo Li 0074, Linbo Qing
Appl. Intell.3
2022 A nonlocal HEVC in-loop filter using CNN-based compression noise estimation
Weiheng Sun, Xiaohai He, Honggang Chen, Shuhua Xiong
Appl. Intell.3
2022 Medical visual question answering based on question-type reasoning and semantic space constraint
Xiaohai He, Luping Liu, Linbo Qing, Honggang Chen, Yan Liu 0078, Chao Ren 0002
Artif. Intell. Medicine5
2022 A two-stage deep generative adversarial quality enhancement network for real-world 3D CT images
Honggang Chen, Xiaohai He, Junxi Feng, Qizhi Teng
Expert Syst. Appl.1
2022 Deep dual-domain semi-blind network for compressed image quality enhancement
Jingbo He, Xiaohai He, Mozhi Zhang, Shuhua Xiong, Honggang Chen
Knowl. Based Syst.5
2022 Fact-based visual question answering via dual-process system
Luping Liu, Xiaohai He, Linbo Qing, Honggang Chen
Knowl. Based Syst.5
2022 A quality enhancement network with coding priors for constant bit rate video coding
Weiheng Sun, Xiaohai He, Chao Ren 0002, Shuhua Xiong, Honggang Chen
Knowl. Based Syst.5
2022 Weakly-supervised contrastive learning-based implicit degradation modeling for blind image super-resolution
Yongfei Zhang, Ling Dong, Linbo Qing, Xiaohai He, Honggang Chen
Knowl. Based Syst.6
2022 Sequential Enhancement for Compressed Video Using Deep Convolutional Generative Adversarial Network
Xiaohai He, Honggang Chen, Shuhua Xiong
Neural Process. Lett.4
2022 Unsupervised Real-World Image Super-Resolution via Dual Synthetic-to-Realistic and Realistic-to-Synthetic Translations
abstract
Due to the challenges of collecting paired low-resolution (LR) and high-resolution (HR) images in real-world scenarios, most existing deep convolutional neural network (CNN)-based single image super-resolution (SR) models are trained with artificially synthesized LR-HR image pairs. However, the domain gap between the synthetic data for model training and the realistic data for testing degrades SR performance significantly, which discourages the application of SR models in practice. One possible solution is to learn from unpaired real-world LR and HR images for their accessibility. Predominant strategies are mainly based on unsupervised domain translation. Despite great advances, there are still noticeable domain gaps between the realistic-like/synthetic-like images generated by unpaired translation and the true realistic/synthetic ones. To address this problem, this letter proposes an effective unsupervised SR framework based on dual synthetic-to-realistic and realistic-to-synthetic translations, namely DTSR. Specifically, to bridge the domain gap between testing and training data, the SR model is optimized using HR images and their realistic-like LR counterparts produced by the synthetic-to-realistic translation. In turn, we propose to narrow the domain gap further via applying the realistic-to-synthetic translation to realistic LR images prior to super-resolving, which also makes the SR model super-resolve simpler examples in testing relative to model training. Moreover, focal frequency and bilateral filtering losses are particularly introduced into DTSR for better details restoration and artifacts suppression. Extensive experiments show that our DTSR outperforms several state-of-the-art models in terms of both quantitative and qualitative comparisons.
Honggang Chen, Ling Dong, Xiaohai He, Ce Zhu
IEEE Signal Process. Lett.1
2022 A Feature-Enriched Deep Convolutional Neural Network for JPEG Image Compression Artifacts Reduction and its Applications
abstract
The amount of multimedia data, such as images and videos, has been increasing rapidly with the development of various imaging devices and the Internet, bringing more stress and challenges to information storage and transmission. The redundancy in images can be reduced to decrease data size via lossy compression, such as the most widely used standard Joint Photographic Experts Group (JPEG). However, the decompressed images generally suffer from various artifacts (e.g., blocking, banding, ringing, and blurring) due to the loss of information, especially at high compression ratios. This article presents a feature-enriched deep convolutional neural network for compression artifacts reduction (FeCarNet, for short). Taking the dense network as the backbone, FeCarNet enriches features to gain valuable information via introducing multi-scale dilated convolutions, along with the efficient 1 ×1 convolution for lowering both parameter complexity and computation cost. Meanwhile, to make full use of different levels of features in FeCarNet, a fusion block that consists of attention-based channel recalibration and dimension reduction is developed for local and global feature fusion. Furthermore, short and long residual connections both in the feature and pixel domains are combined to build a multi-level residual structure, thereby benefiting the network training and performance. In addition, aiming at reducing computation complexity further, pixel-shuffle-based image downsampling and upsampling layers are, respectively, arranged at the head and tail of the FeCarNet, which also enlarges the receptive field of the whole network. Experimental results show the superiority of FeCarNet over state-of-the-art compression artifacts reduction approaches in terms of both restoration capacity and model complexity. The applications of FeCarNet on several computer vision tasks, including image deblurring, edge detection, image segmentation, and object detection, demonstrate the effectiveness of FeCarNet further.
Honggang Chen, Xiaohai He, Linbo Qing, Qizhi Teng
IEEE Trans. Neural Networks Learn. Syst.1
2021 Towards Efficient Co-audit of Privacy-Preserving Data on Consortium Blockchain via Group Key Agreement
abstract
Blockchain is well known for its storage consistency, decentralization and tamper-proof, but the privacy disclosure and difficulty in auditing discourage the innovative application of blockchain technology. As compared to public blockchain and private blockchain, consortium blockchain is widely used across different industries and use cases due to its privacy-preserving ability, auditability and high transaction rate. However, the present co-audit of privacy-preserving data on consortium blockchain is inefficient. Private data is usually encrypted by a session key before being published on a consortium blockchain for privacy preservation. The session key is shared with transaction parties and auditors for their access. For decentralizing auditorial power, multiple auditors on the consortium blockchain jointly undertake the responsibility of auditing. The distribution of the session key to an auditor requires individually encrypting the session key with the public key of the auditor. The transaction initiator needs to be online when each auditor asks for the session key, and one encryption of the session key for each auditor consumes resources. This work proposes GAChain and applies group key agreement technology to efficiently co-audit privacy-preserving data on consortium blockchain. Multiple auditors on the consortium blockchain form a group and utilize the blockchain to generate a shared group encryption key and their respective group decryption keys. The session key is encrypted only once by the group encryption key and stored on the consortium blockchain together with the encrypted private data. Auditors then obtain the encrypted session key from the chain and decrypt it with their respective group decryption key for co-auditing. The group key generation is involved only when the group forms or group membership changes, which happens very infrequently on the consortium blockchain. We implement the prototype of GAChain based on Hyperledger Fabric framework. Our experimental studies demonstrate that GAChain improves the co-audit efficiency of transactions containing private data on Fabric, and its incurred overhead is moderate.
Xiaoyan Hu 0007, Xiaoyi Song, Guang Cheng 0001, Honggang Chen, Zhichao Liang
MSN6
2021 Bi-directional skip connection feature pyramid network and sub-pixel convolution for high-quality object detection
Shuqi Xiong, Honggang Chen, Linbo Qing, Xiaohai He
Neurocomputing3
2021 Single depth map super-resolution via joint non-local self-similarity modeling and local multi-directional gradient-guided regularization
Chao Ren 0002, Honggang Chen, Ce Zhu, Kai Liu 0012
Signal Process. Image Commun.3
2021 Enhanced Separable Convolution Network for Lightweight JPEG Compression Artifacts Reduction
abstract
JPEG images are usually corrupted by various undesirable compression artifacts resulted from block-wise coarse quantization on discrete cosine transform coefficients. In recent years, deep convolutional neural networks (CNNs) have made spectacular achievements in compression artifacts reduction. However, most deep CNNs are difficult to be implemented on mobile devices due to their large number of parameters and operations. In this letter, we propose a novel deep CNN called ESCNet for lightweight JPEG compression artifacts reduction, in which enhanced separable convolution (ESConv) is carefully designed to make full use of image multi-scale information for better dense pixel value predictions. Specifically, ESConv consists of a grouped multi-scale dual depth-wise convolution (GMDDConv) and a wide-activated dual point-wise convolution (WDPConv). GMDDConv is dedicated to efficiently extracting abundant image multi-scale spatial features, which will be sent to WDPConv for effective non-linear feature fusion. The experimental results on benchmark datasets show that compared with state-of-the-art methods, our ESCNet not only achieves better performance in both objective indices and subjective quality but also greatly reduces network parameters and operations.
Zhengxin Chen, Xiaohai He, Chao Ren 0002, Honggang Chen, Tingrong Zhang
IEEE Signal Process. Lett.4
2020 Single depth map super-resolution via joint non-local and local modeling
abstract
Depth maps are widely used in 3D imaging techniques because of the appearance of the consumer depth cameras. However, the practical application of the depth map is limited by the poor image quality. In this paper, we propose a novel framework for the single depth map super-resolution via joint the local and non-local constraints simultaneously in the depth map. For the non-local constraint, we use the group-based sparse representation to explore the non-local self-similarity of the depth map. For the local constraint, we first estimate gradient images in different directions of the desired high-resolution (HR) depth map, and then build a multi-directional gradient guided regularizer using these estimated gradient images to describe depth gradients with different orientations. Finally, the two complementary regularizers are cast into a unified optimization framework to obtain the desired HR image. The experimental results show that the proposed method can achieve better depth super-resolution performance than state-of-the-art methods.
Chao Ren 0002, Honggang Chen, Ce Zhu
MMSP3
2020 A quality enhancement framework with noise distribution characteristics for high efficiency video coding
Weiheng Sun, Xiaohai He, Honggang Chen, Ray E. Sheriff, Shuhua Xiong
Neurocomputing3
2020 Adaptive image coding efficiency enhancement using deep convolutional neural networks
Honggang Chen, Xiaohai He, Cheolhong An, Truong Q. Nguyen
Inf. Sci.1
2019 Single image super-resolution incorporating example-based gradient profile estimation and weighted adaptive p-norm
Tao Li 0014, Xiucheng Dong, Honggang Chen
Neurocomputing3
2019 Deep Wide-Activated Residual Network Based Joint Blocking and Color Bleeding Artifacts Reduction for 4: 2: 0 JPEG-Compressed Images
abstract
Blocking and color bleeding are two well-known artifacts for 4:2:0 JPEG-compressed images. Blocking mainly results from the block-level quantization of the luma component, while color bleeding is mainly caused by the subsampling and quantization of chroma components. Restoring luma can reduce blocking distortion, but with little influence on color bleeding. On the contrary, color bleeding can be removed via chroma components restoration. This letter proposes a deep wide-activated residual network for reducing blocking and color bleeding artifacts simultaneously, in which the luma and chroma components are jointly restored. Chroma components usually suffer from more severe distortion than the luma component due to subsampling and coarse quantization. Thus, we use the luma component to guide the restoration of chroma components. Moreover, we reduce blocking and color bleeding artifacts in low-resolution space via pixel shuffle-based decimation and assembling, which allows to obtain high restoration speed. Experimental results show that the proposed approach achieves state-of-the-art performance on joint blocking and color bleeding artifacts reduction.
Honggang Chen, Xiaohai He, Cheolhong An, Truong Q. Nguyen
IEEE Signal Process. Lett.1
2018 CISRDCNN: Super-resolution of compressed images using deep convolutional neural networks
Honggang Chen, Xiaohai He, Chao Ren 0002, Linbo Qing, Qizhi Teng
Neurocomputing1
2018 SGCRSR: Sequential gradient constrained regression for single image super-resolution
Honggang Chen, Xiaohai He, Linbo Qing, Qizhi Teng, Chao Ren 0002
Signal Process. Image Commun.1
2018 An Iterative Framework of Cascaded Deblocking and Superresolution for Compressed Images
abstract
Superresolution (SR) of compressed images is chall-enging due to the combination of resolution loss and compression artifacts. To solve these intertwined problems, the conventional cascading framework splits the solution into independent deblocking and SR subprocesses, where some existing high-frequency (HF) components are often oversmoothed during deblocking and information exchange between cascaded deblocking and SR remains untouched. In this paper, we propose an iterative cascading framework after analyzing the correlation between the two subprocesses. Deblocking is provided with a shape-adaptive low-rank prior to well preserve edges and an extra prior to restore the lost HF components. The latter prior represents an important feedback link from SR to deblocking, which is a novel design in this framework. To provide an accurate and noise-robust feedback of the extra prior, an SR method via singular value decomposition projection is also developed. The extensive experimental results demonstrate the superior performance of the proposed method.
Tao Li 0014, Xiaohai He, Linbo Qing, Qizhi Teng, Honggang Chen
IEEE Trans. Multim.5
2017 Single Image Super-Resolution via Adaptive Transform-Based Nonlocal Self-Similarity Modeling and Learning-Based Gradient Regularization
abstract
Single image super-resolution (SISR) is a challenging work, which aims to recover the missing information in an observed low-resolution (LR) image and generate the corresponding high-resolution (HR) version. As the SISR problem is severely ill-conditioned, effective prior knowledge of HR images is necessary to well pose the HR estimation. In this paper, an effective SISR method is proposed via the local structure-adaptive transform-based nonlocal self-similarity modeling and learning-based gradient regularization (LSNSGR). The LSNSGR exploits both the natural and learned priors of HR images, thus integrating the merits of conventional reconstruction-based and learning-based SISR algorithms. More specifically, on the one hand, we characterize nonlocal self-similarity prior (natural prior) in transform domain by using the designed local structure-adaptive transform; on the other hand, the gradient prior (learned prior) is learned via the jointly optimized regression model. The former prior is effective in suppressing visual artifacts, while the latter performs well in recovering sharp edges and fine structures. By incorporating the two complementary priors into the maximum a posteriori-based reconstruction framework, we optimize a hybrid L1- and L2-regularized minimization problem to achieve an estimation of the desired HR image. Extensive experimental results suggest that the proposed LSNSGR produces better HR estimations than many state-of-the-art works in terms of both perceptual and quantitative evaluations.
Honggang Chen, Xiaohai He, Linbo Qing, Qizhi Teng
IEEE Trans. Multim.1
2016 Single image super resolution using local smoothness and nonlocal self-similarity priors
Honggang Chen, Xiaohai He, Qizhi Teng, Chao Ren 0002
Signal Process. Image Commun.1