EDBT 2026 Demo / reviewers in the wild / expert
Xiaohai He
dblp:132/5827
· DBLP profile ↗
117ranked-venue papers
0as first author
76since 2021 · last 2026
0000-0001-8399-3172ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 57 · 29 since 2021Artificial intelligence and machine learning · 52 · 41 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A sliced-Wasserstein and neural network framework for statistically controllable 3D microstructure reconstruction
Zhenchuan Ma, Qizhi Teng, Pengcheng Yan, Lindong Li, Kirill M. Gerke, Marina V. Karsanina, Xiaohai He |
Comput. Aided Des. | 7 |
| 2026 | Subjective and objective evaluation of visual security in perceptually encrypted images
Xiaodong Bi, Xiaohai He, Zeming Zhao, Haitao Wei, Shuhua Xiong, Zheng Liu 0002, Ray E. Sheriff |
Expert Syst. Appl. | 2 |
| 2026 | Enhancing arbitrary-scale super-resolution with scale-aware multiscale nonlocal feature extraction and local structure-adaptive upsampling
Honggang Chen, Shuhua Xiong, Xiaohai He |
Neurocomputing | 5 |
| 2026 | DSRIR: Dynamic spatial refinement learning for progressive all-in-one image restoration
Xiao Liu 0022, Yutong Yang, Zhengyong Wang, Xiaohai He, Honggang Chen, Yi Li 0069, Pingyu Wang |
Inf. Process. Manag. | 5 |
| 2026 | Blind visual security assessment using a simple parallel dual-stream network
Xiaodong Bi, Xiaohai He, Zeming Zhao, Shuhua Xiong, Honggang Chen, Ray E. Sheriff |
Signal Process. Image Commun. | 2 |
| 2026 | Degradation-Aware Contrastive Learning for Blind Image Quality AssessmentabstractImages affected by the same distortion type and level usually have consistent statistical characteristics, while different distortions exhibit significant discriminability. Inspired by this observation, this paper proposes a Degradation-aware Contrast Learning (DCL) framework to explicitly model degradation properties for Blind Image Quality Assessment (BIQA). First, a Latent Degradation Space (LDS) is constructed via self-supervised contrastive learning to effectively capture degradation features from distorted images. Then, deep semantic features are extracted using a pre-trained model and fused with the degradation features to provide complementary bias information. Finally, the fused features are mapped to perceptual quality scores by a regression model. The experimental results show that the proposed method outperforms existing state-of-the-art BIQA methods in terms of prediction accuracy, robustness, and generalization ability. Xiaodong Bi, Xiaohai He, Shuhua Xiong, Zheng Liu 0002, Ray E. Sheriff |
IEEE Signal Process. Lett. | 2 |
| 2026 | Human-Machine Vision Collaboration Based Rate Control Scheme for VVCabstractWith the widespread adoption of smart terminals, compressed video is increasingly utilized in the receiver for purposes beyond human vision. Conventional video coding standards are optimized primarily for human visual perception and often fail to accommodate the distinct requirements of machine vision. To simultaneously satisfy the perceptual needs and the analytical demands, we propose a novel rate control scheme based on Versatile Video Coding (VVC) for human-machine vision collaborative video coding. Specifically, we employ the You Only Look Once (YOLO) network to extract task-relevant features for machine vision and formulate a detection feature weight based on these features. Leveraging the feature weight and the spatial location information of Coding Tree Units (CTUs), we propose a region classification algorithm that partitions a frame into machine vision-sensitive region (MVSR) and machine vision non-sensitive region (MVNR). Subsequently, we develop an enhanced and refined bit allocation strategy that performs region-level and CTU-level bit allocation, thereby improving the precision and effectiveness of the rate control. Experimental results demonstrate that the scheme improves machine task detection accuracy while preserving perceptual quality for human observers, effectively meeting the dual encoding requirements of human and machine vision. Zeming Zhao, Xiaohai He, Xiaodong Bi, Shuhua Xiong |
IEEE Signal Process. Lett. | 2 |
| 2026 | Spatial-Temporal Correlation Information-Based Rate Control for Versatile Video CodingabstractAlthough lambda-domain-based rate control is widely used in video encoders, developing an efficient rate control scheme for Coding Tree Units (CTUs) under the rate-distortion (R-D) principle remains a significant challenge. In this paper, we propose a spatial-temporal correlation information-based rate control scheme for Versatile Video Coding (VVC), aiming to improve coding performance. We introduce a weight estimation network to establish a CTU-level bit allocation strategy that fully exploits spatial-temporal contextual information. Moreover, the CTU-level coding parameter λ is adaptively optimized based on a dependency factor derived from distortion dependency information in both the spatial and temporal domains. Experimental results demonstrate that, compared to the default VVC rate control, the proposed scheme achieves BD-Rate savings of 6.48%, 17.33% and 13.75% in terms of the Peak Signal-to-Noise Ratio (PSNR), the Multi-Scale Structural Similarity Index (MS-SSIM) and the Video Multimethod Assessment Fusion (VMAF), respectively, under the Low Delay_P (LDP) configuration in the VVC Test Model (VTM) 19.0. Furthermore, the proposed method outperforms other state-of-the-art rate control schemes. Zeming Zhao, Xiaohai He, Shuhua Xiong, Meng Wang 0017, Shiqi Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Efficient Coding Parameters Optimization for Rate Control in Versatile Video CodingabstractIn mainstream video encoders, rate control is crucial in scenarios with limited bandwidth. Within the existing R-$\lambda$model-based rate control scheme, the coding parameters ($\alpha$and$\beta$) are directly involved in the calculation of$\lambda$, and their values have a significant impact on the calculation results. By selecting precise and appropriate coding parameters, enhanced rate control and rate-distortion performance can be realized. However, during the mapping of target bits to$\lambda$,$\alpha$and$\beta$often inadequately consider the rate-distortion attributes and the content characteristics of the coding units. This paper introduces a parameter optimization algorithm for rate control in Versatile Video Coding (VVC), aimed at enhancing coding efficiency. Utilizing the actual coded contexts of the coded Coding Tree Unit (CTU) alongside pre-coding information, we establish a parameters relationship model to deliver better coding parameters according to the rate-distortion attributes of the current CTU. Furthermore, leveraging coded contexts from spatially adjacent CTUs and the feature complexity, we propose a spatial coupling strategy to further improve the preceding coding parameters, considering the content characteristics of CTU. The proposed coding parameter optimization algorithm is implemented in the rate control of the VVC test model (VTM). Experimental findings indicate that this optimization algorithm accomplishes BD-rate savings concerning Peak Signal-to-Noise Ratio (PSNR) as well as the Multiscale Structural Similarity Index Metric (MS-SSIM) across various configurations. In addition, a more stable buffer status and enhanced visual quality are visible, which highlights the benefits of the proposed algorithm. Zeming Zhao, Xiaohai He, Xiaodong Bi, Qizhi Teng, Shuhua Xiong |
IEEE Trans. Multim. | 2 |
| 2026 | Rate Control for 360$^{\circ }$ Versatile Video Coding Based on Visual Gaze MechanismabstractIn the past few years, 360° video has started to infiltrate various aspects of daily life. Although there have been significant developments in 360° video coding technology, understanding of the human visual gaze mechanism has been somewhat overlooked. In this paper, we propose a rate control scheme for 360° Versatile Video Coding (VVC) based on a human visual gaze mechanism, targeting at improving the coding performance and bitrate accuracy. More specifically, based on the Equi-rectangular Projection (ERP) format, latitude information is systematically analyzed and a stripe-level bit allocation scheme is established, to better mitigate the projection distortion. Subsequently, the Lagrange parameter λ is further optimized with distortion dependency and identification of the visual gaze guided key Coding Tree Units (CTUs). The proposed rate control scheme is implemented on the VVC Test Model for 360° video. Experimental results show that the proposed rate control scheme can achieve BD-rate savings in terms of Weighted to Spherically uniform-Peak Signal-to-Noise Ratio (WS-PSNR) and Sphere-Peak Signal-to-Noise Ratio (S-PSNR) under the various configurations, respectively. Meanwhile, a healthier buffer status and better visual quality can be observed, further demonstrating the advantages of the proposed scheme. Zeming Zhao, Meng Wang 0017, Xiangjie Sui, Peilin Chen 0001, Xiaohai He, Shiqi Wang 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | CTU-Level Rate Control with λ Optimization Based on Visual Gaze Mechanism for 360-Degree Versatile Video CodingabstractUnderstanding the human visual gaze mechanism is crucial for enhancing 360° video coding technology. This paper presents a Coding Tree Unit (CTU)-level rate control scheme with λ optimization strategy for 360° Versatile Video Coding (VVC), with the aim of enhancing rate-distortion performance and bitrate accuracy. Specifically, the Lagrange parameter λ is optimized with consideration of distortion dependency and the identification of key CTUs guided by visual gaze, which are derived from a 360° video path generation network, thoroughly integrating the characteristics of the human visual gaze. Experimental results show that the proposed scheme achieves BD-rate savings in terms of Weighted to Spherically uniform-Peak Signal-to-Noise Ratio (WS-PSNR) and Sphere-Peak Signal-to-Noise Ratio (S-PSNR) across various coding configurations. Zeming Zhao, Meng Wang 0017, Xiangjie Sui, Xiaohai He, Shiqi Wang 0001 |
ICIP | 4 |
| 2025 | PFCPNet: A progressive feature correction and prompt network for robust real-world image denoising
Yizhong Pan, Xiaohai He, Zhengyong Wang, Chao Ren 0002 |
Neurocomputing | 3 |
| 2025 | Real-world blind image super-resolution with mixed and probabilistic scheme based synthetic degradation pipeline
Xiao Liu 0022, Zhengyong Wang, Xiaohai He, Chao Ren 0002 |
Knowl. Based Syst. | 4 |
| 2025 | QP-adaptive compressed video super-resolution with coding priors
Tingrong Zhang, Zhengxin Chen, Xiaohai He, Chao Ren 0002, Qizhi Teng |
Signal Process. | 3 |
| 2025 | Enhanced Attention Context Model for Learned Image CompressionabstractRecently, deep learning has witnessed encouraging advances in image compression. An accurate entropy model, which estimates the probability distribution of the latent representation and reduces the bits required for compressing an image, is one of the keys to the success of learned image compression methods. The latent representation presents potential correlations in local, non-local, and cross-channel contexts. However, most entropy models only consider partial correlations, leading to suboptimal entropy estimation. In this letter, we propose a novel enhanced attention context model (EACM) to make full use of various correlations between latent elements for accurate entropy estimation. The proposed EACM contains a local spatial attention block (LSAB), a local channel attention block (LCAB), a global spatial attention block (GSAB), and a global channel attention block (GCAB). LSAB, LCAB, GSAB, and GCAB are carefully designed to adaptively exploit local spatial, local channel, global spatial, and global channel correlations, respectively. The experimental results on benchmark datasets show that our image compression model with the proposed EACM outperforms several state-of-the-art methods quantitatively and qualitatively. Zhengxin Chen, Xiaohai He, Chao Ren 0002, Tingrong Zhang, Shuhua Xiong |
IEEE Signal Process. Lett. | 2 |
| 2025 | Plug-and-Play General Image Registration for Misaligned Multi-Modal Image FusionabstractContemporary works in multi-modal image fusion often excessively rely on aligned source images, resulting in limited practicality when encountering misaligned data. However, there is still a significant gap in developing effective multimodal image registration methods to address this problem. Moreover, existing multi-modal image registration models are largely restricted to specific types of multi-modal data, lacking a general model applicable to diverse multi-modal data types. To address th above issues, this study introduces a novel method named PGMR, which stands as the first plug-and-play general multi-modal image registration model. PGMR comprises three components: Modality Prompt Module (MPM), Universal Registration Framework (URF), and Detail Enhancement Module (DEM). URF serves as the fundamental registration framework, handling both rigid and non-rigid deformations to achieve basic multi-modal image registration. MPM, one core component of this paper, is embedded within URF. Leveraging prompt learning, MPM dynamically integrates modality prompts into the intermediate output of URF, not only alleviating modality discrepancies but also promoting the ability of the registration model across various multi-modal data types. DEM is a detail enhancement module for multi-modal image registration. It can enrich the details of registration results, thereby enhancing the effectiveness of subsequent tasks. We evaluate the performance of PGMR on four multi-modal types and extensive experiments validate the feasibility of PGMR, demonstrating the superiority of our method compared to state-of-the-art alternatives. The code will be available at https://github.com/stwts/PGMR. Tianheng Zheng, Guanglu Dong, Xiaohai He, Chao Ren 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Transformer-Style Convolutional Network for Efficient Natural and Industrial Image SuperresolutionabstractSingle image superresolution (SISR) is a critical task in computer vision with significant applications in both natural and industrial contexts. Although transformer-based approaches for SISR have achieved notable progress due to their exceptional representational capabilities, their quadratic computational complexity poses challenges for deployment on devices with limited resources. Conversely, convolutional networks (ConvNets) are inherently efficient but have difficulty capturing long-range pixel relationships because of their focus on spatial locality. This gives rise to a complementary relationship between the representational power of transformers and the efficiency of ConvNets, both of which are essential for practical applications. Motivated by this, in this article, we introduce TSCN, a novel transformer-style ConvNet. Our analysis highlights the strengths of transformers, including large-range dependencies modeling, two-order features interaction, input self-adaptation, and incorporating advanced components. Based on these insights, we guide the design of ConvNets to fully exploit these characteristics. Specifically, we rethink spatial convolution to enhance the modeling of spatial features and modify the macrostructure of the transformer by replacing self-attention and feed-forward network with the large-range multiorder convolution modulation (LMCM) layer and spatial awareness dynamic feature flow (SADFF) layer. The LMCM integrates reweighting into the large-range convolutional modulation technology, allowing self-adaptive recalibration of input representations using convolutional features as weight matrices and multiorder features interaction. In addition, the SADFF introduces spatial awareness, locality, and dynamic information flow modulation between layers. Experimental results demonstrate that our TSCN outperforms the state-of-the-art method SRFormer on multiple benchmarks by 0.03$\sim$0.17 dB, while using fewer parameters and computations. Xiao Liu 0022, Zhengyong Wang, Xiaohai He, Haosong Gou, Chao Ren 0002 |
IEEE Trans. Ind. Informatics | 4 |
| 2025 | A compressed video quality enhancement algorithm based on CNN and transformer hybrid network
Xiaohai He, Shuhua Xiong, Haibo He, Honggang Chen |
J. Supercomput. | 2 |
| 2024 | An Attention Transformer-Based Method for the Modelling of Functional Connectivity and the Diagnosis of Autism Spectrum Disorder
Linbo Qing, Yanteng Zhang, Xiaohai He, Yonghong Peng |
ICPR (12) | 6 |
| 2024 | Semantic and geometric information propagation for oriented object detection in aerial images
Xiaohai He, Honggang Chen, Linbo Qing, Qizhi Teng |
Appl. Intell. | 2 |
| 2024 | A Multi-Attention Feature Distillation Neural Network for Lightweight Single Image Super-ResolutionabstractIn recent years, remarkable performance improvements have been produced by deep convolutional neural networks (CNN) for single image super-resolution (SISR). Nevertheless, a high proportion of CNN-based SISR models are with quite a few network parameters and high computational complexity for deep or wide architectures. How to more fully utilize deep features to make a balance between model complexity and reconstruction performance is one of the main challenges in this field. To address this problem, on the basis of the well-known information multi-distillation model, a multi-attention feature distillation network termed as MAFDN is developed for lightweight and accurate SISR. Specifically, an effective multi-attention feature distillation block (MAFDB) is designed and used as the basic feature extraction unit in MAFDN. With the help of multi-attention layers including pixel attention, spatial attention, and channel attention, MAFDB uses multiple information distillation branches to learn more discriminative and representative features. Furthermore, MAFDB introduces the depthwise over-parameterized convolutional layer (DO-Conv)-based residual block (OPCRB) to enhance its ability without incurring any parameter and computation increase in the inference stage. The results on commonly used datasets demonstrate that our MAFDN outperforms existing representative lightweight SISR models when taking both reconstruction performance and model complexity into consideration. For example, for × 4 SR on Set5, MAFDN (597K/33.79G) obtains 0.21 dB/0.0037 and 0.10 dB/0.0015 PSNR/SSIM gains over the attention-based SR model AFAN (692K/50.90G) and the feature distillation-based SR model DDistill-SR (675K/32.83G), respectively. Yongfei Zhang, Xinying Lin, Linbo Qing, Xiaohai He, Yi Li 0069, Honggang Chen |
Int. J. Intell. Syst. | 6 |
| 2024 | Blind video quality assessment based on Spatio-Temporal Feature Resolver
Xiaodong Bi, Xiaohai He, Shuhua Xiong, Zeming Zhao, Honggang Chen, Ray E. Sheriff |
Neurocomputing | 2 |
| 2024 | Multi deep invariant feature learning for cross-resolution person re-identification
Weicheng Zhang, Shuhua Xiong, Xiaohai He, Honggang Chen |
Inf. Process. Manag. | 3 |
| 2024 | A channel-wise contextual module for learned intra video compression
Yanrui Zhan, Shuhua Xiong, Xiaohai He, Honggang Chen |
J. Vis. Commun. Image Represent. | 3 |
| 2024 | Fast CU partition strategy based on texture and neighboring partition information for Versatile Video Coding Intra Coding
Ruolan Yang, Xiaohai He, Shuhua Xiong, Zeming Zhao, Honggang Chen |
Multim. Tools Appl. | 2 |
| 2024 | Dual-stage feedback network for lightweight color image compression artifact reduction
Zhengxin Chen, Xiaohai He, Tingrong Zhang, Shuhua Xiong, Chao Ren 0002 |
Neural Networks | 2 |
| 2024 | Depth map super-resolution via learned nonlocal model and enhanced local regularization
Xiaohai He, Honggang Chen, Chao Ren 0002 |
Signal Process. | 2 |
| 2024 | HCT: Chinese Medical Machine Reading Comprehension Question-Answering via Hierarchically Collaborative TransformerabstractChinese medical machine reading comprehension question-answering (cMed-MRCQA) is a critical component of the intelligence question-answering task, focusing on the Chinese medical domain question-answering task. Its purpose enable machines to analyze and understand the given text and question and then extract the accurate answer. To enhance cMed-MRCQA performance, it is essential to possess a profound comprehension and analysis of the context, deduce concealed information from the textual content and, subsequently, precisely determine the answer's span. The answer span has predominantly been defined by language items, with sentences employed in most instances. However, it is worth noting that sentences may not be properly split to varying degrees in various languages, making it challenging for the model to predict the answer zone. To alleviate this issue, this paper presents a novel architecture called HCT based on a Hierarchically Collaborative Transformer. Specifically, we presented a hierarchical collaborative method to locate the boundaries of sentence and answer spans separately. First, we designed a hierarchical encoding module to obtain the local semantic features of the corpus; second, we proposed a sentence-level self-attention module and a fused interaction-attention module to get the global information about the text. Finally, the model is trained by combining loss functions. Extensive experiments were conducted on the public dataset CMedMRC and the reconstruction dataset eMedicine to validate the effectiveness of the proposed method. Experimental results showed that the proposed method performed better than the state-of-the-art methods. Using the F1 metric, our model scored 90.4% on the CMedMRC and 73.2% on eMedicine. Xiaohai He, Luping Liu, Qingmao Fang, Honggang Chen, Yan Liu 0078 |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | Activating More Information in Arbitrary-Scale Image Super-ResolutionabstractSingle-image super-resolution (SISR) has experienced vigorous growth with the rapid development of deep learning. However, handling arbitrary scales (e.g.,integers, non-integers, or asymmetric) using a single model remains a challenging task. Existing super-resolution (SR) networks commonly employ static convolutions during feature extraction, which cannot effectively perceive changes in scales. Moreover, these continuous-scale upsampling modules only utilize the scale factors, without considering the diversity of local features. To activate more information for better reconstruction, two plug-in and compatible modules for fixed-scale networks are designed to perform arbitrary-scale SR tasks. Firstly, we design a Scale-aware Local Feature Adaptation Module (SLFAM), which adaptively adjusts the attention weights of dynamic filters based on the local features and scales. It enables the network to possess stronger representation capabilities. Then we propose a Local Feature Adaptation Upsampling Module (LFAUM), which combines scales and local features to perform arbitrary-scale reconstruction. It allows the upsampling to adapt to local structures. Besides, deformable convolution is utilized letting more information to be activated in the reconstruction, enabling the network to better adapt to the texture features. Extensive experiments on various benchmark datasets demonstrate that integrating the proposed modules into a fixed-scale SR network enables it to achieve satisfactory results with non-integer or asymmetric scales while maintaining advanced performance with integer scales. Yaoqian Zhao, Qizhi Teng, Honggang Chen, Shujiang Zhang, Xiaohai He, Yi Li 0069, Ray E. Sheriff |
IEEE Trans. Multim. | 5 |
| 2024 | DAG-YOLO: A Context-Feature Adaptive fusion Rotating Detection Network in Remote Sensing ImagesabstractObject detection in remote sensing image (RSI) research has seen significant advancements, particularly with the advent of deep learning. However, challenges such as orientation, scale, aspect ratio variations, dense object distribution, and category imbalances remain. To address these challenges, we present DAG-YOLO, a one-stage context-feature adaptive weighted fusion network that incorporates through three innovative parts. First, we integrate 1D Gaussian Angle-coding with YOLOv5 to convert the angle regression task into a classification task, establishing a more robust rotating object detection baseline, GLR-YOLO. Second, we introduce the Dual Branch Context Adaptive Modeling module, which enhances feature extraction capabilities by capturing global context information. Third, we design an adaptive detect head with the Adaptive Global Feature Aggregation and Reweighting (AGFAR) module. AGFAR addresses feature inconsistency among different output layers of the Feature Pyramid Network, retaining useful semantic information and elevating detection accuracy. Extensive experiments on public datasets DOTA-v1.0, DOTA-v1.5, and UCAS-AOD showcase mAP scores of 77.75%, 73.79%, and 90.27%, respectively. Our proposed method has the best performance among the current mainstream SOTA methods, which proves its effectiveness in RSI object detection. Zhenjiang Guo, Xiaohai He, Linbo Qing, Honggang Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | A joint CNN-GNN framework for early diagnosis of AD using multi-source multi-modal dataabstractWith the abundance of medical data, computer-aided AD diagnosis using multi-source and multi-modal data is a hotspot and trend in research, which brings more possibilities for the realization of accurate assessment of cognitive impairment diseases. Currently, the AD diagnosis based on convolutional neural network (CNN) is still the main method. But clinically, part of data exists in non-imaging form, which makes CNNs have a lot of challenges in fusing imaging and non-imaging. Graph neural network (GNN), which extends classical CNN to non-Euclidean space by using graph topology, affords better flexibility for multi-modal data integration. In order to realize AD diagnosis based on multi-source and multi-modality data, this work takes the advantage of CNN in acquiring image features, and further combines image features and non-imaging information via GNN, proposed a joint CNN-GNN diagnostic framework. Through ablation experiments, we further analyzed the effects of MMSE score, and Apoe4 genotype on AD diagnosis on the basis of image and could provide clinical reference. In addition, our proposed method achieved further improvement in diagnostic performance. Yanteng Zhang, Qingyan Cai, Xiaohai He, Xia Ren, Lipei Zhang, Yan Liu 0078 |
BIBM | 3 |
| 2023 | Nonlocal-guided enhanced interaction spatial-temporal network for compressed video super-resolution
Junxiong Cheng, Shuhua Xiong, Xiaohai He, Chao Ren 0002, Tingrong Zhang, Honggang Chen |
Appl. Intell. | 3 |
| 2023 | Image classification based on self-distillation
Linbo Qing, Xiaohai He, Honggang Chen, Qiang Liu 0021 |
Appl. Intell. | 3 |
| 2023 | Self-supervised cycle-consistent learning for scale-arbitrary real-world single image super-resolution
Honggang Chen, Xiaohai He, Yuanyuan Wu 0001, Linbo Qing, Ray E. Sheriff |
Expert Syst. Appl. | 2 |
| 2023 | RestorNet: An efficient network for multiple degradation image restoration
Honggang Chen, Haosong Gou, Zhengyong Wang, Xiaohai He, Linbo Qing, Ray E. Sheriff |
Knowl. Based Syst. | 6 |
| 2023 | An end-to-end multimodal 3D CNN framework with multi-level features for the prediction of mild cognitive impairment
Yanteng Zhang, Xiaohai He, Charlene Zhi Lin Ong, Yan Liu 0078, Qizhi Teng |
Knowl. Based Syst. | 2 |
| 2023 | Multi-relation graph convolutional network for Alzheimer's disease diagnosis using structural MRI
Xiaohai He, Linbo Qing, Xiang Chen 0008, Yan Liu 0078, Honggang Chen |
Knowl. Based Syst. | 2 |
| 2023 | PM-ARNN: 2D-TO-3D reconstruction paradigm for microstructure of porous media via adversarial recurrent neural network
Xiaohai He, Qizhi Teng, Junfang Cui, Xiucheng Dong |
Knowl. Based Syst. | 2 |
| 2023 | Block-correlation-based intra prediction for VVC
Shuhua Xiong, Xiaohai He, Honggang Chen, Chao Ren 0002 |
Multim. Tools Appl. | 3 |
| 2023 | Mixed Entropy Model Enhanced Residual Attention Network for Remote Sensing Image Compression
Junjun Gao, Qizhi Teng, Xiaohai He, Zhengxin Chen, Chao Ren 0002 |
Neural Process. Lett. | 3 |
| 2023 | BDNet: A BERT-based dual-path network for text-to-image cross-modal person re-identification
Qiang Liu 0021, Xiaohai He, Qizhi Teng, Linbo Qing, Honggang Chen |
Pattern Recognit. | 2 |
| 2023 | Dynamically Optimized Human Eyes-to-Face Generation via Attribute VocabularyabstractGenerating face from human eyes, named eyes-to-face generation, is an interesting research topic of face synthesis, which has great potential in the field of public security. One of the main challenges in eyes-to-face generation is the unbalanced information between inputs and outputs, where the outputs are complete facial images while the inputs only contain limited information in the region of eyes. The existing methods generate faces directly from eyes without considering the possibly available facial information (e.g. facial attributes), resulting in inaccurate predictions and high uncertainty in those features less correlated with eyes (e.g. hairstyle, moustache, facial contour). To address this challenge, we propose a two-stage solution (named EA2F-GAN) to dynamically optimize eyes-to-face generation via attribute vocabulary. In addition, a dataset named TEAF is constructed based on the public datasets CelebA and LFW, containing 138,934 triples of eye image, attribute vocabulary, and face image. Sufficient experimental results show that, by incorporating additional facial attributes, our proposed approach can synthesize realistic face with high consistency to the original one, significantly overwhelming state-of-the-art methods. Xiaodong Luo, Xiaohai He, Xiang Chen 0008, Linbo Qing, Honggang Chen |
IEEE Signal Process. Lett. | 2 |
| 2023 | Efficient Rate Control in Versatile Video Coding With Adaptive Spatial-Temporal Bit Allocation and Parameter UpdatingabstractDespite the fact that Versatile Video Coding (VVC) has achieved superior coding performance, two major problems remain for the rate control (RC) model in VVC. First, the regions concerned by human eyes are not clear enough in the coded video due to the deviation between the target bit allocation strategy of the coding tree unit (CTU) in RC and the human visual attention mechanism (HVAM). Second, there are significant quality fluctuations in the coded video frames due to the inappropriate updating speed. To address the above problems, we propose an efficient rate control (ERC) model. Specifically, in order to make the coded video more consistent with the attention of human eyes, we extract texture and motion-based spatial-temporal information to guide the bit allocation at the CTU level. Furthermore, based on the quasi-Newton algorithm and bit error, we propose an adaptive parameter updating (APU) method with the proper updating speed to precisely control the bits per frame. The proposed ERC outperforms the default RC model of VVC Test Model (VTM) 9.1 by saving the average Bjøntegaard Delta Rate (BD-Rate) on full-frame video sequences by 3.60% and 4.94% under low delay P (LDP) and random access (RA) configurations respectively, with higher bitrate accuracy. Moreover, the Peak Signal-to-Noise Ratio (PSNR) and actual coded bits per frame in the video coded by the proposed ERC are more stable. Liqiang He, Xiaohai He, Shuhua Xiong, Zeming Zhao, Honggang Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Real Image Denoising via Guided Residual Estimation and Noise CorrectionabstractDeep learning-based methods have dominated the field of image denoising with their superior performance. Most of them belong to the non-blind denoising approaches assuming that the noise is known at a specific level. However, real-world noise is complex and usually unknown. Since the distribution and level of noise are often unavailable, it will lead to severe performance degradation for non-blind denoising methods. Therefore, introducing noise levels is crucial for the challenging real-world denoising problem. Meanwhile, we observe that noise level mismatch will bring some artifacts to the denoised images. An intuitive solution is using the intermediate denoised images to correct the inaccurate noise level maps. Thus, we introduce an iterative correction scheme, yielding better results than direct noise prediction. We further propose an effective guided feature domain denoising residual network that can promote denoising for various real-world noises using iteratively denoised features, initial image features, and noise level maps. Experimental results on real-world image datasets show that the proposed method can provide excellent visual and objective performance for the real-world denoising task. Yizhong Pan, Chao Ren 0002, Jie Huang 0036, Xiaohai He |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | CasaPuNet: Channel Affine Self-Attention- Based Progressively Updated Network for Real Image DenoisingabstractRecently, the popularity of deep learning has brought broad applications of computer vision technology in industrial information systems. However, the process of image acquisition will inevitably introduce noise, which may heavily degrade image visual quality. Most of the proposed denoising methods are nonblind and they have limited performance in removing real noise with different noise levels. To overcome this problem, we propose a deep convolutional neural network (CNN)-based blind model, i.e., channel affine self-attention (CASA) based progressively updated network (CasaPuNet) for real image denoising. First, we introduce degradation mapping module (DMM) to extract degradation information, which makes the remaining subnetwork of CasaPuNet perform nonblind denoising. Then, CasaPuNet adopts a multistage architecture, which resolves the large gap between the noisy input and clean output into several small gaps and eliminates these small gaps step by step through progressive inference. Finally, a novel CASA is designed to adaptively fuse the features from multiple stages according to input statistics. Specifically, CASA extracts channel information from different features and converts them into channel weights through an affine structure for adaptive adjustment. CASA brings a significant performance gain with a small number of parameters. Extensive experiments demonstrate that CasaPuNet outperforms state-of-the-art denoising methods both quantitatively and visually. Jie Huang 0036, Xiao Liu 0022, Yizhong Pan, Xiaohai He, Chao Ren 0002 |
IEEE Trans. Ind. Informatics | 4 |
| 2022 | UAMNer: uncertainty-aware multimodal named entity recognition in social media posts
Luping Liu, Mozhi Zhang, Linbo Qing, Xiaohai He |
Appl. Intell. | 5 |
| 2022 | A nonlocal HEVC in-loop filter using CNN-based compression noise estimation
Weiheng Sun, Xiaohai He, Honggang Chen, Shuhua Xiong |
Appl. Intell. | 2 |
| 2022 | Medical visual question answering based on question-type reasoning and semantic space constraint
Xiaohai He, Luping Liu, Linbo Qing, Honggang Chen, Yan Liu 0078, Chao Ren 0002 |
Artif. Intell. Medicine | 2 |
| 2022 | A two-stage deep generative adversarial quality enhancement network for real-world 3D CT images
Honggang Chen, Xiaohai He, Junxi Feng, Qizhi Teng |
Expert Syst. Appl. | 2 |
| 2022 | Deep dual-domain semi-blind network for compressed image quality enhancement
Jingbo He, Xiaohai He, Mozhi Zhang, Shuhua Xiong, Honggang Chen |
Knowl. Based Syst. | 2 |
| 2022 | A prior-guided deep network for real image denoising and its applications
Jie Huang 0036, Zhibo Zhao, Chao Ren 0002, Qizhi Teng, Xiaohai He |
Knowl. Based Syst. | 5 |
| 2022 | Feature separation and double causal comparison loss for visible and infrared person re-identification
Qiang Liu 0021, Xiaohai He, Mozhi Zhang, Qizhi Teng, Bo Li 0074, Linbo Qing |
Knowl. Based Syst. | 2 |
| 2022 | Fact-based visual question answering via dual-process system
Luping Liu, Xiaohai He, Linbo Qing, Honggang Chen |
Knowl. Based Syst. | 3 |
| 2022 | CMAFGAN: A Cross-Modal Attention Fusion based Generative Adversarial Network for attribute word-to-face synthesis
Xiaodong Luo, Xiang Chen 0008, Xiaohai He, Linbo Qing, Xinyue Tan |
Knowl. Based Syst. | 3 |
| 2022 | A quality enhancement network with coding priors for constant bit rate video coding
Weiheng Sun, Xiaohai He, Chao Ren 0002, Shuhua Xiong, Honggang Chen |
Knowl. Based Syst. | 2 |
| 2022 | Weakly-supervised contrastive learning-based implicit degradation modeling for blind image super-resolution
Yongfei Zhang, Ling Dong, Linbo Qing, Xiaohai He, Honggang Chen |
Knowl. Based Syst. | 5 |
| 2022 | Cross-modal multi-relationship aware reasoning for image-text matching
Xiaohai He, Linbo Qing, Luping Liu, Xiaodong Luo |
Multim. Tools Appl. | 2 |
| 2022 | DualG-GAN, a Dual-channel Generator based Generative Adversarial Network for text-to-face synthesis
Xiaodong Luo, Xiaohai He, Xiang Chen 0008, Linbo Qing |
Neural Networks | 2 |
| 2022 | Sequential Enhancement for Compressed Video Using Deep Convolutional Generative Adversarial Network
Xiaohai He, Honggang Chen, Shuhua Xiong |
Neural Process. Lett. | 2 |
| 2022 | Deep Feature Fusion Network for Compressed Video Super-Resolution
Xiaohai He, Chao Ren 0002, Tingrong Zhang |
Neural Process. Lett. | 3 |
| 2022 | An effective deep network using target vector update modules for image restoration
Sen Zhai, Chao Ren 0002, Zhengyong Wang, Xiaohai He, Linbo Qing |
Pattern Recognit. | 4 |
| 2022 | Unsupervised Real-World Image Super-Resolution via Dual Synthetic-to-Realistic and Realistic-to-Synthetic TranslationsabstractDue to the challenges of collecting paired low-resolution (LR) and high-resolution (HR) images in real-world scenarios, most existing deep convolutional neural network (CNN)-based single image super-resolution (SR) models are trained with artificially synthesized LR-HR image pairs. However, the domain gap between the synthetic data for model training and the realistic data for testing degrades SR performance significantly, which discourages the application of SR models in practice. One possible solution is to learn from unpaired real-world LR and HR images for their accessibility. Predominant strategies are mainly based on unsupervised domain translation. Despite great advances, there are still noticeable domain gaps between the realistic-like/synthetic-like images generated by unpaired translation and the true realistic/synthetic ones. To address this problem, this letter proposes an effective unsupervised SR framework based on dual synthetic-to-realistic and realistic-to-synthetic translations, namely DTSR. Specifically, to bridge the domain gap between testing and training data, the SR model is optimized using HR images and their realistic-like LR counterparts produced by the synthetic-to-realistic translation. In turn, we propose to narrow the domain gap further via applying the realistic-to-synthetic translation to realistic LR images prior to super-resolving, which also makes the SR model super-resolve simpler examples in testing relative to model training. Moreover, focal frequency and bilateral filtering losses are particularly introduced into DTSR for better details restoration and artifacts suppression. Extensive experiments show that our DTSR outperforms several state-of-the-art models in terms of both quantitative and qualitative comparisons. Honggang Chen, Ling Dong, Xiaohai He, Ce Zhu |
IEEE Signal Process. Lett. | 4 |
| 2022 | An Optimized Rate Control Algorithm in Versatile Video Coding for 360$^\circ$ VideosabstractToday, 360$°$video has become an integral part of people's lives. Despite the fact that the latest generation standard Versatile Video Coding (VVC) demonstrates a significant gain in encoding capacity over High Efficiency Video Coding (HEVC), it still has room for 360$°$video encoding improvements. To further enhance the applicability of 360$°$video coding, an optimized rate control (RC) algorithm in VVC for 360$°$video is proposed in this paper. We present an efficient extraction algorithm for obtaining the video's saliency feature. Furthermore, for the characteristics of 360$°$video, a partitioning algorithm is also proposed to divide a frame into demand and non-demand regions. Additionally, to achieve precise and rational RC, a Coding Tree Unit (CTU)-level bit allocation strategy is proposed based on the saliency feature for the above-mentioned regions. The experimental results show that the proposed RC algorithm can achieve 11.77$\%$bitrate savings and more accurate allocation compared with the default algorithm of VVC. Also, performance enhancement has been observed in comparison to the most advanced algorithm. Zeming Zhao, Xiaohai He, Shuhua Xiong, Liqiang He, Ray E. Sheriff |
IEEE Signal Process. Lett. | 2 |
| 2022 | A video compression artifact reduction approach combined with quantization parameters estimation
Xin Shuai, Linbo Qing, Mozhi Zhang, Weiheng Sun, Xiaohai He |
J. Supercomput. | 5 |
| 2022 | A Feature-Enriched Deep Convolutional Neural Network for JPEG Image Compression Artifacts Reduction and its ApplicationsabstractThe amount of multimedia data, such as images and videos, has been increasing rapidly with the development of various imaging devices and the Internet, bringing more stress and challenges to information storage and transmission. The redundancy in images can be reduced to decrease data size via lossy compression, such as the most widely used standard Joint Photographic Experts Group (JPEG). However, the decompressed images generally suffer from various artifacts (e.g., blocking, banding, ringing, and blurring) due to the loss of information, especially at high compression ratios. This article presents a feature-enriched deep convolutional neural network for compression artifacts reduction (FeCarNet, for short). Taking the dense network as the backbone, FeCarNet enriches features to gain valuable information via introducing multi-scale dilated convolutions, along with the efficient 1 ×1 convolution for lowering both parameter complexity and computation cost. Meanwhile, to make full use of different levels of features in FeCarNet, a fusion block that consists of attention-based channel recalibration and dimension reduction is developed for local and global feature fusion. Furthermore, short and long residual connections both in the feature and pixel domains are combined to build a multi-level residual structure, thereby benefiting the network training and performance. In addition, aiming at reducing computation complexity further, pixel-shuffle-based image downsampling and upsampling layers are, respectively, arranged at the head and tail of the FeCarNet, which also enlarges the receptive field of the whole network. Experimental results show the superiority of FeCarNet over state-of-the-art compression artifacts reduction approaches in terms of both restoration capacity and model complexity. The applications of FeCarNet on several computer vision tasks, including image deblurring, edge detection, image segmentation, and object detection, demonstrate the effectiveness of FeCarNet further. Honggang Chen, Xiaohai He, Linbo Qing, Qizhi Teng |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Adaptive Consistency Prior Based Deep Network for Image DenoisingabstractRecent studies have shown that deep networks can achieve promising results for image denoising. However, how to simultaneously incorporate the valuable achievements of traditional methods into the network design and improve network interpretability is still an open problem. To solve this problem, we propose a novel model-based denoising method to inform the design of our denoising network. First, by introducing a non-linear filtering operator, a reliability matrix, and a high-dimensional feature transformation function into the traditional consistency prior, we propose a novel adaptive consistency prior (ACP). Second, by incorporating the ACP term into the maximum a posteriori framework, a model-based denoising method is proposed. This method is further used to inform the network design, leading to a novel end-to-end trainable and interpretable deep denoising network, called DeamNet. Note that the unfolding process leads to a promising module called dual element-wise attention mechanism (DEAM) module. To the best of our knowledge, both our ACP constraint and DEAM module have not been reported in the previous literature. Extensive experiments verify the superiority of DeamNet on both synthetic and real noisy image datasets. Chao Ren 0002, Xiaohai He, Chuncheng Wang, Zhibo Zhao |
CVPR | 2 |
| 2021 | Deep Deblocker Driven Adaptive Iteration Scheme for Compressed Image RecoveryabstractIt is challenging to propose a flexible and effective framework for various JPEG compressed image recovery (CIR) tasks. In this paper, we propose a novel deep deblocker-driven adaptive iteration scheme, which can quickly and flexibly address various CIR tasks. First, a novel fidelity (NF) is introduced into CIR, and then the CIR problem is divided into inversion and deblocking subproblems by our improved split Bregman iteration (ISBI) algorithm. Next, we design a set of compact yet effective deep deblockers. These deblockers are used as implicit priors and also used for NF in the CIR problem. The convergence of our method is proved as well. To the best of our knowledge, our method is the first work to use deblockers as implicit priors. Extensive experiments demonstrate the effectiveness of our CIR method. Chao Ren 0002, Xiaohai He, Linbo Qing, Yuanzhouhan Cao |
ICME | 2 |
| 2021 | An enhanced siamese angular softmax network with dual joint-attention for person re-identification
Jie Su 0011, Xiaohai He, Linbo Qing, Yongqiang Cheng 0001, Yonghong Peng |
Appl. Intell. | 2 |
| 2021 | Deep recursive network for image denoising with global non-linear smoothness constraint prior
Chuncheng Wang, Chao Ren 0002, Xiaohai He, Linbo Qing |
Neurocomputing | 3 |
| 2021 | Bi-directional skip connection feature pyramid network and sub-pixel convolution for high-quality object detection
Shuqi Xiong, Honggang Chen, Linbo Qing, Xiaohai He |
Neurocomputing | 6 |
| 2021 | Remote sensing image recovery via enhanced residual learning and dual-luminance scheme
Chao Ren 0002, Xiaohai He, Linbo Qing, Yuanyuan Wu 0001, Yi-Fei Pu |
Knowl. Based Syst. | 2 |
| 2021 | Compressed image restoration via deep deblocker driven unified framework
Chao Ren 0002, Qizhi Teng, Xiaohai He, Linbo Qing, Truong Q. Nguyen |
Knowl. Based Syst. | 3 |
| 2021 | An improved R-λ rate control model based on joint spatial-temporal domain information and HVS characteristics
Zeming Zhao, Shuhua Xiong, Weiheng Sun, Xiaohai He, Feiran Zhang |
Multim. Tools Appl. | 4 |
| 2021 | Enhanced wide-activated residual network for efficient and accurate image deblocking
Zhengxin Chen, Xiaohai He, Chao Ren 0002, Pradeep Karn, Shuhua Xiong |
Signal Process. Image Commun. | 2 |
| 2021 | Enhanced Separable Convolution Network for Lightweight JPEG Compression Artifacts ReductionabstractJPEG images are usually corrupted by various undesirable compression artifacts resulted from block-wise coarse quantization on discrete cosine transform coefficients. In recent years, deep convolutional neural networks (CNNs) have made spectacular achievements in compression artifacts reduction. However, most deep CNNs are difficult to be implemented on mobile devices due to their large number of parameters and operations. In this letter, we propose a novel deep CNN called ESCNet for lightweight JPEG compression artifacts reduction, in which enhanced separable convolution (ESConv) is carefully designed to make full use of image multi-scale information for better dense pixel value predictions. Specifically, ESConv consists of a grouped multi-scale dual depth-wise convolution (GMDDConv) and a wide-activated dual point-wise convolution (WDPConv). GMDDConv is dedicated to efficiently extracting abundant image multi-scale spatial features, which will be sent to WDPConv for effective non-linear feature fusion. The experimental results on benchmark datasets show that compared with state-of-the-art methods, our ESCNet not only achieves better performance in both objective indices and subjective quality but also greatly reduces network parameters and operations. Zhengxin Chen, Xiaohai He, Chao Ren 0002, Honggang Chen, Tingrong Zhang |
IEEE Signal Process. Lett. | 2 |
| 2021 | Learning Image Profile Enhancement and Denoising Statistics Priors for Single-Image Super-ResolutionabstractSingle-image super-resolution (SR) has been widely used in computer vision applications. The reconstruction-based SR methods are mainly based on certain prior terms to regularize the SR problem. However, it is very challenging to further improve the SR performance by the conventional design of explicit prior terms. Because of the powerful learning ability, deep convolutional neural networks (CNNs) have been widely used in single-image SR task. However, it is difficult to achieve further improvement by only designing the network architecture. In addition, most existing deep CNN-based SR methods learn a nonlinear mapping function to directly map low-resolution (LR) images to desirable high-resolution (HR) images, ignoring the observation models of input images. Inspired by the split Bregman iteration (SBI) algorithm, which is a powerful technique for solving the constrained optimization problems, the original SR problem is divided into two subproblems: 1) inversion subproblem and 2) denoising subproblem. Since the inversion subproblem can be regarded as an inversion step to reconstruct an intermediate HR image with sharper edges and finer structures, we propose to use deep CNN to capture low-level explicit image profile enhancement prior (PEP). Since the denoising subproblem aims to remove the noise in the intermediate image, we adopt a simple and effective denoising network to learn implicit image denoising statistics prior (DSP). Furthermore, the penalty parameter in SBI is adaptively tuned during the iterations for better performance. Finally, we also prove the convergence of our method. Thus, the deep CNNs are exploited to capture both implicit and explicit image statistics priors. Due to SBI, the SR observation model is also leveraged. Consequently, it bridges between two popular SR approaches: 1) learning-based method and 2) reconstruction-based method. Experimental results show that the proposed method achieves the state-of-the-art SR results. Chao Ren 0002, Xiaohai He, Yi-Fei Pu, Truong Q. Nguyen |
IEEE Trans. Cybern. | 2 |
| 2020 | EyesGAN: Synthesize human face from human eyes
Xiaodong Luo, Xiaohai He, Linbo Qing, Xiang Chen 0008, Luping Liu |
Neurocomputing | 2 |
| 2020 | A quality enhancement framework with noise distribution characteristics for high efficiency video coding
Weiheng Sun, Xiaohai He, Honggang Chen, Ray E. Sheriff, Shuhua Xiong |
Neurocomputing | 2 |
| 2020 | Reduction of JPEG compression artifacts based on DCT coefficients prediction
Mengdi Sun, Xiaohai He, Shuhua Xiong, Chao Ren 0002, Xinglong Li |
Neurocomputing | 2 |
| 2020 | Adaptive image coding efficiency enhancement using deep convolutional neural networks
Honggang Chen, Xiaohai He, Cheolhong An, Truong Q. Nguyen |
Inf. Sci. | 2 |
| 2020 | An experimental study of relative total variation and probabilistic collaborative representation for iris recognition
Pradeep Karn, Xiaohai He, Yanteng Zhang |
Multim. Tools Appl. | 2 |
| 2020 | Zero-shot recognition with latent visual attributes learning
Yurui Xie, Xiaohai He, Xiaodong Luo |
Multim. Tools Appl. | 2 |
| 2020 | Image deblocking via shape-adaptive low-rank prior and sparsity-based detail enhancement
Chao Ren 0002, Xinglong Li, Xiaohai He |
Signal Process. Image Commun. | 5 |
| 2019 | Facial Expression Recognition Based on Group Domain Random Frame Extraction
Yibo Huang 0003, Linbo Qing, Xiaohai He |
ICIG (1) | 6 |
| 2019 | Machine learning-based H.264/AVC to HEVC transcoding via motion information reuse and coding mode similarity analysisabstractHigh‐efficiency video coding (HEVC), which is the latest video coding standard, is expected to have a dominant position in the market in the near future. However, most video resources are now encoded using the H.264/AVC standard. Consequently, there is a growing need for fast H.264/AVC to HEVC transcoders to facilitate the migration to the updated standard. This paper proposes a fast H.264/AVC to HEVC transcoding scheme, which constructs a three‐level classifier using an optimised tree‐augmented Naive Bayesian approach to predict the HEVC coding unit depth. A feature selection method is then proposed to improve prediction accuracy. A motion vector (MV) calculation method is also proposed to reduce the complexity of MV prediction in HEVC by reusing MVs from H.264/AVC. Experimental results show that, compared with other state‐of‐the‐art transcoding algorithms, the proposed algorithm considerably reduces coding complexity while causing only negligible rate‐distortion degradation. Xiaohai He, Linbo Qing, Shan Su, Shuhua Xiong |
IET Image Process. | 2 |
| 2019 | Deep Wide-Activated Residual Network Based Joint Blocking and Color Bleeding Artifacts Reduction for 4: 2: 0 JPEG-Compressed ImagesabstractBlocking and color bleeding are two well-known artifacts for 4:2:0 JPEG-compressed images. Blocking mainly results from the block-level quantization of the luma component, while color bleeding is mainly caused by the subsampling and quantization of chroma components. Restoring luma can reduce blocking distortion, but with little influence on color bleeding. On the contrary, color bleeding can be removed via chroma components restoration. This letter proposes a deep wide-activated residual network for reducing blocking and color bleeding artifacts simultaneously, in which the luma and chroma components are jointly restored. Chroma components usually suffer from more severe distortion than the luma component due to subsampling and coarse quantization. Thus, we use the luma component to guide the restoration of chroma components. Moreover, we reduce blocking and color bleeding artifacts in low-resolution space via pixel shuffle-based decimation and assembling, which allows to obtain high restoration speed. Experimental results show that the proposed approach achieves state-of-the-art performance on joint blocking and color bleeding artifacts reduction. Honggang Chen, Xiaohai He, Cheolhong An, Truong Q. Nguyen |
IEEE Signal Process. Lett. | 2 |
| 2019 | High-Order Statistical Modeling Based on a Decision Tree for Distributed Video CodingabstractAiming at low-complexity encoding, distributed video coding (DVC) based on the Wyner-Ziv theorem has attracted significant attention. However, there is still a compression performance gap between the state-of-the-art DVC and the conventional video coding. One of the most important factors is the efficient estimation of the source correlation statistics. The first-order Laplacian distribution has been widely used for source correlation modeling, but is not effective enough; high-order statistical modeling is necessary, but needs more context features for the estimation of a source symbol's conditional probability. How to analyze the strength of the correlation between the source symbol and the context features in order to utilize the features effectively is crucial for such modeling. In this paper, the estimation of the source statistical distribution is first treated as a classification problem. The symbols of the source can be classified into different classes when the relevant context features are given. Then decision tree learning is introduced to analyze the strength of the correlation between the source symbol and the context features. Specifically, by constructing the decision trees composed of the selected context features, the selected context features can be organized effectively to derive the rules to estimate the current symbol's conditional probability, upon which the high-order statistical modeling is designed. Experimental results show that the proposed model can achieve significant coding gain over existing DVC systems, especially for natural videos with high motion intensity. Linbo Qing, Wenjun Zeng 0001, Xiaohai He |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | Enhanced Non-Local Total Variation Model and Multi-Directional Feature Prediction Prior for Single Image Super ResolutionabstractIt is widely acknowledged that single image super-resolution (SISR) methods play a critical role in recovering the missing high-frequencies in an input low-resolution image. As SISR is severely ill-conditioned, image priors are necessary to regularize the solution spaces and generate the corresponding high-resolution image. In this paper, we propose an effective SISR framework based on the enhanced non-local similarity modeling and learning-based multi-directional feature prediction (ENLTV-MDFP). Since both the modeled and learned priors are exploited, the proposed ENLTV-MDFP method benefits from the complementary properties of the reconstruction-based and learning-based SISR approaches. Specifically, for the non-local similarity-based modeled prior [enhanced non-local total variation, (ENLTV)], it is characterized via the decaying kernel and stable group similarity reliability schemes. For the learned prior [multi-directional feature prediction prior, (MDFP)], it is learned via the deep convolutional neural network. The modeled prior performs well in enhancing edges and suppressing visual artifacts, while the learned prior is effective in hallucinating details from external images. Combining these two complementary priors in the MAP framework, a combined SR cost function is proposed. Finally, the combined SR problem is solved via the split Bregman iteration algorithm. Based on the extensive experiments, the proposed ENLTV-MDFP method outperforms many state-of-the-art algorithms visually and quantitatively. Chao Ren 0002, Xiaohai He, Yi-Fei Pu, Truong Q. Nguyen |
IEEE Trans. Image Process. | 2 |
| 2019 | Improved Low-Bitrate HEVC Video Coding Using Deep Learning Based Super-Resolution and Adaptive Block PatchingabstractGood-quality video coding for low-bitrate applications is essential for narrow bandwidth transmission and limited capacity storage. In this paper, we propose an adaptive downsampling-based coding model to improve the low-bitrate compression efficiency of high-efficiency video coding (HEVC). At the encoder, the video sequence is adaptively divided into key frames (KFs) and nonkey frames (NKFs), which are encoded at the original resolution and at a reduced resolution, respectively. At the decoder, a super-resolution method based on deep learning and gradient transformation is used to upscale the NKFs. To improve the quality of NKFs without additional information during decoding, we use motion estimation to find the most similar blocks between the upscaled NKFs and the associated high-resolution KFs. Then, an adaptive patching-based method is used to warp the low-quality NKF blocks with the high-quality KF blocks. Experimental results indicate that for standard high-definition test video sequences, the maximum improvement in the peak signal-to-noise ratio can reach 3.54 dB, and the critical bitrate can reach 9.89 Mb/s at a low bitrate when compared to HEVC. These results demonstrate significant improvements compared to existing methods. Xiaohai He, Linbo Qing, Qizhi Teng, Songfan Yang |
IEEE Trans. Multim. | 2 |
| 2019 | Adjusted Non-Local Regression and Directional Smoothness for Image RestorationabstractImage restoration (IR) problems are very important in many low-level vision tasks. Due to their ill-posed natures, image priors are widely used to regularize the solution spaces. Recently, patch-based non-local self-similarity has shown great potential in IR problems, leading to many effective non-local priors. Their performance largely depends on whether the non-local self-similarity of the underlying image can be fully exploited. However, most of these priors, including non-local regression (NLR), only utilize the center pixel of each patch to model the non-local feature, which is suboptimal. We propose an effective overlap-based non-local regression (ONLR) to fully exploit the non-local similar patches: first, the concept of overlap-based similar pixels group (OSPG) is introduced; second, for each pixel within an OSPG, the non-local weight is obtained via a novel similarity measurement method; third, based on the consistency assumption, the non-local fitting deviations (NLFDs) by using OSPGs are uniformly constrained. Because of the uniform constraints, the restoration may be poor in regions where OSPGs are not reliable. Consequently, a weighting scheme is proposed to measure the OSPG reliability, leading to a novel adjusted non-local regression (ANLR). In addition, the integral image technique (IIT) is adopted to speed up the similar patches search process. To further boost the ANLR, a local directional smoothness (DS) prior is proposed as a good complement of the non-local feature. Finally, a fast split Bregman iteration algorithm is designed to solve the ANLR-DS minimization problem. Extensive experiments on two typical IR problems, that is, image deblurring and super resolution, demonstrate the superiority of the proposed method compared to many state-of-the-art IR methods. Chao Ren 0002, Xiaohai He, Truong Q. Nguyen |
IEEE Trans. Multim. | 2 |
| 2018 | CISRDCNN: Super-resolution of compressed images using deep convolutional neural networks
Honggang Chen, Xiaohai He, Chao Ren 0002, Linbo Qing, Qizhi Teng |
Neurocomputing | 2 |
| 2018 | Robust distributed video coding for wireless multimedia sensor networks
Linbo Qing, Xiaohai He, Xianfeng Ou |
Multim. Tools Appl. | 3 |
| 2018 | Adaptive Gradient Information and BFGS Based Inter Frame Rate Control for High Efficiency Video Coding
Yuyun Ye, Xiaohai He, Qizhi Teng, Linbo Qing, Dechun Xia |
Multim. Tools Appl. | 2 |
| 2018 | SGCRSR: Sequential gradient constrained regression for single image super-resolution
Honggang Chen, Xiaohai He, Linbo Qing, Qizhi Teng, Chao Ren 0002 |
Signal Process. Image Commun. | 2 |
| 2018 | Nonlocal Similarity Modeling and Deep CNN Gradient Prior for Super ResolutionabstractThis letter presents a novel super-resolution (SR) method via nonlocal similarity modeling and deep convolutional neural network (CNN) gradient prior (GP). Specifically, on the one hand, the group similarity reliability (GSR) strategy is proposed for improving the adaptive high-dimensional nonlocal total variation (AHNLTV) model [statistical prior, GSR-based AHNLTV (GA)], which captures the structures of the underlying high-resolution (HR) image via the image itself. On the other hand, the GP is learned by using the deep CNN (learned prior), which predicts the gradients from external images. Finally, the GA-GP approach is proposed by incorporating the two complementary priors. The results show that GA-GP achieves better performance than other state-of-the-art SR methods. Chao Ren 0002, Xiaohai He, Yi-Fei Pu |
IEEE Signal Process. Lett. | 2 |
| 2018 | An Iterative Framework of Cascaded Deblocking and Superresolution for Compressed ImagesabstractSuperresolution (SR) of compressed images is chall-enging due to the combination of resolution loss and compression artifacts. To solve these intertwined problems, the conventional cascading framework splits the solution into independent deblocking and SR subprocesses, where some existing high-frequency (HF) components are often oversmoothed during deblocking and information exchange between cascaded deblocking and SR remains untouched. In this paper, we propose an iterative cascading framework after analyzing the correlation between the two subprocesses. Deblocking is provided with a shape-adaptive low-rank prior to well preserve edges and an extra prior to restore the lost HF components. The latter prior represents an important feedback link from SR to deblocking, which is a novel design in this framework. To provide an accurate and noise-robust feedback of the extra prior, an SR method via singular value decomposition projection is also developed. The extensive experimental results demonstrate the superior performance of the proposed method. Tao Li 0014, Xiaohai He, Linbo Qing, Qizhi Teng, Honggang Chen |
IEEE Trans. Multim. | 2 |
| 2017 | 3D MRI image super-resolution for brain combining rigid and large diffeomorphic registrationabstractMost of the recent leading multiple magnetic resonance imaging (MRI) super‐resolution techniques for brain are limited to rigid motion. In this study, the authors aim to develop a super‐resolution technique with diffeomorphism mainly for longitudinal brain MRI data. For the images from different time slots, unpredicted deformation may occur. In previous studies, sole rigid registration or traditional non‐rigid registration has been frequently used to achieve multi‐plane super‐resolution. However, non‐rigid motion of two brains from different time slots is difficult to model, since brain contains a wealth of complex structure such as the cerebral cortex. In order to address such problem, rigid and large diffeomorphic registration has been embedded into their super‐resolution framework. In addition, many previous researchers use norm to achieve super‐resolution framework. In this work, norm minimisation and regularisation based on a bilateral prior are adopted. These operations ensure its robustness to the assumed model of data and noise. Their approach is evaluated using Alzheimer datasets from seven different resolutions. Results show that their reconstructions have advantages over rigid and conventional non‐rigid registration‐based super‐resolution, in terms of the root‐mean‐square error and structure similarity. Furthermore, their reconstruction results improve the precision of brain automatic segmentation. Zifei Liang, Xiaohai He, Qizhi Teng, Lingbo Qing |
IET Image Process. | 2 |
| 2017 | Tree-structured Bayesian compressive sensing via generalised inverse Gaussian distributionabstractCompressive sensing (CS) implements signal sampling and compression simultaneously, which significantly alleviates the pressure on the sampling end. However, the reconstruction algorithm is an underdetermined linear inverse problem. To solve this problem, it is crucial to involve prior knowledge regarding the reconstructed signal. In this study, the compressibility of wavelet coefficients is utilised as prior knowledge. Moreover, a generalised inverse Gaussian (GIG) distribution is integrated in the context of tree‐structured Bayesian CS (TSBCS), which also imposes the persistence property between the successive levels. Finally, variational Bayesian inference is used to infer the posterior probability distribution of the model parameters. Due to the overall algorithm is based on TSBCS, the proposal is referred to as TSBCS via a GIG distribution (TSBCS‐GIG). Experimental results show that the authors’ proposed TSBCS‐GIG algorithm outperforms other well‐known algorithms in both peak signal‐to‐noise ratio and visual quality. Maojiao Wang, Xiaohai He, Linbo Qing, Shuhua Xiong |
IET Signal Process. | 2 |
| 2017 | Moving Object Detection With a Freely Moving Camera via Background Motion SubtractionabstractDetection of moving objects in a video captured by a freely moving camera is a challenging problem in computer vision. Most existing methods often assume that the background (BG) can be approximated by dominant single plane/multiple planes or impose significant geometric constraints on BG, or utilize a complex BG/foreground probabilistic model. Instead, we propose a computationally efficient algorithm that is able to detect moving objects accurately and robustly in a general 3D scene. This problem is formulated as a coarse-to-fine thresholding scheme on the particle trajectories in the video sequence. First, a coarse foreground (CFG) region is extracted by performing reduced singular value decomposition on multiple matrices that are built from bundles of particle trajectories. Next, the BG motion of pixels in the CFG region is reconstructed by a fast inpainting method. After subtracting the BG motion, the fine foreground is segmented out by an adaptive thresholding method that is capable of solving multiple-moving-objects scenarios. Finally, the detected foreground is further refined by the mean-shift segmentation method. Extensive simulations and a comparison with the state-of-the-art methods verify the effectiveness of the proposed method. Yuanyuan Wu 0001, Xiaohai He, Truong Q. Nguyen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Single Image Super-Resolution via Adaptive High-Dimensional Non-Local Total Variation and Adaptive Geometric FeatureabstractSingle image super-resolution (SR) is very important in many computer vision systems. However, as a highly ill-posed problem, its performance mainly relies on the prior knowledge. Among these priors, the non-local total variation (NLTV) prior is very popular and has been thoroughly studied in recent years. Nevertheless, technical challenges remain. Because NLTV only exploits a fixed non-shifted target patch in the patch search process, a lack of similar patches is inevitable in some cases. Thus, the non-local similarity cannot be fully characterized, and the effectiveness of NLTV cannot be ensured. Based on the motivation that more accurate non-local similar patches can be found by using shifted target patches, a novel multishifted similar-patch search (MSPS) strategy is proposed. With this strategy, NLTV is extended as a newly proposed super-high-dimensional NLTV (SHNLTV) prior to fully exploit the underlying non-local similarity. However, as SHNLTV is very high-dimensional, applying it directly to SR is very difficult. To solve this problem, a novel statistics-based dimension reduction strategy is proposed and then applied to SHNLTV. Thus, SHNLTV becomes a more computationally effective prior that we call adaptive high-dimensional non-local total variation (AHNLTV). In AHNLTV, a novel joint weight strategy that fully exploits the potential of the MSPS-based non-local similarity is proposed. To further boost the performance of AHNLTV, the adaptive geometric duality (AGD) prior is also incorporated. Finally, an efficient split Bregman iteration-based algorithm is developed to solve the AHNLTV-AGD-driven minimization problem. Extensive experiments validate the proposed method achieves better results than many state-of-the-art SR methods in terms of both objective and subjective qualities. Chao Ren 0002, Xiaohai He, Truong Q. Nguyen |
IEEE Trans. Image Process. | 2 |
| 2017 | Single Image Super-Resolution via Adaptive Transform-Based Nonlocal Self-Similarity Modeling and Learning-Based Gradient RegularizationabstractSingle image super-resolution (SISR) is a challenging work, which aims to recover the missing information in an observed low-resolution (LR) image and generate the corresponding high-resolution (HR) version. As the SISR problem is severely ill-conditioned, effective prior knowledge of HR images is necessary to well pose the HR estimation. In this paper, an effective SISR method is proposed via the local structure-adaptive transform-based nonlocal self-similarity modeling and learning-based gradient regularization (LSNSGR). The LSNSGR exploits both the natural and learned priors of HR images, thus integrating the merits of conventional reconstruction-based and learning-based SISR algorithms. More specifically, on the one hand, we characterize nonlocal self-similarity prior (natural prior) in transform domain by using the designed local structure-adaptive transform; on the other hand, the gradient prior (learned prior) is learned via the jointly optimized regression model. The former prior is effective in suppressing visual artifacts, while the latter performs well in recovering sharp edges and fine structures. By incorporating the two complementary priors into the maximum a posteriori-based reconstruction framework, we optimize a hybrid L1- and L2-regularized minimization problem to achieve an estimation of the desired HR image. Extensive experimental results suggest that the proposed LSNSGR produces better HR estimations than many state-of-the-art works in terms of both perceptual and quantitative evaluations. Honggang Chen, Xiaohai He, Linbo Qing, Qizhi Teng |
IEEE Trans. Multim. | 2 |
| 2016 | Depth-based distributed multi-view video coding with hierarchical Wyner-Ziv framesabstractDistributed multi-view video coding (DMVC) is a new emerging multi-view video coding (MVC) scheme, in which multi-view video are encoded separately and decoded dependently, so the burden of huge computation is shifted from the encoder to the decoder side. However, there is still a large gap between the DMVC and traditional MVC in terms of compression performance. In order to improve the coding performance of DMVC, wavelet domain DMVC framework based on hierarchical Wyner-Ziv frames is proposed. With the introduction of depth map to multi-view, a fusion algorithm based on error correction with the information from depth map and adjacent views is proposed. The experimental results show better quality of the SI for the WZ frames and significant improvement in rate distortion (RD) performance are achieved. Linbo Qing, Xiaohai He, Wenshi Xiong |
VCIP | 3 |
| 2016 | Reconstruction algorithm using exact tree projection for tree-structured compressive sensingabstractTree‐structured compressive sensing (CS) shows that it is possible to recover tree‐sparse signals using fewer measurements compared with conventional CS. However, performance guarantees rely heavily on the premise that an exact tree projection (ETP) algorithm is employed. Nevertheless, for a given sparsity, the condensing sort and select algorithm in the model‐based compressive sampling matching pursuit (CoSaMP) algorithm can only yield an approximate tree projection. Therefore, in order to ensure reconstruction precision, the authors propose the combination of an ETP algorithm with the CoSaMP algorithm. Further, the hierarchical wavelet connected tree is also integrated into the ETP‐CoSaMP algorithm to offset the high computational complexity of the ETP algorithm. Experimental results indicate that the hierarchical ETP based on CoSaMP algorithm (HETP‐CoSaMP algorithm) enhances reconstruction accuracy while retaining reconstruction time that is comparable with that of the model‐based CoSaMP algorithm. Maojiao Wang, Wenhui Jing, Xiaohai He |
IET Signal Process. | 4 |
| 2016 | Rotation expanded dictionary-based single image super-resolution
Tao Li 0014, Xiaohai He, Qizhi Teng |
Neurocomputing | 2 |
| 2016 | Single image super resolution using local smoothness and nonlocal self-similarity priors
Honggang Chen, Xiaohai He, Qizhi Teng, Chao Ren 0002 |
Signal Process. Image Commun. | 2 |
| 2016 | Long-Range Motion Trajectories Extraction of Articulated Human Using Mesh EvolutionabstractThis letter presents a novel approach to extract reliable dense and long-range motion trajectories of articulated human in a video sequence. Compared with existing approaches that emphasize temporal consistency of each tracked point, we also consider the spatial structure of tracked points on the articulated human. We treat points as a set of vertices, and build a triangle mesh to join them in image space. The problem of extracting long-range motion trajectories is changed to the issue of consistency of mesh evolution over time. First, self-occlusion is detected by a novel mesh-based method and an adaptive motion estimation method is proposed to initialize mesh between successive frames. Furthermore, we propose an iterative algorithm to efficiently adjust vertices of mesh for a physically plausible deformation, which can meet the local rigidity of mesh and silhouette constraints. Finally, we compare the proposed method with the state-of-the-art methods on a set of challenging sequences. Evaluations demonstrate that our method achieves favorable performance in terms of both accuracy and integrity of extracted trajectories. Yuanyuan Wu 0001, Xiaohai He, Byeongkeun Kang, Haiying Song, Truong Q. Nguyen |
IEEE Signal Process. Lett. | 2 |
| 2016 | Single Image Super-Resolution Using Local Geometric Duality and Non-Local SimilarityabstractSuper-resolution (SR) from a single image plays an important role in many computer vision applications. It aims to estimate a high-resolution (HR) image from an input low- resolution (LR) image. To ensure a reliable and robust estimation of the HR image, we propose a novel single image SR method that exploits both the local geometric duality (GD) and the non-local similarity of images. The main principle is to formulate these two typically existing features of images as effective priors to constrain the super-resolved results. In consideration of this principle, the robust soft-decision interpolation method is generalized as an outstanding adaptive GD (AGD)-based local prior. To adaptively design weights for the AGD prior, a local non-smoothness detection method and a directional standard-deviation-based weights selection method are proposed. After that, the AGD prior is combined with a variational-framework-based non-local prior. Furthermore, the proposed algorithm is speeded up by a fast GD matrices construction method, which primarily relies on the selective pixel processing. The extensive experimental results verify the effectiveness of the proposed method compared with several state-of-the-art SR algorithms. Chao Ren 0002, Xiaohai He, Qizhi Teng, Yuanyuan Wu 0001, Truong Q. Nguyen |
IEEE Trans. Image Process. | 2 |
| 2015 | A New Framework for Container Code Recognition by Using Segmentation-Based and HMM-Based ApproachesabstractTraditional methods for automatic recognition of container code in visual images are based on segmentation and recognition of isolated characters. However, when the segment fails to separate each character from the others, those methods will not function properly. Sometimes the container code characters are printed or arranged very closely, which makes it a challenge to isolate each character. To address this issue, a new framework for automatic container code recognition (ACCR) in visual images is proposed in this paper. In this framework, code-character regions are first located by applying a horizontal high-pass filter and scan line analysis. Then, character blocks are extracted from the code-character regions and further classified into two categories, i.e. single-character block and multi-character block. Finally, a segmentation-based approach is implemented for recognition of the characters in single-character blocks, and a hidden Markov model (HMM)-based method is proposed for the multi-character blocks. The experimental results demonstrate the effectiveness of the proposed method, which can successfully recognize the container code with closely arranged characters. Wei Wu 0002, Zheng Liu 0002, Zhiming Liu 0009, Xi Wu 0004, Xiaohai He |
Int. J. Pattern Recognit. Artif. Intell. | 6 |
| 2015 | A fast inter-prediction algorithm for HEVC based on temporal and spatial correlation
Guo-Yun Zhong, Xiaohai He, Linbo Qing |
Multim. Tools Appl. | 2 |
| 2015 | Space-time super-resolution with patch group cuts prior
Tao Li 0014, Xiaohai He, Qizhi Teng, Zhengyong Wang, Chao Ren 0002 |
Signal Process. Image Commun. | 2 |
| 2014 | Content-based group-of-picture size control in distributed video coding
Enrico Masala, Yan-Mei Yu, Xiaohai He |
Signal Process. Image Commun. | 3 |
| 2013 | Subframe video synchronization by matching trajectoriesabstractWe propose a novel approach to align unsynchronized video sequences of the same dynamic scene that can be subframe accurate and is applicable for different frame rate problem. The proposed approach relies on matching motion trajectories and it is assumed that the object moves on a planar surface. By exploring the invariants of planar trajectories under projective transformation, the cross ratio as an invariant feature is computed for each point along the trajectories and the similarity between invariant features of different trajectories is measured with a distance that takes into account the statistical properties of the cross ratio. Then the smooth high frame rate trajectory is synthesized for searching subframe temporal displacement under our alignment framework. The experimental results with synthetic and real-world sequences show that our approach achieves fairly accuracy and efficiency in subframe temporal alignment of the multiple unsynchronized video sequences. Yuanyuan Wu 0001, Xiaohai He, Truong Q. Nguyen |
ICASSP | 2 |
| 2013 | Object tracking using firefly algorithmabstractFirefly algorithm (FA) is a new meta‐heuristic optimisation algorithm that mimics the social behaviour of fireflies flying in the tropical and temperate summer sky. In this study, a novel application of FA is presented as it is applied to solve tracking problem. A general optimisation‐based tracking architecture is proposed and the parameters’ sensitivity and adjustment of the FA in tracking system are studied. Experimental results show that the FA‐based tracker can robustly track an arbitrary target in various challenging conditions. The authors compare the speed and accuracy of the FA with three typical tracking algorithms including the particle filter, meanshift and particle swarm optimisation. Comparative results show that the FA‐based tracker outperforms the other three trackers. Mingliang Gao 0001, Xiaohai He, Dai-Sheng Luo, Qizhi Teng |
IET Comput. Vis. | 2 |
| 2013 | Adaptive regularization-based space-time super-resolution reconstruction
Haiying Song, Linbo Qing, Yuanyuan Wu 0001, Xiaohai He |
Signal Process. Image Commun. | 4 |
| 2012 | An automated vision system for container-code recognition
Wei Wu 0002, Zheng Liu 0002, Xiaomin Yang, Xiaohai He |
Expert Syst. Appl. | 5 |
| 2011 | Learning-based super resolution using kernel partial least squares
Wei Wu 0002, Zheng Liu 0002, Xiaohai He |
Image Vis. Comput. | 3 |
| 2011 | Hidden-Markov-Model-Based Segmentation Confidence Applied to Container Code Character ExtractionabstractAutomatic container code recognition (ACCR) has become an indispensable aspect of current intelligent container management systems. In real applications, an ACCR module sometimes faces the problem of missing characters, i.e., not all the 11 container code characters (CCCs) appear in the input image. However, a few of the present methods can process container code images with missing characters. Therefore, a method is proposed to extract the CCCs for both the situation wherein all the 11 CCCs appear in an image and the situation wherein some CCCs are missing. In this method, hidden Markov model (HMM)-based segmentation confidence is proposed to describe the probability of the segmented characters belonging to the container code. Based on the segmentation confidence, the segmented characters are determined whether they belong to the container code or not, and if there are some characters missing, the positions of these characters can be estimated. Various container code images have been used to test the proposed method. The results of the tests show that the method is effective. Wei Wu 0002, Xiaomin Yang, Xiaohai He |
IEEE Trans. Intell. Transp. Syst. | 4 |