VLDB 2026 Research / reviewers in the wild / expert
Zhuhong Shao
dblp:137/6960
· DBLP profile ↗
53ranked-venue papers
14as first author
39since 2021 · last 2026
0000-0002-4847-282XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 11 first-author · 23 since 2021Artificial intelligence and machine learning · 20 · 3 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dynamic Neuro-Symbolic Adapter for Efficient Fine-Grained Visual Recognition
Zhuhong Shao |
ICPR (10) | 5 |
| 2026 | Multi-source Pseudo-Label Generation for Weakly Supervised Salient Object Detection
Handan Zhang, Zhuhong Shao |
ICPR (8) | 5 |
| 2026 | HIM-PyraNet: Hierarchical Attention and Region-Focused Lightweight Network for Micro-Expression RecognitionabstractABSTRACT Micro‐expression recognition (MER) is a technology that infers emotions or intentions by analyzing subtle changes in facial expressions. It is often used to identify deception or assist interpersonal communication and holds significant value. Currently, MER technology is advancing rapidly; however, several challenges remain. To address the insufficient capture of local features and high computational complexity in MER, this paper proposes a novel lightweight network architecture. We introduce a Hierarchical Interactive Multi‐scale Pyramid Network (HIM‐PyraNet) that identifies key facial muscle movement regions while considering inherent spatial relationships between facial landmarks, differing from most existing studies using self‐attention mechanisms. HIM‐PyraNet comprises two main components: a Cross‐Region Interaction Attention (CRIA) module focusing on local temporal features and a Multi‐scale Feature Pyramid Fusion (MFPF) module integrating local and global semantics. Specifically, the face is divided into four distinct regions: left eye, right eye, left lip, and right lip. The CRIA module captures local micro‐muscle movements with region‐specific self‐attention, while the MFPF module learns interactions between eye and lip regions. This strategy effectively reduces the processing area for facial MER, successfully builds facial regional collaboration, and retains the local detailed features of expressions while reducing computational parameters. Experiments on SAMM, CASME II, and SMIC datasets show that with only 0.25M parameters, our model achieves 85.31% average recognition accuracy and 75.3 frames per second (FPS) inference speed. Compared with existing methods, this architecture demonstrates higher accuracy with a low parameter count, providing an innovative solution for real‐time micro‐expression analysis on edge devices. Fangjie Xue, Zhuhong Shao, Wanlong Cui, Zehao Yuan |
Concurr. Comput. Pract. Exp. | 3 |
| 2026 | CMTNet: A collaborative mamba-transformer network with spatial-temporal cross-fusion for speech emotion recognition
Shihe Dong, Jiajun Wei, Yibing Zhu, Zhuhong Shao, Mingyue Niu, Xiaohui Tan, Yinan Jiang, Rongyin Qin |
Pattern Recognit. | 6 |
| 2026 | Audio-Visual Feature Disentanglement and Fusion Network for Automatic Depression Severity PredictionabstractIn order to achieve early screening and assist clinical decision-making, automatic depression assessment based on multimodal data are highly anticipated. However, the existed methods often suffer from semantic gap and information redundancy due to heterogeneity among modalities. To address this challenge, this paper investigates a novel Feature Disentanglement and Fusion Network (FDFNet) for predicting depression severity from audio-visual cues. Firstly, we design the shared and private encoders to disentangle modality-shared and modalityprivate representations. The former representation that acquires joint information is subjected by similarity constraints between modalities to ensure their distributions as close as possible. The latter that can capture unique features of each modality is restrained by independence constraints for keeping their distributions distinct. The decoder is then developed to reconstruct unimodal representation with constraints to minimize information loss. Finally, an efficient fusion strategy through addition and concatenation is ultilized for aggregating information. Experimental results on four benchmark datasets demonstrate that the proposed FDFNet consistently outperforms several stateof-the-art methods, with the competitive MAE/RMSE values of 6.22/7.58 on AVEC2013, 5.21/6.49 on AVEC2014, 4.25/5.34 on DAIC-WOZ, and 4.41/5.10 on E-DAIC, indicating that multimodal deep learning based on audio-visual is an attractive solution for objectively evaluating the depression severity. Zhuhong Shao, Rongyin Qin, Yongzhen Huang, Peipeng Liang, Yinan Jiang, Yanhe Deng, Xiaohui Tan |
IEEE Trans. Affect. Comput. | 2 |
| 2026 | DSTC: A Multimodal Network for Depression Emotion Recognition and Sentiment AnalysisabstractDepression is a common mental disorder that affects physical and mental health. Recently, automatic recognition of multimodal depression based on deep learning has attracted widespread attention. Current methods for automatic depression recognition based on audio, video and text lack consideration of modality differences, and the feature fusion between modalities is not sufficient. To address these challenges, this study proposes a multimodal depression emotion recognition and sentiment analyses method based on Temporal ProbSparse Attention Transformer (TPAT) and Position-guided Multimodal Cross-Fusion (PMC) module, named DSTC. Specifically, due to the commonalities between audio and video modalities, this study introduces a Temporal ProbSparse Attention Transformer designed to facilitate modality interaction while effectively capturing long-range semantic information. In addition, to fully utilize the positional information in the text modality and integrate the text with audio-video modalities, this study employs a position-guided Multimodal Cross-Fusion module. The proposed method achieved MAE scores of 4.22 and 4.85 on the validation and test sets of the E-DAIC depression dataset, indicating its significant improvements over existing depression recognition methods. To assess the generalization performance, additional experiments were conducted on the sentiment analysis datasets CMU-MOSI and CMU-MOSEI, where Acc-7 scores of 41.6% and 53.0% were achieved, respectively. These results highlight the superior performance of the proposed method compared to state-of-the-art sentiment analysis algorithms. Jiaxi Lu, Zhuhong Shao |
IEEE Trans. Affect. Comput. | 3 |
| 2026 | Multimodal Local Global Interaction Networks for Automatic Depression Severity EstimationabstractPhysiological studies have shown that differences between depressed and healthy individuals are manifested in the audio and video modalities. Hence, some researchers have combined local and global information from audio or video modality to obtain the unimodal representation. Attention mechanisms or Multi-Layer Perceptrons (MLPs) are then used to complete the fusion of different representations. However, attention mechanisms or MLPs is essentially a linear aggregation manner, and lacks the ability to explore the element-wise interaction between local and global representations within and across modalities, which affects the accuracy of estimating the depression severity. To this end, we propose a Representation Interaction (RI) module, which uses the mutual linear adjustment to achieve element-wise interaction between representations. Thus, the RI module can be seen as an mutual observation of two representations, which helps to achieve complementary advantages and improve the model’s ability to characterize depression cues. Furthermore, since the interaction process generates multiple representations, we propose a Multi-representation Prediction (MP) module. This module implements multi-representation vectorization in a hierarchical manner from summarizing a single representation to aggregating multiple representations, and adopts the attention mechanism to obtain the estimation of an individual depression severity. In this way, we use the RI and MP modules to construct the Multimodal Local Global Interaction (MLGI) network. The experimental performance on AVEC 2013 and AVEC 2014 depression datasets demonstrates the effectiveness of our method. Mingyue Niu, Zhuhong Shao, Yongjun He 0002, Jianhua Tao 0001, Björn W. Schuller |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | ODQ-UNet: A Quaternion-based U-Net Integrating OD and RGB Space for Nuclear SegmentationabstractNuclear segmentation is a critical task in computational histopathology, providing quantitative metrics essential for cancer diagnosis and treatment. However, accurately segmenting nuclei in histopathological images remains a significant challenge due to several factors, including wide variations in staining pro-tocols, high cell density, and frequent overlaps. To address these issues, we propose ODQ-UNet, a novel segmentation framework that integrates color space transformation and quaternion-based convolution. Our method integrates RGB and Optical Density (OD) spaces, with OD offering a stain-invariant representation via a logarithmic transformation. This fusion enriches the feature representation and significantly improves our model's robustness to staining variations. We then introduce Quaternion Convolution to unify the multi-channel information from both OD and RGB spaces. This approach treats the combined four channels as a single quaternion entity, allowing the model to capture complex, holistic correlations among the color channels more effectively than traditional convolution. The ODQ-UNet is built upon a U - Net architecture with residual connections and an attention mechanism for improved feature extraction. To further validate our approach, we employed Grad-CAM for interpretability, demonstrating that our model effectively focuses on relevant nuclear structures. Experimental results on two publicly available benchmark datasets, MoNuSeg and CPM17, show that ODQ-UNet achieves superior performance with Dice coefficients of 86.01 % and 92.76%, respectively, outperforming state-of-the-art methods. Junhui Xin, Jingyi Weng, Jierui Zhao, Zhuhong Shao |
BIBM | 6 |
| 2025 | MFMamba: A Multimodal Fusion State Space Model for Depression RecognitionabstractDepression is a severe mental illness, and extracting emotional information from video-audio signals for multimodal depression recognition is a challenging problem. Recent methods use the self-attention (SA) mechanism from Transformers to capture the dynamic relationships between different modalities. However, the quadratic computational complexity of SA reduces its effectiveness in modeling long sequences, making it insufficient for capturing complex intra-modal and inter-modal complementarity. To address this issue, this work proposes a Multimodal Fusion Mamba (MFMamba) framework, which is attention-free and purely focuses on using state space models (SSMs) for long-sequence modeling. Specifically, we devise Video Spatio-Temporal Mamba (VSTMamba) and Audio Temporal Mamba (ATMamba) for video-audio feature extraction. To fully capture the correlations among multimodal features and eliminate information redundancy, we introduce Fusion Mamba (FMamba) to integrate various features effectively. In experiments on AVEC 2013 and AVEC 2014 datasets, our method achieved competitive results. Zhuhong Shao, Jiaxi Lu |
ICASSP | 4 |
| 2025 | K-CMorph: Integrating K-space Consistency and Complex-Valued Processing for Improved MRI Deformable Registration
Zhuhong Shao, Jingbing Yang |
ICIC (3) | 6 |
| 2025 | A Multi-level and Multi-scale Context Refinement Network for Video-based Depression RecognitionabstractDepression recognition based on artificial intelligence is a new research topic in the field of affective computing, which is of great value for depression screening and auxiliary diagnosis. Recent studies have made significant progress in video-based depression recognition through CNNs and self-attention mechanisms. However, challenges include limited understanding of contextual relationships, loss of detailed information, and overfitting. To address these limitations, this work proposes the Multilevel and Multi-scale Context Refinement Network (MM-CRN) to model the relationships across multi-scale features, refine detail features, and fuse multi-level features. Technically, we devise the Video Context Refinement (ViCR) Block, comprising the Global Context Encoding Module (GCEM) and the Local Refinement Encoding Module (LREM), which encode contextual relationships and local detail information independently. The multi-level attention mechanism is introduced to capture feature correlations across different levels, reduce redundant information, and relieve overfitting. Experiments on the AVEC 2013 and AVEC 2014 datasets demonstrate that the proposed model achieves highly competitive results, with MAE/RMSE values of 6.04/7.61 and 5.99/7.59. Moreover, the visualization results applying gradient-weighted class activation mapping (Grad-CAM) demonstrate the effectiveness of the proposed method. Jiaxi Lu, Zhuhong Shao |
IJCNN | 5 |
| 2025 | Multimodal Interpretable Depression Analysis Using Visual, Physiological, Audio and Textual DataabstractMotivated by depression's significant impact on global health, this work proposes MultiDepNet, a novel multi-modal interpretable depression detection system integrating visual, physiological, audio, and textual data. Through ded-icated feature extraction methods (MTCNN for video, TS-CAN for physiological, ResNet-18 for audio, and RoBERTa for text modalities) and a strategic fusion of modality-specific networks including CNN-RNN, Transformer, MLP, and ResNet-18, it achieves significant advancements in depression detection. Its performance, evaluated across four benchmark datasets (AVEC 2013, AVEC 2014, DAIC, and E-DAIC), demonstrates average MAE of 5.64, RMSE of 7.15, accuracy of 74.19%, precision of 0.7373, re-call of 0.7378, and F1 of 0.7376. It also implements a Multiviz-based interpretability mechanism that computes each modality's contribution to the model's performance. The results reveal the visual modality to be the most signifi-cant, contributing 37.88% towards depression detection. Puneet Kumar 0003, Shreshtha Misra, Zhuhong Shao, Balasubramanian Raman |
WACV | 3 |
| 2025 | MSCA-Sp R-CNN: a segmentation algorithm for pneumonia small lesions integrating multi-scale channel attention and sub-pixel upsampling
Zhuhong Shao, Rongyin Qin, Guoping Huo |
Multim. Syst. | 4 |
| 2025 | TTFNet: Temporal-Frequency Features Fusion Network for Speech Based Automatic Depression Recognition and AssessmentabstractRelated studies have revealed that the phonological features of depressed patients are different from those of healthy individuals. With the increasing prevalence of depression, an objective and convenient approach for early screening is necessary. To this end, we propose an automatic depression detection method based on hybrid speech features extracted by deep learning, dubbed as TTFNet. Firstly, to effectively excavate the intrinsic relationship among multidimensional dynamic features in the frequency domain, the log-Mel spectrogram of raw speech and its related derivatives are encoded into quaternion representation. Then, the innovatively designed quaternion VisionLSTM is utilized to capture their synergistic effects. Simultaneously, we integrate sLSTM with the pre-trained wav2vec 2.0 model to fully acquire the temporal features. In addition, to further exploit the complementarity between temporal and frequency features, we design an XConformer block for cross-sequence interactions, which ingeniously combines self-attention mechanisms and convolutional modules. Based on this block, the dual-path fusion module closely utilizes the mutual promotion of features from different domains, thereby enhancing generalization capability of the proposed model. Extensive experiments conducted on the AVEC 2013, AVEC 2014, DAIC-WOZ and E-DAIC datasets demonstrate that our method outperforms current state-of-the-art methods in both depression recognition and severity prediction tasks. Xiyuan Chen 0004, Zhuhong Shao, Yinan Jiang, Runsen Chen, Bicao Li, Mingyue Niu, Hongguang Chen, Jiasong Wu |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | LMS-VDR: Integrating Landmarks into Multi-scale Hybrid Net for Video-Based Depression Recognition
Zhuhong Shao |
PRCV (10) | 4 |
| 2024 | Federated learning with comparative learning-based dynamic parameter updating on glioma whole slide images
Longjian Huang, Lizhi Shao, Meiling Bao, Changsong Guo, Zhuhong Shao, Xiazi Huang, Mingjing Wang, Xiaoming Jiang, Shengzhou Hu |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | Spatial-Temporal Attention Network for Depression Recognition from facial videos
Zhuhong Shao, Guodong Guo |
Expert Syst. Appl. | 4 |
| 2024 | Pyramid quaternion discrete cosine transform based ConvNet for cancelable face recognition
Zhuhong Shao, Leding Li, Xuanyi Li, Bicao Li |
Image Vis. Comput. | 1 |
| 2024 | Stereo image encryption using vector decomposition and symmetry of 2D-DFT in quaternion gyrator domain
Zhuhong Shao, Leding Li, Xiaoxu Zhao, Bicao Li, Xilin Liu 0003 |
Multim. Tools Appl. | 1 |
| 2024 | Cancelable color face recognition using trinion gyrator transform and randomized nonlinear PCANet
Zhuhong Shao, Bicao Li, Junlin Ouyang |
Multim. Tools Appl. | 1 |
| 2024 | Color image encryption based on discrete trinion Fourier transform and compressive sensing
Zhuhong Shao, Bicao Li, Xilin Liu 0003 |
Multim. Tools Appl. | 2 |
| 2024 | Cancelable face recognition using phase retrieval and complex principal component analysis network
Zhuhong Shao, Leding Li, Bicao Li, Xilin Liu 0003 |
Mach. Vis. Appl. | 1 |
| 2024 | Integrating Deep Facial Priors Into Landmarks for Privacy Preserving Multimodal Depression RecognitionabstractAutomatic depression diagnosis is a challenging problem, that requires integrating spatial-temporal information and extracting features from audio-visual signals. In terms of privacy protection, the development trend of recognition algorithms based on facial landmarks has created additional challenges and difficulties. In this paper, we propose an audio-visual attention network (AVA-DepressNet) for depression recognition. It is a novel multimodal framework with facial privacy protection, and uses attention-based modules to enhance audio-visual spatial and temporal features. In addition, an adversarial multistage (AMS) training strategy is developed to optimize the encoder-decoder structure. Additionally, facial structure prior knowledge is creatively used in AMS training. Our AVA-DepressNet is evaluated on popular audio-visual depression datasets: AVEC 2013, AVEC 2014, and AVEC 2017. The results show that our approach reaches the state-of-the-art performance or competitive results for depression recognition. Zhuhong Shao, Guodong Guo |
IEEE Trans. Affect. Comput. | 3 |
| 2023 | Full Quaternion Matrix-Based Multiscale Principal Component Analysis Network for Facial Expression Recognition
Zhuhong Shao |
PRCV (5) | 3 |
| 2023 | CDBIFusion: A Cross-Domain Bidirectional Interaction Fusion Network for PET and MRI Images
Bicao Li, Bei Wang 0005, Zhuhong Shao, Jie Huang 0037, Jiaxi Lu |
PRCV (13) | 4 |
| 2023 | A semi-fragile watermarking tamper localization method based on QDFT and multi-view fusion
Junlin Ouyang, Jingtao Huang, Xingzi Wen, Zhuhong Shao |
Multim. Tools Appl. | 4 |
| 2023 | Trinion discrete cosine transform with application to color image encryption
Zhuhong Shao, Yadong Tang |
Multim. Tools Appl. | 1 |
| 2023 | Correction to: Trinion discrete cosine transform with application to color image encryption
Zhuhong Shao, Yadong Tang |
Multim. Tools Appl. | 1 |
| 2023 | Residual shuffle attention network for image super-resolution
Xuanyi Li, Zhuhong Shao, Bicao Li, Jiasong Wu, Yuping Duan |
Mach. Vis. Appl. | 2 |
| 2023 | Randomized nonlinear two-dimensional principal component analysis network for object recognition
Zhijian Sun, Zhuhong Shao, Bicao Li, Jiasong Wu |
Mach. Vis. Appl. | 2 |
| 2023 | LQGDNet: A Local Quaternion and Global Deep Network for Facial Depression RecognitionabstractRecent visual-based depression recognition methods mostly use hand-crafted features with information lost in color channels, or deep network features with a limited performance from the finite data. In this paper, we propose a method called Local Quaternion and Global Deep Network (LQGDNet) which can combine advantages from hand-crafted and deep features. Specifically, the Quaternion XOR Asymmetrical Regional Local Gradient Coding (XOR-AR-LGC) is first designed, which encodes the facial images with local textures in the quaternion domain to keep the dependence of color channels, and integrated into the Quaternion Feature Extractor (QFE). To the best of our knowledge, it is the first attempt to use a quaternion-based method for facial depression recognition. Second, we design the Local Quaternion Representation Module (LQRM) composed of Local Deep Feature Extractor (LDFE) and QFE to output local quaternion facial features. Third, global deep facial features are encoded from the Global Deep Representation Module (GDRM) with the deep convolutional neural network. Finally, the LQGDNet integrates LQRM and GDRM with the local quaternion and global deep features and predicts the depression score. The experimental results on AVEC 2013 and AVEC 2014 show the superiority of our method compared to the state-of-the-art approaches. Zhuhong Shao, Guodong Guo |
IEEE Trans. Affect. Comput. | 4 |
| 2022 | DCAN: A Dual Cascade Attention Network for Fusing Pet and MRI ImagesabstractTraditional fusion approaches and most deep learning-based methods usually generate the intermediate decision map, resulting in detail loss of source images or fusion results. In this work, to enhance the detailed features and structured information from source images, we propose a dual cascade attention network (DCAN) to obtain a more informative fusion image for PET and MRI images. In our approach, channel attention is employed to improve the ability of features representation and spatial attention can highlight informative regions in the proposed fusion network. Additionally, channel and spatial attention are sequential arrangement in channel-first. Moreover, to achieve good performance in the procedure of feature extraction and image reconstruction, two-stage training strategy is adopted to train our fusion model. Experimental results demonstrate that the proposed approach achieves remarkable performance for PET and MRI images fusion. Bicao Li, Zhoufeng Liu, Chunlei Li 0002, Zhuhong Shao, Zongmin Wang |
ICIP | 5 |
| 2022 | DMF-CL: Dense Multi-scale Feature Contrastive Learning for Semantic Segmentation of Remote-Sensing Images
Mengxing Song, Bicao Li, Pingjun Wei, Zhuhong Shao, Jing Wang 0080, Jie Huang 0037 |
PRCV (4) | 4 |
| 2022 | Color image watermarking based on singular value decomposition and generalized regression neural network
Xilin Liu 0003, Yongfei Wu, Peiting Gao, Junlin Ouyang, Zhuhong Shao |
Multim. Tools Appl. | 5 |
| 2022 | Multiple color image encryption based on cascaded quaternion gyrator transforms
Zhuhong Shao, Yan Zhang 0094, Gouenou Coatrieux |
Signal Process. Image Commun. | 3 |
| 2021 | Double image encryption based on symmetry of 2D-DFT and equal modulus decomposition
Zhuhong Shao, Yadong Tang, Mingxian Liang |
Multim. Tools Appl. | 1 |
| 2021 | Multiple-image encryption based on cascaded gyrator transforms and high-dimensional chaotic system
Xiaoni Sun, Zhuhong Shao, Mingxian Liang, Fengjian Yang |
Multim. Tools Appl. | 2 |
| 2021 | Fusing structure and color features for cancelable face recognition
Zhuhong Shao, Bicao Li |
Multim. Tools Appl. | 2 |
| 2021 | Robust multiple color images encryption using discrete Fourier transforms and chaotic map
Yadong Tang, Zhuhong Shao, Xiaoxu Zhao |
Signal Process. Image Commun. | 2 |
| 2020 | Color image encryption based on discrete trinion Fourier transform and random-multiresolution singular value decomposition
Qijun Yao, Zhuhong Shao, Xilin Liu 0003, Qingbin Tong |
Multim. Tools Appl. | 2 |
| 2020 | The modified generic polar harmonic transforms for image representation
Xilin Liu 0003, Yongfei Wu, Zhuhong Shao, Jiasong Wu |
Pattern Anal. Appl. | 3 |
| 2020 | Multiple-image encryption based on chaotic phase mask and equal modulus decomposition in quaternion gyrator domain
Zhuhong Shao, Xilin Liu 0003, Qijun Yao, Na Qi |
Signal Process. Image Commun. | 1 |
| 2018 | Double-image cryptosystem using chaotic map and mixture amplitude-phase retrieval in gyrator domain
Zhuhong Shao, Xiaoyan Fu, Huimei Yuan, Huazhong Shu |
Multim. Tools Appl. | 1 |
| 2018 | Multiple color image encryption and authentication based on phase retrieval and partial decryption in quaternion gyrator domain
Zhuhong Shao, Qingbin Tong, Xiaoxu Zhao, Xiaoyan Fu |
Multim. Tools Appl. | 1 |
| 2018 | Automated Depression Diagnosis Based on Deep Networks to Encode Facial Appearance and DynamicsabstractAs a severe psychiatric disorder disease, depression is a state of low mood and aversion to activity, which prevents a person from functioning normally in both work and daily lives. The study on automated mental health assessment has been given increasing attentions in recent years. In this paper, we study the problem of automatic diagnosis of depression. A new approach to predict the Beck Depression Inventory II (BDI-II) values from video data is proposed based on the deep networks. The proposed framework is designed in a two stream manner, aiming at capturing both the facial appearance and dynamics. Further, we employ joint tuning layers that can implicitly integrate the appearance and dynamic information. Experiments are conducted on two depression databases, AVEC2013 and AVEC2014. The experimental results show that our proposed approach significantly improve the depression prediction performance, compared to other visual-based approaches. Yu Zhu 0006, Zhuhong Shao, Guodong Guo |
IEEE Trans. Affect. Comput. | 3 |
| 2017 | Fast single image dehazing based on a regression model
Zhong Luan, Xiuzhuang Zhou, Zhuhong Shao, Guodong Guo, Xiaoming Liu 0002 |
Neurocomputing | 4 |
| 2016 | Retinal vessel enhancement using multi-dictionary and sparse codingabstractA novel retinal vessel enhancement method based on multi-dictionary and sparse coding is proposed in this paper. Two dictionaries are utilized to gain the retinal vascular structures and details, one is the representation dictionary (RD) generated from the original retinal images, and another is the enhancement dictionary (ED) extracted from the corresponding label images. The proposed method represents the input image with RD to get the sparse coefficients via a sparse coding process. Then the enhanced retinal vessel image is obtained from the solved sparse coefficients and ED. Experimental results performed on the DRIVE and STARE databases indicate that the proposed method not only can effectively improve the image contrast but also enhance the details of the retinal vessels. Yang Chen 0008, Zhuhong Shao, Limin Luo 0001 |
ICASSP | 3 |
| 2016 | Blood vessel enhancement via multi-dictionary and sparse coding: Application to retinal vessel enhancing
Yang Chen 0008, Zhuhong Shao, Limin Luo 0001 |
Neurocomputing | 3 |
| 2016 | Color image classification via quaternion principal component analysis network
Jiasong Wu, Zhuhong Shao, Yang Chen 0008, Beijing Chen, Lotfi Senhadji, Huazhong Shu |
Neurocomputing | 3 |
| 2016 | Robust watermarking using orthogonal Fourier-Mellin moments and chaotic map for double images
Zhuhong Shao, Xilin Liu 0003, Guodong Guo |
Signal Process. | 1 |
| 2016 | Robust watermarking scheme for color image based on quaternion-type moment invariants and visual cryptography
Zhuhong Shao, Huazhong Shu, Gouenou Coatrieux, Jiasong Wu |
Signal Process. Image Commun. | 1 |
| 2014 | Quaternion Bessel-Fourier moments and their invariant descriptors for object reconstruction and recognition
Zhuhong Shao, Huazhong Shu, Jiasong Wu, Beijing Chen, Jean-Louis Coatrieux |
Pattern Recognit. | 1 |
| 2013 | Quaternion gyrator transform and its application to color image encryptionabstractThe gyrator transform has been proposed in optics a few years ago. By using the theory of quaternion numbers, this paper presents the quaternion gyrator transform (QGT). It is shown that the QGT can be computed via the left-side type of quaternion Fourier transforms. The new transform is applied to color image encryption for validation, where the rotation angles are used as encryption keys making it more secure compared to a recent method using discrete quaternion Fourier transforms (DQFTs). Experimental results show that the proposed encryption algorithm for color image performs as well as the DQFTs method in terms of noise robustness, so that it could be a useful tool for color image encryption. Zhuhong Shao, Jiasong Wu, Jean-Louis Coatrieux, Gouenou Coatrieux, Huazhong Shu |
ICIP | 1 |