Beiji Zou 0001

dblp:37/2732-1 · DBLP profile ↗
← Back
115ranked-venue papers
21as first author
63since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 47 · 10 first-author · 17 since 2021Artificial intelligence and machine learning · 45 · 5 first-author · 31 since 2021Applied, interdisciplinary, general and emerging computing · 24 · 5 first-author · 15 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 7 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 DCFANet: Merging dynamic context clustering mamba and context-to-focus attention for medical image segmentation
Xiaoyan Kui, Zhipeng Hu, Zexin Ji, Shen Jiang, Qianmu Xiao, Ziwei Zou, Qinsong Li, Yang Li 0111, Beiji Zou 0001, Liming Chen 0002
Neurocomputing9
2026 TFKAN: Time-frequency KAN for long-term time series forecasting
Xiaoyan Kui, Canwei Liu, Qinsong Li, Zhipeng Hu, Yangyang Shi, Weixin Si, Beiji Zou 0001
Neurocomputing7
2026 Tri-HGNet: A feature-driven dynamic hypergraph framework for medical image segmentation
Xiaoyan Kui, Lingxiao Liu, Qinsong Li, Haonan Yan, Weixin Si, Zuheng Ming, Beiji Zou 0001
Neurocomputing7
2026 Improving neural radiance fields with depth-aware optimization for novel view synthesis
Junyao Li, Beiji Zou 0001
Image Vis. Comput.4
2026 A comprehensive survey on magnetic resonance image reconstruction
Xiaoyan Kui, Zijie Fan, Zexin Ji, Qinsong Li, Chengtao Liu, Weixin Si, Beiji Zou 0001
Image Vis. Comput.7
2026 RMViM-Net: Residual multi-path vision mamba with graph interaction attention for medical image segmentation
Shen Jiang, Xiaoyan Kui, Xingzhuo Bao, Qinsong Li, Zhipeng Hu, Beiji Zou 0001
Knowl. Based Syst.6
2026 STL-SATVNet: STL Decomposition and self-attention-based time-varying neural network for multi-scale forecasting of multivariate time-series
Hongbo Xiao, Beiji Zou 0001, Xiaoyan Kui
Knowl. Based Syst.2
2026 Alzheimer's disease classification based on multimodal consistent distribution and trusted fusion
Xiaoyan Kui, Yulan Dai, Beiji Zou 0001, Chengzhang Zhu, Yang Li 0111, Zexin Ji, Liming Chen 0002, Miguel Bordallo López
Neural Networks3
2026 Cascading size-dependent deep propagation (CADP): Addressing over-smoothing in graph few-shot dermatology classification
Abdulrahman Noman, Beiji Zou 0001, Chengzhang Zhu, Mohammed Alhabib, Ahmed Alasri
Neural Networks2
2026 Global and local Mamba network for multi-modality medical image super-resolution
Zexin Ji, Beiji Zou 0001, Xiaoyan Kui, Sébastien Thureau, Su Ruan
Pattern Recognit.2
2026 Integrating frequency-aware mamba with diffusion for 4D volumetric image synthesis
Yangyang Shi, Beiji Zou 0001, Xiaonian Deng, Yucong Zhang, Zehua Liu, Xiaoyan Kui, Weixin Si
Pattern Recognit.2
2025 Aligning Medical Images and Language Through Multimodal Medical Rationales
abstract
Large vision-language models (LVLMs) have gained widespread attention in the medical field for their outstanding capability in handling image-text representations. However, the misalignment between medical images and clinical text presents factuality challenges for Medical LVLMs (Med-LVLMs), often resulting in hallucinations. Multimodal Chain-of-Thought (MCoT) can reduce factual errors of Med-LVLMs by encouraging explicit step-by-step reasoning, but it poses two major challenges. First, factually accurate medical rationales are crucial for aligning medical images with the corresponding clinical texts, yet existing Med-LVLMs struggle to generate such rationales. Second, if the model's initial prediction is correct, its inherent knowledge can be disrupted by an over-reliance on the generated medical rationale, resulting in an incorrect answer. To address these challenges, we propose MMReT, a novel multimodal medical reasoning tuning approach designed to improve the factual accuracy of Med-LVLMs. First, we employ carefully designed prompts to guide GPT-4o in generating high-quality medical rationales, which are then used to fine-tune the original Med-LVLM. Second, to address errors stemming from excessive reliance on generated rationales, we introduce medical reasoning preference fine-tuning, which encourages the model to maintain an appropriate balance between leveraging its inherent knowledge and incorporating generated medical rationale. Experimental results show that MMReT substantially enhances the factuality of Med-LVLMs, outperforming previous methods with average improvements of 8.0% on VQA-RAD and 11.6% on SLAKE in factual accuracy.
Zhi Chen 0015, Beiji Zou 0001, Xiaoyan Kui, Ziwei Zou, Jinming Duan 0001
BIBM2
2025 Distance-Aware and Knowledge-Driven Vision Mamba U-Net for Radiotherapy Dose Prediction
abstract
Dose planning is essential in radiotherapy for cancer patients, yet current practice relies on iterative manual optimization, underscoring the need for automated prediction. Existing deep learning approaches remain limited because they often ignore the 3D spatial relationships between tumors and surrounding organs at risk (OARs), and clinical priors on safe dose thresholds. To overcome these limitations, we propose DKVMU-Net, a distance-aware and knowledge-driven Vision Mamba U-Net for automated dose prediction. Our framework incorporates Vision Mamba blocks to capture global, long-range dependencies from CT scans and OAR signed distance field (SDF) maps, which naturally encode spatial information. Additionally, we introduce a deformable dynamic feature enhancement module (DDFEM) for texture refinement, followed by a linear crossattention fusion module to improve cross-modality integration. A customized loss function is also designed to incorporate prior knowledge of OAR dose constraints, ensuring optimal target coverage and OAR protection. To alleviate the scarcity of doseplanning datasets, we collect an in-house radiotherapy lung cancer dataset (RLCD), consisting of CT volumes, OAR masks, and corresponding SDF maps from 116 patients. We evaluate our DKVMU-Net on both the in-house dataset and public available OpenKBP dataset. Compared with the sate-of-the-art method, our approach achieves an 11.6 % improvement in dose score (1.641 vs. 1.857) and 26.3 % in DVH score (6.481 vs. 8.799) on RLCD, and a 7.8 % improvement in dose score (2.421 vs. 2.626) and 13.9 % in DVH score (1.057 vs. 1.227) on OpenKBP. These results demonstrate the robustness and effectiveness of our approach.
Yangyang Shi, Xiaoyan Kui, Yucong Zhang, Shihao Zou, Zuheng Ming, Weixin Si, Azeddine Beghdadi, Beiji Zou 0001
BIBM8
2025 From Global to Local: Mamba-Based Hierarchical Registration for Respiratory Lung Deformation
abstract
Deformable image registration is essential in medical applications, as accurately estimating organ displacements across respiratory phases enables precise radiation dose planning in dynamic environments, mitigates damage to organs at risk (OARs), and thus improves patients' health-related quality of life. Although current learning-based methods have achieved impressive performance in small deformation registration, challenges remain due to their limited ability to capture large deformations occurring during respiration. To address this issue, we propose a novel Mamba-based hierarchical registration framework that effectively extracts both global and local features for accurate deformation prediction. Specifically, given a pair of source and target 3DCT volumes, we incorporate a foundation model pretrained on medical image registration tasks to enhance alignment accuracy. We further propose a directional-deformable Mamba scheme to facilitate global context extraction and local motion awareness. The directional Mamba component scans input features from multiple orientations to achieve broad contextual perception, while the deformable Mamba module employs adaptive directional scanning strategies to capture dynamic local variations. To overcome the scarcity of annotated respiratory data, we also collect a new respiratory lung cancer dataset comprising 100 annotated phases from 20 patients. Experimental results on our in-house dataset demonstrate that our method outperforms state-of-the-art approaches, achieving a 1.3 % improvement in overall Dice accuracy and a 1.6 dB increase in PSNR, underscoring its strong potential for clinical deployment. Code and test data are available at: https://github.com/yangyangshi806/Mamba_based_Registration.
Yangyang Shi, Yucong Zhang, Beiji Zou 0001, Xiaoyan Kui, Zexin Ji, Zuheng Ming, Azeddine Beghdadi, Weixin Si
BIBM3
2025 Mamba Based Feature Extraction and Adaptive Multilevel Feature Fusion for 3D Tumor Segmentation from Multi-modal Medical Image
Zexin Ji, Beiji Zou 0001, Xiaoyan Kui, Hua Li 0003, Pierre Vera, Su Ruan
ICIC (28)2
2025 Robust Multimodal Representation Learning with Information Bottleneck and Balanced Fusion for Alzheimers Disease Classification
abstract
Given the capability of multimodal data to provide information from multiple perspectives, it is beneficial for improving the accuracy of Alzheimer’s disease (AD) classification. However, during practical multimodal learning, there is a phenomenon where certain modalities dominate the decision-making, leading to insufficient learning from other modalities. Moreover, redundant information within multimodal data can also hinder accurate classification decisions. Therefore, we propose a robust multimodal representation learning method for AD classification. Specifically, we first construct dedicated encoders for each multimodal data, including structural Magnetic Resonance Imaging (sMRI) images, Positron Emission Tomography (PET) images, and Mini-Mental State Examination (MMSE) scores, to extract their respective representations. Then, we employ the information bottleneck (IB) theory to guide the model to retain classification-related information in multimodal representations while reducing redundancy among modalities. Furthermore, to promote a balanced fusion of multimodal data, we redefine the classification confidence of each modality’s representation using an orthogonal weight classifier and then introduce a regularization term to amplify the prediction score differences for modalities with lower confidence. The experimental results on the Alzheimer’s Disease Neuroimaging Initiative (ADNI) dataset demonstrate that our method enhances the robustness of multimodal representations and achieves promising performance in AD-related classification tasks.
Yulan Dai, Beiji Zou 0001, Xiaoyan Kui, Zexin Ji, Chengzhang Zhu
ICIP2
2025 Flip Distribution Alignment VAE for Multi-phase MRI Synthesis
Xiaoyan Kui, Qianmu Xiao, Qinsong Li, Zexin Ji, Jielin Zhang, Beiji Zou 0001
MICCAI (14)6
2025 Improving cancer driver genes identifying based on graph embedding hypergraph and hierarchical synergy dominance model
Zhipeng Hu, Xiaoyan Kui, Canwei Liu, Zanbo Sun, Shen Jiang, Kai Zhu 0009, Beiji Zou 0001
Expert Syst. Appl.8
2025 Utilizing Multihead Attention-Based Graph Convolution Networks for Traffic Speed Prediction
abstract
Accurate traffic speed prediction holds immense significance in mitigating traffic congestion and enhancing traffic safety. However, traffic data exhibit distinct patterns across different cycles (such as weekdays, weekends, and holidays), making it challenging for traditional models to effectively capture this multiperiod heterogeneity in traffic data. Furthermore, most existing research on traffic speed prediction struggles to efficiently capture the spatiotemporal characteristics of dynamic traffic data simultaneously. To tackle these challenges, this paper first introduces spatiotemporal‐aware position encoding (STAPE) technology, which addresses the multiperiod heterogeneity in traffic data by integrating temporal cycle information with spatial position information. Second, a multilevel spatiotemporal feature extraction architecture is designed, leveraging graph convolutional network (GCN) to capture the topological structure and spatial features of the traffic road network. By applying gated recurrent unit (GRU) to capture the temporal dependencies of traffic data, and combining GCN and GRU in multiple stages, this architecture deeply explores the spatiotemporal features of traffic data. Additionally, this paper integrates a multihead attention mechanism, which, in conjunction with the parallelized attention channel adaptive mechanism and the multilevel spatiotemporal feature extraction architecture, enhances the model’s ability to adaptively model different spatiotemporal patterns dynamically, thereby efficiently capturing the dynamically changing spatiotemporal features. Extensive performance evaluation experiments conducted on the METR‐LA and PEMS‐BAY datasets demonstrate that the predictive performance of the proposed model surpasses that of nine other baseline methods.
Hongbo Xiao, Beiji Zou 0001, Xiaoyan Kui, Lilian Yuan
Int. J. Intell. Syst.2
2025 Gl-MambaNet: A global-local hybrid Mamba network for medical image segmentation
Xiaoyan Kui, Shen Jiang, Qinsong Li, Yifei Peng, Zhipeng Hu, Beiji Zou 0001
Neurocomputing6
2025 PK-Net: A prior knowledge-driven dual-path network for enhanced glaucoma screening
Xiaoyan Kui, Zeru Hai, Beiji Zou 0001, Yang Li 0111, Wei Liang 0005, Zuheng Ming, Liming Chen 0002
Knowl. Based Syst.3
2025 WinGraphUNet: Advanced windowed graph modeling with remixed contextual learning for efficient medical image segmentation
Xiaoyan Kui, Haonan Yan, Qinsong Li, Lingxiao Liu, Weixin Si, Wei Liang 0005, Beiji Zou 0001
Knowl. Based Syst.7
2025 Generation of super-resolution for medical image via a self-prior guided Mamba network with edge-aware constraint
Zexin Ji, Beiji Zou 0001, Xiaoyan Kui, Hua Li 0003, Pierre Vera, Su Ruan
Pattern Recognit. Lett.2
2025 Predicting Driver Genes From Multi-Omics Data Using Hierarchical Multi-Feature Synergy Model
abstract
Cancer is an extremely complex disease, whose occurrence and development are influenced by a multitude of factors, among which the abnormal activity of cancer driver genes plays a crucial role in the pathological process. Identifying these genes allows researchers to understand pathogenic mechanisms and biological functions of cancer, facilitating the development of targeted therapies. Current methods for identifying driver genes often ignore the synergism among genes and the importance of features, thereby affecting identification accuracy. In this paper, we propose a cancer driver genes identification method called HMFS, which is based on the hierarchical multi-feature synergy model. Firstly, a hypergraph is constructed using Node2vec and K-means algorithm. By analyzing the topological feature and mutual exclusion degree of genes in each hyperedge, the Mutation Aggregation Coefficient is extracted. Then, based on the functional expression mechanism of genes, differential expression analysis is performed using miRNA and mRNA expression data. Finally, by analyzing the importance among features, the Hierarchical Multi-Feature Synergy is proposed for features fusion. In this paper, experiments are conducted on three real cancer datasets. Compared with seven representative methods, HMFS has the best performance on all evaluation indicators.
Zhipeng Hu, Xiaoyan Kui, Canwei Liu, Shen Jiang, Ziwei Zou, Beiji Zou 0001
IEEE Trans. Comput. Biol. Bioinform.7
2025 ChebMixer: Efficient Graph Representation Learning With MLP Mixer
abstract
Graph neural networks (GNNs) have achieved remarkable success in learning graph representations, especially graph Transformers, which have recently shown superior performance on various graph mining tasks. However, the graph Transformer generally treats nodes as tokens, which results in quadratic complexity regarding the number of nodes during self-attention computation. The graph multilayer perceptron (MLP) mixer addresses this challenge using the efficient MLP Mixer technique from computer vision. However, the time-consuming process of extracting graph tokens limits its performance. In this article, we present a novel architecture named ChebMixer, a newly proposed graph MLP Mixer that uses fast Chebyshev polynomials-based spectral filtering to extract a sequence of tokens. First, we produce multiscale representations of graph nodes via fast Chebyshev polynomial-based spectral filtering. Next, we consider each node's multiscale representations as a sequence of tokens and refine the node representation with an effective MLP Mixer. Finally, we aggregate the multiscale representations of nodes through Chebyshev interpolation. Owing to the powerful representation capabilities and fast computational properties of the MLP Mixer, we can quickly extract more informative node representations to improve the performance of downstream tasks. The experimental results prove our significant improvements in various scenarios, ranging from homogeneous and heterophilic graph node classification to medical image segmentation. Compared with NAGphormer, the average performance improved by 1.45% on homogeneous graphs and 4.15% on heterophilic graphs. And the average performance improved by 1.39% on medical image segmentation tasks compared with VM-UNet. We will release the source code after this article is accepted.
Xiaoyan Kui, Haonan Yan, Qinsong Li, Liming Chen 0002, Beiji Zou 0001
IEEE Trans. Neural Networks Learn. Syst.6
2024 Strong Multimodal Representation Learner through Cross-domain Distillation for Alzheimer's Disease Classification
abstract
Vision-language foundational models have achieved commendable results on related tasks. However, their application to medical tasks is still limited due to issues arising from data biases. Currently, leveraging existing foundational models to improve medical tasks remains a challenge. To this end, this paper proposes a strong multimodal representation learning method based on cross-domain distillation handling structural Magnetic Resonance Imaging (sMRI), Positron Emission Computed Tomograph (PET) images, and mini-mental state examination (MMSE) score for Alzheimer’s disease (AD) classification. Specifically, we establish a text-to-image cross-domain distillation learning framework, enabling a text encoder pre-trained on general visual recognition tasks to guide the training of sMRI and PET image feature extractors. Simultaneously, positional encoding is used to extract the magnitude features of MMSE scores. Based on the multimodal representations extracted from sMRI, PET images, and MMSE scores, we perform a self-attention operation equipped with a gating mechanism for multimodal feature fusion. This mechanism controls the contribution of each modality representation to the classification decision, dynamically strengthening or weakening specific modality representations and helping construct stronger fused features for AD classification. Our method undergoes 5-fold cross-validation on the widely used ADNI dataset, and comparative experimental results demonstrate that our method achieves advanced performance in two AD-related binary classification tasks.
Yulan Dai, Beiji Zou 0001, Xiaoyan Kui, Qinsong Li, Wei Zhao 0040, Jun Liu 0075, Miguel Bordallo López
BIBM2
2024 TTCNet: Transformer and Tubular Convolution Feature Attention Network for OCTA vessel segmentation
abstract
Optical coherence tomography angiography (OCTA) is a non-invasive imaging technique that can reveal blood flow and the detailed structure of the retinal and choroidal vasculature. However, uneven image quality and the high complexity of blood vessel distribution bring great challenges to accurate blood vessel segmentation. In this paper, we propose an OCTA blood vessel segmentation network based on the classic encoder-decoder architecture called TTCNet, which integrates a dual-branch encoder to capture global context information effectively. In addition, Snake Detail Attention Block is designed to enhance the perception of the tubular structure of blood vessels and provide richer feature representation in the network. We also design a module named Pyramid Feature Attention Module in each level of skip connection to better extract feature information at different scales. Experiments are conducted on two public datasets, OCTA_6M and OCTA_3M. G-mean Score and Area Under the Curve of our network on these two datasets are 93.05, 93.18, and 95.21, 94.95%, respectively, achieving better results than other state-of-the-art methods.
Chengzhang Zhu, Zhangzheng Yang, Yalong Xiao, Beiji Zou 0001, Haoze Zhou
BIBM5
2024 Self-prior Guided Mamba-UNet Networks for Medical Image Super-Resolution
Zexin Ji, Beiji Zou 0001, Xiaoyan Kui, Pierre Vera, Su Ruan
ICPR (11)2
2024 An Effective Dual-Scale Hybrid Encoder Network for Medical Image Segmentation
abstract
Medical image segmentation’s accuracy is crucial for clinical analysis and diagnosis. Despite progress with U-Net-inspired models, they often underuse multi-scale encoding layers crucial for enhancing detailing visual features and overlooking the importance of merging multi-scale features within the channel dimension to enhance decoder complexity. To address these limitations, we introduce a dual-scale hybrid encoder network DSENet for medical image segmentation. Our network design is characterized by the strategic employment of dual-scale convolutional kernels at each encoder level, integrating the robust feature extraction capabilities of CNNs with the contextual awareness of Transformer models. This synergy enables the precise capture of both granular and broader context features throughout the encoding phase. We further enhance the model by integrating a channel attention fusion (CAF) mechanism within the skip connections. This innovation effectively integrates dual-scale features, subsequently integrating them with the same-level decoder feature map, thereby reinforcing the feature representation. To refine the predicted segmentation, we employ a novel strategy that merges dual-scale feature maps from the initial encoder stage with the segmentation map through a cascade operation. The output fusion feature map is then processed by a self-attention Transformer structure, ensuring a meticulous refinement of the segmentation output, preserving essential details, and enhancing segmentation accuracy. Our proposed DSENet has been evaluated on three distinct medical image datasets, and the experimental results demonstrate that it achieves more accurate segmentation performance and adaptability to varying target segmentation, making it more competitive compared to existing SOTA methods.
Chengzhang Zhu, Renmao Zhang, Yalong Xiao, Beiji Zou 0001, Xian Chai, Zhangzheng Yang, Lan Hua, Xuanchu Duan
IJCNN4
2024 Deform-Mamba Network for MRI Super-Resolution
Zexin Ji, Beiji Zou 0001, Xiaoyan Kui, Pierre Vera, Su Ruan
MICCAI (7)2
2024 PFFNet: A pyramid feature fusion network for microaneurysm segmentation in fundus images
abstract
Abstract Retinal microaneurysm (MA) is a definite earliest clinical sigh of diabetic retinopathy (DR). Its automatic segmentation is key to realizing intelligent screening for early DR, which can significantly reduce the risk of visual impairment in patients. However, the minute scale and subtle contrast of MAs against the background pose challenges for segmentation. This paper focuses on automatic MA segmentation in fundus images. A novel pyramid feature fusion network (PFFNet) that progressively develops and fuses rich contextual information by integrating two pyramid modules is proposed. Multiple global pyramid scene parsing (GPSP) modules are introduced between the encoder and decoder to provide diverse global contextual information for the decoder through reconstructing skip connections. Additionally, a spatial scale‐aware pyramid (SSAP) module is introduced to dynamically fuse multi‐scale contextual information. This rich contextual information will help to identify MAs from low‐contrast background. Furthermore, to mitigate issue related to category imbalance, a combo loss function is introduced. Finally, to validate the effectiveness of the proposed method, experiments are conducted on two publicly available datasets, IDRiD and DDR, and PFFNet is compared with several state‐of‐the‐art models. The experimental results demonstrate the superiority of our PFFNet in the MA segmentation task.
Beiji Zou 0001, Xiaoxia Xiao, Qinghua Peng, Junfeng Yan, Wensheng Zhang 0002, Kejuan Yue
IET Image Process.2
2024 Enhancing cervical cancer diagnosis: Integrated attention-transformer system with weakly supervised learning
Ashfaque Khowaja, Beiji Zou 0001, Xiaoyan Kui
Image Vis. Comput.2
2024 Deep learning-based magnetic resonance image super-resolution: a survey
Zexin Ji, Beiji Zou 0001, Xiaoyan Kui, Jun Liu 0075, Wei Zhao 0040, Chengzhang Zhu, Peishan Dai, Yulan Dai
Neural Comput. Appl.2
2024 Structure-aware neural radiance fields without posed camera
Yaxin Xu, Beiji Zou 0001
Pattern Recognit.4
2023 Wavelet-aware Transformer Network for Multi-contrast Knee MRI Super-resolution
abstract
In this paper, we propose a wavelet-aware transformer network (WATNet) for multi-contrast knee MRI super-resolution. Unlike conventional image domain-based super-resolution methods that can not explicitly model the lost high-frequency information, our WATNet endeavors to adaptively fuse the complementary frequency information of the multi-contrast image in the wavelet domain and further refine it in the image domain. The proposed WATNet consists of the multi-scale wavelet transformation (MSWT) module, wavelet-aware transformer (WAT) module, and reconstruction (Rec) module. Specifically, the MSWT module learns to transform the MR image to multi-scale wavelet domain features by the wavelet transformation. The WAT module can adaptively search and transfer similar wavelet domain reference information to the low-resolution one. The Rec module can restore high-quality images in the image domain. To further capture more high-frequency details, we also design the wavelet-based high-frequency loss. The qualitative and quantitative experimental results indicate that our proposed WATNet outperforms most state-of-the-art methods.
Zexin Ji, Xiaoyan Kui, Chengzhang Zhu, Yang Li 0111, Yulan Dai, Beiji Zou 0001
BIBM7
2023 Prior-knowledge-based self-attention network for 3D human pose estimation
Yaxin Xu, Beiji Zou 0001
Expert Syst. Appl.3
2023 PA-LBF: Prefix-Based and Adaptive Learned Bloom Filter for Spatial Data
abstract
The recently proposed learned bloom filter (LBF) opens a new perspective on how to reconstruct bloom filters with machine learning. However, the LBF has a massive time cost and does not apply to multidimensional spatial data. In this paper, we propose a prefix‐based and adaptive learned bloom filter (PA‐LBF) for spatial data, which efficiently supports the insertion and deletion. The proposed PA‐LBF is divided into three parts: (1) the prefix‐based classification. The Z‐order space‐filling curve is used to extract data, prefix it, and classify it. (2) The adaptive learning process. The multiple independent adaptive sub‐LBFs are designed to train the suffixes of data, combined with part 1, to reduce the false positive rate (FPR), query, and learning process time consumption. (3) The backup filter uses CBF. Two kinds of backup CBF are constructed to meet the situation of different insertion and deletion frequencies. Experimental results prove the validity of the theory and show that the PA‐LBF reduces the FPR by 84.87%, 79.53%, and 43.01% with the same memory usage compared with the LBF on three real‐world spatial datasets. Moreover, the time consumption of PA‐LBF can be reduced to 5× and 2.05× that of the LBF on the query and learning process, respectively.
Meng Zeng, Beiji Zou 0001, Xiaoyan Kui, Chengzhang Zhu, Ling Xiao 0003, Zhi Chen 0015, Jingyu Du
Int. J. Intell. Syst.2
2023 Two-layer partitioned and deletable deep bloom filter for large-scale membership query
Meng Zeng, Beiji Zou 0001, Wensheng Zhang 0002, Xuebing Yang, Guilan Kong, Xiaoyan Kui, Chengzhang Zhu
Inf. Syst.2
2023 Anomaly detection for streaming data based on grid-clustering and Gaussian distribution
Beiji Zou 0001, Kangkang Yang, Xiaoyan Kui, Jun Liu 0075, Wei Zhao 0040
Inf. Sci.1
2023 A deep retinal image quality assessment network with salient structure priors
Ziwen Xu, Beiji Zou 0001, Qing Liu 0003
Multim. Tools Appl.2
2023 Recognition of abnormal human behavior in dual-channel convolutional 3D construction site based on deep learning
Lingzi Jiang, Beiji Zou 0001, Shu Liu 0002, Enquan Huang
Neural Comput. Appl.2
2023 AC-E Network: Attentive Context-Enhanced Network for Liver Segmentation
abstract
Segmentation of liver from CT scans is essential in computer-aided liver disease diagnosis and treatment. However, the 2DCNN ignores the 3D context, and the 3DCNN suffers from numerous learnable parameters and high computational cost. In order to overcome this limitation, we propose an Attentive Context-Enhanced Network (AC-E Network) consisting of 1) an attentive context encoding module (ACEM) that can be integrated into the 2D backbone to extract 3D context without a sharp increase in the number of learnable parameters; 2) a dual segmentation branch including complemental loss making the network attend to both the liver region and boundary so that getting the segmented liver surface with high accuracy. Extensive experiments on the LiTS and the 3D-IRCADb datasets demonstrate that our method outperforms existing approaches and is competitive to the state-of-the-art 2D-3D hybrid method on the equilibrium of the segmentation precision and the number of model parameters.
Yang Li 0111, Beiji Zou 0001, Peishan Dai, Miao Liao, Harrison X. Bai, Zhicheng Jiao
IEEE J. Biomed. Health Informatics2
2023 ESDedup: An efficient and secure deduplication scheme based on data similarity and blockchain for cloud-assisted medical storage systems
Ling Xiao 0003, Beiji Zou 0001, Chengzhang Zhu, Fanbo Nie
J. Supercomput.2
2023 Combining external attention GAN with deep convolutional neural networks for real-fake identification of luxury handbags
Jianbiao Peng, Beiji Zou 0001, Chengzhang Zhu
Vis. Comput.2
2022 A dynamic multi-modal fusion network for ovarian tumor differentiation
abstract
Accurate ovarian tumor differentiation is a challenging task where the benign and malignant tumors share similar T1C and T2WI MRI appearances. Therefore, it is necessary to leverage additional multi-modal data, e.g., the age, CA125level, and other clinical information, which are helpful but rarely exploited. In this paper, we propose a dynamic fusion network that can adaptively make full use of multi-modal data, including MRI and clinical information, to realize precise ovarian tumor differentiation. Specifically, we design a dynamic nonlinear module (D-Non-L module) on the top of the image representation. The D-Non-L module is formulated as an iterative nonlinear projection parameterized by the learned features of the patient-wise clinical information. With the help of this module, the interaction between clinical features and image features could be achieved to adaptively improve the discrimination of visual representations. Moreover, we construct a dual-path-based architecture to fully exploit the complementary information from T1C and T2WI MRIs. Extensive experimental results on the locally organized ovarian tumor dataset demonstrate that our methods are superior to the single-modal and single-path-based methods. And the proposed dynamic non-linear module obtains the best performance compared with other multi-modal fusion strategies.
Yang Li 0111, Beiji Zou 0001, Yulan Dai, Harrison X. Bai, Zhicheng Jiao
BIBM2
2022 MeDA: Using Blockchain for Patient-Controlled Medical Data Auditing in Institutions
abstract
Hospital Information Systems (HIS) are now widely used in various medical institutions due to the rapid advancement of information technology. However, owing to the sensitivity of medical data, the data is usually managed solely by the relevant department of the institution, leaving patients and other departments unable to verify its integrity and authenticity. This unilateral data management by the department erodes the credi-bility of the data and complicates building trust between doctors and patients. To address this issue effectively, we propose a scheme of patient-controlled medical data auditing in institutions by leveraging blockchain technology. We establish a permissioned blockchain network among medical departments and program smart contracts that allow patients to authorize changes to their data to build trust in the data while maintaining patient ownership. More specifically, we propose different evidence on-chain strategies that can be adapted to meet various data integrity and recoverability requirements for different data characteristics. Additionally, we present an adaptive hash method that ensures outstanding computational efficiency across a wide range of file sizes. Therefore, MeDA enables patient-controlled data auditing within the medical institution, guaranteeing data integrity and authenticity. The analysis and evaluation demonstrate that the proposed strategies are able to audit various types of data securely and efficiently.
Beiji Zou 0001, Fanbo Nie, Ling Xiao 0003, Chengzhang Zhu
CCGRID1
2022 GRVT: Toward Effective Grocery Recognition via Vision Transformer
Shu Liu 0002, Chengzhang Zhu, Beiji Zou 0001
CGI4
2022 Toward Efficient Image Denoising: A Lightweight Network with Retargeting Supervision Driven Knowledge Distillation
Beiji Zou 0001, Shu Liu 0002
CGI1
2022 A Learned Prefix Bloom Filter for Spatial Data
Beiji Zou 0001, Meng Zeng, Chengzhang Zhu, Ling Xiao 0003, Zhi Chen 0015
DEXA (1)1
2022 Alps: An Adaptive Load Partitioning Scaling Solution for Stream Processing System on Skewed Stream
Beiji Zou 0001, Chengzhang Zhu, Ling Xiao 0003, Meng Zeng, Zhi Chen 0015
DEXA (2)1
2022 Entity-level Attention Pooling and Information Gating for Document-level Relation Extraction
abstract
Document-level relation extraction intends to extract relation facts among different entity pairs in the entire document. Previously proposed methods make numerous efforts for this task, however, these researches neglect to treat the two entities in an entity pair as an organic unity for relation extraction, which leads to the lack of valuable information about the entity pairs and insufficient information interaction. To tackle the above problem, we propose a framework, named Entity-level Attention Pooling and Information Gating (EAPIG), for document-level relation extraction. Specifically, we first utilize an encoder module to capture the long-distance dependencies of entities in the document, and then we propose two modules: the Entity-level Attention Pooling module obtains the local information of entity pairs, and the Information Gating module introduces the global information of entity pairs and promotes sufficient interaction between the local information and the global information. Experimental results on the benchmark dataset DocRED show that our approach can efficiently capture and combine the abundant information from the entity pairs to achieve better performance than the previous baselines.
Beiji Zou 0001, Zhi Chen 0015, Chengzhang Zhu, Ling Xiao 0003, Meng Zeng
ICPR1
2022 Parameter-Free Latent Space Transformer for Zero-Shot Bidirectional Cross-modality Liver Segmentation
Yang Li 0111, Beiji Zou 0001, Yulan Dai, Chengzhang Zhu, Fan Yang 0054, Xin Li 0079, Harrison X. Bai, Zhicheng Jiao
MICCAI (4)2
2022 OVS-Net: An effective feature extraction network for optical coherence tomography angiography vessel segmentation
abstract
Abstract Optical coherence tomography angiography (OCTA), as a noninvasive imaging modality, has been widely used in clinical ophthalmology. However, the segmentation of retinal vessels in OCTA is under‐studied due to OCTA is a relatively new technology. In this article, an effective feature extraction network, OVS‐Net, is proposed for OCTA vessel segmentation. The OVS‐Net is divided into coarse stage and refine stage which structures are basically the same. In each stage, we utilize OctaveResBlock as the basic block to better extract the hierarchical multifrequency features of OCTA and capture the multiscale semantic features of the vessels. In order to improve the feature characterization, feature enhanced attention block is introduced into the network, which is proved to be more conducive for microvessel segmentation in our experiments. Multiscale feature blocks are introduced into the network to promote the deep integration of semantic features at different scales. Experiments on OCTA‐SS and OCTA‐500 datasets show that our proposed OVS‐Net achieves more competitive segmentation results than the existing methods, especially for microvessel segmentation.
Chengzhang Zhu, Han Wang 0064, Yalong Xiao, Yulan Dai, Beiji Zou 0001
Comput. Animat. Virtual Worlds6
2022 Single image depth estimation based on sculpture strategy
Zhengdong Pu, Jianquan Ouyang 0001, Beiji Zou 0001
Knowl. Based Syst.5
2022 SkeletonPose: Exploiting human skeleton constraint for 3D human pose estimation
Yaxin Xu, Zhengdong Pu, Jianquan Ouyang 0001, Beiji Zou 0001
Knowl. Based Syst.5
2022 Fixing Defect of Photometric Loss for Self-Supervised Monocular Depth Estimation
abstract
View-synthesis-based methods have shown very promising results for the task of unsupervised depth estimation in single images. Most existing approaches synthesize a new image and employ it as the supervision signal for depth and pose prediction. There are two problems in these approaches: 1) There are many combinations of pose and depth that can synthesize a certain new image; therefore, reconstructing the depth and pose based on the view-synthesis method from only two images is an inherently ill-posed problem; 2) The model is trained under the photometric consistency assumption that the brightness or gradient is constant when applied to the video sequences. However, this assumption is easily violated in realistic scenes due to light changes, reflective surfaces and occlusions. To overcome the first drawback, we exploit the point cloud consistency constraint to eliminate ambiguity. To overcome the second drawback, we use threshold masks to filter dynamic and occluded points and introduce matching point constraints that implicitly encode the geometry relationship between two matched points to improve the precision of depth prediction. In addition, we employ epipolar constraints to compensate for the instability of the photometric error in textureless regions and varying illumination conditions. The experimental results on the KITTI, Cityscapes and NYUv2 datasets show that the method can improve the accuracy of depth prediction and enhance the robustness of the model in handling textureless regions and illumination changes. The code and data are available athttps://github.com/XTUPRLAB/FixUnDepth.
Zhengdong Pu, Beiji Zou 0001
IEEE Trans. Circuits Syst. Video Technol.4
2021 Character Flow Detection and Rectification for Scene Text Spotting
Beiji Zou 0001, Kai-Wen Li, Enquan Huang, Shu Liu 0002
CGI1
2021 Multi-branch Multi-task 3D-CNN for Alzheimer's Disease Detection
Junhu Li, Beiji Zou 0001, Ziwen Xu, Qing Liu 0003
PRCV (3)2
2021 A Dark and Bright Channel Prior Guided Deep Network for Retinal Image Quality Assessment
Ziwen Xu, Beiji Zou 0001, Qing Liu 0003
PRCV (3)2
2021 Ground truth free retinal vessel segmentation by learning from simple pixels
abstract
Abstract Retinal vessel segmentation is fundamental for the automatic retinal image analysis and ocular disease screening. This paper aims to learn a ground truth free feature aggregation strategy for the vessel segmentation. Five vesselness maps modelling the vessels'profile, appearance, and shape are first generated. Together, the histogram of the local binary pattern and the green colour are extracted. In each vesselness map, the pixels with large vesselness values are regarded as simple positive samples. The pixels with small vesselness values are regarded as simple negative samples, and the pixels with mediocre values are treated as difficult pixels. The simple positive samples and simple negative samples near the difficult pixels consist of the training dataset while the rest vesselness maps together with the local binary pattern histogram, and green colour channel are used as the features to learn a strong classifier. Then, without leveraging any ground truth, multiple kernel boosting is used to combine four support vector machine kernels to learn a strong vessel model for each image. Applying this learnt model to the pixels with mediocre values in the single vesselness map, their label will be determined. Totally, five strong vessel models are learnt. Finally, pixels with the majority supports from the strong vessel models are labelled as vessel pixels. The proposed method achieves accuracy of 94.83%, sensitivity of 72.59%, and specificity of 98.11% on DRIVE dataset, and accuracy of 95.51%, sensitivity of 78.09%, and specificity of 97.56% on STARE. It outperforms the state‐of‐the‐art unsupervised methods and achieves comparable performances to the supervised methods.
Beiji Zou 0001, Hongpu Fu, Zailiang Chen 0001, Qing Liu 0003
IET Image Process.1
2021 Robust and discriminative zero-watermark scheme based on invariant features and similarity-based retrieval to protect large-scale DIBR 3D videos
Xiyao Liu 0001, Yifan Wang 0008, Ziqiang Sun, Lei Wang 0017, Rongchang Zhao, Yuesheng Zhu, Beiji Zou 0001, Hui Fang 0003
Inf. Sci.7
2021 A Character Flow Framework for Multi-Oriented Scene Text Detection
Beiji Zou 0001, Kai-Wen Li, Shu Liu 0002
J. Comput. Sci. Technol.2
2021 Multi-Label Classification Scheme Based on Local Regression for Retinal Vessel Segmentation
abstract
Segmenting small retinal vessels with width less than 2 pixels in fundus images is a challenging task. In this paper, in order to effectively segment the vessels, especially the narrow parts, we propose a local regression scheme to enhance the narrow parts, along with a novel multi-label classification method based on this scheme. We consider five labels for blood vessels and background in particular: the center of big vessels, the edge of big vessels, the center as well as the edge of small vessels, the center of background, and the edge of background. We first determine the multi-label by the local de-regression model according to the vessel pattern from the ground truth images. Then, we train a convolutional neural network (CNN) for multi-label classification. Next, we perform a local regression method to transform the previous multi-label into binary label to better locate small vessels and generate an entire retinal vessel image. Our method is evaluated using two publicly available datasets and compared with several state-of-the-art studies. The experimental results have demonstrated the effectiveness of our method in segmenting retinal vessels.
Beiji Zou 0001, Yulan Dai, Qi He 0008, Chengzhang Zhu
IEEE ACM Trans. Comput. Biol. Bioinform.1
2020 A Deep Gradient Boosting Network for Optic Disc and Cup Segmentation
abstract
Segmentation of optic disc (OD) and optic cup (OC) is critical in automated fundus image analysis system. Existing state-of-the-arts focus on designing deep neural networks with one or multiple dense prediction branches. Such kind of designs ignore connections among prediction branches and their learning capacity is limited. To build connections among prediction branches, this paper introduces gradient boosting framework to deep classification model and proposes a gradient boosting network called BoostNet. Specifically, deformable side-output unit and aggregation unit with deep supervisions are proposed to learn base functions and expansion coefficients in gradient boosting framework. By stacking aggregation units in a deep-to-shallow manner, models’ performances are gradually boosted along deep to shallow stages. BoostNet achieves superior results to existing deep OD and OC segmentation networks on the public dataset ORIGA.
Qing Liu 0003, Beiji Zou 0001, Yixiong Liang
ICASSP2
2020 Single image dehazing based on fusion strategy
Fan Guo 0001, Jin Tang 0002, Hui Peng 0001, Lijue Liu, Beiji Zou 0001
Neurocomputing6
2020 Non-rigid retinal image registration using an unsupervised structure-driven regression network
Beiji Zou 0001, Zhiyou He, Rongchang Zhao, Chengzhang Zhu, Wangmin Liao, Shuo Li 0001
Neurocomputing1
2020 Clinical Interpretable Deep Learning Model for Glaucoma Diagnosis
abstract
Despite the potential to revolutionise disease diagnosis by performing data-driven classification, clinical interpretability of ConvNet remains challenging. In this paper, a novel clinical interpretable ConvNet architecture is proposed not only for accurate glaucoma diagnosis but also for the more transparent interpretation by highlighting the distinct regions recognised by the network. To the best of our knowledge, this is the first work of providing the interpretable diagnosis of glaucoma with the popular deep learning model. We propose a novel scheme for aggregating features from different scales to promote the performance of glaucoma diagnosis, which we refer to as M-LAP. Moreover, by modelling the correspondence from binary diagnosis information to the spatial pixels, the proposed scheme generates glaucoma activations, which bridge the gap between global semantical diagnosis and precise location. In contrast to previous works, it can discover the distinguish local regions in fundus images as evidence for clinical interpretable glaucoma diagnosis. Experimental results, performed on the challenging ORIGA datasets, show that our method on glaucoma diagnosis outperforms state-of-the-art methods with the highest AUC (0.88). Remarkably, the extensive results, optic disc segmentation (dice of 0.9) and local disease focus localization based on the evidence map, demonstrate the effectiveness of our methods on clinical interpretability.
Wangmin Liao, Beiji Zou 0001, Rongchang Zhao, Yuanqiong Chen, Zhiyou He, Mengjie Zhou
IEEE J. Biomed. Health Informatics2
2019 Weakly-Supervised Simultaneous Evidence Identification and Segmentation for Automated Glaucoma Diagnosis
abstract
Evidence identification, optic disc segmentation and automated glaucoma diagnosis are the most clinically significant tasks for clinicians to assess fundus images. However, delivering the three tasks simultaneously is extremely challenging due to the high variability of fundus structure and lack of datasets with complete annotations. In this paper, we propose an innovative Weakly-Supervised Multi-Task Learning method (WSMTL) for accurate evidence identification, optic disc segmentation and automated glaucoma diagnosis. The WSMTL method only uses weak-label data with binary diagnostic labels (normal/glaucoma) for training, while obtains pixel-level segmentation mask and diagnosis for testing. The WSMTL is constituted by a skip and densely connected CNN to capture multi-scale discriminative representation of fundus structure; a well-designed pyramid integration structure to generate high-resolution evidence map for evidence identification, in which the pixels with higher value represent higher confidence to highlight the abnormalities; a constrained clustering branch for optic disc segmentation; and a fully-connected discriminator for automated glaucoma diagnosis. Experimental results show that our proposed WSMTL effectively and simultaneously delivers evidence identification, optic disc segmentation (89.6% TP Dice), and accurate glaucoma diagnosis (92.4% AUC). This endows our WSMTL a great potential for the effective clinical assessment of glaucoma.
Rongchang Zhao, Wangmin Liao, Beiji Zou 0001, Zailiang Chen 0001, Shuo Li 0001
AAAI3
2019 Multi-index Optic Disc Quantification via MultiTask Ensemble Learning
Rongchang Zhao, Zailiang Chen 0001, Xiyao Liu 0001, Beiji Zou 0001, Shuo Li 0001
MICCAI (1)4
2019 A novel glaucomatous representation method based on Radon and wavelet transform
abstract
BACKGROUND: Glaucoma is an irreversible eye disease caused by the optic nerve injury. Therefore, it usually changes the structure of the optic nerve head (ONH). Clinically, ONH assessment based on fundus image is one of the most useful way for glaucoma detection. However, the effective representation for ONH assessment is a challenging task because its structural changes result in the complex and mixed visual patterns. METHOD: We proposed a novel feature representation based on Radon and Wavelet transform to capture these visual patterns. Firstly, Radon transform (RT) is used to map the fundus image into Radon domain, in which the spatial radial variations of ONH are converted to a discrete signal for the description of image structural features. Secondly, the discrete wavelet transform (DWT) is utilized to capture differences and get quantitative representation. Finally, principal component analysis (PCA) and support vector machine (SVM) are used for dimensionality reduction and glaucoma detection. RESULTS: The proposed method achieves the state-of-the-art detection performance on RIMONE-r2 dataset with the accuracy and area under the curve (AUC) at 0.861 and 0.906, respectively. CONCLUSION: In conclusion, we showed that the proposed method has the capacity as an effective tool for large-scale glaucoma screening, and it can provide a reference for the clinical diagnosis on glaucoma.
Beiji Zou 0001, Changlong Chen, Rongchang Zhao, Ping-Bo Ouyang, Chengzhang Zhu, Qilin Chen, Xuanchu Duan
BMC Bioinform.1
2019 Glaucoma screening pipeline based on clinical measurements and hidden features
abstract
Glaucoma refers to a chronic disease of the eye that leads to vision loss that is irreversible, which is called ‘silent theft of sight’. Thus, an automatic glaucoma screening pipeline from optic disc (OD) localisation to glaucoma risk prediction is proposed in this study. The proposed pipeline consists of three main phases. Firstly, the OD is localised by morphological processing and sliding window methods. Secondly, a novel neural network which is in U‐shape and convolutional introduces concatenating path and fusion loss function is developed to split OD and optic cup (OC) at the same time. Thirdly, both clinical measurements including optic cup‐to‐disc ratio (CDR), neuroretinal rim related features, and hidden features including statistical moments, entropy and energy are combined to train glaucoma classifiers. According to the results of the experiment, the proposed segmentation network achieves the best performance on both OD and OC segmentation and the proposed CDR calculation method is capable of achieving the performance similar to that of ophthalmologist on CDR measurement. Besides, the authors’ glaucoma classification model can obtain the best performance on sensitivity and area under the curve score in comparison with the existing methods.
Fan Guo 0001, Yuxiang Mai, Jin Tang 0002, Xuanchu Duan, Beiji Zou 0001, Lingzi Jiang
IET Image Process.6
2019 A spatial-aware joint optic disc and cup segmentation method
Qing Liu 0003, Xiaopeng Hong, Shuo Li 0001, Zailiang Chen 0001, Guoying Zhao 0001, Beiji Zou 0001
Neurocomputing6
2019 Automatic Diabetic Retinopathy Screening via Cascaded Framework Based on Image- and Lesion-Level Features Fusion
Chengzhang Zhu, Beiji Zou 0001, Rongchang Zhao, Changlong Chen, Yalong Xiao
J. Comput. Sci. Technol.3
2019 Robust hybrid image watermarking scheme based on KAZE features and IWT-SVD
Xiyao Liu 0001, Yifan Wang 0008, Jingyu Du, Jieting Lou, Beiji Zou 0001
Multim. Tools Appl.6
2019 Distributed and Efficient Object Detection via Interactions Among Devices, Edge, and Cloud
abstract
With the rapid development of Internet-of-Things and communication techniques, media transmission in surveillance applications is gradually relying on wireless networks. Meanwhile, the emergence of edge computing has pushed the media data analysis from the cloud to the edge of the network to achieve fast response for delay-sensitive media processing tasks. Object detection is a representative delay-sensitive image processing task in surveillance applications, but faces significant challenges in this context. For example, how to compress images for transmission in wireless environment without compromising the detection accuracy, and how to integrate and update local inference models online in an edge computing-based object detection system. In this paper, we propose an object detection architecture based on edge computing to achieve distributed and efficient object detection for surveillance applications. Under this architecture, we develop an adaptive Region-of-Interest-based image compression scheme for end devices to efficiently compress their captured images for wireless transmission but not to sacrifice the object detection accuracy of edge servers. Furthermore, we carefully design distributed and communication-efficient interactions among end devices, edge servers, and the cloud to dynamically optimize the object detection accuracy online. Extensive simulation results demonstrate that our proposed architecture not only achieves a competitive detection accuracy to traditional cloud-based objective detection solution with reduced response delay but also significantly improves the image transmission efficiency with adaptive image compression ratio.
Yun-Di Guo, Beiji Zou 0001, Ju Ren 0001, Yaoxue Zhang
IEEE Trans. Multim.2
2018 Parameter Selection of Image Fog Removal Using Artificial Fish Swarm Algorithm
Fan Guo 0001, Gonghao Lan, Xiaoming Xiao, Beiji Zou 0001
ICIC (1)4
2018 An Approach for Glaucoma Detection Based on the Features Representation in Radon Domain
Beiji Zou 0001, Qilin Chen, Rongchang Zhao, Ping-Bo Ouyang, Chengzhang Zhu, Xuanchu Duan
ICIC (2)1
2018 Multi-Label Classification Scheme Based on Local Regression for Retinal Vessel Segmentation
abstract
The segmentation of small blood vessels whose width is less than 2 pixels in retinal images is a challenging problem. Existed methods rarely focus on the differences between small vessels and big vessels when doing segmentation. Therefore, previous methods are not accurate enough on small blood vessel segmentation. To effectively segment small blood vessels in retinal images including big vessels, we proposed a novel multi-label classification scheme for retinal vessel segmentation. In our proposed scheme, a local de-regression model is designed for multi-labeling and a convolutional neural network is used for multi-label classification. At addition, a local regression method is utilized to transform multi-label into binary label for locating small vessels. The experimental results show that our method achieves prominent performance for automatic retinal vessel segmentation, especially for small blood vessels.
Qi He 0008, Beiji Zou 0001, Chengzhang Zhu, Xiyao Liu 0001, Hongpu Fu, Lei Wang 0017
ICIP2
2018 Automatic Measurement of Cup-to-Disc Ratio for Retinal Images
Fan Guo 0001, Beiji Zou 0001, Xiyao Liu 0001, Rongchang Zhao
PRCV (1)3
2018 Classified optic disc localization algorithm based on verification model
Beiji Zou 0001, Changlong Chen, Chengzhang Zhu, Xuanchu Duan, Zailiang Chen 0001
Comput. Graph.1
2018 Localisation and segmentation of optic disc with the fractional-order Darwinian particle swarm optimisation algorithm
abstract
Automatic optic disc (OD) localisation and segmentation is still a great challenge in computer‐aided diagnosis and screening system. Here, a new OD segmentation algorithm is proposed based on the distinct features of OD in terms of its intensity and shape. The algorithm includes four stages: image preprocessing, image segmentation, ellipse fitting, and OD localisation and segmentation. In the preprocessing stage, the blood vessel in the input retinal image is removed by using the morphological operation and median filtering in HSL (hue–saturation–lightness) colour space. In the image segmentation and ellipse fitting stages, the fractional‐order Darwinian particle swarm optimisation algorithm is used to extract the brightest region, and the least‐squares optimisation is adopted to detect elliptical OD shape. Finally, the smooth OD borders are generated in the last stage. The proposed method is evaluated by the centroid difference, overlapping ratio, overlap score, and success indexes. Experimental results on the retinal images from DRION, MESSIDOR, ORIGA, and many other public databases demonstrate that the proposed method has superior performance, and may be a suitable tool for automated retinal image analysis.
Fan Guo 0001, Hui Peng 0001, Beiji Zou 0001, Rongchang Zhao, Xiyao Liu 0001
IET Image Process.3
2018 Improved multi-scale line detection method for retinal blood vessel segmentation
abstract
Changes of retinal blood vessel are precursors of many serious diseases such as diabetic retinopathy, hypertension and cardiovascular diseases. Automatic segmentation of retinal blood vessels in the fundus image can better assist in the diagnosis of these diseases and has been studied by many researchers. However, the segmentation of pale vessel pixels remains a problem because of their low contrasts with surrounding pixels. This study proposes an improved multi‐scale line detector to segment retinal vessels. It computes the line responses of vessels in multi‐scale windows and takes the maximum as the response value, which can enhance the responses of pale vessel pixels near strong vessels or dark background pixels. Experimental results on the publicly available database DRIVE demonstrate that the proposed method can detect pale vessel pixels better. It achieves 75.28% in sensitivity and 94.47% in accuracy, which outperforms the state‐of‐the‐art unsupervised methods. Compared with the supervised methods it also gets better sensitivity and comparable accuracy.
Kejuan Yue, Beiji Zou 0001, Zailiang Chen 0001, Qing Liu 0003
IET Image Process.2
2018 3D Filtering by Block Matching and Convolutional Neural Network for Image Denoising
Beiji Zou 0001, Yun-Di Guo, Qi He 0008, Ping-Bo Ouyang, Zailiang Chen 0001
J. Comput. Sci. Technol.1
2018 Distinguishable zero-watermarking scheme with similarity-based retrieval for digital rights Management of Fundus Image
Beiji Zou 0001, Jingyu Du, Xiyao Liu 0001, Yifan Wang 0008
Multim. Tools Appl.1
2017 A robust and synthesized-unseen watermarking for the DRM of DIBR-based 3D video
Xiyao Liu 0001, Fangfang Li 0004, Jingyu Du, Yang Guan, Yuesheng Zhu, Beiji Zou 0001
Neurocomputing6
2017 Automatic Retinal Image Registration Using Blood Vessel Segmentation and SIFT Feature
abstract
Automatic retinal image registration is still a great challenge in computer aided diagnosis and screening system. In this paper, a new retinal image registration method is proposed based on the combination of blood vessel segmentation and scale invariant feature transform (SIFT) feature. The algorithm includes two stages: retinal image segmentation and registration. In the segmentation stage, the blood vessel is segmented by using the guided filter to enhance the vessel structure and the bottom-hat transformation to extract blood vessel. In the registration stage, the SIFT algorithm is adopted to detect the feature of vessel segmentation image, complemented by using a random sample consensus (RANSAC) algorithm to eliminate incorrect matches. We evaluate our method from both segmentation and registration aspects. For segmentation evaluation, we test our method on DRIVE database, which provides manually labeled images from two specialists. The experimental results show that our method achieves 0.9562 in accuracy (Acc), which presents competitive performance compare to other existing segmentation methods. For registration evaluation, we test our method on STARE database, and the experimental results demonstrate the superior performance of the proposed method, which makes the algorithm a suitable tool for automated retinal image analysis.
Fan Guo 0001, Beiji Zou 0001, Yixiong Liang
Int. J. Pattern Recognit. Artif. Intell.3
2017 Automatic Anterior Lamina Cribrosa Surface Depth Measurement Based on Active Contour and Energy Constraint
Zailiang Chen 0001, Beiji Zou 0001, Hailan Shen, Rongchang Zhao
J. Comput. Sci. Technol.3
2017 Supervised Vessels Classification Based on Feature Selection
Beiji Zou 0001, Chengzhang Zhu, Zailiang Chen 0001, Zi-Qian Zhang
J. Comput. Sci. Technol.1
2017 Orientation Histogram-Based Center-Surround Interaction: An Integration Approach for Contour Detection
abstract
Contour is a critical feature for image description and object recognition in many computer vision tasks. However, detection of object contour remains a challenging problem because of disturbances from texture edges. This letter proposes a scheme to handle texture edges by implementing contour integration. The proposed scheme integrates structural segments into contours while inhibiting texture edges with the help of the orientation histogram-based center-surround interaction model. In the model, local edges within surroundings exert a modulatory effect on central contour cues based on the co-occurrence statistics of local edges described by the divergence of orientation histograms in the local region. We evaluate the proposed scheme on two well-known challenging boundary detection data sets (RuG and BSDS500). The experiments demonstrate that our scheme achieves a high [Formula: see text]-measure of up to 0.74. Results show that our scheme achieves integrating accurate contour while eliminating most of texture edges, a novel approach to long-range feature analysis.
Rongchang Zhao, Min Wu 0002, Xiyao Liu 0001, Beiji Zou 0001, Fangfang Li 0004
Neural Comput.4
2017 Novel robust zero-watermarking scheme for digital rights management of 3D videos
Xiyao Liu 0001, Rongchang Zhao, Fangfang Li 0004, Yipeng Ding, Beiji Zou 0001
Signal Process. Image Commun.6
2017 Hierarchical Contour Closure-Based Holistic Salient Object Detection
abstract
Most existing salient object detection methods compute the saliency for pixels, patches, or superpixels by contrast. Such fine-grained contrast-based salient object detection methods are stuck with saliency attenuation of the salient object and saliency overestimation of the background when the image is complicated. To better compute the saliency for complicated images, we propose a hierarchical contour closure-based holistic salient object detection method, in which two saliency cues, i.e., closure completeness and closure reliability, are thoroughly exploited. The former pops out the holistic homogeneous regions bounded by completely closed outer contours, and the latter highlights the holistic homogeneous regions bounded by averagely highly reliable outer contours. Accordingly, we propose two computational schemes to compute the corresponding saliency maps in a hierarchical segmentation space. Finally, we propose a framework to combine the two saliency maps, obtaining the final saliency map. Experimental results on three publicly available datasets show that even each single saliency map is able to reach the state-of-the-art performance. Furthermore, our framework, which combines two saliency maps, outperforms the state of the arts. Additionally, we show that the proposed framework can be easily used to extend existing methods and further improve their performances substantially.
Qing Liu 0003, Xiaopeng Hong, Beiji Zou 0001, Jie Chen 0001, Zailiang Chen 0001, Guoying Zhao 0001
IEEE Trans. Image Process.3
2017 Scene text detection using adaptive color reduction, adjacent character model and hybrid verification strategy
Beiji Zou 0001, Jianjing Guo
Vis. Comput.2
2016 Automatic segmentation for cell images based on bottleneck detection and ellipse fitting
Miao Liao, Yuqian Zhao 0001, Xiang-hua Li, Peishan Dai, Xiao-wen Xu, Jun-kai Zhang, Beiji Zou 0001
Neurocomputing7
2016 Natural scene text detection by multi-scale adaptive color clustering and non-text filtering
Beiji Zou 0001, Zailiang Chen 0001, Chengzhang Zhu, Jianjing Guo
Neurocomputing2
2016 Saliency detection using boundary information
Beiji Zou 0001, Qing Liu 0003, Zailiang Chen 0001, Shijian Liu
Multim. Syst.1
2016 An automatic video text detection method based on BP-adaboost
Beiji Zou 0001, Hongpu Fu
Multim. Tools Appl.2
2015 An Improved Retinal Vessel Segmentation Method Based on Supervised Learning
abstract
This paper propose an improved supervised method for retinal vessel segmentation based on Extreme Learning Machine (ELM). Firstly, a 36-D feature vector is extracted for each pixel of the fund us image consisting of local features, morphological features and divergence of vector fields. Then a matrix for pixels of the training set using the feature vector and the manual segmentation is constructed as the input of the ELM. Finally a classifier is obtained to segment the retinal vessels. The method is evaluated with the DRIVE database and the average accuracy is 0.9581. And the running time is greatly decreased by using ELM. It is applicable for computer-aided diagnosis and disease screening.
Chengzhang Zhu, Beiji Zou 0001, Yao Xiang, Jinkai Cui
CAD/Graphics2
2015 Surroundedness based multiscale saliency detection
Beiji Zou 0001, Qing Liu 0003, Zailiang Chen 0001, Hongpu Fu, Chengzhang Zhu
J. Vis. Commun. Image Represent.1
2014 Motion recognition for 3D human motion capture data using support vector machines with rejection determination
Meiling Cai, Beiji Zou 0001, Huanzhi Gao, Juan Song
Multim. Tools Appl.2
2013 A novel particle filter with implicit dynamic model for irregular motion tracking
Beiji Zou 0001, Lingzhi Li 0006
Mach. Vis. Appl.2
2012 Considering direct interaction of artificial ant colony foraging simulation and animation
abstract
Direct interaction between ants exists in nature, but is rarely seen in ant colony foraging simulations. In this article, we propose a novel ant colony foraging simulation model named Direct Ant Colony Foraging (DACF), which combines the direct interaction mechanism with the indirect pheromone interaction mechanism. Based on Wilensky's ant colony foraging model and Panait's model, we develop DACF2 and DACF3, separately. These models not only well describe the ant colony foraging behaviours in nature, but also give inspiration to the research of the emergent behaviours in complex systems. The simulation results demonstrate not only the superiority of DACF2 and DACF3, but also the validity and effectiveness of the direct interaction mechanism. Also, ant colony foraging animation, which is run with DACF, demonstrates the model in another way.
Zhigang Meng, Beiji Zou 0001
J. Exp. Theor. Artif. Intell.2
2011 Exploring regularized feature selection for person specific face verification
abstract
In this paper, we explore the regularized feature selection method for person specific face verification in unconstrained environments. We reformulate the generalization of the single-task sparsity-enforced feature selection method to multi-task cases as a simultaneous sparse approximation problem. We also investigate two feature selection strategies in the multi-task generalization based on the positive and negative feature correlation assumptions across different persons. Simultaneous orthogonal matching pursuit (SOMP) is adopted and modified to solve the corresponding optimization problems. We further proposed a named simultaneous subspace pursuit (SSP) methods which generalize the subspace pursuit method to solve the corresponding optimization problems. The performance of different feature selection strategies and different solvers for face verification are compared on the challenging LFW face database. Our experimental results show that 1) the selected subsets based on positive correlation assumption are more effective than those based on the negative correlation assumption; 2) the OMP-based solvers outperform SP-based solvers in terms of feature selection and 3) the regularized methods with OMP-based solvers can outperform state-of-the-art feature selection methods.
Yixiong Liang, Lei Wang 0017, Beiji Zou 0001
ICCV4
2011 Multi-task GLOH feature selection for human age estimation
abstract
In this paper, we propose a novel age estimation method based on gradient location and orientation histogram (GLOH) descriptor and multi-task learning (MTL). The GLOH, one of the state-of-the-art local descriptor, is used to capture the age- related local and spatial information of face image. As the extracted GLOH features are often redundant, MTL is designed to select the most informative GLOH bins for age estimation problem, while the corresponding weights are determined by ridge regression. This approach largely reduces the dimensions of feature, which can not only improve performance but also decrease the computational burden. Experiments on the public available FG-NET database show that the proposed method can achieve comparable performance over previous approaches while using much fewer features.
Yixiong Liang, Lingbo Liu, Yao Xiang, Beiji Zou 0001
ICIP5
2011 Feature selection via simultaneous sparse approximation for person specific face verification
abstract
There is an increasing use of some imperceivable and redundant local features for face recognition. While only a relatively small fraction of them is relevant to the final recognition task, the feature selection is a crucial and necessary step to select the most discriminant ones to obtain a compact face representation. In this paper, we investigate the sparsity-enforced regularization-based feature selection methods and propose a multi-task feature selection method for building person specific models for face verification. We assume that the person specific models share a common subset of features and novelly reformulated the common subset selection problem as a simultaneous sparse approximation problem. The effectiveness of the proposed methods is verified with the challenging LFW face databases.
Yixiong Liang, Lei Wang 0017, Beiji Zou 0001
ICIP4
2011 Practical craniofacial surgery simulator based on GPU accelerated lattice shape matching
abstract
Abstract This paper presents an intuitive and practical craniofacial surgery simulation system, which is suitable for daily clinical practice. The key component of the system is a GPU accelerated lattice shape matching method, to estimate individual patient post‐operative appearance interactively. The lattice model can be set up from individual CT data in an easy and robust way, incorporating CT mapping physical inhomogeneous material as well as orthotropic behavior of soft tissue. Instead of simulating dynamic behavior, an iterative optimization is used for direct computation of soft tissue deformation. In addition, a GPU acceleration framework for the lattice shape matching method is further exploited based on the lattice regularity, thus the simulation operations can be performed at interactive or real time. Finally, a craniofacial surgery simulation system is developed for daily clinical practice, which is capable of simulating a variety of surgeries including osteotomies, bone fragment repositioning, and insertion of implants. The treatment of more than 50 patients was found to provide a good correlation between simulation and post‐operative outcome. Copyright © 2011 John Wiley & Sons, Ltd.
Yixiong Liang, Lingzhi Li 0006, Beiji Zou 0001, Xing-Hao Zhu
Comput. Animat. Virtual Worlds4
2010 Secure and Efficient Data Aggregation for Wireless Sensor Networks
abstract
This paper addresses the secure data aggregation for wireless sensor networks (WSNs) with both static tree architecture and dynamic cluster-based architecture. For WSNs with static tree architecture, we propose the Leaf Node Representation (LNR) scheme to solve the Id problem and make the key stream-based encrypted data aggregation feasible and practical for large scale networks. For WSNs with dynamic cluster-based architectures, we propose the Delayed Hop-by-hop Authentication (DHA) scheme to provide hop-by-hop data integrity and data freshness only using individual keys. Analytical results show that the proposed scheme can reduce the communication overhead significantly compared to a well known existing scheme.
Xiaoyan Wang 0003, Jie Li 0002, Xiaoning Peng, Beiji Zou 0001
VTC Fall4
2010 Reconstruction of intersecting curved solids from 2D orthographic views
Zi-Gang Fu, Beiji Zou 0001
Comput. Aided Des.2
2010 Enhanced Hexagonal-Based Search Using Direction-Oriented Inner Search for Motion Estimation
abstract
The newly developed enhanced hexagonal-based search using point-oriented inner search (EHS-POIS) enormously speeds up hexagon-based search (HS). From a different perspective, an inherent correlation between distortion and spatial direction through statistical analysis is found. Based on the observed distortion distribution, a novel enhanced hexagonal-based search with direction-oriented inner search (EHS-DIOS) is proposed to avoid real distortion calculation and thus reduce high computation. Experimental results show that, the proposed algorithm is faster than EHS-POIS by achieving two times improvement in terms of inner search speed, and as compared with previous works, it makes a better tradeoff between speed and decoded image quality.
Beiji Zou 0001, Cao Shi, Canhui Xu 0001
IEEE Trans. Circuits Syst. Video Technol.1
2009 Curvature normal vector driven interpolatory subdivision
abstract
We present an intrinsically nonlinear interpolatory subdivision scheme with geometric information and some free parameters via discrete curvatures normal vector. Our scheme can produce fair G1-continuous curves, which can avoid the potential pitfalls and unacceptable cases appeared in the four-point subdivision scheme. Furthermore, with the proper parameter choice, the proposed scheme is convexity-preserving, and reproduces the conic curve. Finally, the experimental results show our scheme is effective.
Huanxi Zhao, Xia Qiu, Luming Liang, Beiji Zou 0001
Shape Modeling International5
2009 Automatic reconstruction of 3D human motion pose from uncalibrated monocular video sequences based on markerless human motion tracking
Beiji Zou 0001, Cao Shi, Umugwaneza Marie Providence
Pattern Recognit.1
2009 Selective color transfer with multi-source images
Yao Xiang, Beiji Zou 0001
Pattern Recognit. Lett.2
2008 Multi-source color transfer for natural images
abstract
Considering a "target" image with complex scene, a novel method for regionally transferring colors from a set of "source" color images to the complex "target" image is presented. An improved EM (expectation-maximization) algorithm is first adopted to approximately model the regional color distribution of the target image by using Gaussian mixture model (GMM); trained by this target GMM, color pixels in each source image are then assigned into source regions corresponding to the target ones, and a color selection mechanism is proposed to choose the best- matched source colors for re-coloring each target region. A new image quality metric, colorfulness, is adopted for objective performance comparisons with previous methods. Experiments show that the proposed method can generate convincing results with well performance.
Yao Xiang, Beiji Zou 0001, Maddy Hui Wang
ICIP2
2008 A note on the paper "Normal based subdivision scheme for curve design" by Xunnian Yang
Luming Liang, Huanxi Zhao, Beiji Zou 0001
Comput. Aided Geom. Des.3
2007 A New Algorithm for Trademark Image Retrieval Based on Sub-block of Polar Coordinates
Beiji Zou 0001
ICEC1
2005 Bi-Objective Model for Test-Suite Reduction Based on Modified Condition/Decision Coverage
abstract
It is evidence that modified condition/decision coverage (MC/DC) is an effective verification method and can help to detect safety faults despite of its expensive cost. In regression testing, it is quite costly to rerun all of test cases in test suite because new test cases are added to test suite as the software evolves. Therefore, it is necessary to reduce the test suite to improve test efficiency and save test cost. Many existing test-suite reduction techniques are not effective to reduce MC/DC test suite. This paper proposes a new test-suite reduction technique for MC/DC: a bi-objective model that considers both the coverage degree of test case for test requirements and the capability of test cases to reveal error. Our experiment results show that the technique both reduces the size of test suite and better ensures the effectiveness of test suite to reveal error.
Lili Pan 0002, Beiji Zou 0001, Hao Chen 0051
PRDC3