Dongming Zhou 0001

dblp:92/2875-1 · DBLP profile ↗
← Back
65ranked-venue papers
3as first author
44since 2021 · last 2026
0000-0003-0139-9415ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 39 · 2 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 14 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Systems, architecture and hardware · 3 · 3 since 2021
YearPublicationVenuePosition
2026 RAMR: A role-adaptive modality recalibration network for RGBT tracking
Zhao Gao, Dongming Zhou 0001, Yisong Liu, Qingqing Shan
Expert Syst. Appl.2
2026 Learning heterogeneous biological interactions via meta-relation-guided dual-channel graph transformer for circular ribonucleic acid function prediction
Yanbu Guo, Haokun Zhu, Jinde Cao, Rui Chen 0005, Dongming Zhou 0001
Knowl. Based Syst.5
2026 Cross-model and attribute-driven dual-stage knowledge distillation for multimodal medical image fusion
Yanyu Liu, Chunxue Liu, Ruichao Hou, Zhaisheng Ding, Kangjian He, Dongming Zhou 0001
Multim. Syst.6
2026 Learning frequency and memory-aware prompts for multi-modal object tracking
Boyue Xu, Ruichao Hou, Tongwei Ren, Dongming Zhou 0001, Gangshan Wu, Jinde Cao
Pattern Recognit.4
2026 Gradient-guided multi-task framework with genetic optimization for medical image segmentation and classification
Dongming Zhou 0001, Jinde Cao, Weina Zhu
Pattern Recognit.2
2026 HyPSAM: Hybrid Prompt-Driven Segment Anything Model for RGB-Thermal Salient Object Detection
abstract
RGB-thermal salient object detection (RGB-T SOD) aims to identify prominent objects by integrating complementary information from RGB and thermal modalities. However, learning the precise boundaries and complete objects remains challenging due to the intrinsic insufficient feature fusion and the extrinsic limitations of data scarcity. In this paper, we propose a novel hybrid prompt-driven segment anything model (HyPSAM), which leverages the zero-shot generalization capabilities of the segment anything model (SAM) for RGB-T SOD. Specifically, we first propose a dynamic fusion network (DFNet) that generates high-quality initial saliency maps as visual prompts. DFNet employs dynamic convolution and multi-branch decoding to facilitate adaptive cross-modality interaction, overcoming the limitations of fixed-parameter kernels and enhancing multi-modal feature representation. Moreover, we propose a plug-and-play refinement network (P2RNet) which serves as a general optimization strategy to guide SAM in refining saliency maps by using hybrid prompts. The text prompt ensures reliable modality input, while the mask and box prompts enable precise salient object localization. Extensive experiments on three public datasets demonstrate that our method achieves state-of-the-art performance. Notably, HyPSAM has remarkable versatility, seamlessly integrating with different RGB-T SOD methods to achieve significant performance gains, thereby highlighting the potential of prompt engineering in this field. The code and results of our method are available at: https://github.com/milotic233/HyPSAM.
Ruichao Hou, Tongwei Ren, Dongming Zhou 0001, Gangshan Wu, Jinde Cao
IEEE Trans. Circuits Syst. Video Technol.4
2025 EGLC: Enhancing Global Localization Capability for medical image segmentation
Yulong Wan, Dongming Zhou 0001
Comput. Vis. Image Underst.2
2025 Deep gate information bottleneck-based prediction model for complex disease-related micro-ribonucleic acids via heterogeneous biological networks
Yanbu Guo, Yiyang Xin, Jinde Cao, Yaoli Xu, Dongming Zhou 0001
Eng. Appl. Artif. Intell.5
2025 MCINet: Multimodal context-aware network for RGBT tracking
Zhao Gao, Dongming Zhou 0001, Yisong Liu, Qingqing Shan
Knowl. Based Syst.2
2025 MKFTracker: An RGBT tracker via multimodal knowledge embedding and feature interaction
Weidai Xia, Dongming Zhou 0001, Jinde Cao
Knowl. Based Syst.3
2025 Two-stage Unidirectional Fusion Network for RGBT tracking
Yisong Liu, Zhao Gao, Yang Cao 0003, Dongming Zhou 0001
Knowl. Based Syst.4
2025 SCDFuse: A semantic complementary distillation framework for joint infrared and visible image fusion and denoising
Shidong Xie, Yongsheng Zang, Jinde Cao, Dongming Zhou 0001, Mingchuan Tan, Zhaisheng Ding, Guanbo Wang
Knowl. Based Syst.5
2025 ACL-Net: Attribute-Aware Contrastive Learning Network for Medical Image Fusion
abstract
Medical image fusion aims to integrate multi-sensor source images into a unified representation, providing comprehensive and diagnostically enriched information to support clinical decision-making. However, the scarcity of labeled data presents significant challenges in effectively learning complementary features across modalities. In this paper, we propose a novel attribute-aware contrastive learning network, called ACL-Net, boosting medical image fusion performance. Specifically, we introduce the attribute transformation strategy to simulate variations in pixel intensity and structural patterns, guiding the model to focus on critical cross-modal information. In this way, it enhances contrastive learning by generating diverse negative pairs, thereby mitigating the scarcity of negative samples in unsupervised fusion scenarios. Extensive experiments demonstrate that our method achieves superior performance compared to state-of-the-art medical image fusion methods.
Yanyu Liu, Ruichao Hou, Zhaisheng Ding, Dongming Zhou 0001, Jinde Cao
IEEE Signal Process. Lett.4
2025 Pathological Image Segmentation of Breast Cancer via Template Matching
abstract
Accurate pathological image segmentation is crucial for the clinical diagnosis of breast cancer. However, existing methods of pathological segmentation face challenges due to the variability and complexity of breast cancer on pathological images. To address these issues, we propose a novel segmentaion network called template-matching pathological segmentation network. Our method incorporates an innovative template matching strategy inspired by the diagnostic process of pathologists. The template matching strategy is to utilize visual transformer to establish a correlative relationship between cancer lesions and corresponding templates. To improve feature utilization of pathological images, PSVTNet introduces detailed information attention and information entropy attention. Detailed information attention aims to exploit detailed information by serving as the path connecting shallow-layer and deep-layer features. Meanwhile, information entropy attention can redistribute feature weights to high-entropy regions according to the information-entropy attention map. Additionally, this work releases a comprehensive pathological dataset that comprises labeled pathological images. These images are collected from breast and stomach cancers with hematoxylin&eosin and human epidermal growth factor receptor-2 staining. Extensive experiments demonstrate that PSVTNet significantly outperforms state-of-the-art methods on pathologic images of breast cancer, but can also process pathologic images of stomach cancer carrying with same diagnosed features as the breast cancer.
Kaixiang Yan, Yanyu Liu, Jinde Cao, Dongming Zhou 0001
IEEE J. Biomed. Health Informatics5
2025 $\hbox {KD}^{3}$mt: knowledge distillation-driven dynamic mixer transformer for medical image fusion
Zhaijuan Ding, Yanyu Liu, Kangjian He, Dongming Zhou 0001
Vis. Comput.5
2024 Multimodal Sentiment Analysis Based on 3D Stereoscopic Attention
abstract
In the multimodal (text, audio, and visual) sentiment analysis, the current methods generally consider the bi-modal sentiment interaction, resulting in inadequate mining and fusion of relations between modalities. In this paper, we propose the concept of multimodal 3D (3-Dimensional) stereoscopic attention for the first time, which constructs the tri-modal stereoscopic attention with temporal sequences simultaneously to adequately structure the sentiment interaction. To solve the problems of stereoscopic attention construction such as the increased complexity of algorithms caused by rising dimensions, we propose a progressive construction method with 2D attention as an intermediate process. To implement sentiment relations based on stereoscopic attention to integrating modal information sufficiently, a forward propagation mechanism is proposed, which optimizes the representations of each modality with multimodal modulation. The results on two public datasets confirm the superiority of the proposed method in all metrics to the baselines.
Dongming Zhou 0001, Zhengpeng Zhao, Dan Xu 0001, Jinde Cao
ICASSP3
2024 Robust multi-focus image fusion using focus property detection and deep image matting
Changcheng Wang, Yongsheng Zang, Dongming Zhou 0001, Jiatian Mei, Rencan Nie, Lifen Zhou
Expert Syst. Appl.3
2024 Dynamic hypergraph convolutional network for multimodal sentiment analysis
Dongming Zhou 0001, Jinde Cao, Jinjing Gu, Zhengpeng Zhao, Dan Xu 0001
Neurocomputing3
2024 Co-space Representation Interaction Network for multimodal sentiment analysis
Zhengpeng Zhao, Dongming Zhou 0001, Dan Xu 0001, Jinde Cao
Knowl. Based Syst.5
2024 Two-subnet network for real-world image denoising
Lianmin Zhou, Dongming Zhou 0001, Hao Yang 0040, Shaoliang Yang
Multim. Tools Appl.2
2024 SiamMGT: robust RGBT tracking via graph attention and reliable modality weight learning
Lizhi Geng, Dongming Zhou 0001, Kerui Wang, Yisong Liu, Kaixiang Yan
J. Supercomput.2
2024 Context-Aware Poly(A) Signal Prediction Model via Deep Spatial-Temporal Neural Networks
abstract
Polyadenylation [Poly(A)] is an essential process during messenger RNA (mRNA) maturation in biological eukaryote systems. Identifying Poly(A) signals (PASs) from the genome level is the key to understanding the mechanism of translation regulation and mRNA metabolism. In this work, we propose a deep dual-dynamic context-aware Poly(A) signal prediction model, called multiscale convolution with self-attention networks (MCANet), to adaptively uncover the spatial-temporal contextual dependence information. Specifically, the model automatically learns and strengthens informative features from the temporalwise and the spatialwise dimension. The identity connectivity performs contextual feature maps of Poly(A) data by direct connections from previous layers to subsequent layers. Then, a fully parametric rectified linear unit (FP-RELU) with dual-dynamic coefficients is devised to make the training of the model easier and enhance the generalization ability. A cross-entropy loss (CL) function is designed to make the model focus on samples that are easy to misclassify. Experiments on different Poly(A) signals demonstrate the superior performance of the proposed MCANet, and an ablation study shows the effectiveness of the network design for the feature learning and prediction of Poly(A) signals.
Yanbu Guo, Dongming Zhou 0001, Chaoyang Li 0001, Jinde Cao
IEEE Trans. Neural Networks Learn. Syst.2
2024 A novel highland and freshwater-circumstance dataset: advancing underwater image enhancement
Kaixiang Yan, Dongming Zhou 0001, Changcheng Wang, Jiarui Quan
Vis. Comput.3
2023 A robust infrared and visible image fusion framework via multi-receptive-field attention and color visual perception
Zhaisheng Ding, Dongming Zhou 0001, Yanyu Liu, Ruichao Hou
Appl. Intell.3
2023 EDAfuse: A encoder-decoder with atrous spatial pyramid network for infrared and visible image fusion
abstract
Abstract Infrared and visible images come from different sensors, and they have their advantages and disadvantages. In order to make the fused images contain as much salience information as possible, a practical fusion method, termed EDAfuse, is proposed in this paper. In EDAfuse, the authors introduce an encoder–decoder with the atrous spatial pyramid network for infrared and visible image fusion. The authors use the encoding network which includes three convolutional neural network (CNN) layers to extract deep features from input images. Then the proposed atrous spatial pyramid model is utilized to get five different scale features. The same scale features from the two original images are fused by our fusion strategy with the attention model and information quantity model. Finally, the decoding network is utilized to reconstruct the fused image. In the training process, the authors introduce a loss function with saliency loss to improve the ability of the model for extracting salient features from original images. In the experiment process, the authors use the average values of seven metrics for 21 fused images to evaluate the proposed method and the other seven existing methods. The results show that our method has four best values and three second‐best values. The subjective assessment also demonstrates that the proposed method outperforms the state‐of‐the‐art fusion methods.
Cairen Nie, Dongming Zhou 0001, Rencan Nie
IET Image Process.2
2023 Collaborative fine-grained interaction learning for image-text sentiment analysis
Xingwang Xiao, Dongming Zhou 0001, Jinde Cao, Jinjing Gu, Zhengpeng Zhao, Dan Xu 0001
Knowl. Based Syst.3
2023 An interactive deep model combined with Retinex for low-light visible and infrared image fusion
Changcheng Wang, Yongsheng Zang, Dongming Zhou 0001, Rencan Nie, Jiatian Mei
Neural Comput. Appl.3
2023 Variational gated autoencoder-based feature extraction model for inferring disease-miRNA associations based on multiview features
Yanbu Guo, Dongming Zhou 0001, Xiaoli Ruan, Jinde Cao
Neural Networks2
2023 RGBT Tracking via Multi-stage Matching Guidance and Context integration
Kaixiang Yan, Changcheng Wang, Dongming Zhou 0001
Neural Process. Lett.3
2023 CDMC-Net: Context-Aware Image Deblurring Using a Multi-scale Cascaded Network
Qian Zhao 0014, Dongming Zhou 0001, Hao Yang 0040
Neural Process. Lett.2
2023 An Improved Hybrid Network With a Transformer Module for Medical Image Fusion
abstract
Medical image fusion technology is an essential component of computer-aided diagnosis, which aims to extract useful cross-modality cues from raw signals to generate high-quality fused images. Many advanced methods focus on designing fusion rules, but there is still room for improvement in cross-modal information extraction. To this end, we propose a novel encoder-decoder architecture with three technical novelties. First, we divide the medical images into two attributes, namely pixel intensity distribution attributes and texture attributes, and thus design two self-reconstruction tasks to mine as many specific features as possible. Second, we propose a hybrid network combining a CNN and a transformer module to model both long-range and short-range dependencies. Moreover, we construct a self-adaptive weight fusion rule that automatically measures salient features. Extensive experiments on a public medical image dataset and other multimodal datasets show that the proposed method achieves satisfactory performance.
Yanyu Liu, Yongsheng Zang, Dongming Zhou 0001, Jinde Cao, Rencan Nie, Ruichao Hou, Zhaisheng Ding, Jiatian Mei
IEEE J. Biomed. Health Informatics3
2023 External-attention dual-modality fusion network for RGBT tracking
Kaixiang Yan, Jiatian Mei, Dongming Zhou 0001, Lifen Zhou
J. Supercomput.3
2023 RainFormer: a pyramid transformer for single image deraining
Hao Yang 0040, Dongming Zhou 0001, Jinde Cao, Qian Zhao 0014
J. Supercomput.2
2023 A two-stage network with wavelet transformation for single-image deraining
Hao Yang 0040, Dongming Zhou 0001, Qian Zhao 0014
Vis. Comput.2
2022 Deep multi-scale Gaussian residual networks for contextual-aware translation initiation site recognition
Yanbu Guo, Dongming Zhou 0001, Weihua Li 0006, Jinde Cao
Expert Syst. Appl.2
2022 CIRNet: An improved RGBT tracking via cross-modality interaction and re-identification
Weidai Xia, Dongming Zhou 0001, Jinde Cao, Yanyu Liu, Ruichao Hou
Neurocomputing2
2022 Gated residual neural networks with self-normalization for translation initiation site recognition
Yanbu Guo, Dongming Zhou 0001, Jinde Cao, Rencan Nie, Xiaoli Ruan, Yanyu Liu
Knowl. Based Syst.2
2022 Context-aware dynamic neural computational models for accurate Poly(A) signal prediction
Yanbu Guo, Chaoyang Li 0001, Dongming Zhou 0001, Jinde Cao, Hui Liang 0004
Neural Networks3
2022 Rethinking Low-Light Enhancement via Transformer-GAN
abstract
Images and videos shot in low light are often accompanied by severe image degradation, such as color noise, chromatic aberrations and loss of details. Most existing convolutional neural network (CNN)-based low-light enhancement methods focus on decomposing the image into illumination and reflection parts via the Retinex model, but these methods often fail to adequately consider controlling noise during enhancement and perform poorly in the face of complex lighting environments. In this letter, we propose a powerful Vision Transformer-based Generative Adversarial Network (Transformer-GAN) for enhancing low-light images. Transformer-GAN consists of two subnets as follows: (1) the feature extraction is achieved by an iterative multi-branch network in the feature extraction subnet, and (2) the enhancement is completed in the image reconstruction subnet. The innovative core works are multi-head multi-covariance self-attention (MHMCA) and Light feature-forward module structures (LFFM) in Transformer-GAN. Experiments demonstrate that our method outperforms state-of-the-art low-light enhancement methods on popular low-light datasets.
Shaoliang Yang, Dongming Zhou 0001, Jinde Cao, Yanbu Guo
IEEE Signal Process. Lett.2
2022 How to Analyze the Neurodynamic Characteristics of Pulse-Coupled Neural Networks? A Theoretical Analysis and Case Study of Intersecting Cortical Model
abstract
The intersecting cortical model (ICM), initially designed for image processing, is a special case of the biologically inspired pulse-coupled neural-network (PCNN) models. Although the ICM has been widely used, few studies concern the internal activities and firing conditions of the neuron, which may lead to an invalid model in the application. Furthermore, the lack of theoretical analysis has led to inappropriate parameter settings and consequent limitations on ICM applications. To address this deficiency, we first study the continuous firing condition of ICM neurons to determine the restrictions that exist between network parameters and the input signal. Second, we investigate the neuron pulse period to understand the neural firing mechanism. Third, we derive the relationship between the continuous firing condition and the neural pulse period, and the relationship can prove the validity of the continuous firing condition and the neural pulse period as well. A solid understanding of the neural firing mechanism is helpful in setting appropriate parameters and in providing a theoretical basis for widespread applications to use the ICM model effectively. Extensive experiments of numerical tests with a common image reveal the rationality of our theoretical results.
Xin Jin 0005, Dongming Zhou 0001, Xing Chu, Shaowen Yao 0001, Keqin Li 0001, Wei Zhou 0011
IEEE Trans. Cybern.2
2022 A Total Variation With Joint Norms For Infrared and Visible Image Fusion
abstract
A single infrared image or visible image for the same scene is usually insufficient to simultaneously reveal the infrared objects and the scene details. Thus, image fusion techniques play an important role in producing a single image from the images captured by infrared and visible sensors. In this paper, we propose a novel total variation (TV)-based fusion for infrared and visible images. In our model, a weighted fidelity term is employed to fuse both the infrared objects in the infrared image and the salient scenes in the visible image. To this end, a weight estimation method is developed based on the global luminance contrast-based saliency. Also, to overcome the over-fitting, two constraints are further introduced to merge more details from the visible image and prevent the luminance degradation for the fused result, respectively. Moreover, joint norms are exploited to produce a better result.${{\boldsymbol{l}}_{2,1,{\boldsymbol{rc}}}}$provides the structural group sparseness for the fidelity term, whereas${{\boldsymbol{l}}_{1/2}}$presents the better gradient sparse for the detail preserving term and${{\boldsymbol{l}}_2}$is utilized for the luminance degradation preventing term. Experimental results indicate that the proposed method can give state-of-the-art performances both in visual perception and quantitative scores than other methods.
Rencan Nie, Chaozhen Ma, Jinde Cao, Hongwei Ding 0001, Dongming Zhou 0001
IEEE Trans. Multim.5
2022 AEMS: an attention enhancement network of modules stacking for lowlight image enhancement
Dongming Zhou 0001, Rencan Nie, Yanyu Liu, Yixue Wei
Vis. Comput.3
2021 AMBCR: Low-light image enhancement via attention guided multi-branch construction and Retinex theory
abstract
Abstract Due to different lighting environments and equipment limitations, low‐light images have high noise, low contrast and unobvious colours. The main purpose of low‐light image enhancement is to preserve the details and suppress noise as much as possible while improving the contrast of the image. Here, different networks are first combined to construct a multi‐branch module for features extraction, and use the module and Retinex theory to extract the reflection map of the image. Then an attention mechanism is introduced into the multi‐branch construction to balance the feature weight of each branch, and get the final result by the reconstruction module. The Retinex theory is used to calculate the L 1 loss and the gradient loss for the intermediate feature map of the entire model to train our framework. The entire process is completed in an end‐to‐end‐way, which avoids the hand‐crafted reconstruction rules and reduces the workload. What's more, a large number of experiments demonstrate that the proposed framework performs better results than state‐of‐the‐art algorithms in both quantitative and qualitative evaluations of image enhancement.
Dongming Zhou 0001, Rencan Nie, Shidong Xie, Yanyu Liu
IET Image Process.2
2021 Multi-Source Information Exchange Encoding With PCNN for Medical Image Fusion
abstract
Multimodal medical image fusion (MMIF) is to merge multiple images for better imaging quality with preserving different specific features, which could be more informative for efficient clinical diagnosis. In this paper, a novel fusion framework is proposed for multimodal medical images based on multi-source information exchange encoding (MIEE) by using Pulse Coupled Neural Network (PCNN). We construct an MIEE model by using two types of PCNN, such that the information of an image can be exchanged and encoded to another image. Then the fusion contributions for each pixel are estimated qualitatively according to a logical comparison of exchanged information. Further, the exchanged information is nonlinearly transformed using an exponential function with a functional parameter. Finally, quantitative fusion contributions are produced through a reverse-proportional operator to the exchanged information. Also, particle swarm optimization-based derivative-free optimization and a total vibration-based derivative optimization are used to optimize the PCNN and functional transform parameters, respectively. Experiments demonstrate that our method gives the best results than other state-of-the-art fusion approaches.
Rencan Nie, Jinde Cao, Dongming Zhou 0001, Wenhua Qian
IEEE Trans. Circuits Syst. Video Technol.3
2020 Construction of high dynamic range image based on gradient information transformation
abstract
This study proposes a fusion method for high dynamic range images based on gradient information transformation. In the proposed work, the authors first measure the three exposure weights of the source images, namely, local contrast, luminance and spatial structure. Then, the exposure weights are merged through a multi‐scale Laplacian pyramid scheme. For the weight maps measurement, the dense scale‐invariant feature transform method is used to calculate the local contrast around each pixel location, rather than a single pixel. The image luminance levels are computed in the gradient domain to get more visual information and the authors leverage the dictionary learning to effectively extract the luminance of images. Additionally, to better preserve the spatial structure of the source images, the just‐noticeable‐distortion technique is employed. By comparing the experimental results both subjectively and objectively, it is evident that the proposed method represents an improvement over some exciting methods.
Yanyu Liu, Dongming Zhou 0001, Rencan Nie, Ruichao Hou, Zhaisheng Ding
IET Image Process.2
2020 DeepANF: A deep attentive neural framework with distributed representation for chromatin accessibility prediction
Yanbu Guo, Dongming Zhou 0001, Rencan Nie, Xiaoli Ruan, Weihua Li 0006
Neurocomputing2
2020 Attentive gated neural networks for identifying chromatin accessibility
Yanbu Guo, Dongming Zhou 0001, Weihua Li 0006, Rencan Nie, Ruichao Hou, Chengli Zhou
Neural Comput. Appl.2
2019 DeepACLSTM: deep asymmetric convolutional long short-term memory neural models for protein secondary structure prediction
abstract
BACKGROUND: Protein secondary structure (PSS) is critical to further predict the tertiary structure, understand protein function and design drugs. However, experimental techniques of PSS are time consuming and expensive, and thus it's very urgent to develop efficient computational approaches for predicting PSS based on sequence information alone. Moreover, the feature matrix of a protein contains two dimensions: the amino-acid residue dimension and the feature vector dimension. Existing deep learning based methods have achieved remarkable performances of PSS prediction, but the methods often utilize the features from the amino-acid dimension. Thus, there is still room to improve computational methods of PSS prediction. RESULTS: We propose a novel deep neural network method, called DeepACLSTM, to predict 8-category PSS from protein sequence features and profile features. Our method efficiently applies asymmetric convolutional neural networks (ACNNs) combined with bidirectional long short-term memory (BLSTM) neural networks to predict PSS, leveraging the feature vector dimension of the protein feature matrix. In DeepACLSTM, the ACNNs extract the complex local contexts of amino-acids; the BLSTM neural networks capture the long-distance interdependencies between amino-acids. Furthermore, the prediction module predicts the category of each amino-acid residue based on both local contexts and long-distance interdependencies. To evaluate performances of DeepACLSTM, we conduct experiments on three publicly available datasets: CB513, CASP10 and CASP12. Results indicate that the performance of our method is superior to the state-of-the-art baselines on three publicly datasets. CONCLUSIONS: Experiments demonstrate that DeepACLSTM is an efficient predication method for predicting 8-category PSS and has the ability to extract more complex sequence-structure relationships between amino-acid residues. Moreover, experiments also indicate the feature vector dimension contains the useful information for improving PSS prediction.
Yanbu Guo, Weihua Li 0006, Bingyi Wang, Huiqing Liu, Dongming Zhou 0001
BMC Bioinform.5
2019 Infrared and visible images fusion using visual saliency and optimized spiking cortical model in non-subsampled shearlet transform domain
Ruichao Hou, Rencan Nie, Dongming Zhou 0001, Jinde Cao, Dong Liu 0023
Multim. Tools Appl.3
2019 Multi-focus image fusion combining focus-region-level partition and pulse-coupled neural network
Kangjian He, Dongming Zhou 0001, Xuejie Zhang 0002, Rencan Nie, Xin Jin 0005
Soft Comput.2
2019 FuseGAN: Learning to Fuse Multi-Focus Image via Conditional Generative Adversarial Network
abstract
We study the problem of multi-focus image fusion, where the key challenge is detecting the focused regions accurately among multiple partially focused source images. Inspired by the conditional generative adversarial network (cGAN) to image-to-image task, we propose a novel FuseGAN to fulfill the images-to-image for multi-focus image fusion. To satisfy the requirement of dual input-to-one output, the encoder of the generator in FuseGAN is designed as a Siamese network. The least square GAN objective is employed to enhance the training stability of FuseGAN, resulting in an accurate confidence map for focus region detection. Also, we exploit the convolutional conditional random fields technique on the confidence map to reach a refined final decision map for better focus region detection. Moreover, due to the lack of a large-scale standard dataset, we synthesize a large enough multi-focus image dataset based on a public natural image dataset PASCAL VOC 2012, where we utilize a normalized disk point spread function to simulate the defocus and separate the background and foreground in the synthesis for each image. We conduct extensive experiments on two public datasets to verify the effectiveness of the proposed method. Results demonstrate that the proposed method presents accurate decision maps for focus regions in multi-focus images, such that the fused images are superior to 11 recent state-of-the-art algorithms, not only in visual perception, but also in quantitative analysis in terms of five metrics.
Xiaopeng Guo 0001, Rencan Nie, Jinde Cao, Dongming Zhou 0001, Liye Mei, Kangjian He
IEEE Trans. Multim.4
2018 Multi-focus: Focused region finding and multi-scale transform for image fusion
Kangjian He, Dongming Zhou 0001, Xuejie Zhang 0002, Rencan Nie
Neurocomputing2
2018 A Regularized Locality Projection-Based Sparsity Discriminant Analysis for Face Recognition
abstract
Manifold learning and classifiers based on sparse representation are widely used in pattern recognition. Most of the conventional manifold learning methods are subjected to the choice of parameters. In this paper, we present a Regularized Locality Projection based on Sparsity Discriminant Analysis (RLPSD) method for Feature Extraction (FE) to understand the high-dimensional data such as face images. In RLPSD, firstly, we show the sparse representation of training samples by collaborative representation-based classification (CRC). Secondly, the idea of part optimization based on sparse representation is used to ensure the within-class compactness which combines with the labels of measurements and the weights of sparse presentation can be as small as possible. Finally, whole optimization can be directly obtained without the iteration of local optimization. Meanwhile, the separability information of between-class can be well discriminated by scatter matrix which is similar to Fisher linear discriminant analysis (LDA). The great recognition performance of the proposed method is verified by comparing with the popular algorithms on Yale, ORL, AR and Extended YaleB face databases and Oxford 102 flowers dataset.
Chuanbo Yu, Rencan Nie, Dongming Zhou 0001
Int. J. Pattern Recognit. Artif. Intell.3
2018 A lightweight scheme for multi-focus image fusion
Xin Jin 0005, Jingyu Hou 0001, Rencan Nie, Shaowen Yao 0001, Dongming Zhou 0001, Kangjian He
Multim. Tools Appl.5
2018 Fully Convolutional Network-Based Multifocus Image Fusion
abstract
As the optical lenses for cameras always have limited depth of field, the captured images with the same scene are not all in focus. Multifocus image fusion is an efficient technology that can synthesize an all-in-focus image using several partially focused images. Previous methods have accomplished the fusion task in spatial or transform domains. However, fusion rules are always a problem in most methods. In this letter, from the aspect of focus region detection, we propose a novel multifocus image fusion method based on a fully convolutional network (FCN) learned from synthesized multifocus images. The primary novelty of this method is that the pixel-wise focus regions are detected through a learning FCN, and the entire image, not just the image patches, are exploited to train the FCN. First, we synthesize 4500 pairs of multifocus images by repeatedly using a gaussian filter for each image from PASCAL VOC 2012, to train the FCN. After that, a pair of source images is fed into the trained FCN, and two score maps indicating the focus property are generated. Next, an inversed score map is averaged with another score map to produce an aggregative score map, which take full advantage of focus probabilities in two score maps. We implement the fully connected conditional random field (CRF) on the aggregative score map to accomplish and refine a binary decision map for the fusion task. Finally, we exploit the weighted strategy based on the refined decision map to produce the fused image. To demonstrate the performance of the proposed method, we compare its fused results with several start-of-the-art methods not only on a gray data set but also on a color data set. Experimental results show that the proposed method can achieve superior fusion performance in both human visual quality and objective assessment.
Xiaopeng Guo 0001, Rencan Nie, Jinde Cao, Dongming Zhou 0001, Wenhua Qian
Neural Comput.4
2018 Multimodal sensor medical image fusion based on nonsubsampled shearlet transform and S-PCNNs in HSV space
Xin Jin 0005, Jingyu Hou 0001, Dongming Zhou 0001, Shaowen Yao 0001
Signal Process.5
2018 Multi-focus image fusion method using S-PCNN optimized by particle swarm optimization
Xin Jin 0005, Dongming Zhou 0001, Shaowen Yao 0001, Rencan Nie, Kangjian He
Soft Comput.2
2017 Global asymptotic stability by complex-valued inequalities for complex-valued neural networks with delays on period time scales
Zhengqiu Zhang, Dangli Hao, Dongming Zhou 0001
Neurocomputing3
2014 Novel LMI-Based Condition on Global Asymptotic Stability for a Class of Cohen-Grossberg BAM Networks With Extended Activation Functions
abstract
This paper is concerned with global asymptotic stability of a class of Cohen-Grossberg bidirectional associative memory (BAM) neural networks with delays. Under the assumptions that the activation functions only satisfy the so-called extended global Lipschitz condition and the behaved functions only satisfy global Lipschitz condition, we apply linear matrix inequality (LMI) method and homeomorphism theory to propose a new LMI-based sufficient condition for global asymptotic stability of the concerned neural networks. In our results, the extended global Lipschitz condition on the activation functions is less conservative than the assumptions for boundedness and monotonicity and is weaker than the assumption for the general global Lipschitz condition, the global Lipschitz condition on the behaved functions is also less conservative than the assumptions for monotonicity and differentiability in existing papers.
Zhengqiu Zhang, Jinde Cao, Dongming Zhou 0001
IEEE Trans. Neural Networks Learn. Syst.3
2013 New LMI-based conditions for global exponential stability to a class of Cohen-Grossberg BAM networks with delays
Dongming Zhou 0001, Shenghua Yu, Zhengqiu Zhang
Neurocomputing1
2012 Global asymptotic stability to a generalized Cohen-Grossberg BAM neural networks of neutral type delays
Zhengqiu Zhang, Dongming Zhou 0001
Neural Networks3
2009 Global robust exponential stability for second-order Cohen-Grossberg neural networks with multiple delays
Zhengqiu Zhang, Dongming Zhou 0001
Neurocomputing2
2009 Analysis of autowave characteristics for competitive pulse coupled neural network and its application
Dongming Zhou 0001, Rencan Nie, Dongfeng Zhao
Neurocomputing1
2008 A New Algorithm for Finding the Shortest Path Tree Using Competitive Pulse Coupled Neural Network
Dongming Zhou 0001, Rencan Nie, Dongfeng Zhao
ICIC (2)1
1998 Stability analysis of delayed cellular neural networks
Jinde Cao, Dongming Zhou 0001
Neural Networks2