Jing Zhang 0041

dblp:05/3499-41 · DBLP profile ↗
← Back
45ranked-venue papers
31as first author
28since 2021 · last 2026
0000-0001-6270-7771ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 19 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 12 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 Anomaly-aware mutual promotion network for medical visual question answering
Jiong Teng, Li Xi, Feihong Luo, Jing Zhang 0041
Knowl. Based Syst.4
2026 Improving episodic few-shot visual question answering via spatial and frequency domain dual-calibration
Jing Zhang 0041, Yunzuo Hu, Zhe Wang 0002
Pattern Recognit.1
2026 Fine-Grained Emotion Adaptive Alignment Network for Image Emotion Distribution Transfer
Jing Zhang 0041, Jixiang Zhu, Yumo Kang, Dongdong Li 0003, Zhe Wang 0002
IEEE Trans. Affect. Comput.1
2025 Subtype-Former: A Deep Learning Approach for Cancer Subtype Discovery with Multi-Omics Data
abstract
Cancer is heterogeneous, affecting the precise approach to personalized treatment. Accurate subtyping can lead to better survival rates for cancer patients. High-throughput technologies provide multiple omics data for cancer subtyping. This study proposed Subtype-Former, a deep learning method based on MLP and Transformer Block, to extract the lowdimensional representation of the multi-omics data. K-means and Consensus Clustering are also used to achieve accurate subtyping results. We compared Subtype-Former with the other state-of-the-art subtyping methods across the TCGA 10 cancer types. We found that Subtype-Former can perform better on the benchmark datasets of more than 5000 tumors based on the survival analysis. In addition, Subtype-Former also achieved outstanding results in pan-cancer subtyping, which can help analyze the commonalities and differences across various cancer types at the molecular level. Finally, we applied Subtype-Former to the TCGA 10 types of cancers. We identified 50 essential biomarkers, which can be used to study targeted cancer drugs and promote the development of cancer treatments in the era of precision medicine.
Hanwen Huang, Yuhang Sheng, Dongdong Li 0003, Jing Zhang 0041, Hai Yang 0002
BIBM5
2025 Cross-modal heterogeneous graph reasoning network for visual question answering
Jing Zhang 0041, Jiong Teng, Weichao Ding, Zhe Wang 0002
Neural Comput. Appl.1
2025 Saccade and purify: Task adapted multi-view feature calibration network for few shot learning
Jing Zhang 0041, Yunzuo Hu, Xinzhou Zhang, Mingzhe Chen, Zhe Wang 0002
Neural Networks1
2025 Trans-Driver: A Deep Learning Approach for Cancer Driver Gene Discovery With Multi-Omics Data
abstract
Driver genes play a crucial role in the growth of cancer cells. Accurate identification of cancer driver genes is essential for deepening our understanding of cancer pathogenesis and facilitating the development of cancer therapies and drug-targeted driver genes. However, the diversity and complexity of multi-omics data still make cancer driver identification highly challenging. In this study, we propose Transformer-Driver (Trans-Driver), a deep supervised learning method based on a novel transformer architecture, which integrates multi-omics data to learn the differences and associations between different omics modalities for cancer driver discovery. Trans-Driver introduces a kernel-based multi-head self-attention mechanism with gated residual connections, as well as a Dynamic Tanh (DyT) normalization function, to enhance the integration and modeling of heterogeneous multi-omics features. Compared with other state-of-the-art driver gene identification methods, Trans-Driver achieved excellent performance on TCGA, CGC, and PCAWG datasets. Among approximately 20,000 protein-coding genes, Trans-Driver reported 269 candidate driver genes, of which 132 genes (about 49.1%) were included in the gold standard CGC dataset. Feature contribution analysis further demonstrated that integrating multi-omics data improved performance compared to using somatic mutation data alone. Finally, detailed analysis revealed that the candidate drivers are clinically meaningful, demonstrating the practical value of Trans-Driver.
Hai Yang 0002, Zhenbei Yang, Lei Zhang 0224, Yijing Yang, Dongdong Li 0003, Jing Zhang 0041, Zhe Wang 0002
IEEE Trans. Comput. Biol. Bioinform.7
2025 Deep Reciprocal Learning for Image Captioning
abstract
The current training strategies based on knowledge distillation for image captioning assume that each learning model possesses complete learning value, lacking review and guidance mechanisms among the interactive process of models. To address this problem, we propose a novel Captioner with Deep Reciprocal Learning (CaDReL) for image captioning inspired by the social learning theory, which realizes interactive learning between models controlled by salient semantic evaluation. In CaDReL, we analyze the semantic saliency of each learning network to better control the parameter transfer in knowledge distillation by cyclically alternately freezing and unfreezing two learning networks with identical review mechanisms. We also propose a novel cascade bridging diffusion module, which fuses feature information from different levels of visual information and attention ranges in the encoder by a cascade diffusion mechanism to capture rich image details and contextual information. Meanwhile, an attention guided knowledge augmentation module is proposed to guide knowledge transferring by the attention maps from the respective encoders of two peer joint learning modules for improving the robustness of the whole training strategy. Experimental results illustrated that the proposed CaDReL achieves excellent performance on the MSCOCO dataset, and outperforms most state-of-the-art methods. Codes are available athttps://github.com/ZJ-VIP-Lab/Deep-Reciprocal-Learning-for-Image-Captioning.
Jing Zhang 0041, Yingshuai Xie, Zhe Wang 0002
IEEE Trans. Circuits Syst. Video Technol.1
2025 Scale-Wise Semantic Alignment Enhanced Multigrained Adaptive Fusion for Virtual Try-On
abstract
Image-based virtual try-on aims to fit garments onto a target person accurately and naturally while preserving the textural details of the garment. Inspired by the dynamic perception process of the human visual system, which transitions from global perception to local details, we propose a novel multigrained adaptive fusion network for virtual try-on framework named MA-VITON. MA-VITON precisely aligns clothing semantic features with human body parts across different scales, reduces unrealistic textures caused by garment distortion, and employs coarse-to-fine clothing features to progressively guide the generation of try-on results. To achieve this, we introduce a scale-wise semantic alignment (SSA) module that extracts local features of clothing and the target person at various scales using flexible query strategies. It learns semantic correspondences between garments and the human body in the latent space through parallel bidirectional interactions, ensuring accurate feature alignment. Additionally, we propose a multigrained adaptive fusion (MAF) module, which identifies critical garment regions using a polyscale attention mechanism and allocates more tokens to adaptively preserve intricate textural details. Extensive experiments on multiple widely used public datasets demonstrate that MA-VITON achieves outstanding performance and surpasses state-of-the-art methods. The code is publicly available at https://github.com/Max-Teapot/MA-VITON.
Jing Zhang 0041, Yumo Kang, Zhe Wang 0002
IEEE Trans. Neural Networks Learn. Syst.1
2025 Latent Attention Network With Position Perception for Visual Question Answering
abstract
For exploring the complex relative position relationships among multiobject with multiple position prepositions in the question, we propose a novel latent attention (LA) network for visual question answering (VQA), in which LA with position perception is extracted by a novel LA generation module (LAGM) and encoded along with absolute and relative position relations by our proposed position-aware module (PAM). The LAGM reconstructs original attention into LA by capturing the tendency of visual attention shifting according to the position prepositions in the question. The LA accurately captures the complex relative position features of multiple objects and helps the model locate the attention to the correct object or region. The PAM adopts latent state and relative position relations to enhance the capability of comprehending the multiobject correlations. In addition, we also propose a novel gated counting module (GCM) to strengthen the sensitivity of quantitative knowledge for effectively improving the performance of counting questions. Extensive experiments demonstrate that our proposed method achieves excellent performance on VQA and outperforms state-of-the-art methods on the widely used datasets VQA v2 and VQA v1.
Jing Zhang 0041, Zhe Wang 0002
IEEE Trans. Neural Networks Learn. Syst.1
2024 Cross-Modal Feature Distribution Calibration for Few-Shot Visual Question Answering
abstract
Few-shot Visual Question Answering (VQA) realizes few-shot cross-modal learning, which is an emerging and challenging task in computer vision. Currently, most of the few-shot VQA methods are confined to simply extending few-shot classification methods to cross-modal tasks while ignoring the spatial distribution properties of multimodal features and cross-modal information interaction. To address this problem, we propose a novel Cross-modal feature Distribution Calibration Inference Network (CDCIN) in this paper, where a new concept named visual information entropy is proposed to realize multimodal features distribution calibration by cross-modal information interaction for more effective few-shot VQA. Visual information entropy is a statistical variable that represents the spatial distribution of visual features guided by the question, which is aligned before and after the reasoning process to mitigate redundant information and improve multi-modal features by our proposed visual information entropy calibration module. To further enhance the inference ability of cross-modal features, we additionally propose a novel pre-training method, where the reasoning sub-network of CDCIN is pretrained on the base class in a VQA classification paradigm and fine-tuned on the few-shot VQA datasets. Extensive experiments demonstrate that our proposed CDCIN achieves excellent performance on few-shot VQA and outperforms state-of-the-art methods on three widely used benchmark datasets.
Jing Zhang 0041, Mingzhe Chen, Zhe Wang 0002
AAAI1
2024 Object aroused emotion analysis network for image sentiment analysis
Jing Zhang 0041, Jiangpei Liu, Weichao Ding, Zhe Wang 0002
Knowl. Based Syst.1
2024 MCPL: Multi-model co-guided progressive learning for multimodal aspect-based sentiment analysis
Jing Zhang 0041, Jiaqi Qu, Jiangpei Liu, Zhe Wang 0002
Knowl. Based Syst.1
2024 BMPCN: A Bigraph Mutual Prototype Calibration Net for few-shot classification
Jing Zhang 0041, Mingzhe Chen, Yunzuo Hu, Xinzhou Zhang, Zhe Wang 0002
Pattern Recognit.1
2024 Adaptive Semantic-Enhanced Transformer for Image Captioning
abstract
In the research on image captioning, rich semantic information is very important for generating critical caption words as guiding information. However, semantic information from offline object detectors involves many semantic objects that do not appear in the caption, thereby bringing noise into the decoding process. To produce more accurate semantic guiding information and further optimize the decoding process, we propose an end-to-end adaptive semantic-enhanced transformer (AS-Transformer) model for image captioning. For semantic enhancement information extraction, we propose a constrained weaklysupervised learning (CWSL) module, which reconstructs the semantic object's probability distribution detected by the multiple instances learning (MIL) through a joint loss function. These strengthened semantic objects from the reconstructed probability distribution can better depict the semantic meaning of images. Also, for semantic enhancement decoding, we propose an adaptive gated mechanism (AGM) module to adjust the attention between visual and semantic information adaptively for the more accurate generation of caption words. Through the joint control of the CWSL module and AGM module, our proposed model constructs a complete adaptive enhancement mechanism from encoding to decoding and obtains visual context that is more suitable for captions. Experiments on the public Microsoft Common Objects in COntext (MSCOCO) and Flickr30K datasets illustrate that our proposed AS-Transformer can adaptively obtain effective semantic information and adjust the attention weights between semantic and visual information automatically, which achieves more accurate captions compared with semantic enhancement methods and outperforms state-of-the-art methods.
Jing Zhang 0041, Zhongjun Fang, Zhe Wang 0002
IEEE Trans. Neural Networks Learn. Syst.1
2024 Emotion-wise feature interaction analysis-based visual emotion distribution learning
Jing Zhang 0041, Qiuge Qin, Qi Ye 0004, Wen Du
Vis. Comput.1
2023 Improving Image Captioning through Visual and Semantic Mutual Promotion
abstract
Current image captioning methods commonly use semantic attributes extracted by an object detector to guide visual representation, leaving the mutual guidance and enhancement between vision and semantics under-explored. Neurological studies have revealed that the visual cortex of the brain plays a crucial role in recognizing visual objects, while the prefrontal cortex is involved in the integration of contextual semantics. Inspired by the above studies, we propose a novel Visual-Semantic Transformer (VST) to model the neural interaction between vision and semantics, which explores the mechanism of deep fusion and mutual promotion of multimodal information, realizing more accurate image captioning. To better facilitate the complementary strengths between visual objects and semantic contexts, we propose a global position-sensitive co-attention encoder to realize globally associative, position-aware visual and semantic co-interaction through a mutual cross-attention mechanism. In addition, a multimodal mixed attention module is proposed in the decoder, which achieves adaptive multimodal feature fusion for enhancing the decoding capability. Experimental evidence shows that our VST significantly surpasses the state-of-the-art approaches on MSCOCO dataset and reaches the excellent CIDEr score of 142% on the Karpathy test split.
Jing Zhang 0041, Yingshuai Xie
ACM Multimedia1
2023 Multi-feature fusion enhanced transformer with multi-layer fused decoding for image captioning
Jing Zhang 0041, Zhongjun Fang, Zhe Wang 0002
Appl. Intell.1
2023 From multi-omics data to the cancer druggable gene discovery: a novel machine learning-based approach
abstract
The development of targeted drugs allows precision medicine in cancer treatment and optimal targeted therapies. Accurate identification of cancer druggable genes helps strengthen the understanding of targeted cancer therapy and promotes precise cancer treatment. However, rare cancer-druggable genes have been found due to the multi-omics data's diversity and complexity. This study proposes deep forest for cancer druggable genes discovery (DF-CAGE), a novel machine learning-based method for cancer-druggable gene discovery. DF-CAGE integrated the somatic mutations, copy number variants, DNA methylation and RNA-Seq data across ˜10 000 TCGA profiles to identify the landscape of the cancer-druggable genes. We found that DF-CAGE discovers the commonalities of currently known cancer-druggable genes from the perspective of multi-omics data and achieved excellent performance on OncoKB, Target and Drugbank data sets. Among the ˜20 000 protein-coding genes, DF-CAGE pinpointed 465 potential cancer-druggable genes. We found that the candidate cancer druggable genes (CDG) are clinically meaningful and divided the CDG into known, reliable and potential gene sets. Finally, we analyzed the omics data's contribution to identifying druggable genes. We found that DF-CAGE reports druggable genes mainly based on the copy number variations (CNVs) data, the gene rearrangements and the mutation rates in the population. These findings may enlighten the future study and development of new drugs.
Hai Yang 0002, Lipeng Gan, Rui Chen 0021, Dongdong Li 0003, Jing Zhang 0041, Zhe Wang 0002
Briefings Bioinform.5
2023 Hierarchical decoding with latent context for image captioning
Jing Zhang 0041, Yingshuai Xie, Kangkang Li 0003, Zhe Wang 0002, Wen Du
Neural Comput. Appl.1
2023 Cross on Cross Attention: Deep Fusion Transformer for Image Captioning
abstract
Numerous studies have shown that in-depth mining of correlations between multi-modal features can help improve the accuracy of cross-modal data analysis tasks. However, the current image description methods based on the encoder-decoder framework only carry out the interaction and fusion of multi-modal features in the encoding stage or the decoding stage, which cannot effectively alleviate the semantic gap. In this paper, we propose a Deep Fusion Transformer (DFT) for image captioning to provide a deep multi-feature and multi-modal information fusion strategy throughout the encoding to decoding process. We propose a novel global cross encoder to align different types of visual features, which can effectively compensate for the differences between features and incorporate each other’s strengths. In the decoder, a novel cross on cross attention is proposed to realize hierarchical cross-modal data analysis, extending complex cross-modal reasoning capabilities through the multi-level interaction of visual and semantic features. Extensive experiments conducted on the MSCOCO dataset prove that our proposed DFT can achieve excellent performance and outperform state-of-the-art methods. The code is available athttps://github.com/weimingboya/DFT.
Jing Zhang 0041, Yingshuai Xie, Weichao Ding, Zhe Wang 0002
IEEE Trans. Circuits Syst. Video Technol.1
2022 Image sentiment classification via multi-level sentiment region correlation analysis
Jing Zhang 0041, Qi Ye 0004, Zhe Wang 0002
Neurocomputing1
2022 Geometric imbalanced deep learning with feature scaling and boundary sample mining
Zhe Wang 0002, Qida Dong, Wei Guo 0023, Dongdong Li 0003, Jing Zhang 0041, Wenli Du
Pattern Recognit.5
2022 Graph-Based Object Semantic Refinement for Visual Emotion Recognition
abstract
The rich semantic information contained in images is an important clue to explore visual emotions. Therefore, exploring the correlation between visual emotion and the semantic relationship of objects, and extracting more effective semantic features through explicit or implicit modeling is very important for visual emotion analysis. In this paper, a novel Graph-based Object Semantic Refinement (GOSR) model is proposed to extract multi-level semantic features for visual emotion classification, in which graph structures is used to represent the object semantics and their position relationships of an image, and Graph Convolutional Networks (GCN) is used to refine object information by the aggregating neighbor object with their position relationships. The different convolutional layer’s features from GCN are further fused by Gated Recurrent Units (GRU) networks to achieve high-level semantic features. Then a framework with two branches to leverage visual and semantic information for visual sentiment analysis is proposed, which uses convolutional neural networks to extract visual features from images, and collaborates with semantic features from GOSR model to achieve better emotion recognition results. Besides, for alleviating the potentially unreasonable predictions and promote models collaboration, a novel tendency loss function based on the correlations among emotion labels is proposed to adjust the output activation value other than the target label. Extensive experiments on four widely used benchmark datasets show that our proposed method can achieve competitive performance and outperform most of the state-of-the-art methods on visual emotion recognition.
Jing Zhang 0041, Zhe Wang 0002, Hai Yang 0002
IEEE Trans. Circuits Syst. Video Technol.1
2022 Task Encoding With Distribution Calibration for Few-Shot Learning
abstract
Few-shot learning is an extremely challenging task in computer vision that has attracted increased research attention in recent years. However, most recent methods do not fully use the task’s information, and few of the seen samples result in large intraclass differences among the same classes. In this paper, we propose a novel task encoding with distribution calibration (TEDC) model for few-shot learning, which uses the relationships among the feature distributions to reduce intraclass differences. In the TEDC model, an integrated feature extraction module (IFEM) is proposed, which extracts the multiangle visual features of an image and fuses them to obtain more representative features. To effectively utilize the task information, a novel task encoding module (TEM) is proposed, which obtains the task features by fusing all the seen samples’ information and uses them to adjust all the samples’ features for more generalizable task-specific representations. We also propose a distribution calibration module (DCM) to reduce the bias between the distribution of the support features and the query features in the same class. Extensive experiments show that our proposed TEDC model achieves an excellent performance and outperforms the state-of-the-art methods on three widely used few-shot classification benchmarks, specifically miniImageNet, tieredImageNet and CUB-200-2011.
Jing Zhang 0041, Xinzhou Zhang, Zhe Wang 0002
IEEE Trans. Circuits Syst. Video Technol.1
2021 Visual enhanced gLSTM for image captioning
Jing Zhang 0041, Kangkang Li 0003, Zhenkun Wang 0005, Xianwen Zhao, Zhe Wang 0002
Expert Syst. Appl.1
2021 Parallel-fusion LSTM with synchronous semantic and visual information for image captioning
Jing Zhang 0041, Kangkang Li 0003, Zhe Wang 0002
J. Vis. Commun. Image Represent.1
2021 BLSTM and CNN Stacking Architecture for Speech Emotion Recognition
Dongdong Li 0003, Linyu Sun, Xinlei Xu, Zhe Wang 0002, Jing Zhang 0041, Wenli Du
Neural Process. Lett.5
2020 Multiple Universum Empirical Kernel Learning
Zhe Wang 0002, Sisi Hong, Lijuan Yao, Dongdong Li 0003, Wenli Du, Jing Zhang 0041
Eng. Appl. Artif. Intell.6
2020 Object semantics sentiment correlation analysis enhanced image sentiment classification
Jing Zhang 0041, Dongdong Li 0003, Zhe Wang 0002
Knowl. Based Syst.1
2020 Multi-matrices entropy discriminant ensemble learning for imbalanced problem
Zhe Wang 0002, Zhaozhi Chen, Jing Zhang 0041, Wenli Du, Dongdong Li 0003
Neural Comput. Appl.4
2020 Efficient matrixized classification learning with separated solution process
Zonghai Zhu, Zhe Wang 0002, Dongdong Li 0003, Wenli Du, Jing Zhang 0041
Neural Comput. Appl.5
2019 Another Dimension: Towards Multi-subnet Neural Network for Image Sentiment Analysis
abstract
Image sentiment analysis has been studied for many years, and most of algorithms take the image sentiment as independent and discrete labels to predict by machine learning. Actually, as a product of multiple hormone combinations, emotions are generated by mutual suppression signal in brain. Inspired by neural microcircuit in amygdala, we propose a novel Multi-Subnet Neural Network (MSNN) that simulates the human brain mechanism for image sentiment classification. Different from traditional neural network, MSNN extends a new domain channel to imitate the way that images stimulate the brain through different neural circuits and produce sentimental semantic information by multi-subnet and signal reforming network. Experiments show that MSNN is well adapted to multi-class image sentiment classification task, and outperforms other multi-class sentiment classification models.
Jing Zhang 0041, Zhe Wang 0002, Tong Ruan
ICME1
2019 Multi-view learning with fisher kernel and bi-bagging for imbalanced problem
Zhe Wang 0002, Zhaozhi Chen, Jing Zhang 0041, Wenli Du
Appl. Intell.4
2019 Web image annotation based on Tri-relational Graph and semantic context analysis
Jing Zhang 0041, Ti Tao, Yakun Mu, Dongdong Li 0003, Zhe Wang 0002
Eng. Appl. Artif. Intell.1
2019 Cost-sensitive Fuzzy Multiple Kernel Learning for imbalanced problem
Zhe Wang 0002, Bolu Wang, Dongdong Li 0003, Jing Zhang 0041
Neurocomputing5
2019 Image region label refinement using spatial position relation graph
Jing Zhang 0041, Zhenkun Wang 0005, Yakun Mu, Zhe Wang 0002
Knowl. Based Syst.1
2018 Image region annotation based on segmentation and semantic correlation analysis
abstract
The authors propose an image region annotation framework by exploring syntactic and semantic correlations among segmented regions in an image. A texture‐enhanced image segmentation JSEG algorithm is first used to improve the pixel consistency in a segmented image region. Next, each region is represented by a set of image codewords, also known as visual alphabets, with each of them used to characterise certain low‐level image features. A visual lexicon, with its vocabulary items defined as either a codeword or a co‐occurrence of multiple alphabets, is formed and used to model middle‐level semantic concepts. The concept classification models are trained by a maximal figure‐of‐merit algorithm with a collection of training images with multiple correlations, including spatial, syntactic and semantic relationship, between regions and their corresponding concepts. In addition, a region‐semantic correlation model constructed with latent semantic analysis is used to correct the potentially wrong annotations by analysing the relationship between image region positions and labels. When evaluated on the Corel 5K dataset, the proposed image region annotation framework achieves accurate results on image region concept tagging as well as whole image based annotations.
Jing Zhang 0041, Yakun Mu, Sheng-Wei Feng, Kehuang Li, Yubo Yuan 0001
IET Image Process.1
2017 Image retrieval using the extended salient region
Jing Zhang 0041, Sheng-Wei Feng, Yong-Wei Gao, Yubo Yuan 0001
Inf. Sci.1
2016 Automatic image region annotation through segmentation based visual semantic analysis and discriminative classification
abstract
We propose a new framework for automatic image annotation (AIA) of regions through segmentation based semantic analysis and discriminative classification. Given a test image, it is first segmented by a proposed texture-enhanced JSEG algorithm. Then these regions are represented by an extended bag-of-words model in which a feature vector, based on a visual lexicon with its vocabulary consisting of a visual word or a co-occurrence of multiple visual words, is constructed to represent the region content. Finally a concept classifier learned by a maximal figure-of-merit algorithm is used to predict the region labels. These models are discriminatively trained from image regions with multiple associations between regions and concepts. Experiments on a subset of the Corel 5K data set illustrate that our proposed approach to region AIA achieves more accurate annotation results than some sate-of-the-art algorithms.
Jing Zhang 0041, Yong-Wei Gao, Sheng-Wei Feng, Yubo Yuan 0001
ICASSP1
2016 Structure-aware image inpainting using patch scale optimization
Chao Dai, Bin Sheng 0001, Jing Zhang 0041, Weiyao Lin, Yubo Yuan 0001
J. Vis. Commun. Image Represent.5
2016 Image saliency detection using Gabor texture cues
Bin Sheng 0001, Jian-ning Liang, Jing Zhang 0041, Yubo Yuan 0001
Multim. Tools Appl.5
2015 Moving visual focus in salient object segmentation
abstract
Saliency detection plays an important role in image segmentation, object detection and retrieval, which attracts more attention in the field of computer vision recently. Most existing saliency detection algorithms have not considered the influence of visual focus (VF) shifting yet. In this study, a novel algorithm named moving region contrast (MRC) is proposed to analyse image saliency. The algorithm MRC is built on a novel concept of moving VF. The initial VF is defined as the geometric centre of the image. Then the VF is calculated iteratively by focus‐moving technique where a saliency gravitation model is employed to determine the moving direction. The salient region is obtained according to the final VF. The experiments are conducted on the dataset with 1000 images released by Achanta. Experimental results show that the proposed algorithm achieves marked improvements in performance and outperforms other 11 popular algorithms.
Xiao-Long Xiao, Fangli Ying, Jing Zhang 0041, Yubo Yuan 0001
IET Image Process.5
2015 Representation of image content based on RoI-BoW
Jing Zhang 0041, Ya-Xin Zhao, Yubo Yuan 0001
J. Vis. Commun. Image Represent.1
2015 A novel image annotation model based on content representation with multi-layer segmentation
Jing Zhang 0041, Ya-Xin Zhao, Yubo Yuan 0001
Neural Comput. Appl.1