EDBT 2026 Demo / reviewers in the wild / expert
Jing Zhang 0041
dblp:05/3499-41
· DBLP profile ↗
45ranked-venue papers
31as first author
28since 2021 · last 2026
0000-0001-6270-7771ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 19 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 12 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Anomaly-aware mutual promotion network for medical visual question answering
Jiong Teng, Li Xi, Feihong Luo, Jing Zhang 0041 |
Knowl. Based Syst. | 4 |
| 2026 | Improving episodic few-shot visual question answering via spatial and frequency domain dual-calibration
Jing Zhang 0041, Yunzuo Hu, Zhe Wang 0002 |
Pattern Recognit. | 1 |
| 2026 | Fine-Grained Emotion Adaptive Alignment Network for Image Emotion Distribution Transfer
Jing Zhang 0041, Jixiang Zhu, Yumo Kang, Dongdong Li 0003, Zhe Wang 0002 |
IEEE Trans. Affect. Comput. | 1 |
| 2025 | Subtype-Former: A Deep Learning Approach for Cancer Subtype Discovery with Multi-Omics DataabstractCancer is heterogeneous, affecting the precise approach to personalized treatment. Accurate subtyping can lead to better survival rates for cancer patients. High-throughput technologies provide multiple omics data for cancer subtyping. This study proposed Subtype-Former, a deep learning method based on MLP and Transformer Block, to extract the lowdimensional representation of the multi-omics data. K-means and Consensus Clustering are also used to achieve accurate subtyping results. We compared Subtype-Former with the other state-of-the-art subtyping methods across the TCGA 10 cancer types. We found that Subtype-Former can perform better on the benchmark datasets of more than 5000 tumors based on the survival analysis. In addition, Subtype-Former also achieved outstanding results in pan-cancer subtyping, which can help analyze the commonalities and differences across various cancer types at the molecular level. Finally, we applied Subtype-Former to the TCGA 10 types of cancers. We identified 50 essential biomarkers, which can be used to study targeted cancer drugs and promote the development of cancer treatments in the era of precision medicine. Hanwen Huang, Yuhang Sheng, Dongdong Li 0003, Jing Zhang 0041, Hai Yang 0002 |
BIBM | 5 |
| 2025 | Cross-modal heterogeneous graph reasoning network for visual question answering
Jing Zhang 0041, Jiong Teng, Weichao Ding, Zhe Wang 0002 |
Neural Comput. Appl. | 1 |
| 2025 | Saccade and purify: Task adapted multi-view feature calibration network for few shot learning
Jing Zhang 0041, Yunzuo Hu, Xinzhou Zhang, Mingzhe Chen, Zhe Wang 0002 |
Neural Networks | 1 |
| 2025 | Trans-Driver: A Deep Learning Approach for Cancer Driver Gene Discovery With Multi-Omics DataabstractDriver genes play a crucial role in the growth of cancer cells. Accurate identification of cancer driver genes is essential for deepening our understanding of cancer pathogenesis and facilitating the development of cancer therapies and drug-targeted driver genes. However, the diversity and complexity of multi-omics data still make cancer driver identification highly challenging. In this study, we propose Transformer-Driver (Trans-Driver), a deep supervised learning method based on a novel transformer architecture, which integrates multi-omics data to learn the differences and associations between different omics modalities for cancer driver discovery. Trans-Driver introduces a kernel-based multi-head self-attention mechanism with gated residual connections, as well as a Dynamic Tanh (DyT) normalization function, to enhance the integration and modeling of heterogeneous multi-omics features. Compared with other state-of-the-art driver gene identification methods, Trans-Driver achieved excellent performance on TCGA, CGC, and PCAWG datasets. Among approximately 20,000 protein-coding genes, Trans-Driver reported 269 candidate driver genes, of which 132 genes (about 49.1%) were included in the gold standard CGC dataset. Feature contribution analysis further demonstrated that integrating multi-omics data improved performance compared to using somatic mutation data alone. Finally, detailed analysis revealed that the candidate drivers are clinically meaningful, demonstrating the practical value of Trans-Driver. Hai Yang 0002, Zhenbei Yang, Lei Zhang 0224, Yijing Yang, Dongdong Li 0003, Jing Zhang 0041, Zhe Wang 0002 |
IEEE Trans. Comput. Biol. Bioinform. | 7 |
| 2025 | Deep Reciprocal Learning for Image CaptioningabstractThe current training strategies based on knowledge distillation for image captioning assume that each learning model possesses complete learning value, lacking review and guidance mechanisms among the interactive process of models. To address this problem, we propose a novel Captioner with Deep Reciprocal Learning (CaDReL) for image captioning inspired by the social learning theory, which realizes interactive learning between models controlled by salient semantic evaluation. In CaDReL, we analyze the semantic saliency of each learning network to better control the parameter transfer in knowledge distillation by cyclically alternately freezing and unfreezing two learning networks with identical review mechanisms. We also propose a novel cascade bridging diffusion module, which fuses feature information from different levels of visual information and attention ranges in the encoder by a cascade diffusion mechanism to capture rich image details and contextual information. Meanwhile, an attention guided knowledge augmentation module is proposed to guide knowledge transferring by the attention maps from the respective encoders of two peer joint learning modules for improving the robustness of the whole training strategy. Experimental results illustrated that the proposed CaDReL achieves excellent performance on the MSCOCO dataset, and outperforms most state-of-the-art methods. Codes are available athttps://github.com/ZJ-VIP-Lab/Deep-Reciprocal-Learning-for-Image-Captioning. Jing Zhang 0041, Yingshuai Xie, Zhe Wang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Scale-Wise Semantic Alignment Enhanced Multigrained Adaptive Fusion for Virtual Try-OnabstractImage-based virtual try-on aims to fit garments onto a target person accurately and naturally while preserving the textural details of the garment. Inspired by the dynamic perception process of the human visual system, which transitions from global perception to local details, we propose a novel multigrained adaptive fusion network for virtual try-on framework named MA-VITON. MA-VITON precisely aligns clothing semantic features with human body parts across different scales, reduces unrealistic textures caused by garment distortion, and employs coarse-to-fine clothing features to progressively guide the generation of try-on results. To achieve this, we introduce a scale-wise semantic alignment (SSA) module that extracts local features of clothing and the target person at various scales using flexible query strategies. It learns semantic correspondences between garments and the human body in the latent space through parallel bidirectional interactions, ensuring accurate feature alignment. Additionally, we propose a multigrained adaptive fusion (MAF) module, which identifies critical garment regions using a polyscale attention mechanism and allocates more tokens to adaptively preserve intricate textural details. Extensive experiments on multiple widely used public datasets demonstrate that MA-VITON achieves outstanding performance and surpasses state-of-the-art methods. The code is publicly available at https://github.com/Max-Teapot/MA-VITON. Jing Zhang 0041, Yumo Kang, Zhe Wang 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Latent Attention Network With Position Perception for Visual Question AnsweringabstractFor exploring the complex relative position relationships among multiobject with multiple position prepositions in the question, we propose a novel latent attention (LA) network for visual question answering (VQA), in which LA with position perception is extracted by a novel LA generation module (LAGM) and encoded along with absolute and relative position relations by our proposed position-aware module (PAM). The LAGM reconstructs original attention into LA by capturing the tendency of visual attention shifting according to the position prepositions in the question. The LA accurately captures the complex relative position features of multiple objects and helps the model locate the attention to the correct object or region. The PAM adopts latent state and relative position relations to enhance the capability of comprehending the multiobject correlations. In addition, we also propose a novel gated counting module (GCM) to strengthen the sensitivity of quantitative knowledge for effectively improving the performance of counting questions. Extensive experiments demonstrate that our proposed method achieves excellent performance on VQA and outperforms state-of-the-art methods on the widely used datasets VQA v2 and VQA v1. Jing Zhang 0041, Zhe Wang 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Cross-Modal Feature Distribution Calibration for Few-Shot Visual Question AnsweringabstractFew-shot Visual Question Answering (VQA) realizes few-shot cross-modal learning, which is an emerging and challenging task in computer vision. Currently, most of the few-shot VQA methods are confined to simply extending few-shot classification methods to cross-modal tasks while ignoring the spatial distribution properties of multimodal features and cross-modal information interaction. To address this problem, we propose a novel Cross-modal feature Distribution Calibration Inference Network (CDCIN) in this paper, where a new concept named visual information entropy is proposed to realize multimodal features distribution calibration by cross-modal information interaction for more effective few-shot VQA. Visual information entropy is a statistical variable that represents the spatial distribution of visual features guided by the question, which is aligned before and after the reasoning process to mitigate redundant information and improve multi-modal features by our proposed visual information entropy calibration module. To further enhance the inference ability of cross-modal features, we additionally propose a novel pre-training method, where the reasoning sub-network of CDCIN is pretrained on the base class in a VQA classification paradigm and fine-tuned on the few-shot VQA datasets. Extensive experiments demonstrate that our proposed CDCIN achieves excellent performance on few-shot VQA and outperforms state-of-the-art methods on three widely used benchmark datasets. Jing Zhang 0041, Mingzhe Chen, Zhe Wang 0002 |
AAAI | 1 |
| 2024 | Object aroused emotion analysis network for image sentiment analysis
Jing Zhang 0041, Jiangpei Liu, Weichao Ding, Zhe Wang 0002 |
Knowl. Based Syst. | 1 |
| 2024 | MCPL: Multi-model co-guided progressive learning for multimodal aspect-based sentiment analysis
Jing Zhang 0041, Jiaqi Qu, Jiangpei Liu, Zhe Wang 0002 |
Knowl. Based Syst. | 1 |
| 2024 | BMPCN: A Bigraph Mutual Prototype Calibration Net for few-shot classification
Jing Zhang 0041, Mingzhe Chen, Yunzuo Hu, Xinzhou Zhang, Zhe Wang 0002 |
Pattern Recognit. | 1 |
| 2024 | Adaptive Semantic-Enhanced Transformer for Image CaptioningabstractIn the research on image captioning, rich semantic information is very important for generating critical caption words as guiding information. However, semantic information from offline object detectors involves many semantic objects that do not appear in the caption, thereby bringing noise into the decoding process. To produce more accurate semantic guiding information and further optimize the decoding process, we propose an end-to-end adaptive semantic-enhanced transformer (AS-Transformer) model for image captioning. For semantic enhancement information extraction, we propose a constrained weaklysupervised learning (CWSL) module, which reconstructs the semantic object's probability distribution detected by the multiple instances learning (MIL) through a joint loss function. These strengthened semantic objects from the reconstructed probability distribution can better depict the semantic meaning of images. Also, for semantic enhancement decoding, we propose an adaptive gated mechanism (AGM) module to adjust the attention between visual and semantic information adaptively for the more accurate generation of caption words. Through the joint control of the CWSL module and AGM module, our proposed model constructs a complete adaptive enhancement mechanism from encoding to decoding and obtains visual context that is more suitable for captions. Experiments on the public Microsoft Common Objects in COntext (MSCOCO) and Flickr30K datasets illustrate that our proposed AS-Transformer can adaptively obtain effective semantic information and adjust the attention weights between semantic and visual information automatically, which achieves more accurate captions compared with semantic enhancement methods and outperforms state-of-the-art methods. Jing Zhang 0041, Zhongjun Fang, Zhe Wang 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Emotion-wise feature interaction analysis-based visual emotion distribution learning
Jing Zhang 0041, Qiuge Qin, Qi Ye 0004, Wen Du |
Vis. Comput. | 1 |
| 2023 | Improving Image Captioning through Visual and Semantic Mutual PromotionabstractCurrent image captioning methods commonly use semantic attributes extracted by an object detector to guide visual representation, leaving the mutual guidance and enhancement between vision and semantics under-explored. Neurological studies have revealed that the visual cortex of the brain plays a crucial role in recognizing visual objects, while the prefrontal cortex is involved in the integration of contextual semantics. Inspired by the above studies, we propose a novel Visual-Semantic Transformer (VST) to model the neural interaction between vision and semantics, which explores the mechanism of deep fusion and mutual promotion of multimodal information, realizing more accurate image captioning. To better facilitate the complementary strengths between visual objects and semantic contexts, we propose a global position-sensitive co-attention encoder to realize globally associative, position-aware visual and semantic co-interaction through a mutual cross-attention mechanism. In addition, a multimodal mixed attention module is proposed in the decoder, which achieves adaptive multimodal feature fusion for enhancing the decoding capability. Experimental evidence shows that our VST significantly surpasses the state-of-the-art approaches on MSCOCO dataset and reaches the excellent CIDEr score of 142% on the Karpathy test split. Jing Zhang 0041, Yingshuai Xie |
ACM Multimedia | 1 |
| 2023 | Multi-feature fusion enhanced transformer with multi-layer fused decoding for image captioning
Jing Zhang 0041, Zhongjun Fang, Zhe Wang 0002 |
Appl. Intell. | 1 |
| 2023 | From multi-omics data to the cancer druggable gene discovery: a novel machine learning-based approachabstractThe development of targeted drugs allows precision medicine in cancer treatment and optimal targeted therapies. Accurate identification of cancer druggable genes helps strengthen the understanding of targeted cancer therapy and promotes precise cancer treatment. However, rare cancer-druggable genes have been found due to the multi-omics data's diversity and complexity. This study proposes deep forest for cancer druggable genes discovery (DF-CAGE), a novel machine learning-based method for cancer-druggable gene discovery. DF-CAGE integrated the somatic mutations, copy number variants, DNA methylation and RNA-Seq data across ˜10 000 TCGA profiles to identify the landscape of the cancer-druggable genes. We found that DF-CAGE discovers the commonalities of currently known cancer-druggable genes from the perspective of multi-omics data and achieved excellent performance on OncoKB, Target and Drugbank data sets. Among the ˜20 000 protein-coding genes, DF-CAGE pinpointed 465 potential cancer-druggable genes. We found that the candidate cancer druggable genes (CDG) are clinically meaningful and divided the CDG into known, reliable and potential gene sets. Finally, we analyzed the omics data's contribution to identifying druggable genes. We found that DF-CAGE reports druggable genes mainly based on the copy number variations (CNVs) data, the gene rearrangements and the mutation rates in the population. These findings may enlighten the future study and development of new drugs. Hai Yang 0002, Lipeng Gan, Rui Chen 0021, Dongdong Li 0003, Jing Zhang 0041, Zhe Wang 0002 |
Briefings Bioinform. | 5 |
| 2023 | Hierarchical decoding with latent context for image captioning
Jing Zhang 0041, Yingshuai Xie, Kangkang Li 0003, Zhe Wang 0002, Wen Du |
Neural Comput. Appl. | 1 |
| 2023 | Cross on Cross Attention: Deep Fusion Transformer for Image CaptioningabstractNumerous studies have shown that in-depth mining of correlations between multi-modal features can help improve the accuracy of cross-modal data analysis tasks. However, the current image description methods based on the encoder-decoder framework only carry out the interaction and fusion of multi-modal features in the encoding stage or the decoding stage, which cannot effectively alleviate the semantic gap. In this paper, we propose a Deep Fusion Transformer (DFT) for image captioning to provide a deep multi-feature and multi-modal information fusion strategy throughout the encoding to decoding process. We propose a novel global cross encoder to align different types of visual features, which can effectively compensate for the differences between features and incorporate each other’s strengths. In the decoder, a novel cross on cross attention is proposed to realize hierarchical cross-modal data analysis, extending complex cross-modal reasoning capabilities through the multi-level interaction of visual and semantic features. Extensive experiments conducted on the MSCOCO dataset prove that our proposed DFT can achieve excellent performance and outperform state-of-the-art methods. The code is available athttps://github.com/weimingboya/DFT. Jing Zhang 0041, Yingshuai Xie, Weichao Ding, Zhe Wang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Image sentiment classification via multi-level sentiment region correlation analysis
Jing Zhang 0041, Qi Ye 0004, Zhe Wang 0002 |
Neurocomputing | 1 |
| 2022 | Geometric imbalanced deep learning with feature scaling and boundary sample mining
Zhe Wang 0002, Qida Dong, Wei Guo 0023, Dongdong Li 0003, Jing Zhang 0041, Wenli Du |
Pattern Recognit. | 5 |
| 2022 | Graph-Based Object Semantic Refinement for Visual Emotion RecognitionabstractThe rich semantic information contained in images is an important clue to explore visual emotions. Therefore, exploring the correlation between visual emotion and the semantic relationship of objects, and extracting more effective semantic features through explicit or implicit modeling is very important for visual emotion analysis. In this paper, a novel Graph-based Object Semantic Refinement (GOSR) model is proposed to extract multi-level semantic features for visual emotion classification, in which graph structures is used to represent the object semantics and their position relationships of an image, and Graph Convolutional Networks (GCN) is used to refine object information by the aggregating neighbor object with their position relationships. The different convolutional layer’s features from GCN are further fused by Gated Recurrent Units (GRU) networks to achieve high-level semantic features. Then a framework with two branches to leverage visual and semantic information for visual sentiment analysis is proposed, which uses convolutional neural networks to extract visual features from images, and collaborates with semantic features from GOSR model to achieve better emotion recognition results. Besides, for alleviating the potentially unreasonable predictions and promote models collaboration, a novel tendency loss function based on the correlations among emotion labels is proposed to adjust the output activation value other than the target label. Extensive experiments on four widely used benchmark datasets show that our proposed method can achieve competitive performance and outperform most of the state-of-the-art methods on visual emotion recognition. Jing Zhang 0041, Zhe Wang 0002, Hai Yang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Task Encoding With Distribution Calibration for Few-Shot LearningabstractFew-shot learning is an extremely challenging task in computer vision that has attracted increased research attention in recent years. However, most recent methods do not fully use the task’s information, and few of the seen samples result in large intraclass differences among the same classes. In this paper, we propose a novel task encoding with distribution calibration (TEDC) model for few-shot learning, which uses the relationships among the feature distributions to reduce intraclass differences. In the TEDC model, an integrated feature extraction module (IFEM) is proposed, which extracts the multiangle visual features of an image and fuses them to obtain more representative features. To effectively utilize the task information, a novel task encoding module (TEM) is proposed, which obtains the task features by fusing all the seen samples’ information and uses them to adjust all the samples’ features for more generalizable task-specific representations. We also propose a distribution calibration module (DCM) to reduce the bias between the distribution of the support features and the query features in the same class. Extensive experiments show that our proposed TEDC model achieves an excellent performance and outperforms the state-of-the-art methods on three widely used few-shot classification benchmarks, specifically miniImageNet, tieredImageNet and CUB-200-2011. Jing Zhang 0041, Xinzhou Zhang, Zhe Wang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Visual enhanced gLSTM for image captioning
Jing Zhang 0041, Kangkang Li 0003, Zhenkun Wang 0005, Xianwen Zhao, Zhe Wang 0002 |
Expert Syst. Appl. | 1 |
| 2021 | Parallel-fusion LSTM with synchronous semantic and visual information for image captioning
Jing Zhang 0041, Kangkang Li 0003, Zhe Wang 0002 |
J. Vis. Commun. Image Represent. | 1 |
| 2021 | BLSTM and CNN Stacking Architecture for Speech Emotion Recognition
Dongdong Li 0003, Linyu Sun, Xinlei Xu, Zhe Wang 0002, Jing Zhang 0041, Wenli Du |
Neural Process. Lett. | 5 |
| 2020 | Multiple Universum Empirical Kernel Learning
Zhe Wang 0002, Sisi Hong, Lijuan Yao, Dongdong Li 0003, Wenli Du, Jing Zhang 0041 |
Eng. Appl. Artif. Intell. | 6 |
| 2020 | Object semantics sentiment correlation analysis enhanced image sentiment classification
Jing Zhang 0041, Dongdong Li 0003, Zhe Wang 0002 |
Knowl. Based Syst. | 1 |
| 2020 | Multi-matrices entropy discriminant ensemble learning for imbalanced problem
Zhe Wang 0002, Zhaozhi Chen, Jing Zhang 0041, Wenli Du, Dongdong Li 0003 |
Neural Comput. Appl. | 4 |
| 2020 | Efficient matrixized classification learning with separated solution process
Zonghai Zhu, Zhe Wang 0002, Dongdong Li 0003, Wenli Du, Jing Zhang 0041 |
Neural Comput. Appl. | 5 |
| 2019 | Another Dimension: Towards Multi-subnet Neural Network for Image Sentiment AnalysisabstractImage sentiment analysis has been studied for many years, and most of algorithms take the image sentiment as independent and discrete labels to predict by machine learning. Actually, as a product of multiple hormone combinations, emotions are generated by mutual suppression signal in brain. Inspired by neural microcircuit in amygdala, we propose a novel Multi-Subnet Neural Network (MSNN) that simulates the human brain mechanism for image sentiment classification. Different from traditional neural network, MSNN extends a new domain channel to imitate the way that images stimulate the brain through different neural circuits and produce sentimental semantic information by multi-subnet and signal reforming network. Experiments show that MSNN is well adapted to multi-class image sentiment classification task, and outperforms other multi-class sentiment classification models. Jing Zhang 0041, Zhe Wang 0002, Tong Ruan |
ICME | 1 |
| 2019 | Multi-view learning with fisher kernel and bi-bagging for imbalanced problem
Zhe Wang 0002, Zhaozhi Chen, Jing Zhang 0041, Wenli Du |
Appl. Intell. | 4 |
| 2019 | Web image annotation based on Tri-relational Graph and semantic context analysis
Jing Zhang 0041, Ti Tao, Yakun Mu, Dongdong Li 0003, Zhe Wang 0002 |
Eng. Appl. Artif. Intell. | 1 |
| 2019 | Cost-sensitive Fuzzy Multiple Kernel Learning for imbalanced problem
Zhe Wang 0002, Bolu Wang, Dongdong Li 0003, Jing Zhang 0041 |
Neurocomputing | 5 |
| 2019 | Image region label refinement using spatial position relation graph
Jing Zhang 0041, Zhenkun Wang 0005, Yakun Mu, Zhe Wang 0002 |
Knowl. Based Syst. | 1 |
| 2018 | Image region annotation based on segmentation and semantic correlation analysisabstractThe authors propose an image region annotation framework by exploring syntactic and semantic correlations among segmented regions in an image. A texture‐enhanced image segmentation JSEG algorithm is first used to improve the pixel consistency in a segmented image region. Next, each region is represented by a set of image codewords, also known as visual alphabets, with each of them used to characterise certain low‐level image features. A visual lexicon, with its vocabulary items defined as either a codeword or a co‐occurrence of multiple alphabets, is formed and used to model middle‐level semantic concepts. The concept classification models are trained by a maximal figure‐of‐merit algorithm with a collection of training images with multiple correlations, including spatial, syntactic and semantic relationship, between regions and their corresponding concepts. In addition, a region‐semantic correlation model constructed with latent semantic analysis is used to correct the potentially wrong annotations by analysing the relationship between image region positions and labels. When evaluated on the Corel 5K dataset, the proposed image region annotation framework achieves accurate results on image region concept tagging as well as whole image based annotations. Jing Zhang 0041, Yakun Mu, Sheng-Wei Feng, Kehuang Li, Yubo Yuan 0001 |
IET Image Process. | 1 |
| 2017 | Image retrieval using the extended salient region
Jing Zhang 0041, Sheng-Wei Feng, Yong-Wei Gao, Yubo Yuan 0001 |
Inf. Sci. | 1 |
| 2016 | Automatic image region annotation through segmentation based visual semantic analysis and discriminative classificationabstractWe propose a new framework for automatic image annotation (AIA) of regions through segmentation based semantic analysis and discriminative classification. Given a test image, it is first segmented by a proposed texture-enhanced JSEG algorithm. Then these regions are represented by an extended bag-of-words model in which a feature vector, based on a visual lexicon with its vocabulary consisting of a visual word or a co-occurrence of multiple visual words, is constructed to represent the region content. Finally a concept classifier learned by a maximal figure-of-merit algorithm is used to predict the region labels. These models are discriminatively trained from image regions with multiple associations between regions and concepts. Experiments on a subset of the Corel 5K data set illustrate that our proposed approach to region AIA achieves more accurate annotation results than some sate-of-the-art algorithms. Jing Zhang 0041, Yong-Wei Gao, Sheng-Wei Feng, Yubo Yuan 0001 |
ICASSP | 1 |
| 2016 | Structure-aware image inpainting using patch scale optimization
Chao Dai, Bin Sheng 0001, Jing Zhang 0041, Weiyao Lin, Yubo Yuan 0001 |
J. Vis. Commun. Image Represent. | 5 |
| 2016 | Image saliency detection using Gabor texture cues
Bin Sheng 0001, Jian-ning Liang, Jing Zhang 0041, Yubo Yuan 0001 |
Multim. Tools Appl. | 5 |
| 2015 | Moving visual focus in salient object segmentationabstractSaliency detection plays an important role in image segmentation, object detection and retrieval, which attracts more attention in the field of computer vision recently. Most existing saliency detection algorithms have not considered the influence of visual focus (VF) shifting yet. In this study, a novel algorithm named moving region contrast (MRC) is proposed to analyse image saliency. The algorithm MRC is built on a novel concept of moving VF. The initial VF is defined as the geometric centre of the image. Then the VF is calculated iteratively by focus‐moving technique where a saliency gravitation model is employed to determine the moving direction. The salient region is obtained according to the final VF. The experiments are conducted on the dataset with 1000 images released by Achanta. Experimental results show that the proposed algorithm achieves marked improvements in performance and outperforms other 11 popular algorithms. Xiao-Long Xiao, Fangli Ying, Jing Zhang 0041, Yubo Yuan 0001 |
IET Image Process. | 5 |
| 2015 | Representation of image content based on RoI-BoW
Jing Zhang 0041, Ya-Xin Zhao, Yubo Yuan 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2015 | A novel image annotation model based on content representation with multi-layer segmentation
Jing Zhang 0041, Ya-Xin Zhao, Yubo Yuan 0001 |
Neural Comput. Appl. | 1 |