Qiaohong Chen

dblp:128/0390 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Interactive edge awareness network for salient object detection in optical remote sensing images
Xian Fang, Qiaohong Chen, Gongyang Li
Expert Syst. Appl.3
2025 Selective Guidance Network with edge and texture awareness for polyp segmentation
Qiaohong Chen, Xian Fang
Expert Syst. Appl.1
2025 Edge and semantic collaboration framework with cross coordination attention for co-saliency detection
Qiaohong Chen, Xian Fang, Jinchao Zhu
Knowl. Based Syst.1
2025 MGCNet: Multiple group-wise correlation network with hierarchical contrastive learning for co-salient object detection
Xian Fang, Jinchao Zhu, Qiaohong Chen, Zuofan Chen
Knowl. Based Syst.4
2025 CaVMamba: convolution-augmented VMamba for medical image segmentation
Qiaohong Chen, Xian Fang
Vis. Comput.1
2025 CTHFNet: contrastive translation and hierarchical fusion network for text-video-audio sentiment analysis
Qiaohong Chen, Shufan Xie, Xian Fang
Vis. Comput.1
2025 DAMAF: dual attention network with multi-level adaptive complementary fusion for medical image segmentation
Yueqian Pan, Qiaohong Chen, Xian Fang
Vis. Comput.2
2024 DFEDC: Dual fusion with enhanced deformable convolution for medical image segmentation
Xian Fang, Yueqian Pan, Qiaohong Chen
Image Vis. Comput.3
2024 Global information regulation network for multimodal sentiment analysis
Shufan Xie, Qiaohong Chen, Xian Fang
Image Vis. Comput.2
2024 Dual triple attention guided CNN-VMamba for medical image segmentation
Qiaohong Chen, Xian Fang
Multim. Syst.1
2024 Sub-pixel multi-scale fusion network for medical image segmentation
Qiaohong Chen, Xian Fang
Multim. Tools Appl.2
2023 Improving Image Captioning with Feature Filtering and Injection
Qiaohong Chen, Xian Fang, Jia Bao, Shenxiang Xiang
ICANN (2)2
2023 Improving Visual Question Answering by Multimodal Gate Fusion Network
abstract
Visual question answering (VQA) is a difficult multimodal task that requires answering questions about images. It requires a fine-grained level of understanding of both the visual content of the image and the textual content of the question. However, most of the existing models perform weakly in filtering noisy information and are unable to fuse features from multiple modalities effectively. To resolve the above restriction, we propose a novel multimodal gate fusion network (MGFN), which consists of an attention-on-attention interaction module (AoAIM) and a multimodal gate fusion module (MGFM). The role of AoAIM is to capture intra-modal and inter-modal dependencies and to filter out some irrelevant attention. The proposed MGFM can effectively fuse textual and visual features based on the relative importance of textual and visual modalities. We have performed many ablation experiments on the VQA-v2 dataset to validate the effectiveness of AoAIM and MGFM. The ablation experiments demonstrate that both AoAIM and MGFM play a key role in improving the performance of the model. By embedding these two modules, MGFN performs better than the previous state-of-the-art (SOTA) model on the VQA-v2 dataset. Particularly, the MGFN achieves an overall accuracy of 71.68% on the test-dev set and 72.12% on the test-std set.
Shenxiang Xiang, Qiaohong Chen, Xian Fang
IJCNN2
2023 Gesture image recognition method based on DC-Res2Net and a feature fusion attention module
Qiuhong Tian, Wenxuan Sun, Lizao Zhang, Qiaohong Chen, Jialu Wu
J. Vis. Commun. Image Represent.5
2022 Global-Local Enhancement Network for Short Text Classification
abstract
Because of the limited context information, it is a challenging task to classify short texts. Most existing methods only focus on extracting high-quality local features or global features from text to construct text representations, which is not comprehensive enough. This article proposes a global–local enhancement network (GLEN), which can construct a high-quality text representation by integrating the global and local features of the text. To improve the quality of global and local features, a group-wise enhancement mechanism is introduced, which can effectively enhance the important features while weakening the unimportant features. For the information-loss problem of traditional pooling operation, we designed a global–local pooling mechanism, which can filter out local features that are more relevant to the whole information of the text. Seven benchmark datasets of text classification were selected to test the performance of the model, and GLEN obtained the best results on most datasets. The ablation experiment on GLEN shows that both the group-wise enhancement mechanism and the global–local pooling mechanism can effectively improve the performance of the model.
Qiaohong Chen, Ji Wang 0009, Yubo Jia
IEEE Trans. Comput. Soc. Syst.1
2021 Visual-Semantic Dual Channel Network for Visual Question Answering
abstract
Recently, the existing visual question answering (VQA) models based on the attention mechanism have achieved state-of-art results. However, attention-based networks only rely on question guidance to capture relevant image features, ignoring the high-level semantic information of the image. Therefore, the key challenge of the VQA task lies in obtaining effective semantic embedding and fine-grained visual understanding during the reasoning process. In this research, we propose a novel visual-semantic dual channel network to answer related questions from both visual and semantic perspectives. Specifically, the visual channel uses the relational reasoning method with an attention mechanism to capture visual objects and their relations, while the semantic channel can capture high-level semantic information from the global and local image by the semantic attention module. We confirmed the effectiveness of the proposed model and each module through extensive experiments on two versions of VQA datasets. Interpretability shows that the visual-semantic dual channel network can dynamically model to infer the most relevant answer to the question.
Qiaohong Chen, Yubo Jia
IJCNN2
2020 Visual Relational Reasoning for Image Caption
abstract
Recently, various attention-based networks have achieved state-of-art results on image captioning tasks. However, this simple mechanism is insufficient to modelling and reasoning the relationships between the visual regions required for scene understanding. In this research, we propose a visual relational reasoning module to implicit learning semantic and spatial relationships between pairs of relevant visual objects and infers the feature output that is most relevant to the currently generated word. Furthermore, a context gate is introduced to dynamically control the contribution of visual region attention modules and visual relational reasoning module which allows predicting different words according to different type of features (visual or visual relationship). We evaluate our model on the MSCOCO dataset and achieved state-of-the-art results. Qualitative analysis shows that our visual relational reasoning model can dynamically model and reason the most relevant features of different types of generated words and improve the quality of the caption.
Haolei Pei, Qiaohong Chen, Ji Wang 0009, Yubo Jia
IJCNN2
2020 Dynamic Global-Local Attention Network Based On Capsules for Text Classification
abstract
Text classification requires a comprehensive consideration of global and local information for the text. However, most methods only treat the global and local features of the text as two separate parts and ignore the relationship between them. In this paper, we propose a Dynamic Global-Local Attention Network based on Capsules (DGLA) that can use global features to dynamically adjust the importance of local features (e.g., sentence-level features or phrase-level features). The global features of the text are extracted by the capsule network, which can capture the mutual positional relationship of the input features to mine more hidden information. Furthermore, we have designed two global-local attention mechanisms within DGLA to measure the importance of two different local features and effectively leverage the advantages of these two attention mechanisms through the residual network. The performance of the model was evaluated on seven benchmark text classification datasets, and DGLA achieved the highest accuracy on all datasets. Ablation experiments show that the global-local attention mechanism can significantly improve the performance of the model.
Ji Wang 0009, Qiaohong Chen, Haolei Pei, Yubo Jia
IJCNN2
2020 Hybrid Neural Network for Sina Weibo Sentiment Analysis
abstract
Sina Weibo sentiment analysis technology provides the methods to survey public emotion about the related events or products in China. Most of the current works in sentiment analysis are to apply neural networks, such as convolution neural network (CNN), long short-term memory (LSTM), or C-LSTM. In this article, a novel structure of a hybrid neural network model is proposed to deal with the polysemy phenomena of words and topic confusion with Sina Weibo. First, the embeddings from language models (ELMo) and some statistical methods based on the corpus and sentiment lexicon are employed to extract the features. This method uses latent semantic relationships in different linguistic contexts and cooccurrence statistical features between words in Weibo. Second, for the classification model, unlike traditional C-LSTM which feeds CNN's output into LSTM, we employ several filters with variable window sizes to extract a sequence of high-level word representation in different granularity distributions of text data in multichannel CNN. At the same time, obtain the sentence representation in Bi-LSTM. Then, concatenate the outputs of multichannel CNN and Bi-LSTM. In conclusion, the results indicate that the proposed model performs better on the precision, recall, and F1-score for Weibo sentiment analysis.
Mingjie Ling, Qiaohong Chen, Yubo Jia
IEEE Trans. Comput. Soc. Syst.2