Gongpeng Song

dblp:380/7347 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Context-Aware Enhancement and Transformer Network for Image-Text Retrieval
abstract
Image-text matching is an important and challenging task in the field of multimedia analysis, aimed at bridging the semantic gap between visual content and language descriptions. Although this field has significant implications for enhancing cross-modal interaction, most previous work still faces challenges in accurately aligning images with text descriptions, especially when dealing with images containing rich semantic information. To address this issue, we propose a novel context-aware image-text matching model that extracts and summarizes visual region information aligned with multiple text descriptions from a single image. Specifically, we designed an adaptive context-aware self-attention module to extract representations of visual regions and text words. By controlling the complementary semantic relationships within each modality, our model can adaptively capture contextual information for each modality. We then introduce a Transformer-based encoder layer to extract features at multiple levels and aggregate region-level features into image-level features. Finally, fine-grained word-region alignment is conducted to match image features with corresponding text features. To evaluate the effectiveness of our method, we conducted extensive experiments on two benchmark datasets, Flickr30K and MS-COCO. The experimental results show that our model outper-forms several state-of-the-art baselines in image-text matching tasks, demonstrating the effectiveness and practicality of our model in handline cross-modal retrieval tasks.
Ruijia Zhang, Gongpeng Song
CSCWD4
2025 Hierarchical Text Classification Method Driven by Dynamic Hierarchical Fusion and Prompt Embedding
abstract
Hierarchical Text Classification, as a critical task in natural language processing, has broad applications in scenarios with complex label structures and limited samples. However, existing methods still face challenges in capturing hierarchical label information and text interactions, resulting in suboptimal classification performance. This study proposes a few-shot hierarchical text classification method based on prompt embeddings and dynamically fused hierarchical features. Building on pre-trained language models, the method encodes hierarchical structure information as prompts through an adaptive prompt embedding model and dynamically fuses hierarchical features to generate more expressive text representations. Additionally, a positive sample generation strategy based on contrastive learning is introduced, enabling the model to effectively distinguish hierarchical features across different categories, further enhancing classification performance. Experimental results demonstrate that this method significantly improves classification accuracy on multiple public datasets, exhibiting stronger generalization ability and robustness compared to mainstream methods. This approach offers a novel and effective solution for hierarchical text classification tasks.
Yiyun Xing, Kaili Zhou, Gongpeng Song
CSCWD4
2025 Multimodal Knowledge Graph Completion Method Based on Integrated Modality Adversarial Training and Relation-Enhanced Attention Mechanism
abstract
Negative sampling (NS) is widely used in knowledge graph completion (KGC) to generate negative triples for contrastive learning during training. However, existing NS methods are not suitable when multimodal information is incorporated into KGC models. Due to their complex design, these methods are also inefficient. In this paper, we propose the integration of Modality-Aware Adversarial Training (IMAT) to generate higher-quality negative samples for Multimodal Knowledge Graph Completion (MMKGC), and introduce a Relation-Enhanced Cross-modal Attention (RECA) mechanism to evaluate bidirectional attention weights between multimodal features using relational information, thereby improving the model's ability to identify hard negative samples. Our approach represents a joint design of MMKGC models and training strategies, surpassing 16 recent MMKGC methods and achieving new state-of-the-art results on three public MMKGC benchmarks.
Ruijia Zhang, Gongpeng Song
CSCWD4
2024 A Novel Text Matching Model Based on Multilayer Coding and Feature Enhancement
abstract
With the rapid development of Natural Language Processing (NLP), text matching has become the basis of many downstream tasks in NLP, and the study of text matching is of great research significance for solving tasks such as question and answer and information retrieval in NLP. Most of the current text matching methods tend to have problems such as mismatch of grammatical structures and insufficient interaction information in sentences. In order to solve the problems of insufficient interaction information and lack of features and ability to capture keyword and sequence information in text matching, this paper proposes a text matching method based on multi-layer coding and soft attention mechanism. The method first embeds the text and sends it to a gating module for processing, then sends the processed result to a module containing a combination of multilayer coding and soft attention mechanism for further operations such as multiple alignment, and finally sends it to a classifier containing a three-layer fully-connected network for predicting whether the input text pairs match or not. The two modules proposed in this paper are practically feasible, and comparison and ablation experiments have been conducted on the publicly available datasets LCQMC dataset and BQ dataset, and the experimental results show that the two modules improve the accuracy of text matching by 2.44% and 1.02%, respectively.
Gongpeng Song
CSCWD3
2024 SCTAR: A Multi-Layer BiLSTM-Based Chinese Short Text Similarity Computation Model with Attention Mechanism
abstract
Text semantic similarity is a crucial research area in Natural Language Processing (NLP). However, traditional methods for calculating the similarity of short Chinese texts often fall short in accuracy and other crucial aspects. The Bidirectional Long Short-Term Memory Network (BiLSTM) has demonstrated remarkable performance in computing the similarity of short Chinese texts by effectively capturing long-range dependencies and semantic information. Additionally, the multi-head attention mechanism considers the interaction between different text locations and semantic information, enhancing the model’s representative capacity. Building upon this foundation, our paper proposes an enhanced model known as SCTbilstmAttRdrop (SCTAR). This model is constructed based on a multilayer BiLSTM architecture and incorporates an SE-gated convolutional module and a convolutional multi-head attention-aware model. We conducted extensive evaluations using two Chinese short text datasets, Chinese-SNLI and CCKS2018_Task3. The experimental results unequivocally demonstrate that our SCTAR model surpasses other common methods in terms of accuracy, precision, recall, and F1 scores when tasked with computing the similarity of Chinese short texts.
Yiyun Xing, Gongpeng Song
CSCWD4
2024 FERI: Feature Enhancement and Relational Interaction for Image-text Matching
abstract
Image-text matching is an important problem at the intersection of computer vision and natural language processing. It aims to establish the semantic link between image and text to achieve high-quality semantic alignment between the two modalities. However, the existing methods have the problem that the meaning expressed in the image or the complex narrative in the text cannot be fully understood due to insufficient feature extraction. Moreover, due to the essential modal differences between images and texts, how to effectively and accurately align the semantic contents in images and texts has become the key of research. In order to solve the above problems, this paper proposes a method based on feature enhancement and relationship interaction. When processing images, the proposed method fuses labeled features, region features and location features to represent images. When processing text, a combination of Bi-GRU and self-attention mechanism is used to represent the text. In order to further align the semantic content in images and texts accurately, this paper improves two relational interaction mechanisms by identifying connection relationships and learning association relationships. Thus, the relation enhanced embedding is obtained. Finally, it calculated the similarity of the enhanced embedding to judge the matching degree of the image and text. Extensive experiments on the public datasets Flickr30K and MSCOCO demonstrate the effectiveness of our method.
Gongpeng Song
ICPADS3
2024 MFFLEN: Multi-Label Text Classification Based on Multi-Feature Fusion and Label Embedding
abstract
To address the challenges associated with insufficiently extracting and utilizing features at different levels, overlooking the connection between label meanings and text, and facing problems of over-compression or information loss when extracting global information using recurrent neural networks in the field of multi-label text categorization, this paper introduces an innovative model known as MFFLEN (MultiFeature Fusion and Label Embedding Neural Network). First, a back-translated enhanced label set is constructed by back-translated splicing enhancement of the original label set. This set, together with the text, is then input into the embedding layer, which consists of the pre-trained model of bert-baseChinese, thus establishing the initial connection between the text and the labels within the same vector space. Then, to comprehensively extract multi-level semantic features, the model uses a convolutional layer to extract local features and an embedding layer to extract sentence-level features. A bidirectional attention embedded GRU (BAE-GRU) layer is used to extract hybrid finegrained features, which are then fed into the attention layer to further extract hybrid labeled features based on labeling information. Finally, these three different types of features are fused and multi-label text classification results are obtained using a classifier. The experiments proved that the MFFLEN model achieved 73.82% and 88.44% macro-F1 and 88.00% and 88.86% micro-F1 on the two datasets CAIL 2018 Small and CAIL 2018 Split, respectively, which is better than other baseline models.
Qiliang Gu, Gongpeng Song
SMC4
2024 SMGC-SBERT: A Multi-Feature Fusion Chinese Short Text Similarity Computation Model Based on Optimised SBERT
abstract
Chinese short text similarity computation stands as a pivotal task within natural language processing, garnering significant attention. However, existing models grapple with limitations in handling intricate semantic relationships, such as the challenge of discerning subtle semantic nuances in text, inadequacies in effectively integrating diverse levels of semantic information, and the struggle to capture polysemous meanings accurately. In addressing these issues, this paper introduces an innovative Chinese short text similarity computation model, SMGC-SBERT. This model addresses shortcomings of the existing models by employing a multi-module fusion strategy, thereby enabling a more precise measurement of semantic similarity between texts. Primarily, the model incorporates SAT embedding to acquire phrase-level semantic information and leverages the MS-BERT model to encode text, improving the model's comprehension of textual polysemy and obtaining richer semantic representation. Then, the fusion of module features, including multi-branch convolutional networks and mix pooling, enables the extraction of textual features from varied levels, bolstering the model's representational capacity. Additionally, to further reduce overfitting risks while improving accuracy and other performance, a multi-layer feature adjustment network is utilized for short text similarity calculation. The final resultant experimental findings showcase the superiority of the SMGC-SBERT model over other neural network models, demonstrating significant advancements across both Chinese-SNLI and CCKS2018_Task3 Chinese short text datasets.
Qiliang Gu, Gongpeng Song
SMC4