Shouqiang Liu

dblp:30/3361 · DBLP profile ↗
← Back
18ranked-venue papers
1as first author
16since 2021 · last 2026
0000-0002-2795-0685ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Systems, architecture and hardware · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CoGrad: Profiling Competence Boundaries for Efficient Mathematical Alignment in LLMs
Youhui Zhang, WenLiang Wu, Xifan Liu, Shouqiang Liu
ICIC (22)5
2026 Evol-Preference: Automatic evolution of preference data for safety alignment
Youyou Huang, Shouqiang Liu, Xiongwen Pang
Expert Syst. Appl.3
2025 AdaptiCare-RL: Adaptive Multi-Agent Collaborative Counseling Dialogues Regulated by Reinforcement Learning
Linying Su, Shouqiang Liu
IEEE Big Data3
2025 Feedback-Integrated Short Answer Grading with Data Augmentation and Large Language Models
Zebiao Chen, Shouqiang Liu
IEEE Big Data3
2025 TransEdge: Leveraging Transformer and EfficientKAN with Edge Sensitivity for Advanced Medical Image Segmentation
Jingguang Liao, Linying Su, JingJing Liang, Shouqiang Liu
ICIC (28)4
2025 MAMR: Modality-Aligned Multi-modal Retrieval
abstract
Retrieval-Augmented Generation (RAG) is a mainstream approach to tackling complex Visual Question Answering (VQA) tasks, such as Knowledge-Based Visual Question Answering (KB-VQA). However, previous multimodal retrieval methods typically rely on pre-trained textual and visual models to generate feature representations separately and retrieve knowledge bases by concatenating vectors directly or using simple networks for processing. While these methods are intuitive and effective, they often overlook the deep semantic correlations between text and image, limiting their robustness in complex scenarios. Moreover, the spatial dimensions of visual feature vectors are usually significantly higher than those of textual feature vectors, resulting in considerable redundancy in visual representations. This imbalance between modalities further undermines the effectiveness of modality alignment. To address these issues, we propose a novel multimodal retrieval approach with enhanced modality alignment capabilities, named Modality-Aligned Multimodal Retrieval (MAMR). By enabling image and text to learn relevant information from each other mutually, MAMR significantly improves modality alignment and fusion capabilities, achieving superior retrieval performance. Furthermore, to enhance the robustness of the model and reduce redundancy in visual features, we introduce a multi-level feature segmentation and fusion method. This approach reduces the spatial dimensions of image features while integrating multi-scale local information, thereby strengthening the representation of global visual features. Experimental results on the OKVQA and EVQA datasets demonstrate that MAMR outperforms numerous methods and achieves state-of-the-art (SOTA) performance across most evaluation metrics.
YangZhen Xu, JiaWei Mo, FangZhen Lin, Zebiao Chen, ChengFeng Chen, Shouqiang Liu
IJCNN6
2025 SAE: A Multimodal Sentiment Analysis Large Language Model
Xiyin Zeng, Qianyi Zhou, Shouqiang Liu
IUI3
2025 A noise-robust framework for multi-step wind power forecasting via self-supervised joint optimization
Gangliang Li, Chengfeng Chen, Shouqiang Liu
Eng. Appl. Artif. Intell.5
2024 MCQG: Reading Comprehension Multiple Choice Questions Generation Based on Pre-trained Language Models
Zebiao Chen, Runfeng Lin, Sujun Zhong, Shouqiang Liu
PRICAI (2)6
2024 Anatomy-Aware Enhancement and Cross-Modal Disease Representation Fusion for Medical Report Generation
abstract
The objective of medical report generation is to convert medical images into structured text reports. The limitations of existing methods primarily include the lack of efficient local modeling and fusion algorithms, as well as severe data bias in medical datasets. To address these issues, we propose a novel method for report generation based on multi-feature fusion, comprising Anatomy-Aware Fusion Enhancement (AAFE) and Cross-modal Disease Representation Fusion (CDRF). This method first localizes specific anatomical structures and then models the interdependencies among anatomy-aware tokens through AAFE, achieving deep integration of global and anatomical tokens. Simultaneously, CDRF dynamically fuses cross-modal clinical indications, visual features, and disease states to obtain a high-order representation rich in visual and disease information, enabling the generator to generate fine-grained medical reports based on rich multi-feature embedding. Furthermore, we introduce a cross-modal cosine similarity loss mechanism to further promote the extraction of disease text features with anatomy-aware information. Experimental results on the IU-Xray and MIMIC-CXR datasets demonstrate that our method outperforms existing advanced models in the medical report generation task and exhibits outstanding performance.
Runfeng Lin, Dacheng Xu, Shouqiang Liu
SMC7
2024 DGRC: An Effective Fine-Tuning Framework for Distractor Generation in Chinese Multi-Choice Reading Comprehension
abstract
When evaluating a learner's knowledge proficiency, the multiple-choice question is an efficient and widely used format in standardized tests. Nevertheless, generating these questions, particularly plausible distractors (incorrect options), poses a considerable challenge. Generally, the distractor generation can be classified into cloze-style distractor generation (CDG) and natural questions distractor generation (NQDG). In contrast to the CDG, utilizing pre-trained language models (PLMs) for NQDG presents three primary challenges: (1) PLMs are typically trained to generate “correct” content, like answers, while rarely trained to generate “plausible” content, like distractors; (2) PLMs often struggle to produce content that aligns well with specific knowledge and the style of exams; (3) NQDG necessitates the model to produce longer, context-sensitive, and question-relevant distractors. In this study, we introduce a fine-tuning framework named DGRC for NQDG in Chinese multi-choice reading comprehension from authentic examinations. DGRC comprises three major components: hard chain-of-thought, multi-task learning, and generation mask patterns. The experiment results demonstrate that DGRC significantly enhances generation performance, achieving a more than 2.5-fold improvement in BLEU scores.
Runfeng Lin, Dacheng Xu, Huijiang Wang, Zebiao Chen, Shouqiang Liu
SMC6
2024 ROP: Exploring Image Captioning with Retrieval-Augmented and Object Detection of Prompt
abstract
Image captioning models aim to bridge the modalities of vision and language by generating natural language descriptions that match the content of an input image. Existing methods for generating image captions do so by integrating visual encoders with language models, and by training the model using large volumes of image-text pairs. However, this approach leads to a significant increase in the model's parameter size due to the need to store extensive visual concepts and detailed textual descriptions. With substantial investments in datasets and computational resources, the quality of image caption generation has notably improved, though these advancements come at a high cost. In this paper, we introduce a novel method, termed ROP, that combines retrieval and detection to address this challenge. This approach significantly reduces the number of trainable parameters while preserving the accuracy of the descriptions. Through the retrieval module, the model can find the k most similar sentences to the input image, utilizing this rich contextual information to enhance the overall understanding of the image content. Simultaneously, the detection module enhances the model's interpretation of fine details by identifying prominent regions within the image. Our method has proven its effectiveness and feasibility in experiments.
Shouqiang Liu
SMC3
2024 DMPA: A Compact and Effective Pipeline for Detecting Multiple Phishing Attacks
abstract
In recent years, deep learning has significantly advanced the field of phishing defense, but most methods are typically effective only against specific types of attacks, such as phishing emails. Additionally, the large and complex nature of these models often makes them difficult to integrate efficiently into devices. To address these issues, we propose a pipeline for Detecting Multiple Phishing Attacks (DMPA) that combines fine-tuning of a pretrained model with adversarial training. We first apply a unique fine-tuning strategy to optimize resource usage, and then employ adversarial training across multiple types of phishing text attacks. By introducing perturbations to the model’s embedding vectors, we generate a large number of adversarial examples to enhance the accuracy and robustness of our model. We constructed a dataset of 100K high-quality samples and compared our method against the latest approaches from various aspects. The results show that our method surpasses all existing state-of-the-art models in overall performance. Ablation experiments further demonstrate that adversarial training significantly improves the model’s performance, confirming the effectiveness of our approach. Additionally, we have packaged the trained model into a complete pipeline, making it easier to integrate into devices. The code is made available at https://github.com/donghuang-cyber/PHISHINGDENFESE/tree/main
Gangliang Li, Chengfeng Chen, Shouqiang Liu
TrustCom4
2024 Research on the application of knowledge mapping and knowledge structure construction based on adaptive learning model
Xiyin Zeng, Shouqiang Liu
Expert Syst. Appl.2
2024 An application study on multimodal fake news detection based on Albert-ResNet50 Model
Mingyue Jiang, Chang Jing, Shouqiang Liu
Multim. Tools Appl.5
2023 Deepfake detection based on remote photoplethysmography
Qingzhen Xu, Han Qiao, Shouqiang Liu
Multim. Tools Appl.4
2020 Research of animals image semantic segmentation based on deep learning
abstract
Summary It is imperative for us to develop the technology of image semantic segmentation with the increasing demand in the image processing. Nowadays, the development of deep learning is of great significance to the improvement of image segmentation. Furthermore, the paper discussed the relationship between image semantic segmentation and animal image research based on the actual situation, and found that animal image processing technology plays a more important role in the field of protecting precious animals. The end‐to‐end network training of this paper is consisted of Fully Convolutional Network (FCN) for the front end and Conditional Random Fields as Recurrent Neural Networks (CRF‐RNN) for the back end via comparing a variety of research methods. The experiments achieved desired outcome for the semantic segmentation of animal images by utilizing Caffe deep learning framework and explained the implementation details from the aspects of training and testing.
Shouqiang Liu, Qingzhen Xu
Concurr. Comput. Pract. Exp.1
2020 A new algorithm of stock data mining in Internet of Multimedia Things
Jinfei Yang, Shouqiang Liu
J. Supercomput.3