Peiwu Qin

dblp:256/7630 · DBLP profile ↗
← Back
15ranked-venue papers
0as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 7 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Brain network construction and analysis for epilepsy: A methodology review
Yuge Yang, Duanpo Wu, Tiejia Jiang, Chenggang Yan 0001, Yixuan Yuan, Samaneh Kashi, Peiwu Qin
Neural Networks9
2025 ViTCM-LLM: A Multimodal RAG Framework for Advanced TCM Clinical Decision Support
abstract
Traditional Chinese Medicine (TCM) diagnosis, particularly methods like tongue diagnosis, faces significant challenges in subjectivity and scalability. The application of Large Language Models (LLMs) to fundamental TCM tasks, such as syndrome differentiation and prescription generation, is significantly hampered by the difficulty of integrating visual tongue data with clinical text, and by the scarcity of suitable public datasets. To overcome these barriers, we introduce ViTCM-LLM, a novel framework that emulates an expert's diagnostic process by integrating multimodal language modeling with Retrieval-Augmented Generation (RAG). Employing a dual-component architecture, ViTCM-LLM integrates a fine-tuned Qwen2.5-VL model for visual analysis with a Qwen3-based RAG for clinical reasoning. The framework was developed and validated using MedTCM, a new large-scale multimodal dataset that we introduce specifically for advanced TCM research. To properly evaluate our framework's clinical accuracy, which existing metrics fail to capture, we also developed TDEU, a domain-specific evaluation metric. Evaluated using this metric alongside standard benchmarks, our comprehensive experiments on MedTCM demonstrate that ViTCM-LLM significantly outperforms leading models, including GPT-4o and Gemini 2.5 Flash. These findings not only establish the feasibility of a generalizable tongue diagnosis model but also validate the critical role of integrating visual data for advanced TCM diagnostics. The code and data can be found at https://github.com/jw-chae/ViTCM_LLM.
Lihui Luo, Joongwon Chae, Yang Liu 0471, Igor Pantic, Vladan Devedzic, Zhumei Sun, Zelin Zeng, Sanatbek Matlatipov, Xiaoming Yin, Peiwu Qin
BIBM10
2025 Enhancing Dual-Stream Attention Network: A Novel Multimodal Fusion Approach for Cognitive Assessment in Cerebral Small Vessel Disease
abstract
Cerebral small vessel disease (CSVD) is a major contributor to cognitive impairment and dementia in aging populations, necessitating early and precise diagnosis for effective clinical intervention. While existing multimodal imaging approaches address CSVD-related cognitive deficits, they often lack deep semantic integration of complementary modalities and struggle with class imbalance in heterogeneous data. We present the Enhancing Dual-Stream Attention Network (EDSAN), a novel framework that synergistically combines T1-weighted and diffusion tensor imaging (DTI) through multi-stage preprocessing for modality consistency. Our architecture introduces bidirectional cross-attention fusion and heterogeneous encoder optimization to achieve dynamically aligned multimodal feature learning. A unified optimization strategy employing gated weight learning and knowledge distillation streamlines cross-modal interactions while preserving complementary information under modality-specific constraints. The framework incorporates diffusion models and GAN-based auxiliary modules to enhance feature robustness and cross-modal distribution modeling. By integrating Focal Loss, cross-entropy, WGAN-GP, and diffusion reconstruction losses into a hybrid objective function, EDSAN demonstrates state-of-the-art performance in classifying CSVD-induced cognitive impairment. This modular paradigm establishes an extensible foundation for multimodal medical imaging analysis, adaptable to early diagnostic applications across neurological disorders.
Sen Zeng, Chenying Lu, Jiansong Ji, Peiwu Qin
BIBM9
2025 TAGMO: Temporal Control Audio Generation for Multiple Visual Objects Without Training
abstract
With the great popularity of Sora, video-based audio generation has become indispensable. While numerous video-to-audio generation models have emerged, they frequently face difficulties including semantic incompatibilities and synchronization problems, especially in situations with multiple objects. To address these difficulties, we introduce TAGMO, a novel training-free audio generation method that offers precise time control for multi-object video scenarios. Our approach first employs object detection to obtain the class labels and temporal labels of each object, which are then structured and utilized as control conditions within a latent diffusion model (LDM) to generate multi-object audio. Additionally, we innovatively design a time mask based on the corresponding temporal labels and integrate it into the denoising process of the pre-trained audio generation model to achieve accurate temporal control. Experimental results demonstrate that our method enhances temporal alignment accuracy and semantic consistency. Audio demonstrations are available at https://coco-create.github.io/.
Keyu Fan, Yingshan Liang, Jiasheng Lu, Zhicheng Du, Qingyang Shi, Peiwu Qin
ICASSP8
2025 HCMA-UNet: A Hybrid CNN-Mamba UNet with Axial Self-Attention for Efficient Breast Cancer Segmentation
abstract
Breast cancer lesion segmentation in DCE-MRI remains challenging due to heterogeneous tumor morphology and indistinct boundaries. To address these challenges, this study proposes a novel hybrid segmentation network, HCMA-UNet, for lesion segmentation of breast cancer. Our network consists of a lightweight CNN backbone and a Multi-view Axial Self-Attention Mamba (MISM) module. The MISM module integrates Visual State Space Block (VSSB) and Axial Self-Attention (ASA) mechanism, effectively reducing parameters through Asymmetric Split Channel (ASC) strategy to achieve efficient tri-directional feature extraction. Our lightweight model achieves superior performance with 2.87M parameters and 126.44 GFLOPs. A Feature-guided Region-aware loss function (FRLoss) is proposed to enhance segmentation accuracy. Extensive experiments on one private and two public DCE-MRI breast cancer datasets demonstrate that our approach achieves state-of-the-art performance while maintaining computational efficiency. FRLoss also exhibits good cross-architecture generalization capabilities. The source code is available at https://github.com/Haoxuanli-Thu/HCMA-UNet.
Peiwu Qin, Xi Yuan, Zhenglin Chen
ICME3
2025 DBF-UNet: A Two-Stage Framework for Carotid Artery Segmentation with Pseudo-Label Generation
abstract
Medical image analysis faces significant challenges due to limited annotation data, particularly in three-dimensional carotid artery segmentation tasks, where existing datasets exhibit spatially discontinuous slice annotations with only a small portion of expert-labeled slices in complete 3D volumetric data. To address this challenge, we propose a two-stage segmentation framework. First, we construct continuous vessel centerlines by interpolating between annotated slice centroids and propagate labels along these centerlines to generate interpolated annotations for unlabeled slices. The slices with expert annotations are used for fine-tuning SAM-Med2D, while the interpolated labels on unlabeled slices serve as prompts to guide segmentation during inference. In the second stage, we propose a novel Dense Bidirectional Feature Fusion UNet (DBF-UNet). This lightweight architecture achieves precise segmentation of complete 3D vascular structures. The network incorporates bidirectional feature fusion in the encoder and integrates multi-scale feature aggregation with dense connectivity for effective feature reuse. Experimental validation on public datasets demonstrates that our proposed method effectively addresses the sparse annotation challenge in carotid artery segmentation while achieving superior performance compared to existing approaches. The source code is available at https://github.com/Haoxuanli-Thu/DBF-UNet.
Aofan Liu, Peiwu Qin
IJCNN4
2025 AdaDocVQA: Adaptive Framework for Long Document Visual Question Answering in Low-Resource Settings
abstract
Document Visual Question Answering (Document VQA) faces significant challenges when processing long documents in low-resource environments due to context limitations and insufficient training data. This paper presents AdaDocVQA, a unified adaptive framework addressing these challenges through three core innovations: a hybrid text retrieval architecture for effective document segmentation, an intelligent data augmentation pipeline that automatically generates high-quality reasoning question-answer pairs with multi-level verification, and adaptive ensemble inference with dynamic configuration generation and early stopping mechanisms. Experiments on Japanese document VQA benchmarks demonstrate substantial improvements with 83.04% accuracy on Yes/No questions, 52.66% on factual questions, and 44.12% on numerical questions in JDocQA, and 59% accuracy on LAVA dataset. Ablation studies confirm meaningful contributions from each component, and our framework establishes new state-of-the-art results for Japanese document VQA while providing a scalable foundation for other low-resource languages and specialized domains. Our code available at: https://github.com/Haoxuanli-Thu/AdaDocVQA.
Aofan Liu, Peiwu Qin
ACM Multimedia4
2024 External Prompt Features Enhanced Parameter-Efficient Fine-Tuning for Salient Object Detection
Wen Liang, Peipei Ran, Mengchao Bai, Bilha Githinji, Peiwu Qin
ICPR (23)7
2024 Object detection for caries or pit and fissure sealing requirement in children's first permanent molars
abstract
Abstract Dental caries, a common oral disease, poses serious risks if untreated, necessitating effective preventive measures like pit and fissure sealing. However, the reliance on experienced dentists for pit and fissures or caries detection limits accessibility, potentially leading to missed treatment opportunities, especially among children. To bridge this gap, we leverage deep learning in object detection to develop a method for autonomously identifying caries and determining pit and fissure sealing requirements using smartphone oral photos. We test several detection models and adopt a tiling strategy to reduce information loss during image pre‐processing. Our implementation achieves 72.3 mAP.5 with the YOLOXs model and tiling strategy. We enhance accessibility by deploying the pre‐trained network as a WeChat applet on mobile devices, enabling in‐home detection by parents or guardians. In addition, our data set of children's first permanent molars will also aid in the broader study of pediatric oral disease.
Chenyao Jiang, Shiyao Zhai, Hengrui Song, Yuqing Ma, Yachen Fan, Yancheng Fang, Dongmei Yu, Canyang Zhang, Sanyang Han, Runming Wang, Zhenglin Chen, Peiwu Qin
Comput. Intell.14
2024 Deep learning model for human-intuitive shoeprint reconstruction
Yan Wang 0028, Di Wang 0004, Wei Pang 0001, Daixi Li, You Zhou 0008, Dong Xu 0002, Sami Ur Rahman, Amin ur Rahman, Ahmed Ameen Fateh, Peiwu Qin
Expert Syst. Appl.11
2024 Unfolding Explainable AI for Brain Tumor Segmentation
abstract
Brain tumor segmentation (BTS) has been studied from handcrafted engineered features to conventional machine learning (ML) methods, followed by the cutting-edge deep learning approaches. Each recent approach has attempted to overcome the challenges of previous methods and brought conveniences in efficacy, throughput, computation, explainability, investigation, and interpretability. Recently, deep learning (DL) algorithms show excellent performance regarding diverse fields, including image process, computer vision, health analytics, autonomous vehicles, and natural language processes; however, ultimately impediment in making the artificial intelligence explainable and interpretable to clinicians while dealing with critical health informatics and radiomics. Besides the sophisticated deep learning models for brain tumor segmentation, notorious notions like explainability, investigation, trust, and interpretability of DL raised significant concerns for clinicians in their domains. Among many DL methods, the neuro-symbolic learning (NSL) concept has gained more attention as it can contribute to explainable and interpretable AI. In the current study, we survey the prominent approaches, from handcrafted engineering conventional ML to deep learning algorithms, highlight the challenges in DL algorithms, and propose NSL architectures for BTS. Compared to existing surveys, our study not only outlines handcrafted to DL methods for BTS but also proposed explainable and interpretable pipelines appropriate for clinical practices. Our study can better facilitate novice learners in explainable AI and propose efficient, robust, interpretable DL models to facilitate the diagnosis, prognosis, and treatment of BTS.
Ahmed Ameen Fateh, Jieqiong Lin, Yijiang Zhuang, Guisen Lin, Hairui Xiong, You Zhou 0008, Peiwu Qin, Hongwu Zeng
Neurocomputing8
2022 Multidimensional Hypergraph on Delineated Retinal Features for Pathological Myopia Task
Bilha Githinji, Lin An, Yuhan Dong, Wen B. Wei, Peiwu Qin
MICCAI (2)11
2021 VRBT: A Non-pharmacological VR approach towards hypertension
abstract
Hypertension is a prevalent disease that is known to affect the vascular system especially to the people with poor living habits and lifestyles. Virtual reality (VR) is effective to interact with people to release their pressure and cheer them up, which however is less conducted towards manipulating blood pressure and hypertension. In this paper, we consider how hypertension can be treated with VR devices and design virtual reality river bathing therapy (VRBT) with respect to a combination of traditional methods through sensory stimulation, audio interventions, and motor training.
Yui Lo, Qinglan Shan, Jie Xu 0010, Peiwu Qin, Yuhan Dong
VRST6
2021 Stroke prediction from electrocardiograms by deep neural network
Yifeng Xie, Hongnan Yang, Xi Yuan, Ruitao Zhang, Qianyun Zhu, Zhenhai Chu, Chengming Yang, Peiwu Qin, Chenggang Yan 0001
Multim. Tools Appl.9
2020 Weighted Convolutional Motion-Compensated Frame Rate Up-Conversion Using Deep Residual Network
abstract
Frame rate up-conversion (FRUC) usually suffers from unreliable motion vectors due to the absence of the current frame to be interpolated. In addition, since the majority of video sequences are usually compressed by various coding standards to reduce the data volume, the quality of the generated frames in the FRUC will be further impaired. To address this problem, we proposed two FRUC algorithms based on deep residual network. We first present a deep residual network for the FRUC (DRNFRUC), which consists of feature extraction, feature recursive analysis, and image restoration parts with a skip connection between the input and the output of the network. The proposed DRNFRUC takes the result of an arbitrary existing FRUC method as the input and is able to significantly reduce the edge blurring and blocking artifacts when the motion of the block is violent. In addition, we proposed a deep residual network with weighted convolutional motion compensation (DRNWCMC) for the FRUC, where the convolution operations can be embedded into the motion compensation interpolation (MCI) in any existing MCI-based FRUC method. In DRNWCMC, we first devise two convolutional neural networks corresponding to the forward and backward motion compensated frames, respectively. And then, the adaptive interpolation coefficients for motion compensation are designed as two$1\times1$convolutional kernels. Finally, the interpolation result of WCMC is fed into another convolutional neural network to further improve the performance. All the parameters involved in the DRNWCMC are trained simultaneously under the same cost function. The experimental results show that the two proposed algorithms can remarkably improve both the objective and subjective quality of the interpolated frames.
Yongbing Zhang 0002, Lixin Chen, Chenggang Yan 0001, Peiwu Qin, Xiangyang Ji, Qionghai Dai
IEEE Trans. Circuits Syst. Video Technol.4