EDBT 2026 Demo / reviewers in the wild / expert
Yaping Wu
dblp:07/10021
· DBLP profile ↗
12ranked-venue papers
2as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 8 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A hierarchical prompt and prototype learning framework for brain disorder classification
Kaicong Sun, Yaping Wu, Weilin Zhou, Haoyue Yuan, Xintong Wu, Yichu He, Qingxia Wu, Zeng-Yang Che, Yiqiang Zhan, Sean Zhou, Dijia Wu, Feng Shi 0001, Dinggang Shen |
Medical Image Anal. | 3 |
| 2025 | STARNeT: Multidimensional spatial-temporal attention recall network for accurate encrypted traffic classification
Xinjie Guan, Shuyan Zhu, Xili Wan, Yaping Wu |
J. Netw. Comput. Appl. | 4 |
| 2025 | Disentangle and Then Fuse: A Cross-Modal Network for Synthesizing Gadolinium-Enhanced Brain MR ImagesabstractDespite the widespread use of gadolinium-based contrast agents in clinical MRI examinations due to their significant advantages in structural localization and tumor identification, there is a risk of brain deposition and nephrogenic systemic fibrosis. Cross-modal image synthesis methods offer a new alternative, yet lesion synthesis remains challenging. On one hand, brain lesions vary significantly in location, shape, and size. On the other hand, the high background ratio associated with brain lesions makes their synthesis more difficult. To address these issues, we first introduce a Multi-Objective Local Perception Module (M-OLPM), which utilizes edge generation and lesion segmentation tasks to prioritize local lesions from the disentangled local perceptual feature subspaces. To better extend to multi-objective local perception, we propose a ‘Disentangle and Then Fuse’ learning strategy, including a Feature Disentanglement Module (FDM) and a Global Fusion Module (GFM). The FDM decouples multimodal deep features into low-frequency semantic features and high-frequency edge features, alleviating feature conflicts from weakly related perception tasks. To enhance feature interaction among multiple perception tasks, the GFM progressively integrates these local perceptual features and underlying detail features through an attention mechanism, further refining the global image quality. Evaluated on the publicly available BRaTS2020, BRaTS2021 datasets, and the private HPPH dataset, our method significantly outperforms the existing technology in both visual and quantitative assessments of gadolinium-enhanced MRI images in global and localized lesion areas, providing a safe alternative to gadolinium enhancement. The source code is publicly available athttps://github.com/zengyangche/DTF-Net. Zeng-Yang Che, Zheng Zhang 0006, Yaping Wu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Automatic Brain Segmentation for PET/MR Dual-Modal Images Through a Cross-Fusion MechanismabstractThe precise segmentation of different brain regions and tissues is usually a prerequisite for the detection and diagnosis of various neurological disorders in neuroscience. Considering the abundance of functional and structural dual-modality information for positron emission tomography/magnetic resonance (PET/MR) images, we propose a novel 3D whole-brain segmentation network with a cross-fusion mechanism introduced to obtain 45 brain regions. Specifically, the network processes PET and MR images simultaneously, employing UX-Net and a cross-fusion block for feature extraction and fusion in the encoder. We test our method by comparing it with other deep learning-based methods, including 3DUXNET, SwinUNETR, UNETR, nnFormer, UNet3D, NestedUNet, ResUNet, and VNet. The experimental results demonstrate that the proposed method achieves better segmentation performance in terms of both visual and quantitative evaluation metrics and achieves more precise segmentation in three views while preserving fine details. In particular, the proposed method achieves superior quantitative results, with a Dice coefficient of 85.73% 0.01%, a Jaccard index of 76.68% 0.02%, a sensitivity of 85.00% 0.01%, a precision of 83.26% 0.03% and a Hausdorff distance (HD) of 4.4885 14.85%. Moreover, the distribution and correlation of the SUV in the volume of interest (VOI) are also evaluated (PCC > 0.9), indicating consistency with the ground truth and the superiority of the proposed method. In future work, we will utilize our whole-brain segmentation method in clinical practice to assist doctors in accurately diagnosing and treating brain diseases. Hongyan Tang, Zhenxing Huang, Yaping Wu, Jianmin Yuan, Yang Yang 0186, Harry Qin, Hairong Zheng, Dong Liang 0001, Zhanli Hu |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | Prompt-Agent-Driven Integration of Foundation Model Priors for Low-Count PET ReconstructionabstractLow-count Positron Emission Tomography reconstruction is critical for maintaining high imaging quality while minimizing tracer doses and radiation exposure. Although integrating structural information from CT and MR data has been shown to enhance PET reconstruction, this typically requires simultaneous PET and CT/MRI scans, complicating workflows and increasing radiation exposure. Recent advancements in foundation models offer a promising alternative to in-person CT/MRI imaging, potentially overcoming these limitations. However, the use of foundation models' segmentation masks as semantic guides has been observed to introduce erroneous structures in low-count PET reconstructions. To address this challenge, this work introduces an innovative prompting agent-based framework that dynamically interacts with the foundation model to retrieve and refine priors, minimizing undue influence on the reconstruction process. Specifically, a box agent is designed for single-instance local area information retrieval, while a point agent is introduced to progressively prompt broader semantic structures globally, utilizing history point prompts. Additionally, an MDP paradigm has been developed to address the challenges of utilizing historical point prompts while maintaining the independence required by MDPs. Evaluated on both simulated and real datasets, the proposed method demonstrates superior qualitative and quantitative performance compared to state-of-the-art methods, even those leveraging in-person CT/MRI priors. Xingyu Xie, Mu Nan, Yaping Wu, Hairong Zheng, Dong Liang 0001, Zhanli Hu |
IEEE Trans. Medical Imaging | 5 |
| 2024 | Accurate Whole-Brain Image Enhancement for Low-Dose Integrated PET/MR Imaging Through Spatial Brain TransformationabstractPositron emission tomography/magnetic resonance imaging (PET/MRI) systems can provide precise anatomical and functional information with exceptional sensitivity and accuracy for neurological disorder detection. Nevertheless, the radiation exposure risks and economic costs of radiopharmaceuticals may pose significant burdens on patients. To mitigate image quality degradation during low-dose PET imaging, we proposed a novel 3D network equipped with a spatial brain transform (SBF) module for low-dose whole-brain PET and MR images to synthesize high-quality PET images. The FreeSurfer toolkit was applied to derive the spatial brain anatomical alignment information, which was then fused with low-dose PET and MR features through the SBF module. Moreover, several deep learning methods were employed as comparison measures to evaluate the model performance, with the peak signal-to-noise ratio (PSNR), structural similarity (SSIM) and Pearson correlation coefficient (PCC) serving as quantitative metrics. Both the visual results and quantitative results illustrated the effectiveness of our approach. The obtained PSNR and SSIM were 41.96 ± 4.91 dB (p < 0.01) and 0.9654 ± 0.0215 (p < 0.01), which achieved a 19% and 20% improvement, respectively, compared to the original low-dose brain PET images. The volume of interest (VOI) analysis of brain regions such as the left thalamus (PCC = 0.959) also showed that the proposed method could achieve a more accurate standardized uptake value (SUV) distribution while preserving the details of brain structures. In future works, we hope to apply our method to other multimodal systems, such as PET/CT, to assist clinical brain disease diagnosis and treatment. Zhenxing Huang, Yaping Wu, Yun Dong 0002, Yongfeng Yang, Hairong Zheng, Dong Liang 0001, Zhanli Hu |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | Quad-Net: Quad-Domain Network for CT Metal Artifact ReductionabstractMetal implants and other high-density objects in patients introduce severe streaking artifacts in CT images, compromising image quality and diagnostic performance. Although various methods were developed for CT metal artifact reduction over the past decades, including the latest dual-domain deep networks, remaining metal artifacts are still clinically challenging in many cases. Here we extend the state-of-the-art dual-domain deep network approach into a quad-domain counterpart so that all the features in the sinogram, image, and their corresponding Fourier domains are synergized to eliminate metal artifacts optimally without compromising structural subtleties. Our proposed quad-domain network for MAR, referred to as Quad-Net, takes little additional computational cost since the Fourier transform is highly efficient, and works across the four receptive fields to learn both global and local features as well as their relations. Specifically, we first design a Sinogram-Fourier Restoration Network (SFR-Net) in the sinogram domain and its Fourier space to faithfully inpaint metal-corrupted traces. Then, we couple SFR-Net with an Image-Fourier Refinement Network (IFR-Net) which takes both an image and its Fourier spectrum to improve a CT image reconstructed from the SFR-Net output using cross-domain contextual information. Quad-Net is trained on clinical datasets to minimize a composite loss function. Quad-Net does not require precise metal masks, which is of great importance in clinical practice. Our experimental results demonstrate the superiority of Quad-Net over the state-of-the-art MAR methods quantitatively, visually, and statistically. The Quad-Net code is publicly available at https://github.com/longzilicart/Quad-Net. Zilong Li 0001, Yaping Wu, Chuang Niu, Junping Zhang, Ge Wang 0001, Hongming Shan |
IEEE Trans. Medical Imaging | 3 |
| 2024 | Non-Invasive Quantification of the Brain [¹⁸F]FDG-PET Using Inferred Blood Input Function Learned From Total-Body Data With Physical ConstraintabstractFull quantification of brain PET requires the blood input function (IF), which is traditionally achieved through an invasive and time-consuming arterial catheter procedure, making it unfeasible for clinical routine. This study presents a deep learning based method to estimate the input function (DLIF) for a dynamic brain FDG scan. A long short-term memory combined with a fully connected network was used. The dataset for training was generated from 85 total-body dynamic scans obtained on a uEXPLORER scanner. Time-activity curves from 8 brain regions and the carotid served as the input of the model, and labelled IF was generated from the ascending aorta defined on CT image. We emphasize the goodness-of-fitting of kinetic modeling as an additional physical loss to reduce the bias and the need for large training samples. DLIF was evaluated together with existing methods in terms of RMSE, area under the curve, regional and parametric image quantifications. The results revealed that the proposed model can generate IFs that closer to the reference ones in terms of shape and amplitude compared with the IFs generated using existing methods. All regional kinetic parameters calculated using DLIF agreed with reference values, with the correlation coefficient being 0.961 (0.913) and relative bias being 1.68±8.74% (0.37±4.93%) for [Formula: see text] ( [Formula: see text]. In terms of the visual appearance and quantification, parametric images were also highly identical to the reference images. In conclusion, our experiments indicate that a trained model can infer an image-derived IF from dynamic brain PET data, which enables subsequent reliable kinetic modeling. Yaping Wu, Zeheng Xia, Dong Liang 0001, Hairong Zheng, Yongfeng Yang, Shanshan Wang 0002, Tao Sun 0025 |
IEEE Trans. Medical Imaging | 2 |
| 2023 | C2SFormer: Rethinking the Local-Global Design for Efficient Visual Recognition ModelabstractVision Transformers are born with the property of data-dependent and long-range dependencies, accomplishing a number of astonishing results against their contemporary competitor CNNs. To alleviate the excessive computational burden, previous methods apply the local operation (e.g., convolution, local attention) in the high-resolution stages. Although these designs are efficient for local relations learning, especially for the high redundancy stages, they inevitably lead to the losses of non-locality and are constrained by the limited receptive field. In this paper, we present an effective hybrid-style vision backbone that is explicitly built with dynamic convolution and self-attention to respectively undertake both local and global interaction, dubbed C2SFormer. We adopt two homogeneous modules whose structure follows the typical Transformers. For local relations learning, we take the parallel multi-scale design and additive aggregation as simple but effective ideas, named MS-SCDC. For the global context modeling, we leverage the efficient factorized self-attention mechanism proposed in CoaT and apply it with the MS-SCDC in a cross-stacking manner over the high-resolution stages. Additionally, we further introduce a general approach for multi-scale learning of transformer-based modules, named MS-MHSA. The experiments conducted on a variety of general-purpose vision tasks demonstrate the superiority of the proposed model. Xili Wan, Yaping Wu, Xinjie Guan, Aichun Zhu |
IJCNN | 3 |
| 2023 | PGTCN: A novel password-guessing model based on temporal convolution network
Yaping Wu, Xili Wan, Xinjie Guan, Tingxiang Ji, Feng Ye 0002 |
J. Netw. Comput. Appl. | 1 |
| 2023 | CSTSUNet: A Cross Swin Transformer-Based Siamese U-Shape Network for Change Detection in Remote Sensing ImagesabstractChange detection (CD) in remote sensing images is a critical task that has achieved significant success by deep learning. Current networks often employ pixel-based differencing, proportion, classification-based, or feature concatenation methods to represent changes of interest. However, these methods fail to effectively detect the desired changes, as they are highly sensitive to factors such as atmospheric conditions, lighting variations, and phenological variations, resulting in detection errors. Inspired by the Transformer structure, we adopt a cross-attention mechanism to more robustly extract feature differences between bitemporal images. The motivation of the method is based on the assumption that if there is no change between image pairs, the semantic features from one temporal image can well be represented by the semantic features from another temporal image. Conversely if there is a change, there are significant reconstruction errors. Therefore, a Cross Swin Transformer based Siamese U-shaped network namely CSTSUNet is proposed for remote sensing change detection. CSTSUnet consists of encoder, difference feature extraction, and decoder. The encoder is based on a hierarchical Resnet with the Siamese U-net structure, allowing parallel processing of bitemporal images and extraction of multi-scale features. The difference feature extraction consists of four difference feature extraction modules that compute difference feature at multiple scales. In this module, Cross Swin Transformer is employed in each difference feature extraction module to communicate the information of bitemporal images. The decoder takes in the multi-scale difference features as input, injects details and boundaries iteratively level by level, and makes the change map more and more accurate. We conduct experiments on three public datasets, and the experimental results demonstrate that the proposed CSTSUNet outperforms other state-of-the-art methods in terms of both qualitative and quantitative analyses. Our code is available at https://github.com/l7170/CSTSUNet.git. Yaping Wu, Lu Li 0005, Nan Wang 0038, Wei Li 0032, Junfang Fan, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | A Two-Branch Neural Network for Short-Axis PET Image Quality EnhancementabstractThe axial field of view (FOV) is a key factor that affects the quality of PET images. Due to hardware FOV restrictions, conventional short-axis PET scanners with FOVs of 20 to 35 cm can acquire only low-quality PET (LQ-PET) images in fast scanning times (2-3 minutes). To overcome hardware restrictions and improve PET image quality for better clinical diagnoses, several deep learning-based algorithms have been proposed. However, these approaches use simple convolution layers with residual learning and local attention, which insufficiently extract and fuse long-range contextual information. To this end, we propose a novel two-branch network architecture with swin transformer units and graph convolution operation, namely SW-GCN. The proposed SW-GCN provides additional spatial- and channel-wise flexibility to handle different types of input information flow. Specifically, considering the high computational cost of calculating self-attention weights in full-size PET images, in our designed spatial adaptive branch, we take the self-attention mechanism within each local partition window and introduce global information interactions between nonoverlapping windows by shifting operations to prevent the aforementioned problem. In addition, the convolutional network structure considers the information in each channel equally during the feature extraction process. In our designed channel adaptive branch, we use a Watts Strogatz topology structure to connect each feature map to only its most relevant features in each graph convolutional layer, substantially reducing information redundancy. Moreover, ensemble learning is adopted in our SW-GCN for mapping distinct features from the two well-designed branches to the enhanced PET images. We carried out extensive experiments on three single-bed position scans for 386 patients. The test results demonstrate that our proposed SW-GCN approach outperforms state-of-the-art methods in both quantitative and qualitative evaluations. Minghan Fu, Yaping Wu, Na Zhang 0001, Yongfeng Yang, Fang-Xiang Wu, Hairong Zheng, Dong Liang 0001, Zhanli Hu |
IEEE J. Biomed. Health Informatics | 3 |