Pengyu Wang 0005

dblp:233/6773-5 · DBLP profile ↗
← Back
20ranked-venue papers
9as first author
16since 2021 · last 2024
0000-0003-0997-9887ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2024 Medical Image Synthesis via Fine-Grained Image-Text Alignment and Anatomy-Pathology Prompting
Wenting Chen, Pengyu Wang 0005, Hui Ren 0001, Lichao Sun 0001, Quanzheng Li, Yixuan Yuan, Xiang Li 0001
MICCAI (12)2
2024 F2TNet: FMRI to T1w MRI Knowledge Transfer Network for Brain Multi-phenotype Prediction
Wuyang Li, Yu Jiang 0013, Zhihao Peng 0002, Pengyu Wang 0005, Xiang Li 0001, Tianming Liu 0001, Junwei Han 0001, Yixuan Yuan
MICCAI (11)5
2024 GBT: Geometric-Oriented Brain Transformer for Autism Diagnosis
Zhihao Peng 0002, Yu Jiang 0013, Pengyu Wang 0005, Yixuan Yuan
MICCAI (12)4
2024 fTSPL: Enhancing Brain Analysis with FMRI-Text Synergistic Prompt Learning
Pengyu Wang 0005, Huaqi Zhang, Zhihao Peng 0002, Yixuan Yuan
MICCAI (12)1
2024 Graph-based multi-source domain adaptation with contrastive and collaborative learning for image deraining
Pengyu Wang 0005, Hongqing Zhu, Huaqi Zhang, Ning Chen 0007, Suyi Yang
Eng. Appl. Artif. Intell.1
2024 LRB-T: local reasoning back-projection transformer for the removal of bad weather effects in images
Pengyu Wang 0005, Hongqing Zhu, Huaqi Zhang, Suyi Yang
Neural Comput. Appl.1
2024 MHD-Net: Memory-Aware Hetero-Modal Distillation Network for Thymic Epithelial Tumor Typing With Missing Pathology Modality
abstract
Fusing multi-modal radiology and pathology data with complementary information can improve the accuracy of tumor typing. However, collecting pathology data is difficult since it is high-cost and sometimes only obtainable after the surgery, which limits the application of multi-modal methods in diagnosis. To address this problem, we propose comprehensively learning multi-modal radiology-pathology data in training, and only using uni-modal radiology data in testing. Concretely, a Memory-aware Hetero-modal Distillation Network (MHD-Net) is proposed, which can distill well-learned multi-modal knowledge with the assistance of memory from the teacher to the student. In the teacher, to tackle the challenge in hetero-modal feature fusion, we propose a novel spatial-differentiated hetero-modal fusion module (SHFM) that models spatial-specific tumor information correlations across modalities. As only radiology data is accessible to the student, we store pathology features in the proposed contrast-boosted typing memory module (CTMM) that achieves type-wise memory updating and stage-wise contrastive memory boosting to ensure the effectiveness and generalization of memory items. In the student, to improve the cross-modal distillation, we propose a multi-stage memory-aware distillation (MMD) scheme that reads memory-aware pathology features from CTMM to remedy missing modal-specific information. Furthermore, we construct a Radiology-Pathology Thymic Epithelial Tumor (RPTET) dataset containing paired CT and WSI images with annotations. Experiments on the RPTET and CPTAC-LUAD datasets demonstrate that MHD-Net significantly improves tumor typing and outperforms existing multi-modal methods on missing modality situations.
Huaqi Zhang, Jie Liu 0044, Weifan Liu, Zekuan Yu, Yixuan Yuan, Pengyu Wang 0005, Harry Qin
IEEE J. Biomed. Health Informatics7
2024 MCPL: Multi-Modal Collaborative Prompt Learning for Medical Vision-Language Model
abstract
Multi-modal prompt learning is a high-performance and cost-effective learning paradigm, which learns text as well as image prompts to tune pre-trained vision-language (V-L) models like CLIP for adapting multiple downstream tasks. However, recent methods typically treat text and image prompts as independent components without considering the dependency between prompts. Moreover, extending multi-modal prompt learning into the medical field poses challenges due to a significant gap between general- and medical-domain data. To this end, we propose a Multi-modal Collaborative Prompt Learning (MCPL) pipeline to tune a frozen V-L model for aligning medical text-image representations, thereby achieving medical downstream tasks. We first construct the anatomy-pathology (AP) prompt for multi-modal prompting jointly with text and image prompts. The AP prompt introduces instance-level anatomy and pathology information, thereby making a V-L model better comprehend medical reports and images. Next, we propose graph-guided prompt collaboration module (GPCM), which explicitly establishes multi-way couplings between the AP, text, and image prompts, enabling collaborative multi-modal prompt producing and updating for more effective prompting. Finally, we develop a novel prompt configuration scheme, which attaches the AP prompt to the query and key, and the text/image prompt to the value in self-attention layers for improving the interpretability of multi-modal prompts. Extensive experiments on numerous medical classification and object detection datasets show that the proposed pipeline achieves excellent effectiveness and generalization. Compared with state-of-the-art prompt learning methods, MCPL provides a more reliable multi-modal prompt paradigm for reducing tuning costs of V-L models on medical downstream tasks. Our code: https://github.com/CUHK-AIM-Group/MCPL.
Pengyu Wang 0005, Huaqi Zhang, Yixuan Yuan
IEEE Trans. Medical Imaging1
2024 MGIML: Cancer Grading With Incomplete Radiology-Pathology Data via Memory Learning and Gradient Homogenization
abstract
Taking advantage of multi-modal radiology-pathology data with complementary clinical information for cancer grading is helpful for doctors to improve diagnosis efficiency and accuracy. However, radiology and pathology data have distinct acquisition difficulties and costs, which leads to incomplete-modality data being common in applications. In this work, we propose a Memory- and Gradient-guided Incomplete Modal-modal Learning (MGIML) framework for cancer grading with incomplete radiology-pathology data. Firstly, to remedy missing-modality information, we propose a Memory-driven Hetero-modality Complement (MH-Complete) scheme, which constructs modal-specific memory banks constrained by a coarse-grained memory boosting (CMB) loss to record generic radiology and pathology feature patterns, and develops a cross-modal memory reading strategy enhanced by a fine-grained memory consistency (FMC) loss to take missing-modality information from well-stored memories. Secondly, as gradient conflicts exist between missing-modality situations, we propose a Rotation-driven Gradient Homogenization (RG-Homogenize) scheme, which estimates instance-specific rotation matrices to smoothly change the feature-level gradient directions, and computes confidence-guided homogenization weights to dynamically balance gradient magnitudes. By simultaneously mitigating gradient direction and magnitude conflicts, this scheme well avoids the negative transfer and optimization imbalance problems. Extensive experiments on CPTAC-UCEC and CPTAC-PDA datasets show that the proposed MGIML framework performs favorably against state-of-the-art multi-modal methods on missing-modality situations.
Pengyu Wang 0005, Huaqi Zhang, Meilu Zhu, Xi Jiang 0001, Harry Qin, Yixuan Yuan
IEEE Trans. Medical Imaging1
2023 Fire detection in video surveillance using superpixel-based region proposal and ESE-ShuffleNet
Pengyu Wang 0005, Jianmei Zhang, Hongqing Zhu
Multim. Tools Appl.1
2023 M-CBN: Manifold constrained joint image dehazing and super-resolution based on chord boosting network
Pengyu Wang 0005, Hongqing Zhu, Han Zhang 0053, Nan Wang 0003
Pattern Recognit.1
2022 GA-SRN: graph attention based text-image semantic reasoning network for fine-grained image classification and retrieval
Hongqing Zhu, Suyi Yang, Pengyu Wang 0005, Han Zhang 0053
Neural Comput. Appl.4
2022 TMS-GAN: A Twofold Multi-Scale Generative Adversarial Network for Single Image Dehazing
abstract
In recent years, learning-based single image dehazing networks have been comprehensively developed. However, performance improvement is limited due to domain shift between trained synthetic hazy images and untrained real-world hazy images. To alleviate this issue, this paper proposes a real-world dehazing targeted training scheme which nearly realizes paired real-world data training. As a result, a Twofold Multi-scale Generative Adversarial Network (TMS-GAN) consisting of a Haze-generation GAN (HgGAN) and a Haze-removal GAN (HrGAN) is designed. HgGAN attributes real haze properties to synthetic images and HrGAN removes haze from both synthetic and generated fake realistic data under supervision. Thus, the proposed method can better adapt to real-world image dehazing using this cooperative training scheme. Meanwhile, several structural advances of TMS-GAN also improve dehazing performance. Specifically, a haze residual map based on atmospheric scattering model is deduced in HgGAN for fake realistic data generation. The dual-branch generator in HrGAN draws attention to detail restoration by one branch along with another color-branch. A plug-and-play Multi-attention Progressive Fusion Module (MAPFM) is proposed and inserted in both HgGAN and HrGAN. MAPFM incorporates multi-attention mechanism to guide multi-scale feature fusion in a progressive manner, in which Adjacency-attention Block (AAB) can capture contributing features of each level and Self-attention Block (SAB) can establish non-local dependency of feature fusion. Experiments on mainstream benchmarks show that the proposed framework is superior especially on real-world hazy images among single image dehazing methods.
Pengyu Wang 0005, Hongqing Zhu, Han Zhang 0053, Nan Wang 0003
IEEE Trans. Circuits Syst. Video Technol.1
2022 Cross-Boosted Multi-Target Domain Adaptation for Multi-Modality Histopathology Image Translation and Segmentation
abstract
Recent digital pathology workflows mainly focus on mono-modality histopathology image analysis. However, they ignore the complementarity between Haematoxylin & Eosin (H&E) and Immunohistochemically (IHC) stained images, which can provide comprehensive gold standard for cancer diagnosis. To resolve this issue, we propose a cross-boosted multi-target domain adaptation pipeline for multi-modality histopathology images, which contains Cross-frequency Style-auxiliary Translation Network (CSTN) and Dual Cross-boosted Segmentation Network (DCSN). Firstly, CSTN achieves the one-to-many translation from fluorescence microscopy images to H&E and IHC images for providing source domain training data. To generate images with realistic color and texture, Cross-frequency Feature Transfer Module (CFTM) is developed to pertinently restructure and normalize high-frequency content and low-frequency style features from different domains. Then, DCSN fulfills multi-target domain adaptive segmentation, where a dual-branch encoder is introduced, and Bidirectional Cross-domain Boosting Module (BCBM) is designed to implement cross-modality information complementation through bidirectional inter-domain collaboration. Finally, we establish Multi-modality Thymus Histopathology (MThH) dataset, which is the largest publicly available H&E and IHC image benchmark. Experiments on MThH dataset and several public datasets show that the proposed pipeline outperforms state-of-the-art methods on both histopathology image translation and segmentation.
Huaqi Zhang, Jie Liu 0044, Pengyu Wang 0005, Zekuan Yu, Weifan Liu
IEEE J. Biomed. Health Informatics3
2021 Single-image de-raining using joint filter and multi-scale deep alternate-connection dense network
Pengyu Wang 0005, Hongqing Zhu
Neurocomputing1
2021 MASG-GAN: A multi-view attention superpixel-guided generative adversarial network for efficient and simultaneous histopathology image segmentation and classification
Huaqi Zhang, Jie Liu 0044, Zekuan Yu, Pengyu Wang 0005
Neurocomputing4
2020 Order Fulfillment Cycle Time Estimation for On-Demand Food Delivery
abstract
By providing customers with conveniences such as easy access to an extensive variety of restaurants, effortless food ordering and fast delivery, on-demand food delivery (OFD) platforms have achieved explosive growth in recent years. A crucial machine learning task performed at OFD platforms is prediction of the Order Fulfillment Cycle Time (OFCT), which refers to the amount of time elapsed between a customer places an order and he/she receives the meal. The accuracy of predicted OFCT is important for customer satisfaction, as it needs to be communicated to a customer before he/she places the order, and is considered as a service promise that should be fulfilled as well as possible. As a result, the estimated OFCT also heavily influences planning decisions such as dispatching and routing.
Kairong Zhou, Wenxing Feng, Pengyu Wang 0005, Ning Chen 0007, Pei Lee
KDD6
2020 Image clustering algorithm using superpixel segmentation and non-symmetric Gaussian-Cauchy mixture model
abstract
In this study, an unsupervised clustering algorithm is proposed to label superpixel density images. Firstly, the authors propose a novel superpixel segmentation algorithm driven by a modified fuzzy C‐means objective function, Kullback–Leibler (KL) divergence, and an entropy term, which generate superpixels with good boundary adherence and intensity homogeneity. In this model, the logarithm of Gaussian distribution as a new distance metric is used to improve the accuracy of boundary pixel classification, the KL divergence is applied to regularise the fuzzy objective function. Based on this model, the generated superpixel intensity images with a highly distinctive background colour from the colour of the target are obtained. Grouping cues generated by superpixels can affect the performance of image clustering greatly. Next, according to the small amount of clustering data generated by the superpixel intensity images, they construct a non‐symmetric mixture model based on a mixture of Gaussian distribution and Cauchy distribution for implementing image clustering. Thus, clustering of colour images is transformed into clustering of these newly generated data. The advantage of this model is its well adaption to different shapes of observed data. Experimental results on publicly available data sets are provided to demonstrate the effectiveness of the proposed algorithm.
Sifan Ji, Hongqing Zhu, Pengyu Wang 0005, Xiaofeng Ling
IET Image Process.3
2019 Segmentation of Overlapping Cervical Smear Cells Based on U-Net and Improved Level Set
abstract
Full convolution network (FCN) is widely used in medical image segmentation and its performance is better than other conventional techniques. This paper proposes a new fusion algorithm that combined the convolutional neural network U-net with a new modified level set method to segment overlapping cervical smear cells. U-net could provide more excellent segmentation results of nuclei and cytoplasm cluster. Then, a modified level set energy function with distance map and a new shape prior term is applied to extract the contour of cervical cells. Owing to this new level set energy function, the segmentation of every individual cell performed well, especially in overlapping area of cells. The evaluation of results also proves the improvement of our fusion algorithm.
Hongqing Zhu, Pengyu Wang 0005, Deping Dong
SMC3
2018 A Motion Artifact Reduction Method in Cerebrovascular DSA Sequence Images
abstract
Digital Subtraction Angiography (DSA) can be used for diagnosing the pathologies of vascular system including systemic vascular disease, coronary heart disease, arrhythmia, valvular disease and congenital heart disease. Previous studies have provided some image enhancement algorithms for DSA images. However, these studies are not suitable for automated processes in huge amounts of data. Furthermore, few algorithms solved the problems of image contrast corruption after artifact removal. In this paper, we propose a fully automatic method for cerebrovascular DSA sequence images artifact removal based on rigid registration and guided filter. The guided filtering method is applied to fuse the original DSA image and registered DSA image, the results of which preserve clear vessel boundary from the original DSA image and remove the artifacts by the registered procedure. The experimental evaluation with 40 DSA sequence images shows that the proposed method increases the contrast index by 24.1% for improving the quality of DSA images compared with other image enhancement methods, and can be implemented as a fully automatic procedure.
Guanglei Wang 0002, Pengyu Wang 0005, Yan Li 0053, Tianqi Su, Hongrui Wang 0002
Int. J. Pattern Recognit. Artif. Intell.2