VLDB 2026 Research / reviewers in the wild / expert
Wentao Zhu 0002
dblp:117/0354-2
· DBLP profile ↗
21ranked-venue papers
3as first author
20since 2021 · last 2025
0000-0001-9290-1778ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PD-INR: Prior-Driven Implicit Neural Representations for TOF-PET Reconstruction
Yuxuan Long, Hong Wang 0021, Xiaodong Kuang, Hailiang Huang 0001, Fan Rao, Huafeng Liu 0003, Yefeng Zheng 0001, Wentao Zhu 0002 |
MICCAI (3) | 9 |
| 2025 | A Multimodal Contrastive Learning for Detecting Aortic Dissection on 3D Non-contrast CT with Anatomy Simplification
Duoer Zhang, Yuxuan Qiu, Zhan Feng, Hong Wang 0021, Yefeng Zheng 0001, Wentao Zhu 0002 |
MICCAI (7) | 8 |
| 2025 | Dual-Source CBCT for Large FoV Imaging Under Short-Scan TrajectoriesabstractCone-beam CT is extensively used in medical diagnosis and treatment. Despite its large longitudinal field of view (FoV), the horizontal FoV of CBCT systems is severely limited due to the detector width. Certain commercial CBCT systems increase the horizontal FoV by employing the offset detector method. However, this method necessitates 360° full circular scanning trajectory which increases the scanning time and is not compatible with specific CBCT system models. In this paper, we investigate the feasibility of large FoV imaging under short scan trajectories with an additional X-ray source. A dual-source CBCT geometry is proposed as well as two corresponding image reconstruction algorithms. The first one is based on cone-parallel rebinning and the subsequent employs a modified Parker weighting scheme. Theoretical calculations demonstrate that the proposed geometry achieves a wider horizontal FoV than the ${90}\%$ detector offset geometry (radius of ${214}.{83}\textit {mm}$ vs. ${198}.{99}\textit {mm}$ ) with a significantly reduced rotation angle (less than 230° vs. 360°). As demonstrated by experiments, the proposed geometry and reconstruction algorithms obtain comparable imaging qualities within the FoV to conventional CBCT imaging techniques. Implementing the proposed geometry is straightforward and does not substantially increase development expenses. It possesses the capacity to expand CBCT applications even further. Tianling Lyu, Xinyun Zhong, Zhan Wu, Yan Xi, Wei Zhao 0029, Yang Chen 0008, Yuanjing Feng, Wentao Zhu 0002 |
IEEE Trans. Medical Imaging | 9 |
| 2025 | A Novel Spatio-Temporal Hub Identification in Brain Networks by Learning Dynamic Graph Embedding on Grassmannian ManifoldsabstractMounting evidence has revealed that functional brain networks are intrinsically dynamic, undergoing changes over time, even in the resting-state environment. Notably, recent studies have highlighted the existence of a small number of critical brain regions within each functional brain network that exhibit a flexible role in adapting the geometric pattern of brain connectivity over time, referred to as "temporal hub" regions. Therefore, the identification of these temporal hubs becomes pivotal for comprehending the mechanisms that underlie the dynamic evolution of brain connectivity. However, existing spatio-temporal hub identification methods rely on static network-based approaches, wherein each temporal hub region is independently inferred from individual time-segmented networks without considering their temporal consistency and consequently fails to align the evolution of hubs with the dynamic changes in brain states. To address this limitation, we propose a novel spatio-temporal hub identification method that fully leverages dynamic graph embedding to distinguish temporal hubs from peripheral nodes, in which dynamic graph embeddings are learned from both spatial and temporal dimensions. Specifically, to preserve the temporal consistency of evolving networks, we model the dynamic graph embedding as a physical model of time, where the network-to-network transition is mathematically expressed as a total variation of dynamic graph embedding with respect to time. Furthermore, a Grassmannian manifold optimization scheme is introduced to enhance graph embedding learning and capture the time-varying topology of brain networks. Experimental results on both synthetic and real fMRI data demonstrate superior temporal consistency in hub identification, surpassing conventional approaches. Defu Yang, Minghan Chen 0001, Shuai Wang 0003, Jiazhou Chen 0001, Hongmin Cai, Guorong Wu 0001, Wentao Zhu 0002 |
IEEE Trans. Medical Imaging | 9 |
| 2025 | Hierarchical Token-Aware Cross-Modality Reconstruction for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) aims to query the same pedestrian's visible (infrared) images in the gallery set from the infrared (visible) images. VI-ReID not only needs to deal with the challenging factors like pose variation and occlusion, but also requires handling the large modality discrepancy. Previous methods mainly focus on learning single-scale modality-shared features and do not effectively explore the multi-scale features of two modalities from both short-range and long-range perspectives. In order to solve these problems, this paper proposes a novel Hierarchical Token-Aware Cross-Modality Reconstruction (HTCR) network to significantly mitigate the modality discrepancy for effective VI-ReID. The HTCR network consists of two main components, i.e., Hierarchical Token-aware Fusion (HTF) and Cross-modality Feature Reconstruction (CFR). The HTF module first bidirectionally exchanges the short-range and long-range multi-scale modality-shared features with a few learnable tokens to achieve discriminative pedestrian features by making full use of the advantages of both Convolutional Neural Network (CNN) and Transformer. Moreover, the CFR module reconstructs global and local pedestrian features of one modality by using the token sequence of the other modality with multi-scale cues to further explore the relationship between the two distinct modalities and alleviate the modality discrepancy. In addition, the Modality-shared feature Reconstruction (MR) loss is leveraged to reduce the noises between the reconstructed and the target features. Experimental results indicate that the proposed HTCR can significantly improve the VI-ReID performance and outperform the state-of-the-art methods on the cross-modality SYSU-MM01, RegDB, and LLCM datasets. Si Chen 0002, Liuxiang Qiu, Dahan Wang, Wentao Zhu 0002, Yang Hua 0001, Yan Yan 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | MedSegViG: Medical Image Segmentation with a Vision Graph Neural NetworkabstractMedical image segmentation is a crucial step toward automatic clinical diagnosis, which has received growing interest. Although some existing methods based on convolutional neural networks or transformers have achieved remarkable success in this task, they still show limitations in effectively modeling the relationships among different objects in images. In this paper, we propose a novel deep learning based model to address this issue by leveraging a vision graph neural network (ViG). Our model, MedSegViG, mainly consists of a hierarchical ViG encoder and a lightweight convolutional decoder. The hierarchical encoder extracts multi-level features from the image and captures the object relationships with graph neural networks. The lightweight decoder then fuses these features and generates the corresponding segmentation map. Extensive experiments are conducted on seven datasets for three typical medical image segmentation tasks: polyp segmentation, skin lesion segmentation, and retinal vessel segmentation. The results demonstrate the superiority of our MedSegViG over state-of-the-art models across various tasks and datasets. The code is released on https://github.com/Xinhong-Li/MedSegViG. Geng Chen 0001, Yuanfeng Wu, Junqing Yang, Tao Zhou 0002, Yi Zhou 0007, Wentao Zhu 0002 |
BIBM | 7 |
| 2024 | D-NAF: Dynamic Neural Attenuation Fields for 4D CBCT Reconstruction in Pulmonary ImagingabstractFour-dimensional Cone-beam CT (4D-CBCT) has already been integrated into many commercial radiotherapy systems to facilitate image-guided radiotherapy (IGRT). These techniques reconstruct a sequence of three-dimensional (3D) CBCT volumes based on respiratory phases, providing motion-compensated images. Nevertheless, current 4D-CBCT methods still suffer from poor imaging quality due to the limited number of views used for phase-resolved reconstruction, and deep learning-based methods require high-quality training data, which is commonly unavailable. In this paper, we propose a self-supervised 4D-CBCT imaging method based on dynamic neural attenuation fields (D-NAF). The dynamic pulmonary CBCT images are encoded into a 4D implicit representation incorporating both spatial information and temporal information. Moreover, we employed a separable spatial-temporal encoding method to reduce the dimensionality of the solution space, leading to enhanced convergence and reconstruction quality. The results demonstrate that the proposed method yields superior imaging quality in comparison to alternative approaches, with an RMSE value of 59.82 HU, PSNR of 35.33 dB and SSIM of 0.9358. This method is expected to be a new benchmark for self-supervised 4D-CBCT imaging. Yuxuan Long, Tianling Lyu, Fan Rao, Yang Chen 0008, Wentao Zhu 0002 |
BIBM | 6 |
| 2024 | Classification of lung cancer subtypes on CT images with synthetic pathological priors
Wentao Zhu 0002, Gege Ma, Geng Chen 0001, Jan Egger, Shaoting Zhang 0001, Dimitris N. Metaxas |
Medical Image Anal. | 1 |
| 2024 | DA-Tran: Multiphase liver tumor segmentation with a domain-adaptive transformer network
Yangfan Ni, Geng Chen 0001, Zhan Feng, Heng Cui, Dimitris N. Metaxas, Shaoting Zhang 0001, Wentao Zhu 0002 |
Pattern Recognit. | 7 |
| 2024 | SCRN: Single-Cell Gene Regulatory Network Identification in Alzheimer's DiseaseabstractAlzheimer's disease (AD) is the most common neurodegenerative disease, and it consumes considerable medical resources with increasing number of patients every year. Mounting evidence show that the regulatory disruptions altering the intrinsic activity of genes in brain cells contribute to AD pathogenesis. To gain insights into the underlying gene regulation in AD, we proposed a graph learning method, Single-Cell based Regulatory Network (SCRN), to identify the regulatory mechanisms based on single-cell data. SCRN implements the γ-decaying heuristic link prediction based on graph neural networks and can identify reliable gene regulatory networks using locally closed subgraphs. In this work, we first performed UMAP dimension reduction analysis on single-cell RNA sequencing (scRNA-seq) data of AD and normal samples. Then we used SCRN to construct the gene regulatory network based on three well-recognized AD genes (APOE, CX3CR1, and P2RY12). Enrichment analysis of the regulatory network revealed significant pathways including NGF signaling, ERBB2 signaling, and hemostasis. These findings demonstrate the feasibility of using SCRN to uncover potential biomarkers and therapeutic targets related to AD. Wentao Zhu 0002, Ziang Xu 0002, Defu Yang, Minghan Chen 0001, Qianqian Song 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2024 | Interpretable Heterogeneous Teacher-Student Learning Framework for Hybrid-Supervised Pulmonary Nodule DetectionabstractExisting pulmonary nodule detection methods often train models in a fully-supervised setting that requires strong labels (i.e., bounding box labels) as label information. However, manual annotation of bounding boxes in CT images is very time-consuming and labor-intensive. To alleviate the annotation burden, in this paper, we investigate pulmonary nodule detection by leveraging both strong labels and weak labels (i.e., center point labels) for training, and propose a novel hybrid-supervised pulmonary nodule detection (HND) method. The training of HND involves a heterogeneous teacher-student learning framework in two stages. In the first stage, we design a point-based consistency calibration network (PCC-Net) as a teacher, which is pre-trained to generate high-quality pseudo bounding box labels given point-augmented CT images as inputs. In the second stage, we develop an information bottleneck-guided pulmonary nodule detection network (IBD-Net) as a student to perform pulmonary nodule detection. In particular, we introduce information bottleneck to learn reliable pulmonary nodule-specific heatmaps under the guidance of PCC-Net, largely enhancing the model’s interpretability and improving the final detection performance. Based on the above designs, our method can effectively detect pulmonary nodule regions with only a limited number of bounding box labels. Experimental results on the public pulmonary nodule detection dataset LUNA16 show that our HND method achieves an excellent balance between the annotation cost and the detection performance. Guangyu Huang, Yan Yan 0001, Jing-Hao Xue, Wentao Zhu 0002, Xióngbiao Luó |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Anatomically Guided PET Image Reconstruction Using Conditional Weakly-Supervised Multi-Task Learning Integrating Self-AttentionabstractTo address the lack of high-quality training labels in positron emission tomography (PET) imaging, weakly-supervised reconstruction methods that generate network-based mappings between prior images and noisy targets have been developed. However, the learned model has an intrinsic variance proportional to the average variance of the target image. To suppress noise and improve the accuracy and generalizability of the learned model, we propose a conditional weakly-supervised multi-task learning (MTL) strategy, in which an auxiliary task is introduced serving as an anatomical regularizer for the PET reconstruction main task. In the proposed MTL approach, we devise a novel multi-channel self-attention (MCSA) module that helps learn an optimal combination of shared and task-specific features by capturing both local and global channel-spatial dependencies. The proposed reconstruction method was evaluated on NEMA phantom PET datasets acquired at different positions in a PET/CT scanner and 26 clinical whole-body PET datasets. The phantom results demonstrate that our method outperforms state-of-the-art learning-free and weakly-supervised approaches obtaining the best noise/contrast tradeoff with a significant noise reduction of approximately 50.0% relative to the maximum likelihood (ML) reconstruction. The patient study results demonstrate that our method achieves the largest noise reductions of 67.3% and 35.5% in the liver and lung, respectively, as well as consistently small biases in 8 tumors with various volumes and intensities. In addition, network visualization reveals that adding the auxiliary task introduces more anatomical information into PET reconstruction than adding only the anatomical loss, and the developed MCSA can abstract features and retain PET image details. Bao Yang, Kuang Gong, Huafeng Liu 0003, Quanzheng Li, Wentao Zhu 0002 |
IEEE Trans. Medical Imaging | 5 |
| 2023 | Dichotomous Image Segmentation with Frequency PriorsabstractDichotomous image segmentation (DIS) has a wide range of real-world applications and gained increasing research attention in recent years. In this paper, we propose to tackle DIS with informative frequency priors. Our model, called FP-DIS, stems from the fact that prior knowledge in the frequency domain can provide valuable cues to identify fine-grained object boundaries. Specifically, we propose a frequency prior generator to jointly utilize a fixed filter and learnable filters to extract informative frequency priors. Before embedding the frequency priors into the network, we first harmonize the multi-scale side-out features to reduce their heterogeneity. This is achieved by our feature harmonization module, which is based on a gating mechanism to harmonize the grouped features. Finally, we propose a frequency prior embedding module to embed the frequency priors into multi-scale features through an adaptive modulation strategy. Extensive experiments on the benchmark dataset, DIS5K, demonstrate that our FP-DIS outperforms state-of-the-art methods by a large margin in terms of key evaluation metrics. Bo Dong 0001, Yuanfeng Wu, Wentao Zhu 0002, Geng Chen 0001, Yanning Zhang 0001 |
IJCAI | 4 |
| 2023 | Spatiotemporal Hub Identification in Brain Network by Learning Dynamic Graph Embedding on Grassmannian Manifold
Defu Yang, Minghan Chen 0001, Yitian Xue, Shuai Wang 0003, Guorong Wu 0001, Wentao Zhu 0002 |
MICCAI (2) | 7 |
| 2023 | Learning pyramidal multi-scale harmonic wavelets for identifying the neuropathology propagation patterns of Alzheimer's disease
Huan Liu 0017, Hongmin Cai, Defu Yang, Wentao Zhu 0002, Guorong Wu 0001, Jiazhou Chen 0001 |
Medical Image Anal. | 4 |
| 2023 | scENT for Revealing Gene Clusters From Single-Cell RNA-Seq DataabstractRecently, the fast development of single-cell RNA-seq (scRNA-seq) techniques has enabled high-resolution transcriptomic statistical analysis of individual cells in heterogeneous tissues, which can help researchers to explore the relationship between genes and human diseases. The emerging scRNA-seq data results in new analysis methods aiming to identify cell-level clustering and annotations. However, there are few methods developed to gain insights into the gene-level clusters with biological significance. This study proposes a new deep learning-based framework, scENT (single cell gENe clusTer), to identify significant gene clusters from single-cell RNA-seq data. We started with clustering the scRNA-seq data into multiple optimal groups, followed by a gene set enrichment analysis to identify classes of over-represented genes. Considering high-dimensional data with extensive zeros and dropout issues, scENT integrates perturbation in the learning process of clustering scRNA-seq data to improve its robustness and performance. Experimental results show that scENT outperformed other benchmarking methods on simulation data. To validate the biological insights of scENT, we applied it to the public experimental scRNA-seq data profiled from patients with Alzheimer's disease and brain metastasis. scENT successfully identified novel functional gene clusters and associated functions, facilitating the discovery of prospective mechanisms and the understanding of related diseases. Fan Rao, Minghan Chen 0001, Defu Yang, Bess Morrell, Qianqian Song 0002, Wentao Zhu 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2023 | Drop Loss for Person Attribute Recognition With Imbalanced Noisy-Labeled SamplesabstractPerson attribute recognition (PAR) aims to simultaneously predict multiple attributes of a person. Existing deep learning-based PAR methods have achieved impressive performance. Unfortunately, these methods usually ignore the fact that different attributes have an imbalance in the number of noisy-labeled samples in the PAR training datasets, thus leading to suboptimal performance. To address the above problem of imbalanced noisy-labeled samples, we propose a novel and effective loss called drop loss for PAR. In the drop loss, the attributes are treated differently in an easy-to-hard way. In particular, the noisy-labeled candidates, which are identified according to their gradient norms, are dropped with a higher drop rate for the harder attribute. Such a manner adaptively alleviates the adverse effect of imbalanced noisy-labeled samples on model learning. To illustrate the effectiveness of the proposed loss, we train a simple ResNet-50 model based on the drop loss and term it DropNet. Experimental results on two representative PAR tasks (including facial attribute recognition and pedestrian attribute recognition) demonstrate that the proposed DropNet achieves comparable or better performance in terms of both balanced accuracy and classification accuracy over several state-of-the-art PAR methods. Yan Yan 0001, Youze Xu, Jing-Hao Xue, Yang Lu 0009, Hanzi Wang, Wentao Zhu 0002 |
IEEE Trans. Cybern. | 6 |
| 2023 | UPL-SFDA: Uncertainty-Aware Pseudo Label Guided Source-Free Domain Adaptation for Medical Image SegmentationabstractDomain Adaptation (DA) is important for deep learning-based medical image segmentation models to deal with testing images from a new target domain. As the source-domain data are usually unavailable when a trained model is deployed at a new center, Source-Free Domain Adaptation (SFDA) is appealing for data and annotation-efficient adaptation to the target domain. However, existing SFDA methods have a limited performance due to lack of sufficient supervision with source-domain images unavailable and target-domain images unlabeled. We propose a novel Uncertainty-aware Pseudo Label guided (UPL) SFDA method for medical image segmentation. Specifically, we propose Target Domain Growing (TDG) to enhance the diversity of predictions in the target domain by duplicating the pre-trained model's prediction head multiple times with perturbations. The different predictions in these duplicated heads are used to obtain pseudo labels for unlabeled target-domain images and their uncertainty to identify reliable pseudo labels. We also propose a Twice Forward pass Supervision (TFS) strategy that uses reliable pseudo labels obtained in one forward pass to supervise predictions in the next forward pass. The adaptation is further regularized by a mean prediction-based entropy minimization term that encourages confident and consistent results in different prediction heads. UPL-SFDA was validated with a multi-site heart MRI segmentation dataset, a cross-modality fetal brain segmentation dataset, and a 3D fetal tissue segmentation dataset. It improved the average Dice by 5.54, 5.01 and 6.89 percentage points for the three tasks compared with the baseline, respectively, and outperformed several state-of-the-art SFDA methods. Jianghao Wu 0001, Guotai Wang, Ran Gu, Wentao Zhu 0002, Tom Vercauteren, Sébastien Ourselin, Shaoting Zhang 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2022 | MSFL-Net: Multi-Semantic Feature Learning Network for Occluded Person Re-IdentificationabstractRecently, occluded person re-identification (Re-ID) has received significant interest due to its widespread real-world applications. However, most existing occluded person Re-ID methods ignore semantic granularities that indicate different levels of occluded information of the human body, leading to sub-optimal performance. To address this, we propose a Multi-Semantic Feature Learning Network (MSFL-Net) for occluded person Re-ID. Specifically, MSFL-Net involves a backbone network and a Multi-branch Feature Learning sub-network (MFL). MFL consists of two local-global branches and a global branch to learn multisemantic features in a multi-branch deep network architecture. In each local-global branch, we design a local subbranch and a semantic-guided global sub-branch to extract discriminative features at a certain level of feature granularity and semantic granularity. In the global branch, we learn global features at the largest level of semantic granularity. In particular, a patch contrastive loss is developed to explicitly encourage the semantic feature maps to capture the information from specific body parts. By extracting multi-semantic features, our method is effective in dealing with person Re-ID at different occlusion levels. Experimental results on an occluded person Re-ID dataset (Occluded-REID) and two partial person Re-ID datasets (Partial-iLIDS and Partial-REID) show the superiority of our method against state-of-the-art person Re-ID methods. Guangyu Huang, Yan Yan 0001, Si Chen 0002, Wentao Zhu 0002, Hanzi Wang |
IJCB | 5 |
| 2022 | A Graph Convolutional Multiple Instance Learning on a Hypersphere Manifold Approach for Diagnosing Chronic Obstructive Pulmonary Disease in CT ImagesabstractChronic obstructive pulmonary disease (COPD) is a prevalent chronic disease with high morbidity and mortality. The early diagnosis of COPD is vital for clinical treatment, which helps patients to have a better quality of life. Because COPD can be ascribed to chronic bronchitis and emphysema, lesions in a computed tomography (CT) image can present anywhere inside the lung with different types, shapes and sizes. Multiple instance learning (MIL) is an effective tool for solving COPD discrimination. In this study, a novel graph convolutional MIL with the adaptive additive margin loss (GCMIL-AAMS) approach is proposed to diagnose COPD by CT. Specifically, for those early stage patients, the selected instance-level features can be more discriminative if they were learned by our proposed graph convolution and pooling with self-attention mechanism. The AAMS loss can utilize the information of COPD severity on a hypersphere manifold by adaptively setting the angular margins to improve the performance, as the severity can be quantified as four grades by pulmonary function test. The results show that our proposed GCMIL-AAMS method provides superior discrimination and generalization abilities in COPD discrimination, with areas under a receiver operating characteristic curve (AUCs) of 0.960 ± 0.014 and 0.862 ± 0.010 in the test set and external testing set, respectively, in 5-fold stratified cross validation; moreover, it demonstrates that graph learning is applicable to MIL and suggests that MIL may be adaptable to graph learning. Qixing Feng, Xi Yin 0009, Xiangde Min, Defu Yang, Yen-Wei Chen 0001, Daoqiang Zhang, Wentao Zhu 0002 |
IEEE J. Biomed. Health Informatics | 9 |
| 2014 | Patlak Image Estimation From Dual Time-Point List-Mode PET DataabstractWe investigate using dual time-point PET data to perform Patlak modeling. This approach can be used for whole body dynamic PET studies in which we compute voxel-wise estimates of Patlak parameters using two frames of data for each bed position. Our approach directly uses list-mode arrival times for each event to estimate the Patlak parametric image. We use a penalized likelihood method in which the penalty function uses spatially variant weighting to ensure a count independent local impulse response. We evaluate performance of the method in comparison to fractional changes in SUV values (%DSUV) between the two frames using Cramer Rao analysis and Monte Carlo simulation. Receiver operating characteristic (ROC) curves are used to compare performance in differentiating tumors relative to background based on the dynamic data sets. Using area under the ROC curve as a performance metric, we show superior performance of Patlak relative to %DSUV over a range of dynamic data sets and parameters. These results suggest that Patlak analysis may be appropriate for analysis of dual time-point whole body PET data and could lead to superior detection of tumors relative to %DSUV metrics. Wentao Zhu 0002, Quanzheng Li, Peter S. Conti, Richard M. Leahy |
IEEE Trans. Medical Imaging | 1 |