EDBT 2026 Demo / reviewers in the wild / expert
Wentao Zhu 0001
dblp:117/0354-1
· DBLP profile ↗
21ranked-venue papers
8as first author
14since 2021 · last 2026
0000-0002-7505-9512ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 7 first-author · 7 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Restoration Adaptation for Semantic Segmentation on Low Quality ImagesabstractAbstract In real-world scenarios, the performance of semantic segmentation often deteriorates when processing low-quality (LQ) images, which may lack clear semantic structures and high-frequency details. Although image restoration techniques offer a promising direction for enhancing degraded visual content, conventional real-world image restoration (Real-IR) models primarily focus on pixel-level fidelity and often fail to recover task-relevant semantic cues, limiting their effectiveness when directly applied to downstream vision tasks. Conversely, existing segmentation models trained on high-quality data lack robustness under real-world degradations. In this paper, we propose Restoration Adaptation for Semantic Segmentation (RASS), which effectively integrates semantic image restoration into the segmentation process, enabling high-quality semantic segmentation on the LQ images directly. Specifically, we first propose a Semantic-Constrained Restoration (SCR) model, which injects segmentation priors into the restoration model by aligning its cross-attention maps with segmentation masks, encouraging semantically faithful image reconstruction. Then, RASS transfers semantic restoration knowledge into segmentation through LoRA-based module merging and task-specific fine-tuning, thereby enhancing the model’s robustness to LQ images. To validate the effectiveness of our framework, we construct a real-world LQ image segmentation dataset with high-quality annotations, and conduct extensive experiments on both synthetic and real-world LQ benchmarks. The results show that SCR and RASS significantly outperform state-of-the-art methods in segmentation and restoration tasks. Code, models, and datasets will be available at https://github.com/Ka1Guan/RASS.git . Rongyuan Wu, Shuai Li 0014, Wentao Zhu 0001, Wenjun Zeng 0001, Lei Zhang 0006 |
Int. J. Comput. Vis. | 4 |
| 2023 | Towards Comprehensive Monocular Depth Estimation: Multiple Heads are Better Than OneabstractDepth estimation attracts widespread attention in the computer vision community. However, it is still quite difficult to recover an accurate depth map using only one RGB image. We observe a phenomenon that existing methods tend to fail in different cases, caused by differences in network architecture, loss function and so on. In this work, we investigate into the phenomenon and propose to integrate the strengths of multiple weak depth predictor to build a comprehensive and accurate depth predictor, which is critical for many real-world applications, e.g., 3D reconstruction. Specifically, we construct multiple base (weak) depth predictors by utilizing different Transformer-based and convolutional neural network (CNN)-based architectures. Transformer establishes long-range correlation while CNN preserves local information ignored by Transformer due to the spatial inductive bias. Therefore, the coupling of Transformer and CNN contributes to the generation of complementary depth estimates, which are essential to achieve a comprehensive depth predictor. Then, we design mixers to learn from multiple weak predictions and adaptively fuse them into a strong depth estimate. The resultant model, which we refer to as Transformer-assisted depth ensembles (TEDepth). On the standard NYU-Depth-v2 and KITTI datasets, we thoroughly explore how the neural ensembles affect the depth estimation and demonstrate that our TEDepth achieves better results than previous state-of-the-art approaches. To validate the generalizability across cameras, we directly apply the models trained on NYU-Depth-v2 to the SUN RGB-D dataset without any fine-tuning, and the superior results emphasize its strong generalizability. Shuwei Shao, Zhongcai Pei, Zhong Liu 0005, Weihai Chen, Wentao Zhu 0001, Xingming Wu, Baochang Zhang 0001 |
IEEE Trans. Multim. | 6 |
| 2022 | Anti-retroactive Interference for Lifelong Learning
Runqi Wang, Yuxiang Bao, Baochang Zhang 0001, Jianzhuang Liu, Wentao Zhu 0001, Guodong Guo |
ECCV (24) | 5 |
| 2022 | Bandwidth-Aware Adaptive Codec for DNN Inference Offloading in IoT
Xiufeng Xie, Wentao Zhu 0001, Ji Liu 0002 |
ECCV (38) | 3 |
| 2022 | Uncertainty Learning towards Unsupervised Deformable Medical Image RegistrationabstractUncertainty estimation in medical image registration enables surgeons to evaluate the operative risk based on the trustworthiness of the registered image data thus of paramount importance for practical clinical applications. Despite the recent promising results obtained with deep unsupervised learning-based registration methods, reasoning about uncertainty of unsupervised registration models remains largely unexplored. In this work, we propose a predictive module to learn the registration and uncertainty in correspondence simultaneously. Our framework introduces empirical randomness and registration error based uncertainty prediction. We systematically assess the performances on two MRI datasets with different ensemble paradigms. Experimental results highlight that our proposed framework significantly improves the registration accuracy and uncertainty compared with the baseline. Luckyson Khaidem, Wentao Zhu 0001, Baochang Zhang 0001, David S. Doermann |
WACV | 3 |
| 2022 | Self-Supervised monocular depth and ego-Motion estimation in endoscopy: Appearance flow to the rescue
Shuwei Shao, Zhongcai Pei, Weihai Chen, Wentao Zhu 0001, Xingming Wu, Dianmin Sun, Baochang Zhang 0001 |
Medical Image Anal. | 4 |
| 2022 | Cardiac segmentation on late gadolinium enhancement MRI: A benchmark study from multi-sequence cardiac MR segmentation challenge
Xiahai Zhuang, Jiahang Xu, Xinzhe Luo, Chen Chen 0042, Cheng Ouyang, Daniel Rueckert, Víctor M. Campello, Karim Lekadir, Sulaiman Vesal, Nishant Ravikumar, Yashu Liu 0003, Gongning Luo, Jingkun Chen, Hongwei Li 0004, Buntheng Ly, Maxime Sermesant, Holger Roth, Wentao Zhu 0001, Jiexiang Wang, Xinghao Ding, Sen Yang 0006, Lei Li 0020 |
Medical Image Anal. | 18 |
| 2021 | SpeechNAS: Towards Better Trade-Off Between Latency and Accuracy for Large-Scale Speaker VerificationabstractRecently, x-vector [1] has been a successful and popular approach for speaker verification, which employs a time delay neural network (TDNN) and statistics pooling to extract speaker characterizing embedding from variable-length utterances. Improvement upon the x-vector has been an active research area, and enormous neural networks have been elaborately designed based on the x-vector, e.g., extended TDNN (E-TDNN) [2], factorized TDNN (F-TDNN) [3], and densely connected TDNN (D-TDNN) [4]. In this work, we try to identify the optimal architectures from a TDNN based search space employing neural architecture search (NAS), named SpeechNAS. Leveraging the recent advances in the speaker recognition, such as high-order statistics pooling, multi-branch mechanism, D-TDNN and angular additive margin softmax (AAM) loss with a minimum hyper-spherical energy (MHE), SpeechNAS automatically discovers five network architectures, from SpeechNAS-1 to SpeechNAS-5, of various numbers of parameters and GFLOPs on the large-scale text-independent speaker recognition dataset VoxCelebl. Our derived best neural network achieves an equal error rate (EER) of 1.02% on the standard test set of VoxCelebl, which surpasses previous TDNN based state-of-the-art approaches by a large margin. Wentao Zhu 0001, Tianlong Kong, Shun Lu 0001, Feng Deng, Sen Yang 0004, Ji Liu 0002 |
ASRU | 1 |
| 2021 | Test-Time Training for Deformable Multi-Scale Image RegistrationabstractRegistration is a fundamental task in medical robotics and is often a crucial step for many downstream tasks such as motion analysis, intra-operative tracking and image segmentation. Popular registration methods such as ANTs and NiftyReg optimize objective functions for each pair of images from scratch, which are time-consuming for 3D and sequential images with complex deformations. Recently, deep learning-based registration approaches such as VoxelMorph have been emerging and achieve competitive performance. In this work, we construct a test-time training for deep deformable image registration to improve the generalization ability of conventional learning-based registration model. We design multi-scale deep networks to consecutively model the residual deformations, which is effective for high variational deformations. Extensive experiments validate the effectiveness of multi-scale deep registration with test-time training based on Dice coefficient for image segmentation and mean square error (MSE), normalized local cross-correlation (NLCC) for tissue dense tracking tasks. Wentao Zhu 0001, Yufang Huang, Daguang Xu, Wei Fan 0001, Xiaohui Xie |
ICRA | 1 |
| 2021 | Federated Whole Prostate Segmentation in MRI with Personalized Neural Architectures
Holger Roth, Dong Yang 0005, Wenqi Li 0001, Andriy Myronenko, Wentao Zhu 0001, Ziyue Xu 0001, Xiaosong Wang 0001, Daguang Xu |
MICCAI (3) | 5 |
| 2021 | Shifted Chunk Transformer for Spatio-Temporal Representational LearningabstractSpatio-temporal representational learning has been widely adopted in various fields such as action recognition, video object segmentation, and action anticipation.Previous spatio-temporal representational learning approaches primarily employ ConvNets or sequential models, e.g., LSTM, to learn the intra-frame and inter-frame features. Recently, Transformer models have successfully dominated the study of natural language processing (NLP), image classification, etc. However, the pure-Transformer based spatio-temporal learning can be prohibitively costly on memory and computation to extract fine-grained features from a tiny patch. To tackle the training difficulty and enhance the spatio-temporal learning, we construct a shifted chunk Transformer with pure self-attention blocks. Leveraging the recent efficient Transformer design in NLP, this shifted chunk Transformer can learn hierarchical spatio-temporal features from a local tiny patch to a global videoclip. Our shifted self-attention can also effectively model complicated inter-frame variances. Furthermore, we build a clip encoder based on Transformer to model long-term temporal dependencies. We conduct thorough ablation studies to validate each component and hyper-parameters in our shifted chunk Transformer, and it outperforms previous state-of-the-art approaches on Kinetics-400, Kinetics-600,UCF101, and HMDB51. Xuefan Zha, Wentao Zhu 0001, Xun Lv, Sen Yang 0004, Ji Liu 0002 |
NeurIPS | 2 |
| 2021 | Deformable Gabor Feature Networks for Biomedical Image ClassificationabstractIn recent years, deep learning has dominated progress in the field of medical image analysis. We find however, that the ability of current deep learning approaches to represent the complex geometric structures of many medical images is insufficient. One limitation is that deep learning models require a tremendous amount of data, and it is very difficult to obtain a sufficient amount with the necessary detail. A second limitation is that there are underlying features of these medical images that are well established, but the black-box nature of existing convolutional neural networks (CNNs) do not allow us to exploit them. In this paper, we revisit Gabor filters and introduce a deformable Gabor convolution (DGConv) to expand deep networks interpretability and enable complex spatial variations. The features are learned at deformable sampling locations with adaptive Gabor convolutions to improve representitiveness and robustness to complex objects. The DGConv replaces standard convolutional layers and is easily trained end-to-end, resulting in deformable Gabor feature network (DGFN) with few additional parameters and minimal additional training cost. We introduce DGFN for addressing deep multi-instance multi-label classification on the INbreast dataset for mammograms and on the ChestX-ray14 dataset for pulmonary x-ray images. Xin Xia 0005, Wentao Zhu 0001, Baochang Zhang 0001, David S. Doermann, Lian Zhuo |
WACV | 3 |
| 2021 | Federated semi-supervised learning for COVID region segmentation in chest CT using multi-national data from China, Italy, Japan
Dong Yang 0005, Ziyue Xu 0001, Wenqi Li 0001, Andriy Myronenko, Holger Roth, Stephanie A. Harmon, Sheng Xu 0001, Baris Turkbey, Evrim Turkbey, Xiaosong Wang 0001, Wentao Zhu 0001, Gianpaolo Carrafiello, Francesca Patella, Maurizio Cariati, Hirofumi Obinata, Hitoshi Mori, Kaku Tamura, Peng An 0002, Bradford J. Wood, Daguang Xu |
Medical Image Anal. | 11 |
| 2021 | Multi-Domain Image Completion for Random Missing Input DataabstractMulti-domain data are widely leveraged in vision applications taking advantage of complementary information from different modalities, e.g., brain tumor segmentation from multi-parametric magnetic resonance imaging (MRI). However, due to possible data corruption and different imaging protocols, the availability of images for each domain could vary amongst multiple data sources in practice, which makes it challenging to build a universal model with a varied set of input data. To tackle this problem, we propose a general approach to complete the random missing domain(s) data in real applications. Specifically, we develop a novel multi-domain image completion method that utilizes a generative adversarial network (GAN) with a representational disentanglement scheme to extract shared content encoding and separate style encoding across multiple domains. We further illustrate that the learned representation in multi-domain image completion could be leveraged for high-level tasks, e.g., segmentation, by introducing a unified framework consisting of image completion and segmentation with a shared content encoder. The experiments demonstrate consistent performance improvement on three datasets for brain tumor segmentation, prostate segmentation, and facial expression image completion respectively. Liyue Shen, Wentao Zhu 0001, Xiaosong Wang 0001, Lei Xing 0001, John M. Pauly, Baris Turkbey, Stephanie A. Harmon, Thomas Sanford, Sherif Mehralivand, Peter L. Choyke, Bradford J. Wood, Daguang Xu |
IEEE Trans. Medical Imaging | 2 |
| 2020 | Cycle-Consistent Adversarial Autoencoders for Unsupervised Text Style TransferabstractUnsupervised text style transfer is full of challenges due to the lack of parallel data and difficulties in content preservation.In this paper, we propose a novel neural approach to unsupervised text style transfer which we refer to as Cycle-consistent Adversarial autoEncoders (CAE) trained from non-parallel data.CAE consists of three essential components: (1) LSTM autoencoders that encode a text in one style into its latent representation and decode an encoded representation into its original text or a transferred representation into a style-transferred text, (2) adversarial style transfer networks that use an adversarially trained generator to transform a latent representation in one style into a representation in another style, and (3) a cycle-consistent constraint that enhances the capacity of the adversarial style transfer networks in content preservation.The entire CAE with these three components can be trained end-to-end.Extensive experiments and in-depth analyses on two widely-used public datasets consistently validate the effectiveness of proposed CAE in both style transfer and content preservation against several strong baselines in terms of four automatic evaluation metrics and human evaluation. Yufang Huang, Wentao Zhu 0001, Deyi Xiong, Yiye Zhang, Changjian Hu |
COLING | 2 |
| 2020 | LAMP: Large Deep Nets with Automated Model Parallelism for Image Segmentation
Wentao Zhu 0001, Can Zhao 0001, Wenqi Li 0001, Holger Roth, Ziyue Xu 0001, Daguang Xu |
MICCAI (4) | 1 |
| 2020 | NeurReg: Neural Registration and Its Application to Image SegmentationabstractRegistration is a fundamental task in medical image analysis which can be applied to several tasks including image segmentation, intra-operative tracking, multi-modal image alignment, and motion analysis. Popular registration tools such as ANTs and NiftyReg optimize an objective function for each pair of images from scratch which is time-consuming for large images with complicated deformation. Facilitated by the rapid progress of deep learning, learning-based approaches such as VoxelMorph have been emerging for image registration. These approaches can achieve competitive performance in a fraction of a second on advanced GPUs. In this work, we construct a neural registration framework, called NeurReg, with a hybrid loss of displacement fields and data similarity, which substantially improves the current state-of-the-art of registrations. Within the framework, we simulate various transformations by a registration simulator which generates fixed image and displacement field ground truth for training. Furthermore, we design three segmentation frameworks based on the proposed registration framework: 1) atlas-based segmentation, 2) joint learning of both segmentation and registration tasks, and 3) multi-task learning with atlas-based segmentation as an intermediate feature. Extensive experimental results validate the effectiveness of the proposed NeurReg framework based on various metrics: the endpoint error (EPE) of the predicted displacement field, mean square error (MSE), normalized local cross-correlation (NLCC), mutual information (MI), Dice coefficient, uncertainty estimation, and the interpretability of the segmentation. The proposed NeurReg improves registration accuracy with fast inference speed, which can greatly accelerate related medical image analysis tasks. Wentao Zhu 0001, Andriy Myronenko, Ziyue Xu 0001, Wenqi Li 0001, Holger Roth, Yufang Huang, Fausto Milletari, Daguang Xu |
WACV | 1 |
| 2018 | DeepEM: Deep 3D ConvNets with EM for Weakly Supervised Pulmonary Nodule Detection
Wentao Zhu 0001, Yeeleng Scott Vang, Yufang Huang, Xiaohui Xie |
MICCAI (2) | 1 |
| 2018 | DeepLung: Deep 3D Dual Path Nets for Automated Pulmonary Nodule Detection and ClassificationabstractIn this work, we present a fully automated lung computed tomography (CT) cancer diagnosis system, DeepLung. DeepLung consists of two components, nodule detection (identifying the locations of candidate nodules) and classification (classifying candidate nodules into benign or malignant). Considering the 3D nature of lung CT data and the compactness of dual path networks (DPN), two deep 3D DPN are designed for nodule detection and classification respectively. Specifically, a 3D Faster Regions with Convolutional Neural Net (R-CNN) is designed for nodule detection with 3D dual path blocks and a U-net-like encoder-decoder structure to effectively learn nodule features. For nodule classification, gradient boosting machine (GBM) with 3D dual path network features is proposed. The nodule classification subnetwork was validated on a public dataset from LIDC-IDRI, on which it achieved better performance than state-of-the-art approaches and surpassed the performance of experienced doctors based on image modality. Within the DeepLung system, candidate nodules are detected first by the nodule detection subnetwork, and nodule diagnosis is conducted by the classification subnetwork. Extensive experimental results demonstrate that DeepLung has performance comparable to experienced doctors both for the nodule-level and patient-level diagnosis on the LIDC-IDRI dataset. Wentao Zhu 0001, Chaochun Liu, Wei Fan 0001, Xiaohui Xie |
WACV | 1 |
| 2017 | Deep Multi-instance Networks with Sparse Label Assignment for Whole Mammogram Classification
Wentao Zhu 0001, Qi Lou, Yeeleng Scott Vang, Xiaohui Xie |
MICCAI (3) | 1 |
| 2016 | Co-Occurrence Feature Learning for Skeleton Based Action Recognition Using Regularized Deep LSTM NetworksabstractSkeleton based action recognition distinguishes human actions using the trajectories of skeleton joints, which provide a very good representation for describing actions. Considering that recurrent neural networks (RNNs) with Long Short-Term Memory (LSTM) can learn feature representations and model long-term temporal dependencies automatically, we propose an end-to-end fully connected deep LSTM network for skeleton based action recognition. Inspired by the observation that the co-occurrences of the joints intrinsically characterize human actions, we take the skeleton as the input at each time slot and introduce a novel regularization scheme to learn the co-occurrence features of skeleton joints. To train the deep LSTM network effectively, we propose a new dropout algorithm which simultaneously operates on the gates, cells, and output responses of the LSTM neurons. Experimental results on three human action recognition datasets consistently demonstrate the effectiveness of the proposed model. Wentao Zhu 0001, Cuiling Lan, Junliang Xing, Wenjun Zeng 0001, Yanghao Li, Li Shen 0005, Xiaohui Xie |
AAAI | 1 |