VLDB 2026 Research / reviewers in the wild / expert
Zhiwei Wang 0002
dblp:10/293-2
· DBLP profile ↗
56ranked-venue papers
8as first author
44since 2021 · last 2026
0000-0002-1612-8573ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 32 · 3 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 5 first-author · 18 since 2021Artificial intelligence and machine learning · 15 · 2 first-author · 14 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RouterNet: Hierarchical Point Routing Network for Robust Vertebral Landmark Localization on AP X-ray ImagesabstractLocating vertebral landmarks on anteroposterior (AP) X-ray images is challenging due to the tissue overlap. Despite the great progress of heatmap-based methods, they often predict missing/false points, which are intolerable in the downstream applications like scoliosis assessment. In this paper, we instead modernize the classic point-regression scheme, and propose a novel model termed RouterNet to locate the 68 vertebral landmarks completely and accurately. RouterNet starts from an initial root point, and then gradually routes it onto more and more points with finer and finer semantics. RouterNet naturally couples such point routing process with its hierarchical and multi-scale feature learning. That is, lower-scale feature maps are utilized to regress points with coarser semantics, and the regressed points pilot a more focused local feature extraction on the next higher-scale map to route onto their subsequent positions with finer semantics. With this divide-and-conquer, RouterNet alleviates the task difficulty, and can robustly localize by routing from the whole spinal center to 17 vertebral centers, and further to their 68 corner points. Extensive and comprehensive experiments on both public and private datasets demonstrate our superior performance over other state-of-the-arts, by decreasing NMSE by 73.8% for landmark localization, and SMAPE by 14.8% for the downstream scoliosis assessment. Yingjie Guo, Jinxin Lv, Qiang Li 0018, Zhiwei Wang 0002 |
AAAI | 5 |
| 2026 | Pairing-free Group-level Knowledge Distillation for Robust Gastrointestinal Lesion Classification in White-Light EndoscopyabstractWhite-Light Imaging (WLI) is the standard for endoscopic cancer screening, but Narrow-Band Imaging (NBI) offers superior diagnostic details. A key challenge is transferring knowledge from NBI to enhance WLI-only models, yet existing methods are critically hampered by their reliance on paired NBI-WLI images of the same lesion, a costly and often impractical requirement that leaves vast amounts of clinical data untapped. In this paper, we break this paradigm by introducing PaGKD, a novel Pairing-free Group-level Knowledge Distillation framework that that enables effective cross-modal learning using unpaired WLI and NBI data. Instead of forcing alignment between individual, often semantically mismatched image instances, PaGKD operates at the group level to distill more complete and compatible knowledge across modalities. Central to PaGKD are two complementary modules: (1) Group-level Prototype Distillation (GKD-Pro) distills compact group representations by extracting modality-invariant semantic prototypes via shared lesion-aware queries; (2) Group-level Dense Distillation (GKD-Den) performs dense cross-modal alignment by guiding group-aware attention with activation-derived relation maps. Together, these modules enforce global semantic consistency and local structural coherence without requiring image-level correspondence. Extensive experiments on four clinical datasets demonstrate that PaGKD consistently and significantly outperforms state-of-the-art methods, boosting AUC by 3.3%, 1.1%, 2.8%, and 3.2%, respectively, establishing a new direction for cross-modal learning from unpaired data. Qimei Wang, Yingjie Guo, Qiang Li 0018, Zhiwei Wang 0002 |
AAAI | 5 |
| 2026 | FIA-Edit: Frequency-Interactive Attention for Efficient and High-Fidelity Inversion-Free Text-Guided Image EditingabstractText-guided image editing has advanced rapidly with the rise of diffusion models. While flow-based inversion-free methods offer high efficiency by avoiding latent inversion, they often fail to effectively integrate source information, leading to poor background preservation, spatial inconsistencies, and over-editing due to the lack of effective integration of source information. In this paper, we present FIA-Edit, a novel inversion-free framework that achieves high-fidelity and semantically precise edits through a Frequency-Interactive Attention. Specifically, we design two key components: (1) a Frequency Representation Interaction (FRI) module that enhances cross-domain alignment by exchanging frequency components between source and target features within self-attention, and (2) a Feature Injection (FIJ) module that explicitly incorporates source-side queries, keys, values, and text embeddings into the target branch's cross-attention to preserve structure and semantics. Comprehensive and extensive experiments demonstrate that FIA-Edit supports high-fidelity editing at low computational cost (~6s per 512 * 512 image on an RTX 4090) and consistently outperforms existing methods across diverse tasks in visual quality, background fidelity, and controllability. Furthermore, we are the first to extend text-guided image editing to clinical applications. By synthesizing anatomically coherent hemorrhage variations in surgical images, FIA-Edit opens new opportunities for medical data augmentation and delivers significant gains in downstream bleeding classification. Kaixiang Yang 0004, Boyang Shen, Xin Li 0001, Yuchen Dai, Yueran Ma, Qiang Li 0018, Zhiwei Wang 0002 |
AAAI | 9 |
| 2026 | Multi-class segmentation of aortic branches and zones in computed tomography angiography: The AortaSeg24 challenge
Muhammad Imran 0013, Jonathan R. Krebs, Vishal Balaji Sivaraman, Amarjeet Kumar, Walker R. Ueland, Michael J. Fassler, Lisheng Wang, Maximilian Rokuss, Michael Baumgartner 0001, Yannick Kirchhof, Klaus H. Maier-Hein, Fabian Isensee, Shuolin Liu, Bong Thanh Nguyen, Dong-jin Shin, Park Ji-Woo, Matthew Choi, Kwang-Hyun Uhm, Sung-Jea Ko, Chanwoong Lee, Jaehee Chun, Yun Gu, Zhaohong Pan, Xiaokun Liang, Markus Tiefenthaler, Enrique Almar-Munoz, Matthias Schwab, Mikhail Kotyushev, Rostislav Epifanov, Marek Wodzinski, Henning Müller, Abdul Qayyum 0002, Moona Mazher, Steven A. Niederer, Zhiwei Wang 0002, Kaixiang Yang 0004, Jintao Ren, Stine Sofia Korreman, Yuchong Gao, Hongye Zeng, Jinghua Yue, Fugen Zhou, Alexander Cosman, Muxuan Liang, Gilbert R. Upchurch Jr., Yuyin Zhou, Michol A. Cooper, Wei Shao 0008 |
Medical Image Anal. | 46 |
| 2026 | StyleSeg V2: Towards robust single-label-supervised segmentation of brain tissue via optimization-free registration error perception
Chongwei Wu, Xiaoyu Zeng, Tingwei Quan, Jinxin Lv, Qiang Li 0018, Zhiwei Wang 0002 |
Pattern Recognit. | 8 |
| 2026 | Improving 3D Thin Vessel Segmentation in Brain TOF-MRA via a Dual-Space Context-Aware Networkabstract3D cerebrovascular segmentation poses a significant challenge, akin to locating a line within a vast 3D environment. This complexity can be substantially reduced by projecting the vessels onto a 2D plane, enabling easier segmentation. In this paper, we create a vessel-segmentation-friendly space using a clinical visualization technique called maximum intensity projection (MIP). Leveraging this, we propose a Dual-space Context-Aware Network (DCANet) for 3D vessel segmentation, designed to capture even the finest vessel structures accurately. DCANet begins by transforming a magnetic resonance angiography (MRA) volume into a 3D Regional-MIP volume, where each Regional-MIP slice is constructed by projecting adjacent MRA slices. This transformation highlights vessels as prominent continuous curves rather than the small circular or ellipsoidal cross-sections seen in MRA slices. DCANet encodes vessels separately in the MRA and the projected Regional-MIP spaces and introduces the Regional-MIP Image Fusion Block (MIFB) between these dual spaces to selectively integrate contextual features from Regional-MIP into MRA. Following dual-space encoding, DCANet employs a Dual-mask Spatial Guidance TransFormer (DSGFormer) decoder to focus on vessel regions while effectively excluding background areas, which reduces the learning burden and improves segmentation accuracy. We benchmark DCANet on four datasets: two public datasets, TubeTK and IXI-IOP, and two in-house datasets, Xiehe and IXI-HH. The results demonstrate that DCANet achieves superior performance, with improvements in average DSC values of at least 2.26%, 2.17%, 2.62%, and 2.58% for thin vessels, respectively. Wenqi Shan, Qiang Li 0018, Zhiwei Wang 0002 |
IEEE J. Biomed. Health Informatics | 5 |
| 2026 | Cross-Correlation Rectification for Robust Deformable Registration of Brain Tumor MRI Between Preoperative and Postoperative PhasesabstractA key challenge in registering pre- and post-operative brain tumor images lies in the anatomical inconsistencies caused by pathological changes and surgical resections. Recent efforts have addressed this issue by masking affected regions during optimization, but such approaches discard contextual information and rely on CNN backbones that implicitly model deformation, often overfitting to distant normal tissues and failing to capture the severe nonlinear distortions near the tumor. Correlation-based alternatives enhance generalization to diverse deformation patterns by explicitly modeling geometric correspondences, yet they frequently yield unreliable matches in and around tumor regions, disrupting the deformation field. In this paper, we propose Cross-correlation Rectification-based Registration Network (CRRNet), the first framework that introduces an active rectification mechanism specifically for robustness and structurally coherent pre- to post-operative brain tumor image registration. Specifically, CRRNet achieves this through two complementary modules: 1) a Cross-correlation Analysis-Based Inconsistency module that identifies invalid correspondences via bidirectional loop-closure evaluation on cross-correlations, and 2) a Dual-level Correspondence Rectification module that adaptively integrates contextually reliable correlations from local and long-range perspectives to restore structurally coherent matches. This synergistic design retains the strengths of cross-correlation while effectively mitigating correspondence mismatches. Extensive experiments on multiple tumor benchmarks demonstrate the superiority of CRRNet. Specifically, on the BraTS-Reg dataset, it reduces the mean registration errors by 8.18% in near-tumor regions and 3.95% in far-from-tumor regions, surpassing state-of-the-art methods. Chongwei Wu, Xiaoyu Zeng, Shuxian Niu, Qiang Li 0018, Zhiwei Wang 0002 |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | MonoBox: Tightness-Free Box-Supervised Polyp Segmentation Using Monotonicity ConstraintabstractWe propose MonoBox, an innovative box-supervised segmentation method constrained by monotonicity to liberate its training from the user-unfriendly box-tightness assumption. In contrast to conventional box-supervised segmentation, where the box edges must precisely touch the target boundaries, MonoBox leverages imprecisely-annotated boxes to achieve robust pixel-wise segmentation. The 'linchpin' is that, within the noisy zones around box edges, MonoBox discards the traditional misguiding multiple-instance learning loss, and instead optimizes a carefully-designed objective, termed monotonicity constraint. Along directions transitioning from the foreground to background, this new constraint steers responses to adhere to a trend of monotonically decreasing values. Consequently, the originally unreliable learning within the noisy zones is transformed into a correct and effective monotonicity optimization. Moreover, an adaptive label correction is introduced, enabling MonoBox to enhance the tightness of box annotations using predicted masks from the previous epoch and dynamically shrink the noisy zones as training progresses. We verify MonoBox in the box-supervised segmentation task of polyps, where satisfying box-tightness is challenging due to the vague boundaries between the polyp and normal tissues. Experiments on both public synthetic and in-house real noisy datasets demonstrate that MonoBox exceeds other anti-noise state-of-the-arts by improving Dice by at least 5.5% and 3.3%, respectively. Zhenyu Yi, Qiang Li 0018, Zhiwei Wang 0002 |
AAAI | 7 |
| 2025 | TDD: Task-Aware Diffusion Dual-Matching for Robust Electrocardiogram DenoisingabstractElectrocardiograms (ECGs) play a critical role in the diagnosis and monitoring of cardiovascular diseases. However, their signal quality is often compromised by various types of noise which can impair diagnostic accuracy and destabilize downstream automated analysis. Recent deep learning-based denoising methods have demonstrated promising results, but they often focus solely on signal reconstruction accuracy while overlooking consistency with downstream tasks, such as waveform segmentation or disease classification. This misalignment may lead to clean-looking signals that are nonetheless structurally misleading or clinically unreliable. To address these limitations, we propose the Task-aware Diffusion Dual-matching (TDD) method, a novel denoising framework that integrates multi-task learning with a modified diffusion process. Unlike conventional multi-task approaches that treat auxiliary tasks as feature enhancers, TDD leverages auxiliary predictions to directly complement and guide the denoising process. Specifically, TDD uses noisy ECG signals as the diffusion starting point to reduce reconstruction uncertainty, and segmentation outputs provide structure-aware gradients that refine the denoising trajectory, which further encourages more confident, unimodal segmentation predictions. This dual-matching process forms a positive feedback loop: cleaner signals enable more accurate segmentation, which in turn reinforces denoising with semantically meaningful guidance. Notably, segmentation labels are not required during inference. Experiments on multiple public ECG datasets show that TDD outperforms existing denoising methods under various noise conditions. Moreover, it exhibits superior structure-preserving capabilities under complex cardiac rhythms and enhances downstream disease diagnosis performance, leading to more accurate clinical interpretation. Siyi Fang, Zhiwei Wang 0002, Qiang Li 0018, Peng Zhang 0106 |
BIBM | 3 |
| 2025 | Bidirectional Mammogram View Translation with Column-Aware and Implicit 3D Conditional DiffusionabstractDual-view mammography, including craniocaudal (CC) and mediolateral oblique (MLO) projections, offers complementary anatomical views crucial for breast cancer diagnosis. However, in real-world clinical workflows, one view may be missing, corrupted, or degraded due to acquisition errors or compression artifacts, limiting the effectiveness of downstream analysis. View-to-view translation can help recover missing views and improve lesion alignment. Unlike natural images, this task in mammography is highly challenging due to large non-rigid deformations and severe tissue overlap in X-ray projections, which obscure pixel-level correspondences. In this paper, we propose Column-Aware and Implicit 3D Diffusion (CA3D-Diff), a novel bidirectional mammogram view translation framework based on conditional diffusion model. To address cross-view structural misalignment, we first design a column-aware Cross-Attention mechanism that leverages the geometric property that anatomically corresponding regions tend to lie in similar column positions across views. Furthermore, we introduce an implicit 3D structure reconstruction module that back-projects noisy 2D latents into a coarse 3D feature volume based on breast-view projection geometry. Extensive experiments demonstrate that CA3D-Diff achieves superior performance in bidirectional tasks, outperforming state-of-the-art methods in visual fidelity and structural consistency. Furthermore, the synthesized views effectively improve single-view malignancy classification in screening settings, demonstrating the practical value of our method in realworld diagnostics. Our code is available at https://github.com/lixinHUST/CA3D-Diff. Xin Li 0001, Kaixiang Yang 0004, Qiang Li 0018, Zhiwei Wang 0002 |
BIBM | 4 |
| 2025 | TiS-TSL: Image-Label Supervised Surgical Video Stereo Matching via Time-Switchable Teacher-Student LearningabstractStereo matching in minimally invasive surgery (MIS) is essential for next-generation navigation and augmented reality. Yet, dense disparity supervision is nearly impossible due to anatomical constraints, typically limiting annotations to only a few image-level labels acquired before the endoscope enters deep body cavities. Teacher-Student Learning (TSL) offers a promising solution by leveraging a teacher trained on sparse labels to generate pseudo labels and associated confidence maps from abundant unlabeled surgical videos. However, existing TSL methods are confined to image-level supervision, providing only spatial confidence and lacking temporal consistency estimation. This absence of spatio-temporal reliability results in unstable disparity predictions and severe flickering artifacts across video frames. To overcome these challenges, we propose TiS-TSL, a novel time-switchable teacher-student learning framework for video stereo matching under minimal supervision. At its core is a unified model that operates in three distinct modes: ImagePrediction (IP), Forward Video-Prediction (FVP), and Backward Video-Prediction (BVP), enabling flexible temporal modeling within a single architecture. Enabled by this unified model, TiSTSL adopts a two-stage learning strategy. The Image-to-Video (I2V) stage transfers sparse image-level knowledge to initialize temporal modeling. The subsequent Video-to-Video (V2V) stage refines temporal disparity predictions by comparing forward and backward predictions to calculate bidirectional spatio-temporal consistency. This consistency identifies unreliable regions across frames, filters noisy video-level pseudo labels, and enforces temporal coherence. Experimental results on two public datasets demonstrate that TiS-TSL exceeds other image-based state-of-the-arts by improving TEPE and EPE by at least 2.11 % and$\mathbf{4. 5 4 \%}$, respectively. Codes are at https://github.com/wr2167/TiSTSL. Hao Wang 0218, Qiang Li 0018, Zhiwei Wang 0002 |
BIBM | 6 |
| 2025 | First-frame Supervised Video Polyp Segmentation via Propagative and Semantic Dual-teacher NetworkabstractAutomatic video polyp segmentation plays a critical role in gastrointestinal cancer screening, but the cost of frame-by-frame annotations is prohibitively high. While sparse-frame supervised methods have reduced this burden proportionately, the cost remains overwhelming for long-duration videos and large-scale datasets. In this paper, we, for the first time, reduce the annotation cost to just a single frame per polyp video, regardless of the video's length. To this end, we introduce a new task, First-Frame Supervised Video Polyp Segmentation (FSVPS), and propose a novel Propagative and Semantic Dual-Teacher Network (PSDNet). Specifically, PSDNet adopts a teacher-student framework but employs two distinct types of teachers: the propagative teacher and the semantic teacher. The propagative teacher is a universal object tracker that propagates the first-frame annotation to subsequent frames as pseudo labels. However, tracking errors may accumulate over time, gradually degrading the pseudo labels and misguiding the student model. To address this, we introduce the semantic teacher, an exponential moving average of the student model, which produces more stable and time-invariant pseudo labels. PSDNet merges the pseudo labels from both teachers using a carefully-designed back-propagation strategy. This strategy assesses the quality of the pseudo labels by tracking them backward to the first frame. High-quality pseudo labels are more likely to spatially align with the first-frame annotation after this backward tracking, ensuring more accurate teacher-to-student knowledge transfer and improved segmentation performance. Benchmarking on SUN-SEG, the largest VPS dataset, demonstrates the competitive performance of PSDNet compared to fully-supervised approaches, and its superiority over sparse-frame supervised state-of-the-arts with a minimum improvement of 4.5% in Dice score. Codes are at https://github.com/Huster-Hq/PSDNet. Qiang Li 0018, Zhiwei Wang 0002 |
ICASSP | 4 |
| 2025 | SPNet: Sparse-mask Prompt-learning Network for Cerebrovascular SegmentationabstractCerebrovascular segmentation demands a comprehensive understanding of vascular topology from a global perspective on 3D Time-of-Flight Magnetic Resonance Angiography (TOF-MRA) images. However, existing 3D segmentation methods typically rely on patch-wise processing to manage computational costs, where different segments of the same vessel branch may be handled independently in separate patches. This might restrict their ability to capture the full 3D context of the vessel, thus severely impacting segmentation performance. In this paper, we focus on achieving comprehensive 3D vessel segmentation in a single feedforward pass by leveraging Maximum Intensity Projection (MIP). By compressing the original 3D image, the vessels of the entire brain are superimposed and projected onto a 2D MIP image. We propose a Sparse-mask Prompt-learning Network (SPNet) to leverage MIP for precise cerebrovascular segmentation. Specifically, SPNet generates MIP images along three orthogonal directions and performs 2D segmentation on each of these MIP images independently. The segmented 2D vessels of MIPs are then back-projected into 3D space and fused into a sparse vessel mask. SPNet treats this sparse mask as a topological prompt that captures the overall vascular pathways, thereby enhancing feature learning on the original 3D TOF-MRA data, ultimately improving the final cerebrovascular segmentation performance. Experimental results on the ADAM and IXI datasets demonstrate superior performance of SPNet with least improvements of 1.73% and 0.75% on average Dices, 3.04mm and 1.19mm on Hausdorff distances, respectively, over state-of-the-art methods. Codes are available at: https://github.com/shanwq/SPNet. Wenqi Shan, Qiang Li 0018, Zhiwei Wang 0002 |
ICASSP | 3 |
| 2025 | DACAT: Dual-stream Adaptive Clip-aware Time Modeling for Robust Online Surgical Phase RecognitionabstractSurgical phase recognition has become a crucial requirement in laparoscopic surgery, enabling various clinical applications like surgical risk forecasting. Current methods typically identify the surgical phase using individual frame-wise embeddings as the fundamental unit for time modeling. However, this approach is overly sensitive to current observations, often resulting in discontinuous and erroneous predictions within a complete surgical phase. In this paper, we propose DACAT, a novel dual-stream model that adaptively learns clip-aware context information to enhance the temporal relationship. In one stream, DACAT pretrains a frame encoder, caching all historical frame-wise features. In the other stream, DACAT fine-tunes a new frame encoder to extract the frame-wise feature at the current moment. Additionally, a max clip-response read-out (Max-R) module is introduced to bridge the two streams by using the current frame-wise feature to adaptively fetch the most relevant past clip from the feature cache. The clip-aware context feature is then encoded via cross-attention between the current frame and its fetched adaptive clip, and further utilized to enhance the time modeling for accurate online surgical phase recognition. The benchmark results on three public datasets, i.e., Cholec80, M2CAI16, and AutoLaparo, demonstrate the superiority of our proposed DACAT over existing state-of-the-art methods, with improvements in Jaccard scores of at least 4.5%, 4.6%, and 2.7%, respectively. Our code and models have been released at https://github.com/kk42yy/DACAT. Kaixiang Yang 0004, Qiang Li 0018, Zhiwei Wang 0002 |
ICASSP | 3 |
| 2025 | CoStoDet-DDPM: Collaborative Training of Stochastic and Deterministic Models Improves Surgical Workflow Anticipation and RecognitionabstractAnticipating and recognizing surgical workflows are critical for intelligent surgical assistance systems. However, existing methods rely on deterministic decision-making, struggling to generalize across the large anatomical and procedural variations inherent in real-world surgeries.In this paper, we introduce an innovative framework that incorporates stochastic modeling through a denoising diffusion probabilistic model (DDPM) into conventional deterministic learning for surgical workflow analysis. At the heart of our approach is a collaborative co-training paradigm: the DDPM branch captures procedural uncertainties to enrich feature representations, while the task branch focuses on predicting surgical phases and instrument usage.Theoretically, we demonstrate that this mutual refinement mechanism benefits both branches: the DDPM reduces prediction errors in uncertain scenarios, and the task branch directs the DDPM toward clinically meaningful representations. Notably, the DDPM branch is discarded during inference, enabling real-time predictions without sacrificing accuracy.Experiments on the Cholec80 dataset show that for the anticipation task, our method achieves a 16% reduction in eMAE compared to state-of-the-art approaches, and for phase recognition, it improves the Jaccard score by 1.0%. Additionally, on the AutoLaparo dataset, our method achieves a 1.5% improvement in the Jaccard score for phase recognition, while also exhibiting robust generalization to patient-specific variations. Our code and weight are available at https://github.com/kk42yy/CoStoDet-DDPM. Kaixiang Yang 0004, Xin Li 0001, Qiang Li 0018, Zhiwei Wang 0002 |
ICCV | 4 |
| 2025 | Holistic White-Light Polyp Classification via Alignment-Free Dense Distillation of Auxiliary Optical Chromoendoscopy
Qimei Wang, Xuantao Ji, Qiang Li 0018, Zhiwei Wang 0002 |
MICCAI (11) | 7 |
| 2025 | Targeted False Positive Synthesis via Detector-Guided Adversarial Diffusion Attacker for Robust Polyp Detection
Quan Zhou 0011, Gan Luo, Qingyong Zhang, Yinjiao Tian, Qiang Li 0018, Zhiwei Wang 0002 |
MICCAI (11) | 8 |
| 2025 | Joint Holistic and Lesion Controllable Mammogram Synthesis via Gated Conditional Diffusion ModelabstractMammography is the most commonly used imaging modality for breast cancer screening, driving an increasing demand for deep-learning techniques to support large-scale analysis. However, the development of accurate and robust methods is often limited by insufficient data availability and a lack of diversity in lesion characteristics. While generative models offer a promising solution for data synthesis, current approaches often fail to adequately emphasize lesion-specific features and their relationships with surrounding tissues. In this paper, we propose Gated Conditional Diffusion Model (GCDM), a novel framework designed to jointly synthesize holistic mammogram images and localized lesions. GCDM is built upon a latent denoising diffusion framework, where the noised latent image is concatenated with a soft mask embedding that represents breast, lesion, and their transitional regions, ensuring anatomical coherence between them during the denoising process. To further emphasize lesion-specific features, GCDM incorporates a gated conditioning branch that guides the denoising process by dynamically selecting and fusing the most relevant radiomic and geometric properties of lesions, effectively capturing their interplay. Experimental results demonstrate that GCDM achieves precise control over small lesion areas while enhancing the realism and diversity of synthesized mammograms. These advancements position GCDM as a promising tool for clinical applications in mammogram synthesis. Our code is available at https://github.com/lixinHUST/Gated-Conditional-Diffusion-Model/ Xin Li 0001, Kaixiang Yang 0004, Qiang Li 0018, Zhiwei Wang 0002 |
ACM Multimedia | 4 |
| 2025 | SegRap2023: A benchmark of organs-at-risk and gross tumor volume Segmentation for Radiotherapy Planning of Nasopharyngeal Carcinoma
Xiangde Luo, Yunxin Zhong, Shuolin Liu, Mehdi Astaraki, Simone Bendazzoli, Iuliana Toma-Dasu, Yiwen Ye, Ziyang Chen 0003, Yong Xia 0001, Yanzhou Su, Jin Ye 0002, Junjun He, Zhaohu Xing, Hongqiu Wang, Lei Zhu 0003, Kaixiang Yang 0004, Zhiwei Wang 0002, Chan Woong Lee, Sang Joon Park, Jaehee Chun, Constantin Ulrich, Klaus H. Maier-Hein, Nchongmaje Ndipenoch, Alina Dana Miron, Yongmin Li 0001, Chengyang An, Lisheng Wang, Kaiwen Huang 0002, Yunqi Gu, Tao Zhou 0002, Mu Zhou, Shichuan Zhang, Wenjun Liao, Guotai Wang, Shaoting Zhang 0001 |
Medical Image Anal. | 20 |
| 2025 | Toward high-quality pseudo masks from noisy or weak annotations for robust medical image segmentation
Zhiwei Wang 0002, Tianyu Zhao 0005, Xiaohuan Ding, Xin Yang 0008 |
Neural Networks | 2 |
| 2025 | Boosting Few-Shot Semantic Segmentation of 3D Medical Images via Collaborative Slice AlignmentabstractFew-shot semantic segmentation (FSS) of 3D medical images requires finding a 2D slice from the labeled volume as support to 'query' slices of the unlabeled one. Accurately determining support slices is crucial for learning representative prototypical features, thereby enhancing segmentation accuracy. The existing methods typically resort to the true position of the query target to align the query with support slices or simply exploit one key support slice to segment all query slices, which inevitably results in poor practicality and mis-segmentation. In this regard, we seek a practical and efficient solution by proposing a novel Collaborative Slice Alignment (CSA) module, which densely assigns each query slice its own fittest support without knowing the target prior. Concretely, our CSA first estimates the confidence scores of slices from the sorting task to implicitly reflect their physical location in the human body. The estimated scores are considered as spatial references for aligning support slices and query slices so that each matching pair shares the most similar image contents. Moreover, the self-learnable ranking objective allows CSA to transfer internal knowledge into both support and query features to further boost the FSS performance. Additionally, we introduce an Information Reconciliation (InRe) module to mitigate the inconsistent feature distribution caused by the individual differences between support and query images. Experimental results demonstrate that the combination of CSA and InRe achieves an average Dice score improvement of at least 8.61% across three datasets, consistently outperforming other state-of-the-art methods. Jialun Pei, Zhiwei Wang 0002, Qiang Li 0018, Pheng-Ann Heng |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | Fine-Grained Temporal Site Monitoring in EGD Streams via Visual Time-Aware Embedding and Vision-Text Asymmetric CoworkingabstractEsophagogastroduodenoscopy (EGD) requires inspecting plentiful upper gastrointestinal (UGI) sites completely for a precise cancer screening. Automated temporal site monitoring for EGD assistance is thus of high demand, yet often fails if directly applying the existing methods of online action detection. The key challenges are two-fold: 1) the global camera motion dominates, invalidating the temporal patterns derived from the object optical flows, and 2) the UGI sites are fine-grained, yielding highly homogenized appearances. In this paper, we propose an EGD-customized model, powered by two novel designs, i.e., Visual Time-aware Embedding plus Vision-text Asymmetric Coworking (VTE+VAC), for real-time accurate fine-grained UGI site monitoring. Concretely, VTE learns visual embeddings by differentiating frames via classification losses, and meanwhile by reordering the sampled time-agnostic frames to be temporally coherent via a ranking loss. Such joint objective encourages VTE to capture the sequential relation without resorting to the inapplicable object optical flows, and thus to provide the time-aware frame-wise embeddings. In the subsequent analysis, VAC uses a temporal sliding window, and extracts vision-text multimodal knowledge from each frame and its corresponding textualized prediction via the learned VTE and a frozen BERT. The text embeddings help provide more representative cues, but also may cause misdirection due to prediction errors. Thus, VAC randomly drops or replaces historical predictions to increase the error tolerance to avoid collapsing onto the last few predictions. Qualitative and quantitative experiments demonstrate that the proposed method achieves superior performance compared to other state-of-the-art methods, with an average F1-score improvement of at least 7.66%. Hongkuan Shi, Shiquan He, Xinxia Feng, Jiazhi Liao, Qiang Li 0018, Zhiwei Wang 0002 |
IEEE J. Biomed. Health Informatics | 11 |
| 2025 | Single-Slice Semi-Supervised 3D Medical Image Segmentation via Correlation Information Enhancement and Hybrid Pseudo Mask GenerationabstractThree-dimensional (3D) medical image segmentation typically demands extensive labeled training samples, which is prohibitively time-consuming and requires significant expertise. Although this demand can be mitigated by special learning paradigms such as semi-supervised learning, the cost is still high due to the reader-unfriendly 3D data structure. In this paper, we seek a solution of robust 3D segmentation using extremely simplified annotation that delineates only a single slice per each volume for only a subset of the 3D samples. To this end, we propose two innovative modules: a correlation-enhanced 3D segmentation model (CE-Seg) and a hybrid 3D pseudo mask generator (Hy-Gen). CE-Seg aims to comprehensively understand the 3D targets under super-sparse single-slice supervision by maximizing its ability to mine correlations across slices, spaces and scales. Specifically, CE-Seg mimics the radiologist's interpretation by 'seeing' a dynamically scrolling 3D image to enrich the slice-correlated context. It also introduces a drop-then-restoration self-played task to enhance the spatial correlations of features, and uses a bidirectional cascaded attention to interactively fuse features across different scales. To train CS-Seg, Hy-Gen combines learning-based and learning-free strategies to generate reliable pseudo 3D masks as supervisions. Concretely, Hy-Gen first employs a level-set evolution to 'spread' the single annotation to its neighboring slices as initialization. It then builds a teacher-student framework to progressively refine the initialized 3D mask by dynamically merging the predictions of the CS-Seg's teacher-copy. Extensive experiments on three public and one in-house datasets indicate that our method exceeds eight state-of-the-art semi-supervised methods by at least 3$\%$ in dice, and is even on par with the full-supervised counterpart. Quan Zhou 0011, Mingwei Wen, Mingyue Ding, Yixin Su 0002, Zhiwei Wang 0002 |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | Decoupling Feature Representations of Ego and Other Modalities for Incomplete Multi-modal Brain Tumor SegmentationabstractMulti-modal brain tumor segmentation typically involves four magnetic resonance imaging (MRI) modalities, while incomplete modalities significantly degrade performance. Existing solutions employ explicit or implicit modality adaptation, aligning features across modalities or learning a fused feature robust to modality incompleteness. They share a common goal of encouraging each modality to express both itself and the others. However, the two expression abilities are entangled as a whole in a seamless feature space, resulting in prohibitive learning burdens. In this paper, we propose DeMoSeg to enhance the modality adaptation by Decoupling the task of representing the ego and other Modalities for robust incomplete multi-modal Segmentation. The decoupling is super lightweight by simply using two convolutions to map each modality onto four feature sub-spaces. The first sub-space expresses itself (Self-feature), while the remaining sub-spaces substitute for other modalities (Mutual-features). The Self- and Mutual-features interactively guide each other through a carefully-designed Channel-wised Sparse Self-Attention (CSSA). After that, a Radiologist-mimic Cross-modality expression Relationships (RCR) is introduced to have available modalities provide Self-feature and also ‘lend’ their Mutual-features to compensate for the absent ones by exploiting the clinical prior knowledge. The benchmark results on BraTS2020, BraTS2018 and BraTS2015 verify the DeMoSeg’s superiority thanks to the alleviated modality adaptation difficulty. Concretely, for BraTS2020, DeMoSeg increases Dice by at least 0.92%, 2.95% and 4.95% on whole tumor, tumor core and enhanced tumor regions, respectively, compared to other state-of-the-arts. Codes are at https://github.com/kk42yy/DeMoSeg. Kaixiang Yang 0004, Wenqi Shan, Xikai Yang, Xi Wang 0013, Pheng-Ann Heng, Qiang Li 0018, Zhiwei Wang 0002 |
BIBM | 9 |
| 2024 | Improved Self-supervised Monocular Endoscopic Depth Estimation based on Pose Alignment-friendly Dynamic View Selectionabstract3D depth estimation in monocular endoscopic videos typically relies on self-supervised learning due to the absence of in-body ground-truth depth labels. This learning paradigm optimizes predicted 3D depth using adjacent frames by aligning camera poses and minimizing photometric discrepancies. However, previous efforts often suffer from insufficient visual cues caused by minimal camera movement between adjacent frames, resulting in significant pose alignment errors and reduced learning effectiveness. In this paper, we propose DVSMono, a dynamic view selection-based self-supervised monocular depth estimation that, for the first time, leverages pose alignment-friendly views instead of static adjacent ones. DVSMono employs a pretrained pose estimator to calculate poses between the target frame and its surrounding frames over a larger time span. Accurate poses can create highly-unanimous geometric cost volumes across different depth proposals, while errors often appear as outliers. Motivated by this, we design Temporally-Consistent View Scoring (TCVS), which selects candidate frames with cost volumes showing minimal temporal variance. Additionally, we introduce Photometric-Valid Region Computation (PVRC), which filters out unreliable regions between frames by identifying non-overlapping fields of view and bidirectionally inconsistent occlusion pixels. Ultimately, among the TCVS-selected candidates, DVSMono assigns each target frame its most suitable source frame with the largest PVRC-filtered valid regions. This enables the training of a robust depth estimator, benefiting from improved self-supervised learning built upon dynamically-selected, pose alignment-friendly source-target pairs. Benchmark results on three public datasets demonstrate that DVSMono exceeds the existing state-of-the-art methods by reducing Abs Rel by at least 6.35%. Codes are at https://github.com/adam99goat/DVSMono. Shiquan He, Hao Wang 0218, Qiang Li 0018, Zhiwei Wang 0002 |
BIBM | 7 |
| 2024 | SALI: Short-Term Alignment and Long-Term Interaction Network for Colonoscopy Video Polyp Segmentation
Zhenyu Yi, Qiang Li 0018, Zhiwei Wang 0002 |
MICCAI (6) | 7 |
| 2024 | Noise Removed Inconsistency Activation Map for Unsupervised Registration of Brain Tumor MRI Between Pre-operative and Follow-Up Phases
Chongwei Wu, Xiaoyu Zeng, Hao Wang 0218, Qiang Li 0018, Zhiwei Wang 0002 |
MICCAI (2) | 7 |
| 2024 | Co-learning-assisted progressive dense fusion network for cardiovascular disease detection using ECG and PCG signals
Haobo Zhang 0003, Peng Zhang 0106, Fan Lin, Lianying Chao, Zhiwei Wang 0002, Qiang Li 0018 |
Expert Syst. Appl. | 5 |
| 2024 | Multi-Feature Decision Fusion Network for Heart Sound Abnormality Detection and ClassificationabstractThe heart sound reflects the movement status of the cardiovascular system and contains the early pathological information of cardiovascular diseases. Automatic heart sound diagnosis plays an essential role in the early detection of cardiovascular diseases. In this study, we aim to develop a novel end-to-end heart sound abnormality detection and classification method, which can be adapted to different heart sound diagnosis tasks. Specifically, we developed a Multi-feature Decision Fusion Network (MDFNet) composed of a Multi-dimensional Feature Extraction (MFE) module and a Multi-dimensional Decision Fusion (MDF) module. The MFE module extracted spatial features, multi-level temporal features and spatial-temporal fusion features to learn heart sound characteristics from multiple perspectives. Through deep supervision and decision fusion, the MDF module made the multi-dimensional features extracted by the MFE module more discriminative, and fused the decision results of multi-dimensional features to integrate complementary information. Furthermore, attention modules were embedded in the MDFNet to emphasize the fundamental heart sounds containing effective feature information. Finally, we proposed an efficient data augmentation method to circumvent the diagnosis performance degradation caused by the lack of cardiac cycle segmentation in other end-to-end methods. The developed method achieved an overall accuracy of 94.44% and a F1-score of 86.90% on the binary classification task and a F1-score of 99.30% on the five-classification task. Our method outperformed other state-of-the-art methods and had good clinical application prospects. Haobo Zhang 0003, Peng Zhang 0106, Zhiwei Wang 0002, Lianying Chao, Qiang Li 0018 |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | Robust Semi-Supervised 3D Medical Image Segmentation With Diverse Joint-Task Learning and Decoupled Inter-Student LearningabstractSemi-supervised segmentation is highly significant in 3D medical image segmentation. The typical solutions adopt a teacher-student dual-model architecture, and they constrain the two models' decision consistency on the same segmentation task. However, the scarcity of medical samples can lower the diversity of tasks, reducing the effectiveness of consistency constraint. The issue can further worsen as the weights of the models gradually become synchronized. In this work, we have proposed to construct diverse joint-tasks using masked image modelling for enhancing the reliability of the consistency constraint, and develop a novel architecture consisting of a single teacher but multiple students to enjoy the additional knowledge decoupled from the synchronized weights. Specifically, the teacher and student models 'see' varied randomly-masked versions of an input, and are trained to segment the same targets but reconstruct different missing regions concurrently. Such joint-task of segmentation and reconstruction can have the two learners capture related but complementary features to derive instructive knowledge when constraining their consistency. Moreover, two extra students join the original one to perform an inter-student learning. The three students share the same encoding but different decoding designs, and learn decoupled knowledge by constraining their mutual consistencies, preventing themselves from suboptimally converging to the biased predictions of the dictatorial teacher. Experimental on four medical datasets show that our approach performs better than six mainstream semi-supervised methods. Particularly, our approach achieves at least 0.61% and 0.36% higher Dice and Jaccard values, respectively, than the most competitive approach on our in-house dataset. The code will be released at https://github.com/zxmboshi/DDL. Quan Zhou 0011, Mingyue Ding, Zhiwei Wang 0002, Xuming Zhang 0003 |
IEEE Trans. Medical Imaging | 5 |
| 2023 | Robust One-Shot Segmentation of Brain Tissues via Image-Aligned Style TransformationabstractOne-shot segmentation of brain tissues is typically a dual-model iterative learning: a registration model (reg-model) warps a carefully-labeled atlas onto unlabeled images to initialize their pseudo masks for training a segmentation model (seg-model); the seg-model revises the pseudo masks to enhance the reg-model for a better warping in the next iteration. However, there is a key weakness in such dual-model iteration that the spatial misalignment inevitably caused by the reg-model could misguide the seg-model, which makes it converge on an inferior segmentation performance eventually. In this paper, we propose a novel image-aligned style transformation to reinforce the dual-model iterative learning for robust one-shot segmentation of brain tissues. Specifically, we first utilize the reg-model to warp the atlas onto an unlabeled image, and then employ the Fourier-based amplitude exchange with perturbation to transplant the style of the unlabeled image into the aligned atlas. This allows the subsequent seg-model to learn on the aligned and style-transferred copies of the atlas instead of unlabeled images, which naturally guarantees the correct spatial correspondence of an image-mask training pair, without sacrificing the diversity of intensity patterns carried by the unlabeled images. Furthermore, we introduce a feature-aware content consistency in addition to the image-level similarity to constrain the reg-model for a promising initialization, which avoids the collapse of image-aligned style transformation in the first iteration. Experimental results on two public datasets demonstrate 1) a competitive segmentation performance of our method compared to the fully-supervised method, and 2) a superior performance over other state-of-the-art with an increase of average Dice by up to 4.67%. The source code is available at: https://github.com/JinxLv/One-shot-segmentation-via-IST. Jinxin Lv, Xiaoyu Zeng, Sheng Wang 0013, Zhiwei Wang 0002, Qiang Li 0018 |
AAAI | 5 |
| 2023 | Rethinking Low-Dose CT Synthesis: Degrading Normal-Dose CT from Origin for Pairwise Training of CT DenoiserabstractTraining a denoiser for translating low-dose to normal-dose computed tomography (LDCT to NDCT) requires collecting spatially corresponded paired data, yet it is impractical. Existing works aim to ‘add’ noise to real NDCT for synthesizing paired LDCT, but often failed to model the complex noise-structure entanglement, and therefore led to a suboptimal trained denoiser. In this paper, we propose a fresh solution to mimic more reasonable LDCT by modeling the structure-entangled noise patterns. Our motivation is based on a fact that the noise is originally carried by the source sinogram, and then passed to the reconstructed CT and intertwined with the structures via the filtered back-projection (FBP). In light of this, we first convert a NDCT image to the source sinogram by forward projection, and then design a Degrading CT Network (DeCTNet) to learn noise from the origin. Specifically, DeCTNet consists of two sequential networks, i.e., sinogram and reconstruction networks (SinoNet and RecoNet). SinoNet learns to directly add noise to the NDCT-converted sinogram, and RecoNet learns to further process the reconstructed CT image guided by real but unpaired LDCT images. DeCTNet utilizes a differentiable FBP operator to naturally reinvent the sinogram noise to the entangled noise-structure patterns in CT, and thus bridge SinoNet and RecoNet in an end-to-end training. Moreover, adversarial loss and content-fidelity loss are jointly minimized to effectively learn the noise characteristics and content retention in LDCT synthesis. Both quantitative and qualitative evaluations demonstrate that the denoiser trained using DeCTNet-synthesized pairs outperforms those trained using pairs synthesized by other state-of-the-arts, and radiologists are often not able to distinguish our denoised LDCTs from the real NDCTs. Lianying Chao, Taotao Zhang, Wenqi Shan, Qiang Li 0018, Zhiwei Wang 0002 |
BIBM | 6 |
| 2023 | Towards Learning-based Surgical Planning of Glioma Resection via a Contrastively Constrained Siamese Neural NetworkabstractThe risk assessment of craniotomy paths is a vital stage in the automated surgical planning of glioma resection. However, current methods heavily rely on the manually-assigned risk coefficients tied to factors like path length and intersecting brain structures. This overreliance on empirical coefficients poses limitations in terms of their adaptability to broader cohorts. In this paper, we propose an innovative risk assessment approach for automated surgical planning by leveraging deep learning techniques. Specifically, we develop a Siamese Neural Network (SNN) with two weight-sharing branches, and also compile a glioma dataset, containing the extracted brain vessels, regions, and the craniotomy region on the scalp. During training, SNN learns to predict two risk values for a pair of paths, which are sampled inside and outside the craniotomy region, respectively. A contrastive loss is then employed to guide SNN by penalizing the higher predicted risk inside the craniotomy region or the lower risk outside the region. Such contrastive constraint empowers effective learning of SNN in our unique scenario, where the true risk values of surgical paths are unavailable. During inference, the trained SNN generates a risk map on the scalp, and the path is determined by finding the local minimum on the risk map. To the best of our knowledge, this work proposes one of the first attempts to successfully incorporate deep learning techniques into the risk assessment process of glioma resection planning. Quantitative and neurosurgeons’ qualitative evaluations of six testing cases proved the effectiveness and superiority of our learning-based approach. A consensus is reached that our method shows great potential to help surgeons make safer and more accurate plans in practice. Wenqi Shan, Yingjie Guo, Lianying Chao, Haobo Zhang 0003, Qiang Li 0018, Zhiwei Wang 0002 |
BIBM | 8 |
| 2023 | Dual-view Correlation Hybrid Attention Network for Robust Holistic Mammogram ClassificationabstractMammogram image is important for breast cancer screening, and typically obtained in a dual-view form, i.e., cranio-caudal (CC) and mediolateral oblique (MLO), to provide complementary information for clinical decisions. However, previous methods mostly learn features from the two views independently, which violates the clinical knowledge and ignores the importance of dual-view correlation in the feature learning. In this paper, we propose a dual-view correlation hybrid attention network (DCHA-Net) for robust holistic mammogram classification. Specifically, DCHA-Net is carefully designed to extract and reinvent deep feature maps for the two views, and meanwhile to maximize the underlying correlations between them. A hybrid attention module, consisting of local relation and non-local attention blocks, is proposed to alleviate the spatial misalignment of the paired views in the correlation maximization. A dual-view correlation loss is introduced to maximize the feature similarity between corresponding strip-like regions with equal distance to the chest wall, motivated by the fact that their features represent the same breast tissues, and thus should be highly-correlated with each other. Experimental results on the two public datasets, i.e., INbreast and CBIS-DDSM, demonstrate that the DCHA-Net can well preserve and maximize feature correlations across views, and thus outperforms previous state-of-the-art methods for classifying a whole mammogram as malignant or not. Zhiwei Wang 0002, Junlin Xian, Kangyi Liu, Xin Li 0001, Qiang Li 0018, Xin Yang 0008 |
IJCAI | 1 |
| 2023 | PSDP: Pseudo-supervised dual-processing for low-dose cone-beam computed tomography reconstruction
Lianying Chao, Wenqi Shan, Wenting Xu, Haobo Zhang 0003, Zhiwei Wang 0002, Qiang Li 0018 |
Expert Syst. Appl. | 6 |
| 2023 | Accurate Cobb Angle Estimation on Scoliosis X-Ray Images via Deeply-Coupled Two-Stage Network With Differentiable Cropping and Random PerturbationabstractAutomated Cobb angle estimation on X-ray images is crucial to scoliosis diagnosis. The existing efforts are typically two extremes, which either laboriously detect the raw vertebral landmarks or directly regress Cobb angles from the entire image. In this paper, we propose a novel two-stage end-to-end method as a balanced solution, to avoid vulnerability to false landmarks, and to preserve flexibility in clinical usages. Concretely, we cascade two stages sequentially for detecting vertebrae and then regressing their bending directions instead of raw landmarks. In the detection stage, we combine two networks called LocNet and SegNet to robustly localize vertebrae, and meanwhile to suppress the false positives by additionally segmenting the whole spine. In the subsequent stage, we introduce a regression network named RegNet to accurately regress bending directions of localized vertebrae. Furthermore, the vertebra-aligned local regions on LocNet's intermediate features are cropped via RoIAlign-pooling, and RegNet inherits the cropped regions to learn only feature residuals. By doing so, the regression difficulty can be dramatically alleviated, and the two stages are deeply coupled and mutually guided in an end-to-end training. Moreover, a random perturbation on the inherited features further enhances RegNet's robustness. We benchmark our method on both public and private datasets, and the errors are 2.92 $\pm$ 2.34$^{\circ }$ and 6.87 $\pm$ 6.26% in terms of CMAE and SMAPE on the widely-employed AASCE dataset, outperforming other state-of-the-arts by at least 16.81% and 6.15%, respectively. Also, a clinical user study verifies our promising flexibility for allowing convenient rectifications to further decrease errors by a large marge. Yuanhuai Liang, Jinxin Lv, Dun Li, Xin Yang 0008, Zhiwei Wang 0002, Qiang Li 0018 |
IEEE J. Biomed. Health Informatics | 5 |
| 2023 | Bidirectional Semi-Supervised Dual-Branch CNN for Robust 3D Reconstruction of Stereo Endoscopic Images via Adaptive Cross and Parallel SupervisionsabstractSemi-supervised learning via teacher-student network can train a model effectively on a few labeled samples. It enables a student model to distill knowledge from the teacher's predictions of extra unlabeled data. However, such knowledge flow is typically unidirectional, having the accuracy vulnerable to the quality of teacher model. In this paper, we seek to robust 3D reconstruction of stereo endoscopic images by proposing a novel fashion of bidirectional learning between two learners, each of which can play both roles of teacher and student concurrently. Specifically, we introduce two self-supervisions, i.e., Adaptive Cross Supervision (ACS) and Adaptive Parallel Supervision (APS), to learn a dual-branch convolutional neural network. The two branches predict two different disparity probability distributions for the same position, and output their expectations as disparity values. The learned knowledge flows across branches along two directions: a cross direction (disparity guides distribution in ACS) and a parallel direction (disparity guides disparity in APS). Moreover, each branch also learns confidences to dynamically refine its provided supervisions. In ACS, the predicted disparity is softened into a unimodal distribution, and the lower the confidence, the smoother the distribution. In APS, the incorrect predictions are suppressed by lowering the weights of those with low confidence. With the adaptive bidirectional learning, the two branches enjoy well-tuned mutual supervisions, and eventually converge on a consistent and more accurate disparity estimation. The experimental results on four public datasets demonstrate our superior accuracy over other state-of-the-arts with a relative decrease of averaged disparity error by at least 9.76%. Hongkuan Shi, Zhiwei Wang 0002, Dun Li, Xin Yang 0008, Qiang Li 0018 |
IEEE Trans. Medical Imaging | 2 |
| 2022 | Improving the Quality of Sparse-view Cone-Beam Computed Tomography via Reconstruction-Friendly Interpolation Network
Lianying Chao, Wenqi Shan, Haobo Zhang 0003, Zhiwei Wang 0002, Qiang Li 0018 |
ACCV (6) | 5 |
| 2022 | Sparse-view cone beam CT reconstruction using dual CNNs in projection domain and image domain
Lianying Chao, Zhiwei Wang 0002, Haobo Zhang 0003, Wenting Xu, Peng Zhang 0106, Qiang Li 0018 |
Neurocomputing | 2 |
| 2022 | Dual-domain attention-guided convolutional neural network for low-dose cone-beam computed tomography reconstruction
Lianying Chao, Peng Zhang 0106, Zhiwei Wang 0002, Wenting Xu, Qiang Li 0018 |
Knowl. Based Syst. | 4 |
| 2022 | Joint Progressive and Coarse-to-Fine Registration of Brain MRI via Deformation Field Integration and Non-Rigid Feature FusionabstractRegistration of brain MRI images requires to solve a deformation field, which is extremely difficult in aligning intricate brain tissues, e.g., subcortical nuclei, etc. Existing efforts resort to decomposing the target deformation field into intermediate sub-fields with either tiny motions, i.e., progressive registration stage by stage, or lower resolutions, i.e., coarse-to-fine estimation of the full-size deformation field. In this paper, we argue that those efforts are not mutually exclusive, and propose a unified framework for robust brain MRI registration in both progressive and coarse-to-fine manners simultaneously. Specifically, building on a dual-encoder U-Net, the fixed-moving MRI pair is encoded and decoded into multi-scale sub-fields from coarse to fine. Each decoding block contains two proposed novel modules: i) in Deformation Field Integration (DFI), a single integrated deformation sub-field is calculated, warping by which is equivalent to warping progressively by sub-fields from all previous decoding blocks, and ii) in Non-rigid Feature Fusion (NFF), features of the fixed-moving pair are aligned by DFI-integrated deformation field, and then fused to predict a finer sub-field. Leveraging both DFI and NFF, the target deformation field is factorized into multi-scale sub-fields, where the coarser fields alleviate the estimate of a finer one and the finer field learns to make up those misalignments insolvable by previous coarser ones. The extensive and comprehensive experimental results on both private and two public datasets demonstrate a superior registration performance of brain MRI images over progressive registration only and coarse-to-fine estimation only, with an increase by at most 8% in the average Dice. Jinxin Lv, Zhiwei Wang 0002, Hongkuan Shi, Haobo Zhang 0003, Sheng Wang 0013, Yilang Wang, Qiang Li 0018 |
IEEE Trans. Medical Imaging | 2 |
| 2021 | Towards Robust Dual-View Transformation via Densifying Sparse Supervision for Mammography Lesion Matching
Junlin Xian, Zhiwei Wang 0002, Kwang-Ting Cheng, Xin Yang 0008 |
MICCAI (5) | 2 |
| 2021 | Semi-supervised Learning via Improved Teacher-Student Network for Robust 3D Reconstruction of Stereo Endoscopic Imageabstract3D reconstruction of stereo endoscope image, as an enabling technique for varied surgical systems, e.g., medical droids, navigations, etc., suffers from severe overfitting problems due to scarce labels. Semi-supervised learning based on Teacher-Student Network (TSN) is a potential solution, which utilizes a supervised teacher model trained on available labeled data to teach a student model on all images via assigning them pseudo labels. However, TSN often faces a dilemma: if given only few labeled endoscope images, the teacher model will be trained to be defective and induce high-noised pseudo labels, degrading the student model significantly. To solve this, we propose an improved TSN for a robust 3D reconstruction of stereo endoscope image. Specifically, two novel modules are introduced: 1) a semi-supervised teacher model based on adversarial learning to produce mostly correct pseudo labels by forcing a consistency in predictions for both labeled and unlabeled data, and 2) a confidence network to further filter out noisy pseudo labels by estimating a confidence for each prediction of the teacher model. By doing so, the student model is able to distill knowledge from more accurate and noiseless pseudo labels, thus achieving improved performance. Experimental results on two public datasets show that our improved TSN achieves a superior performance than the state-of-the-arts by reducing the averaged disparity error by at least 13.5%. Hongkuan Shi, Zhiwei Wang 0002, Jinxin Lv, Yilang Wang, Peng Zhang 0106, Qiang Li 0018 |
ACM Multimedia | 2 |
| 2021 | Variation-Aware Federated Learning With Multi-Source Decentralized Medical Image DataabstractPrivacy concerns make it infeasible to construct a large medical image dataset by fusing small ones from different sources/institutions. Therefore, federated learning (FL) becomes a promising technique to learn from multi-source decentralized data with privacy preservation. However, the cross-client variation problem in medical image data would be the bottleneck in practice. In this paper, we propose a variation-aware federated learning (VAFL) framework, where the variations among clients are minimized by transforming the images of all clients onto a common image space. We first select one client with the lowest data complexity to define the target image space and synthesize a collection of images through a privacy-preserving generative adversarial network, called PPWGAN-GP. Then, a subset of those synthesized images, which effectively capture the characteristics of the raw images and are sufficiently distinct from any raw image, is automatically selected for sharing with other clients. For each client, a modified CycleGAN is applied to translate its raw images to the target image space defined by the shared synthesized images. In this way, the cross-client variation problem is addressed with privacy preservation. We apply the framework for automated classification of clinically significant prostate cancer and evaluate it using multi-source decentralized apparent diffusion coefficient (ADC) image data. Experimental results demonstrate that the proposed VAFL framework stably outperforms the current horizontal FL framework. As VAFL is independent of deep learning architectures for classification, we believe that the proposed framework is widely applicable to other medical image classification tasks. Zengqiang Yan, Jeffry Wicaksana, Zhiwei Wang 0002, Xin Yang 0008, Kwang-Ting Cheng |
IEEE J. Biomed. Health Informatics | 3 |
| 2020 | Multi-phase and Multi-level Selective Feature Fusion for Automated Pancreas Segmentation from CT Images
Xixi Jiang, Qingqing Luo, Zhiwei Wang 0002, Xin Li 0001, Kwang-Ting Cheng, Xin Yang 0008 |
MICCAI (4) | 3 |
| 2020 | Semi-supervised mp-MRI data synthesis with StitchLayer and auxiliary distance maximization
Zhiwei Wang 0002, Yi Lin 0009, Kwang-Ting Cheng, Xin Yang 0008 |
Medical Image Anal. | 1 |
| 2020 | Bi-Modality Medical Image Synthesis Using Semi-Supervised Sequential Generative Adversarial NetworksabstractIn this paper, we propose a bi-modality medical image synthesis approach based on sequential generative adversarial network (GAN) and semi-supervised learning. Our approach consists of two generative modules that synthesize images of the two modalities in a sequential order. A method for measuring the synthesis complexity is proposed to automatically determine the synthesis order in our sequential GAN. Images of the modality with a lower complexity are synthesized first, and the counterparts with a higher complexity are generated later. Our sequential GAN is trained end-to-end in a semi-supervised manner. In supervised training, the joint distribution of bi-modality images are learned from real paired images of the two modalities by explicitly minimizing the reconstruction losses between the real and synthetic images. To avoid overfitting limited training images, in unsupervised training, the marginal distribution of each modality is learned based on unpaired images by minimizing the Wasserstein distance between the distributions of real and fake images. We comprehensively evaluate the proposed model using two synthesis tasks based on three types of evaluate metrics and user studies. Visual and quantitative results demonstrate the superiority of our method to the state-of-the-art methods, and reasonable visual quality and clinical significance. Code is made publicly available at https://github.com/hust- linyi/Multimodal-Medical-Image-Synthesis. Xin Yang 0008, Yi Lin 0009, Zhiwei Wang 0002, Xin Li 0001, Kwang-Ting Cheng |
IEEE J. Biomed. Health Informatics | 3 |
| 2020 | Multi-Task Siamese Network for Retinal Artery/Vein Separation via Deep Convolution Along VesselabstractVascular tree disentanglement and vessel type classification are two crucial steps of the graph-based method for retinal artery-vein (A/V) separation. Existing approaches treat them as two independent tasks and mostly rely on ad hoc rules (e.g. change of vessel directions) and hand-crafted features (e.g. color, thickness) to handle them respectively. However, we argue that the two tasks are highly correlated and should be handled jointly since knowing the A/V type can unravel those highly entangled vascular trees, which in turn helps to infer the types of connected vessels that are hard to classify based on only appearance. Therefore, designing features and models isolatedly for the two tasks often leads to a suboptimal solution of A/V separation. In view of this, this paper proposes a multi-task siamese network which aims to learn the two tasks jointly and thus yields more robust deep features for accurate A/V separation. Specifically, we first introduce Convolution Along Vessel (CAV) to extract the visual features by convolving a fundus image along vessel segments, and the geometric features by tracking the directions of blood flow in vessels. The siamese network is then trained to learn multiple tasks: i) classifying A/V types of vessel segments using visual features only, and ii) estimating the similarity of every two connected segments by comparing their visual and geometric features in order to disentangle the vasculature into individual vessel trees. Finally, the results of two tasks mutually correct each other to accomplish final A/V separation. Experimental results demonstrate that our method can achieve accuracy values of 94.7%, 96.9%, and 94.5% on three major databases (DRIVE, INSPIRE, WIDE) respectively, which outperforms recent state-of-the-arts. Zhiwei Wang 0002, Xixi Jiang, Jingen Liu, Kwang-Ting Cheng, Xin Yang 0008 |
IEEE Trans. Medical Imaging | 1 |
| 2018 | StitchAD-GAN for Synthesizing Apparent Diffusion Coefficient Images of Clinically Significant Prostate Cancer
Zhiwei Wang 0002, Yi Lin 0009, Chunyuan Liao, Kwang-Ting Cheng, Xin Yang 0008 |
BMVC | 1 |
| 2018 | Monocular Camera Based Real-Time Dense Mapping Using Generative Adversarial NetworkabstractMonocular simultaneous localization and mapping (SLAM) is a key enabling technique for many computer vision and robotics applications. However, existing methods either can obtain only sparse or semi-dense maps in highly-textured image areas or fail to achieve a satisfactory reconstruction accuracy. In this paper, we present a new method based on a generative adversarial network,named DM-GAN, for real-time dense mapping based on a monocular camera. Specifcally, our depth generator network takes a semidense map obtained from motion stereo matching as a guidance to supervise dense depth prediction of a single RGB image. The depth generator is trained based on a combination of two loss functions, i.e. an adversarial loss for enforcing the generated depth maps to reside on the manifold of the true depth maps and a pixel-wise mean square error (MSE) for ensuring the correct absolute depth values. Extensive experiments on three public datasets demonstrate that our DM-GAN signifcantly outperforms the state-of-the-art methods in terms of greater reconstruction accuracy and higher depth completeness. Xin Yang 0008, Zhiwei Wang 0002, Qiaozhe Zhang, Wenyu Liu 0001, Chunyuan Liao, Kwang-Ting Cheng |
ACM Multimedia | 3 |
| 2018 | Automated Detection of Clinically Significant Prostate Cancer in mp-MRI Images Based on an End-to-End Deep Neural NetworkabstractAutomated methods for detecting clinically significant (CS) prostate cancer (PCa) in multi-parameter magnetic resonance images (mp-MRI) are of high demand. Existing methods typically employ several separate steps, each of which is optimized individually without considering the error tolerance of other steps. As a result, they could either involve unnecessary computational cost or suffer from errors accumulated over steps. In this paper, we present an automated CS PCa detection system, where all steps are optimized jointly in an end-to-end trainable deep neural network. The proposed neural network consists of concatenated subnets: 1) a novel tissue deformation network (TDN) for automated prostate detection and multimodal registration and 2) a dual-path convolutional neural network (CNN) for CS PCa detection. Three types of loss functions, i.e., classification loss, inconsistency loss, and overlap loss, are employed for optimizing all parameters of the proposed TDN and CNN. In the training phase, the two nets mutually affect each other and effectively guide registration and extraction of representative CS PCa-relevant features to achieve results with sufficient accuracy. The entire network is trained in a weakly supervised manner by providing only image-level annotations (i.e., presence/absence of PCa) without exact priors of lesions' locations. Compared with most existing systems which require supervised labels, e.g., manual delineation of PCa lesions, it is much more convenient for clinical usage. Comprehensive evaluation based on fivefold cross validation using 360 patient data demonstrates that our system achieves a high accuracy for CS PCa detection, i.e., a sensitivity of 0.6374 and 0.8978 at 0.1 and 1 false positives per normal/benign patient. Zhiwei Wang 0002, Chaoyue Liu 0002, Danpeng Cheng, Liang Wang 0052, Xin Yang 0008, Kwang-Ting Cheng |
IEEE Trans. Medical Imaging | 1 |
| 2017 | Joint Detection and Diagnosis of Prostate Cancer in Multi-parametric MRI Based on Multimodal Convolutional Neural Networks
Xin Yang 0008, Zhiwei Wang 0002, Chaoyue Liu 0002, Hung Le Minh, Kwang-Ting Cheng, Liang Wang 0052 |
MICCAI (3) | 2 |
| 2017 | DeepCADx: Automated Prostate Cancer Detection and Diagnosis in mp-MRI based on Multimodal Convolutional Neural NetworksabstractIn this paper, we present DeepCADx, a computer-aided prostate detection and diagnosis (CADx) system powered by a novel deep convolutional neural networks (CNNs). Specifically, the developed DeepCADx system processes multi-parametric magnetic resonance imaging (mp-MRI) sequences in three major steps: 1) pre-processing which registers images from different modalities and detect prostates, 2) multimodal CNNs which jointly identifies images containing prostate cancers (PCa) and generate cancer response maps (CRM) with each pixel indicating the probability to be cancerous, and 3) post-processing which localize lesion in CRMs and assess the aggressiveness (i.e. Gleason score) of each localized lesion using multimodal CNN features and a 5-class SVM classifier. Zhiwei Wang 0002, Chaoyue Liu 0002, Xiang Bai, Xin Yang 0008 |
ACM Multimedia | 1 |
| 2017 | Joint Face Detection and Initialization for Face Alignment
Zhiwei Wang 0002, Xin Yang 0008 |
MMM (1) | 1 |
| 2017 | V-Head: Face Detection and Alignment for Facial Augmented Reality Applications
Zhiwei Wang 0002, Xin Yang 0008 |
MMM (2) | 1 |
| 2017 | Co-trained convolutional neural networks for automated detection of prostate cancer in multi-parametric MRI
Xin Yang 0008, Chaoyue Liu 0002, Zhiwei Wang 0002, Hung Le Minh, Liang Wang 0052, Kwang-Ting Cheng |
Medical Image Anal. | 3 |