Qiang Li 0018

dblp:72/872-18 · DBLP profile ↗
← Back
42ranked-venue papers
1as first author
42since 2021 · last 2026
0000-0002-9815-4432ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 21 · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 18 since 2021Artificial intelligence and machine learning · 15 · 1 first-author · 15 since 2021
YearPublicationVenuePosition
2026 RouterNet: Hierarchical Point Routing Network for Robust Vertebral Landmark Localization on AP X-ray Images
abstract
Locating vertebral landmarks on anteroposterior (AP) X-ray images is challenging due to the tissue overlap. Despite the great progress of heatmap-based methods, they often predict missing/false points, which are intolerable in the downstream applications like scoliosis assessment. In this paper, we instead modernize the classic point-regression scheme, and propose a novel model termed RouterNet to locate the 68 vertebral landmarks completely and accurately. RouterNet starts from an initial root point, and then gradually routes it onto more and more points with finer and finer semantics. RouterNet naturally couples such point routing process with its hierarchical and multi-scale feature learning. That is, lower-scale feature maps are utilized to regress points with coarser semantics, and the regressed points pilot a more focused local feature extraction on the next higher-scale map to route onto their subsequent positions with finer semantics. With this divide-and-conquer, RouterNet alleviates the task difficulty, and can robustly localize by routing from the whole spinal center to 17 vertebral centers, and further to their 68 corner points. Extensive and comprehensive experiments on both public and private datasets demonstrate our superior performance over other state-of-the-arts, by decreasing NMSE by 73.8% for landmark localization, and SMAPE by 14.8% for the downstream scoliosis assessment.
Yingjie Guo, Jinxin Lv, Qiang Li 0018, Zhiwei Wang 0002
AAAI4
2026 Pairing-free Group-level Knowledge Distillation for Robust Gastrointestinal Lesion Classification in White-Light Endoscopy
abstract
White-Light Imaging (WLI) is the standard for endoscopic cancer screening, but Narrow-Band Imaging (NBI) offers superior diagnostic details. A key challenge is transferring knowledge from NBI to enhance WLI-only models, yet existing methods are critically hampered by their reliance on paired NBI-WLI images of the same lesion, a costly and often impractical requirement that leaves vast amounts of clinical data untapped. In this paper, we break this paradigm by introducing PaGKD, a novel Pairing-free Group-level Knowledge Distillation framework that that enables effective cross-modal learning using unpaired WLI and NBI data. Instead of forcing alignment between individual, often semantically mismatched image instances, PaGKD operates at the group level to distill more complete and compatible knowledge across modalities. Central to PaGKD are two complementary modules: (1) Group-level Prototype Distillation (GKD-Pro) distills compact group representations by extracting modality-invariant semantic prototypes via shared lesion-aware queries; (2) Group-level Dense Distillation (GKD-Den) performs dense cross-modal alignment by guiding group-aware attention with activation-derived relation maps. Together, these modules enforce global semantic consistency and local structural coherence without requiring image-level correspondence. Extensive experiments on four clinical datasets demonstrate that PaGKD consistently and significantly outperforms state-of-the-art methods, boosting AUC by 3.3%, 1.1%, 2.8%, and 3.2%, respectively, establishing a new direction for cross-modal learning from unpaired data.
Qimei Wang, Yingjie Guo, Qiang Li 0018, Zhiwei Wang 0002
AAAI4
2026 FIA-Edit: Frequency-Interactive Attention for Efficient and High-Fidelity Inversion-Free Text-Guided Image Editing
abstract
Text-guided image editing has advanced rapidly with the rise of diffusion models. While flow-based inversion-free methods offer high efficiency by avoiding latent inversion, they often fail to effectively integrate source information, leading to poor background preservation, spatial inconsistencies, and over-editing due to the lack of effective integration of source information. In this paper, we present FIA-Edit, a novel inversion-free framework that achieves high-fidelity and semantically precise edits through a Frequency-Interactive Attention. Specifically, we design two key components: (1) a Frequency Representation Interaction (FRI) module that enhances cross-domain alignment by exchanging frequency components between source and target features within self-attention, and (2) a Feature Injection (FIJ) module that explicitly incorporates source-side queries, keys, values, and text embeddings into the target branch's cross-attention to preserve structure and semantics. Comprehensive and extensive experiments demonstrate that FIA-Edit supports high-fidelity editing at low computational cost (~6s per 512 * 512 image on an RTX 4090) and consistently outperforms existing methods across diverse tasks in visual quality, background fidelity, and controllability. Furthermore, we are the first to extend text-guided image editing to clinical applications. By synthesizing anatomically coherent hemorrhage variations in surgical images, FIA-Edit opens new opportunities for medical data augmentation and delivers significant gains in downstream bleeding classification.
Kaixiang Yang 0004, Boyang Shen, Xin Li 0001, Yuchen Dai, Yueran Ma, Qiang Li 0018, Zhiwei Wang 0002
AAAI8
2026 A dual-graph neural network for glioblastoma progression diagnosis
abstract
Glioblastoma (GBM) is a malignant brain tumor with poor prognosis even after extensive treatment. This poor prognosis is essentially due to the high rate of post-treatment pseudoprogression (PsP), which can be easily confused with true tumor progression (TTP). Thus, there has been a pressing need for developing reliable non-invasive methods to distinguish between these two types of tumor progression in clinical practice. Diffusion tensor imaging has demonstrated a high potential for automated GBM diagnosis. However, performance on this diagnosis task is typically restricted by data insufficiency and the difficulty of fine-grained classification. Hence, we developed a dual-graph neural network with a randomness-constrained strategy, which jointly achieves discriminative feature extraction and stable label inference. Specifically, we designed a dual-graph architecture comprising two structures. One is a multi-instance-learning-based graph embedding, which extracts features of local target regions. The other is a metric graph that accounts for interactions between patients and consequently enables patient-level label inference. Subsequently, we introduced a randomness-constraint strategy to maximize the utilization of historical data and employed data relationships to obtain highly discriminative graph structures. These aspects boosted performance stability on small-sample data. Extensive experiments have demonstrated promising performance of the proposed model for distinguishing between PsP and TTP. Furthermore, its scalability was verified on two public datasets addressing fine-grained classification challenges in different tasks. Overall, the proposed model provides an effective solution for GBM progression assessment and stable fine-grained classification under small-sample conditions. The code will be released at https://github.com/SJTUBME-QianLab/GBM-Dual_GNN .
Qiang Li 0018, Hailiang Tang, Xiaohua Qian
Pattern Recognit.1
2026 StyleSeg V2: Towards robust single-label-supervised segmentation of brain tissue via optimization-free registration error perception
Chongwei Wu, Xiaoyu Zeng, Tingwei Quan, Jinxin Lv, Qiang Li 0018, Zhiwei Wang 0002
Pattern Recognit.7
2026 Improving 3D Thin Vessel Segmentation in Brain TOF-MRA via a Dual-Space Context-Aware Network
abstract
3D cerebrovascular segmentation poses a significant challenge, akin to locating a line within a vast 3D environment. This complexity can be substantially reduced by projecting the vessels onto a 2D plane, enabling easier segmentation. In this paper, we create a vessel-segmentation-friendly space using a clinical visualization technique called maximum intensity projection (MIP). Leveraging this, we propose a Dual-space Context-Aware Network (DCANet) for 3D vessel segmentation, designed to capture even the finest vessel structures accurately. DCANet begins by transforming a magnetic resonance angiography (MRA) volume into a 3D Regional-MIP volume, where each Regional-MIP slice is constructed by projecting adjacent MRA slices. This transformation highlights vessels as prominent continuous curves rather than the small circular or ellipsoidal cross-sections seen in MRA slices. DCANet encodes vessels separately in the MRA and the projected Regional-MIP spaces and introduces the Regional-MIP Image Fusion Block (MIFB) between these dual spaces to selectively integrate contextual features from Regional-MIP into MRA. Following dual-space encoding, DCANet employs a Dual-mask Spatial Guidance TransFormer (DSGFormer) decoder to focus on vessel regions while effectively excluding background areas, which reduces the learning burden and improves segmentation accuracy. We benchmark DCANet on four datasets: two public datasets, TubeTK and IXI-IOP, and two in-house datasets, Xiehe and IXI-HH. The results demonstrate that DCANet achieves superior performance, with improvements in average DSC values of at least 2.26%, 2.17%, 2.62%, and 2.58% for thin vessels, respectively.
Wenqi Shan, Qiang Li 0018, Zhiwei Wang 0002
IEEE J. Biomed. Health Informatics4
2026 Cross-Correlation Rectification for Robust Deformable Registration of Brain Tumor MRI Between Preoperative and Postoperative Phases
abstract
A key challenge in registering pre- and post-operative brain tumor images lies in the anatomical inconsistencies caused by pathological changes and surgical resections. Recent efforts have addressed this issue by masking affected regions during optimization, but such approaches discard contextual information and rely on CNN backbones that implicitly model deformation, often overfitting to distant normal tissues and failing to capture the severe nonlinear distortions near the tumor. Correlation-based alternatives enhance generalization to diverse deformation patterns by explicitly modeling geometric correspondences, yet they frequently yield unreliable matches in and around tumor regions, disrupting the deformation field. In this paper, we propose Cross-correlation Rectification-based Registration Network (CRRNet), the first framework that introduces an active rectification mechanism specifically for robustness and structurally coherent pre- to post-operative brain tumor image registration. Specifically, CRRNet achieves this through two complementary modules: 1) a Cross-correlation Analysis-Based Inconsistency module that identifies invalid correspondences via bidirectional loop-closure evaluation on cross-correlations, and 2) a Dual-level Correspondence Rectification module that adaptively integrates contextually reliable correlations from local and long-range perspectives to restore structurally coherent matches. This synergistic design retains the strengths of cross-correlation while effectively mitigating correspondence mismatches. Extensive experiments on multiple tumor benchmarks demonstrate the superiority of CRRNet. Specifically, on the BraTS-Reg dataset, it reduces the mean registration errors by 8.18% in near-tumor regions and 3.95% in far-from-tumor regions, surpassing state-of-the-art methods.
Chongwei Wu, Xiaoyu Zeng, Shuxian Niu, Qiang Li 0018, Zhiwei Wang 0002
IEEE J. Biomed. Health Informatics6
2025 MonoBox: Tightness-Free Box-Supervised Polyp Segmentation Using Monotonicity Constraint
abstract
We propose MonoBox, an innovative box-supervised segmentation method constrained by monotonicity to liberate its training from the user-unfriendly box-tightness assumption. In contrast to conventional box-supervised segmentation, where the box edges must precisely touch the target boundaries, MonoBox leverages imprecisely-annotated boxes to achieve robust pixel-wise segmentation. The 'linchpin' is that, within the noisy zones around box edges, MonoBox discards the traditional misguiding multiple-instance learning loss, and instead optimizes a carefully-designed objective, termed monotonicity constraint. Along directions transitioning from the foreground to background, this new constraint steers responses to adhere to a trend of monotonically decreasing values. Consequently, the originally unreliable learning within the noisy zones is transformed into a correct and effective monotonicity optimization. Moreover, an adaptive label correction is introduced, enabling MonoBox to enhance the tightness of box annotations using predicted masks from the previous epoch and dynamically shrink the noisy zones as training progresses. We verify MonoBox in the box-supervised segmentation task of polyps, where satisfying box-tightness is challenging due to the vague boundaries between the polyp and normal tissues. Experiments on both public synthetic and in-house real noisy datasets demonstrate that MonoBox exceeds other anti-noise state-of-the-arts by improving Dice by at least 5.5% and 3.3%, respectively.
Zhenyu Yi, Qiang Li 0018, Zhiwei Wang 0002
AAAI6
2025 TDD: Task-Aware Diffusion Dual-Matching for Robust Electrocardiogram Denoising
abstract
Electrocardiograms (ECGs) play a critical role in the diagnosis and monitoring of cardiovascular diseases. However, their signal quality is often compromised by various types of noise which can impair diagnostic accuracy and destabilize downstream automated analysis. Recent deep learning-based denoising methods have demonstrated promising results, but they often focus solely on signal reconstruction accuracy while overlooking consistency with downstream tasks, such as waveform segmentation or disease classification. This misalignment may lead to clean-looking signals that are nonetheless structurally misleading or clinically unreliable. To address these limitations, we propose the Task-aware Diffusion Dual-matching (TDD) method, a novel denoising framework that integrates multi-task learning with a modified diffusion process. Unlike conventional multi-task approaches that treat auxiliary tasks as feature enhancers, TDD leverages auxiliary predictions to directly complement and guide the denoising process. Specifically, TDD uses noisy ECG signals as the diffusion starting point to reduce reconstruction uncertainty, and segmentation outputs provide structure-aware gradients that refine the denoising trajectory, which further encourages more confident, unimodal segmentation predictions. This dual-matching process forms a positive feedback loop: cleaner signals enable more accurate segmentation, which in turn reinforces denoising with semantically meaningful guidance. Notably, segmentation labels are not required during inference. Experiments on multiple public ECG datasets show that TDD outperforms existing denoising methods under various noise conditions. Moreover, it exhibits superior structure-preserving capabilities under complex cardiac rhythms and enhances downstream disease diagnosis performance, leading to more accurate clinical interpretation.
Siyi Fang, Zhiwei Wang 0002, Qiang Li 0018, Peng Zhang 0106
BIBM5
2025 Bidirectional Mammogram View Translation with Column-Aware and Implicit 3D Conditional Diffusion
abstract
Dual-view mammography, including craniocaudal (CC) and mediolateral oblique (MLO) projections, offers complementary anatomical views crucial for breast cancer diagnosis. However, in real-world clinical workflows, one view may be missing, corrupted, or degraded due to acquisition errors or compression artifacts, limiting the effectiveness of downstream analysis. View-to-view translation can help recover missing views and improve lesion alignment. Unlike natural images, this task in mammography is highly challenging due to large non-rigid deformations and severe tissue overlap in X-ray projections, which obscure pixel-level correspondences. In this paper, we propose Column-Aware and Implicit 3D Diffusion (CA3D-Diff), a novel bidirectional mammogram view translation framework based on conditional diffusion model. To address cross-view structural misalignment, we first design a column-aware Cross-Attention mechanism that leverages the geometric property that anatomically corresponding regions tend to lie in similar column positions across views. Furthermore, we introduce an implicit 3D structure reconstruction module that back-projects noisy 2D latents into a coarse 3D feature volume based on breast-view projection geometry. Extensive experiments demonstrate that CA3D-Diff achieves superior performance in bidirectional tasks, outperforming state-of-the-art methods in visual fidelity and structural consistency. Furthermore, the synthesized views effectively improve single-view malignancy classification in screening settings, demonstrating the practical value of our method in realworld diagnostics. Our code is available at https://github.com/lixinHUST/CA3D-Diff.
Xin Li 0001, Kaixiang Yang 0004, Qiang Li 0018, Zhiwei Wang 0002
BIBM3
2025 TiS-TSL: Image-Label Supervised Surgical Video Stereo Matching via Time-Switchable Teacher-Student Learning
abstract
Stereo matching in minimally invasive surgery (MIS) is essential for next-generation navigation and augmented reality. Yet, dense disparity supervision is nearly impossible due to anatomical constraints, typically limiting annotations to only a few image-level labels acquired before the endoscope enters deep body cavities. Teacher-Student Learning (TSL) offers a promising solution by leveraging a teacher trained on sparse labels to generate pseudo labels and associated confidence maps from abundant unlabeled surgical videos. However, existing TSL methods are confined to image-level supervision, providing only spatial confidence and lacking temporal consistency estimation. This absence of spatio-temporal reliability results in unstable disparity predictions and severe flickering artifacts across video frames. To overcome these challenges, we propose TiS-TSL, a novel time-switchable teacher-student learning framework for video stereo matching under minimal supervision. At its core is a unified model that operates in three distinct modes: ImagePrediction (IP), Forward Video-Prediction (FVP), and Backward Video-Prediction (BVP), enabling flexible temporal modeling within a single architecture. Enabled by this unified model, TiSTSL adopts a two-stage learning strategy. The Image-to-Video (I2V) stage transfers sparse image-level knowledge to initialize temporal modeling. The subsequent Video-to-Video (V2V) stage refines temporal disparity predictions by comparing forward and backward predictions to calculate bidirectional spatio-temporal consistency. This consistency identifies unreliable regions across frames, filters noisy video-level pseudo labels, and enforces temporal coherence. Experimental results on two public datasets demonstrate that TiS-TSL exceeds other image-based state-of-the-arts by improving TEPE and EPE by at least 2.11 % and$\mathbf{4. 5 4 \%}$, respectively. Codes are at https://github.com/wr2167/TiSTSL.
Hao Wang 0218, Qiang Li 0018, Zhiwei Wang 0002
BIBM5
2025 First-frame Supervised Video Polyp Segmentation via Propagative and Semantic Dual-teacher Network
abstract
Automatic video polyp segmentation plays a critical role in gastrointestinal cancer screening, but the cost of frame-by-frame annotations is prohibitively high. While sparse-frame supervised methods have reduced this burden proportionately, the cost remains overwhelming for long-duration videos and large-scale datasets. In this paper, we, for the first time, reduce the annotation cost to just a single frame per polyp video, regardless of the video's length. To this end, we introduce a new task, First-Frame Supervised Video Polyp Segmentation (FSVPS), and propose a novel Propagative and Semantic Dual-Teacher Network (PSDNet). Specifically, PSDNet adopts a teacher-student framework but employs two distinct types of teachers: the propagative teacher and the semantic teacher. The propagative teacher is a universal object tracker that propagates the first-frame annotation to subsequent frames as pseudo labels. However, tracking errors may accumulate over time, gradually degrading the pseudo labels and misguiding the student model. To address this, we introduce the semantic teacher, an exponential moving average of the student model, which produces more stable and time-invariant pseudo labels. PSDNet merges the pseudo labels from both teachers using a carefully-designed back-propagation strategy. This strategy assesses the quality of the pseudo labels by tracking them backward to the first frame. High-quality pseudo labels are more likely to spatially align with the first-frame annotation after this backward tracking, ensuring more accurate teacher-to-student knowledge transfer and improved segmentation performance. Benchmarking on SUN-SEG, the largest VPS dataset, demonstrates the competitive performance of PSDNet compared to fully-supervised approaches, and its superiority over sparse-frame supervised state-of-the-arts with a minimum improvement of 4.5% in Dice score. Codes are at https://github.com/Huster-Hq/PSDNet.
Qiang Li 0018, Zhiwei Wang 0002
ICASSP3
2025 SPNet: Sparse-mask Prompt-learning Network for Cerebrovascular Segmentation
abstract
Cerebrovascular segmentation demands a comprehensive understanding of vascular topology from a global perspective on 3D Time-of-Flight Magnetic Resonance Angiography (TOF-MRA) images. However, existing 3D segmentation methods typically rely on patch-wise processing to manage computational costs, where different segments of the same vessel branch may be handled independently in separate patches. This might restrict their ability to capture the full 3D context of the vessel, thus severely impacting segmentation performance. In this paper, we focus on achieving comprehensive 3D vessel segmentation in a single feedforward pass by leveraging Maximum Intensity Projection (MIP). By compressing the original 3D image, the vessels of the entire brain are superimposed and projected onto a 2D MIP image. We propose a Sparse-mask Prompt-learning Network (SPNet) to leverage MIP for precise cerebrovascular segmentation. Specifically, SPNet generates MIP images along three orthogonal directions and performs 2D segmentation on each of these MIP images independently. The segmented 2D vessels of MIPs are then back-projected into 3D space and fused into a sparse vessel mask. SPNet treats this sparse mask as a topological prompt that captures the overall vascular pathways, thereby enhancing feature learning on the original 3D TOF-MRA data, ultimately improving the final cerebrovascular segmentation performance. Experimental results on the ADAM and IXI datasets demonstrate superior performance of SPNet with least improvements of 1.73% and 0.75% on average Dices, 3.04mm and 1.19mm on Hausdorff distances, respectively, over state-of-the-art methods. Codes are available at: https://github.com/shanwq/SPNet.
Wenqi Shan, Qiang Li 0018, Zhiwei Wang 0002
ICASSP2
2025 DACAT: Dual-stream Adaptive Clip-aware Time Modeling for Robust Online Surgical Phase Recognition
abstract
Surgical phase recognition has become a crucial requirement in laparoscopic surgery, enabling various clinical applications like surgical risk forecasting. Current methods typically identify the surgical phase using individual frame-wise embeddings as the fundamental unit for time modeling. However, this approach is overly sensitive to current observations, often resulting in discontinuous and erroneous predictions within a complete surgical phase. In this paper, we propose DACAT, a novel dual-stream model that adaptively learns clip-aware context information to enhance the temporal relationship. In one stream, DACAT pretrains a frame encoder, caching all historical frame-wise features. In the other stream, DACAT fine-tunes a new frame encoder to extract the frame-wise feature at the current moment. Additionally, a max clip-response read-out (Max-R) module is introduced to bridge the two streams by using the current frame-wise feature to adaptively fetch the most relevant past clip from the feature cache. The clip-aware context feature is then encoded via cross-attention between the current frame and its fetched adaptive clip, and further utilized to enhance the time modeling for accurate online surgical phase recognition. The benchmark results on three public datasets, i.e., Cholec80, M2CAI16, and AutoLaparo, demonstrate the superiority of our proposed DACAT over existing state-of-the-art methods, with improvements in Jaccard scores of at least 4.5%, 4.6%, and 2.7%, respectively. Our code and models have been released at https://github.com/kk42yy/DACAT.
Kaixiang Yang 0004, Qiang Li 0018, Zhiwei Wang 0002
ICASSP2
2025 CoStoDet-DDPM: Collaborative Training of Stochastic and Deterministic Models Improves Surgical Workflow Anticipation and Recognition
abstract
Anticipating and recognizing surgical workflows are critical for intelligent surgical assistance systems. However, existing methods rely on deterministic decision-making, struggling to generalize across the large anatomical and procedural variations inherent in real-world surgeries.In this paper, we introduce an innovative framework that incorporates stochastic modeling through a denoising diffusion probabilistic model (DDPM) into conventional deterministic learning for surgical workflow analysis. At the heart of our approach is a collaborative co-training paradigm: the DDPM branch captures procedural uncertainties to enrich feature representations, while the task branch focuses on predicting surgical phases and instrument usage.Theoretically, we demonstrate that this mutual refinement mechanism benefits both branches: the DDPM reduces prediction errors in uncertain scenarios, and the task branch directs the DDPM toward clinically meaningful representations. Notably, the DDPM branch is discarded during inference, enabling real-time predictions without sacrificing accuracy.Experiments on the Cholec80 dataset show that for the anticipation task, our method achieves a 16% reduction in eMAE compared to state-of-the-art approaches, and for phase recognition, it improves the Jaccard score by 1.0%. Additionally, on the AutoLaparo dataset, our method achieves a 1.5% improvement in the Jaccard score for phase recognition, while also exhibiting robust generalization to patient-specific variations. Our code and weight are available at https://github.com/kk42yy/CoStoDet-DDPM.
Kaixiang Yang 0004, Xin Li 0001, Qiang Li 0018, Zhiwei Wang 0002
ICCV3
2025 Holistic White-Light Polyp Classification via Alignment-Free Dense Distillation of Auxiliary Optical Chromoendoscopy
Qimei Wang, Xuantao Ji, Qiang Li 0018, Zhiwei Wang 0002
MICCAI (11)6
2025 Targeted False Positive Synthesis via Detector-Guided Adversarial Diffusion Attacker for Robust Polyp Detection
Quan Zhou 0011, Gan Luo, Qingyong Zhang, Yinjiao Tian, Qiang Li 0018, Zhiwei Wang 0002
MICCAI (11)7
2025 Joint Holistic and Lesion Controllable Mammogram Synthesis via Gated Conditional Diffusion Model
abstract
Mammography is the most commonly used imaging modality for breast cancer screening, driving an increasing demand for deep-learning techniques to support large-scale analysis. However, the development of accurate and robust methods is often limited by insufficient data availability and a lack of diversity in lesion characteristics. While generative models offer a promising solution for data synthesis, current approaches often fail to adequately emphasize lesion-specific features and their relationships with surrounding tissues. In this paper, we propose Gated Conditional Diffusion Model (GCDM), a novel framework designed to jointly synthesize holistic mammogram images and localized lesions. GCDM is built upon a latent denoising diffusion framework, where the noised latent image is concatenated with a soft mask embedding that represents breast, lesion, and their transitional regions, ensuring anatomical coherence between them during the denoising process. To further emphasize lesion-specific features, GCDM incorporates a gated conditioning branch that guides the denoising process by dynamically selecting and fusing the most relevant radiomic and geometric properties of lesions, effectively capturing their interplay. Experimental results demonstrate that GCDM achieves precise control over small lesion areas while enhancing the realism and diversity of synthesized mammograms. These advancements position GCDM as a promising tool for clinical applications in mammogram synthesis. Our code is available at https://github.com/lixinHUST/Gated-Conditional-Diffusion-Model/
Xin Li 0001, Kaixiang Yang 0004, Qiang Li 0018, Zhiwei Wang 0002
ACM Multimedia3
2025 Y-Net-ECG: A Multi-Lead informed and interpretable architecture for ECG segmentation across diverse rhythms
Peng Zhang 0106, Xiaoli Feng, Kaibiao Huang, Yinuo Zhao, Zuoming Fu, Zhigang Ye, Tao Wang 0138, Xiaoyun Yang, Fan Lin, Qiang Li 0018
Expert Syst. Appl.15
2025 Boosting Few-Shot Semantic Segmentation of 3D Medical Images via Collaborative Slice Alignment
abstract
Few-shot semantic segmentation (FSS) of 3D medical images requires finding a 2D slice from the labeled volume as support to 'query' slices of the unlabeled one. Accurately determining support slices is crucial for learning representative prototypical features, thereby enhancing segmentation accuracy. The existing methods typically resort to the true position of the query target to align the query with support slices or simply exploit one key support slice to segment all query slices, which inevitably results in poor practicality and mis-segmentation. In this regard, we seek a practical and efficient solution by proposing a novel Collaborative Slice Alignment (CSA) module, which densely assigns each query slice its own fittest support without knowing the target prior. Concretely, our CSA first estimates the confidence scores of slices from the sorting task to implicitly reflect their physical location in the human body. The estimated scores are considered as spatial references for aligning support slices and query slices so that each matching pair shares the most similar image contents. Moreover, the self-learnable ranking objective allows CSA to transfer internal knowledge into both support and query features to further boost the FSS performance. Additionally, we introduce an Information Reconciliation (InRe) module to mitigate the inconsistent feature distribution caused by the individual differences between support and query images. Experimental results demonstrate that the combination of CSA and InRe achieves an average Dice score improvement of at least 8.61% across three datasets, consistently outperforming other state-of-the-art methods.
Jialun Pei, Zhiwei Wang 0002, Qiang Li 0018, Pheng-Ann Heng
IEEE J. Biomed. Health Informatics5
2025 Fine-Grained Temporal Site Monitoring in EGD Streams via Visual Time-Aware Embedding and Vision-Text Asymmetric Coworking
abstract
Esophagogastroduodenoscopy (EGD) requires inspecting plentiful upper gastrointestinal (UGI) sites completely for a precise cancer screening. Automated temporal site monitoring for EGD assistance is thus of high demand, yet often fails if directly applying the existing methods of online action detection. The key challenges are two-fold: 1) the global camera motion dominates, invalidating the temporal patterns derived from the object optical flows, and 2) the UGI sites are fine-grained, yielding highly homogenized appearances. In this paper, we propose an EGD-customized model, powered by two novel designs, i.e., Visual Time-aware Embedding plus Vision-text Asymmetric Coworking (VTE+VAC), for real-time accurate fine-grained UGI site monitoring. Concretely, VTE learns visual embeddings by differentiating frames via classification losses, and meanwhile by reordering the sampled time-agnostic frames to be temporally coherent via a ranking loss. Such joint objective encourages VTE to capture the sequential relation without resorting to the inapplicable object optical flows, and thus to provide the time-aware frame-wise embeddings. In the subsequent analysis, VAC uses a temporal sliding window, and extracts vision-text multimodal knowledge from each frame and its corresponding textualized prediction via the learned VTE and a frozen BERT. The text embeddings help provide more representative cues, but also may cause misdirection due to prediction errors. Thus, VAC randomly drops or replaces historical predictions to increase the error tolerance to avoid collapsing onto the last few predictions. Qualitative and quantitative experiments demonstrate that the proposed method achieves superior performance compared to other state-of-the-art methods, with an average F1-score improvement of at least 7.66%.
Hongkuan Shi, Shiquan He, Xinxia Feng, Jiazhi Liao, Qiang Li 0018, Zhiwei Wang 0002
IEEE J. Biomed. Health Informatics10
2024 Decoupling Feature Representations of Ego and Other Modalities for Incomplete Multi-modal Brain Tumor Segmentation
abstract
Multi-modal brain tumor segmentation typically involves four magnetic resonance imaging (MRI) modalities, while incomplete modalities significantly degrade performance. Existing solutions employ explicit or implicit modality adaptation, aligning features across modalities or learning a fused feature robust to modality incompleteness. They share a common goal of encouraging each modality to express both itself and the others. However, the two expression abilities are entangled as a whole in a seamless feature space, resulting in prohibitive learning burdens. In this paper, we propose DeMoSeg to enhance the modality adaptation by Decoupling the task of representing the ego and other Modalities for robust incomplete multi-modal Segmentation. The decoupling is super lightweight by simply using two convolutions to map each modality onto four feature sub-spaces. The first sub-space expresses itself (Self-feature), while the remaining sub-spaces substitute for other modalities (Mutual-features). The Self- and Mutual-features interactively guide each other through a carefully-designed Channel-wised Sparse Self-Attention (CSSA). After that, a Radiologist-mimic Cross-modality expression Relationships (RCR) is introduced to have available modalities provide Self-feature and also ‘lend’ their Mutual-features to compensate for the absent ones by exploiting the clinical prior knowledge. The benchmark results on BraTS2020, BraTS2018 and BraTS2015 verify the DeMoSeg’s superiority thanks to the alleviated modality adaptation difficulty. Concretely, for BraTS2020, DeMoSeg increases Dice by at least 0.92%, 2.95% and 4.95% on whole tumor, tumor core and enhanced tumor regions, respectively, compared to other state-of-the-arts. Codes are at https://github.com/kk42yy/DeMoSeg.
Kaixiang Yang 0004, Wenqi Shan, Xikai Yang, Xi Wang 0013, Pheng-Ann Heng, Qiang Li 0018, Zhiwei Wang 0002
BIBM8
2024 Improved Self-supervised Monocular Endoscopic Depth Estimation based on Pose Alignment-friendly Dynamic View Selection
abstract
3D depth estimation in monocular endoscopic videos typically relies on self-supervised learning due to the absence of in-body ground-truth depth labels. This learning paradigm optimizes predicted 3D depth using adjacent frames by aligning camera poses and minimizing photometric discrepancies. However, previous efforts often suffer from insufficient visual cues caused by minimal camera movement between adjacent frames, resulting in significant pose alignment errors and reduced learning effectiveness. In this paper, we propose DVSMono, a dynamic view selection-based self-supervised monocular depth estimation that, for the first time, leverages pose alignment-friendly views instead of static adjacent ones. DVSMono employs a pretrained pose estimator to calculate poses between the target frame and its surrounding frames over a larger time span. Accurate poses can create highly-unanimous geometric cost volumes across different depth proposals, while errors often appear as outliers. Motivated by this, we design Temporally-Consistent View Scoring (TCVS), which selects candidate frames with cost volumes showing minimal temporal variance. Additionally, we introduce Photometric-Valid Region Computation (PVRC), which filters out unreliable regions between frames by identifying non-overlapping fields of view and bidirectionally inconsistent occlusion pixels. Ultimately, among the TCVS-selected candidates, DVSMono assigns each target frame its most suitable source frame with the largest PVRC-filtered valid regions. This enables the training of a robust depth estimator, benefiting from improved self-supervised learning built upon dynamically-selected, pose alignment-friendly source-target pairs. Benchmark results on three public datasets demonstrate that DVSMono exceeds the existing state-of-the-art methods by reducing Abs Rel by at least 6.35%. Codes are at https://github.com/adam99goat/DVSMono.
Shiquan He, Hao Wang 0218, Qiang Li 0018, Zhiwei Wang 0002
BIBM6
2024 SALI: Short-Term Alignment and Long-Term Interaction Network for Colonoscopy Video Polyp Segmentation
Zhenyu Yi, Qiang Li 0018, Zhiwei Wang 0002
MICCAI (6)6
2024 Noise Removed Inconsistency Activation Map for Unsupervised Registration of Brain Tumor MRI Between Pre-operative and Follow-Up Phases
Chongwei Wu, Xiaoyu Zeng, Hao Wang 0218, Qiang Li 0018, Zhiwei Wang 0002
MICCAI (2)6
2024 Co-learning-assisted progressive dense fusion network for cardiovascular disease detection using ECG and PCG signals
Haobo Zhang 0003, Peng Zhang 0106, Fan Lin, Lianying Chao, Zhiwei Wang 0002, Qiang Li 0018
Expert Syst. Appl.7
2024 Multi-Feature Decision Fusion Network for Heart Sound Abnormality Detection and Classification
abstract
The heart sound reflects the movement status of the cardiovascular system and contains the early pathological information of cardiovascular diseases. Automatic heart sound diagnosis plays an essential role in the early detection of cardiovascular diseases. In this study, we aim to develop a novel end-to-end heart sound abnormality detection and classification method, which can be adapted to different heart sound diagnosis tasks. Specifically, we developed a Multi-feature Decision Fusion Network (MDFNet) composed of a Multi-dimensional Feature Extraction (MFE) module and a Multi-dimensional Decision Fusion (MDF) module. The MFE module extracted spatial features, multi-level temporal features and spatial-temporal fusion features to learn heart sound characteristics from multiple perspectives. Through deep supervision and decision fusion, the MDF module made the multi-dimensional features extracted by the MFE module more discriminative, and fused the decision results of multi-dimensional features to integrate complementary information. Furthermore, attention modules were embedded in the MDFNet to emphasize the fundamental heart sounds containing effective feature information. Finally, we proposed an efficient data augmentation method to circumvent the diagnosis performance degradation caused by the lack of cardiac cycle segmentation in other end-to-end methods. The developed method achieved an overall accuracy of 94.44% and a F1-score of 86.90% on the binary classification task and a F1-score of 99.30% on the five-classification task. Our method outperformed other state-of-the-art methods and had good clinical application prospects.
Haobo Zhang 0003, Peng Zhang 0106, Zhiwei Wang 0002, Lianying Chao, Qiang Li 0018
IEEE J. Biomed. Health Informatics6
2023 Robust One-Shot Segmentation of Brain Tissues via Image-Aligned Style Transformation
abstract
One-shot segmentation of brain tissues is typically a dual-model iterative learning: a registration model (reg-model) warps a carefully-labeled atlas onto unlabeled images to initialize their pseudo masks for training a segmentation model (seg-model); the seg-model revises the pseudo masks to enhance the reg-model for a better warping in the next iteration. However, there is a key weakness in such dual-model iteration that the spatial misalignment inevitably caused by the reg-model could misguide the seg-model, which makes it converge on an inferior segmentation performance eventually. In this paper, we propose a novel image-aligned style transformation to reinforce the dual-model iterative learning for robust one-shot segmentation of brain tissues. Specifically, we first utilize the reg-model to warp the atlas onto an unlabeled image, and then employ the Fourier-based amplitude exchange with perturbation to transplant the style of the unlabeled image into the aligned atlas. This allows the subsequent seg-model to learn on the aligned and style-transferred copies of the atlas instead of unlabeled images, which naturally guarantees the correct spatial correspondence of an image-mask training pair, without sacrificing the diversity of intensity patterns carried by the unlabeled images. Furthermore, we introduce a feature-aware content consistency in addition to the image-level similarity to constrain the reg-model for a promising initialization, which avoids the collapse of image-aligned style transformation in the first iteration. Experimental results on two public datasets demonstrate 1) a competitive segmentation performance of our method compared to the fully-supervised method, and 2) a superior performance over other state-of-the-art with an increase of average Dice by up to 4.67%. The source code is available at: https://github.com/JinxLv/One-shot-segmentation-via-IST.
Jinxin Lv, Xiaoyu Zeng, Sheng Wang 0013, Zhiwei Wang 0002, Qiang Li 0018
AAAI6
2023 Rethinking Low-Dose CT Synthesis: Degrading Normal-Dose CT from Origin for Pairwise Training of CT Denoiser
abstract
Training a denoiser for translating low-dose to normal-dose computed tomography (LDCT to NDCT) requires collecting spatially corresponded paired data, yet it is impractical. Existing works aim to ‘add’ noise to real NDCT for synthesizing paired LDCT, but often failed to model the complex noise-structure entanglement, and therefore led to a suboptimal trained denoiser. In this paper, we propose a fresh solution to mimic more reasonable LDCT by modeling the structure-entangled noise patterns. Our motivation is based on a fact that the noise is originally carried by the source sinogram, and then passed to the reconstructed CT and intertwined with the structures via the filtered back-projection (FBP). In light of this, we first convert a NDCT image to the source sinogram by forward projection, and then design a Degrading CT Network (DeCTNet) to learn noise from the origin. Specifically, DeCTNet consists of two sequential networks, i.e., sinogram and reconstruction networks (SinoNet and RecoNet). SinoNet learns to directly add noise to the NDCT-converted sinogram, and RecoNet learns to further process the reconstructed CT image guided by real but unpaired LDCT images. DeCTNet utilizes a differentiable FBP operator to naturally reinvent the sinogram noise to the entangled noise-structure patterns in CT, and thus bridge SinoNet and RecoNet in an end-to-end training. Moreover, adversarial loss and content-fidelity loss are jointly minimized to effectively learn the noise characteristics and content retention in LDCT synthesis. Both quantitative and qualitative evaluations demonstrate that the denoiser trained using DeCTNet-synthesized pairs outperforms those trained using pairs synthesized by other state-of-the-arts, and radiologists are often not able to distinguish our denoised LDCTs from the real NDCTs.
Lianying Chao, Taotao Zhang, Wenqi Shan, Qiang Li 0018, Zhiwei Wang 0002
BIBM5
2023 Towards Learning-based Surgical Planning of Glioma Resection via a Contrastively Constrained Siamese Neural Network
abstract
The risk assessment of craniotomy paths is a vital stage in the automated surgical planning of glioma resection. However, current methods heavily rely on the manually-assigned risk coefficients tied to factors like path length and intersecting brain structures. This overreliance on empirical coefficients poses limitations in terms of their adaptability to broader cohorts. In this paper, we propose an innovative risk assessment approach for automated surgical planning by leveraging deep learning techniques. Specifically, we develop a Siamese Neural Network (SNN) with two weight-sharing branches, and also compile a glioma dataset, containing the extracted brain vessels, regions, and the craniotomy region on the scalp. During training, SNN learns to predict two risk values for a pair of paths, which are sampled inside and outside the craniotomy region, respectively. A contrastive loss is then employed to guide SNN by penalizing the higher predicted risk inside the craniotomy region or the lower risk outside the region. Such contrastive constraint empowers effective learning of SNN in our unique scenario, where the true risk values of surgical paths are unavailable. During inference, the trained SNN generates a risk map on the scalp, and the path is determined by finding the local minimum on the risk map. To the best of our knowledge, this work proposes one of the first attempts to successfully incorporate deep learning techniques into the risk assessment process of glioma resection planning. Quantitative and neurosurgeons’ qualitative evaluations of six testing cases proved the effectiveness and superiority of our learning-based approach. A consensus is reached that our method shows great potential to help surgeons make safer and more accurate plans in practice.
Wenqi Shan, Yingjie Guo, Lianying Chao, Haobo Zhang 0003, Qiang Li 0018, Zhiwei Wang 0002
BIBM7
2023 Dual-view Correlation Hybrid Attention Network for Robust Holistic Mammogram Classification
abstract
Mammogram image is important for breast cancer screening, and typically obtained in a dual-view form, i.e., cranio-caudal (CC) and mediolateral oblique (MLO), to provide complementary information for clinical decisions. However, previous methods mostly learn features from the two views independently, which violates the clinical knowledge and ignores the importance of dual-view correlation in the feature learning. In this paper, we propose a dual-view correlation hybrid attention network (DCHA-Net) for robust holistic mammogram classification. Specifically, DCHA-Net is carefully designed to extract and reinvent deep feature maps for the two views, and meanwhile to maximize the underlying correlations between them. A hybrid attention module, consisting of local relation and non-local attention blocks, is proposed to alleviate the spatial misalignment of the paired views in the correlation maximization. A dual-view correlation loss is introduced to maximize the feature similarity between corresponding strip-like regions with equal distance to the chest wall, motivated by the fact that their features represent the same breast tissues, and thus should be highly-correlated with each other. Experimental results on the two public datasets, i.e., INbreast and CBIS-DDSM, demonstrate that the DCHA-Net can well preserve and maximize feature correlations across views, and thus outperforms previous state-of-the-art methods for classifying a whole mammogram as malignant or not.
Zhiwei Wang 0002, Junlin Xian, Kangyi Liu, Xin Li 0001, Qiang Li 0018, Xin Yang 0008
IJCAI5
2023 PSDP: Pseudo-supervised dual-processing for low-dose cone-beam computed tomography reconstruction
Lianying Chao, Wenqi Shan, Wenting Xu, Haobo Zhang 0003, Zhiwei Wang 0002, Qiang Li 0018
Expert Syst. Appl.7
2023 Accurate Cobb Angle Estimation on Scoliosis X-Ray Images via Deeply-Coupled Two-Stage Network With Differentiable Cropping and Random Perturbation
abstract
Automated Cobb angle estimation on X-ray images is crucial to scoliosis diagnosis. The existing efforts are typically two extremes, which either laboriously detect the raw vertebral landmarks or directly regress Cobb angles from the entire image. In this paper, we propose a novel two-stage end-to-end method as a balanced solution, to avoid vulnerability to false landmarks, and to preserve flexibility in clinical usages. Concretely, we cascade two stages sequentially for detecting vertebrae and then regressing their bending directions instead of raw landmarks. In the detection stage, we combine two networks called LocNet and SegNet to robustly localize vertebrae, and meanwhile to suppress the false positives by additionally segmenting the whole spine. In the subsequent stage, we introduce a regression network named RegNet to accurately regress bending directions of localized vertebrae. Furthermore, the vertebra-aligned local regions on LocNet's intermediate features are cropped via RoIAlign-pooling, and RegNet inherits the cropped regions to learn only feature residuals. By doing so, the regression difficulty can be dramatically alleviated, and the two stages are deeply coupled and mutually guided in an end-to-end training. Moreover, a random perturbation on the inherited features further enhances RegNet's robustness. We benchmark our method on both public and private datasets, and the errors are 2.92 $\pm$ 2.34$^{\circ }$ and 6.87 $\pm$ 6.26% in terms of CMAE and SMAPE on the widely-employed AASCE dataset, outperforming other state-of-the-arts by at least 16.81% and 6.15%, respectively. Also, a clinical user study verifies our promising flexibility for allowing convenient rectifications to further decrease errors by a large marge.
Yuanhuai Liang, Jinxin Lv, Dun Li, Xin Yang 0008, Zhiwei Wang 0002, Qiang Li 0018
IEEE J. Biomed. Health Informatics6
2023 Learn2Reg: Comprehensive Multi-Task Medical Image Registration Challenge, Dataset and Evaluation in the Era of Deep Learning
abstract
Image registration is a fundamental medical image analysis task, and a wide variety of approaches have been proposed. However, only a few studies have comprehensively compared medical image registration approaches on a wide range of clinically relevant tasks. This limits the development of registration methods, the adoption of research advances into practice, and a fair benchmark across competing approaches. The Learn2Reg challenge addresses these limitations by providing a multi-task medical image registration data set for comprehensive characterisation of deformable registration algorithms. A continuous evaluation will be possible at https://learn2reg.grand-challenge.org. Learn2Reg covers a wide range of anatomies (brain, abdomen, and thorax), modalities (ultrasound, CT, MR), availability of annotations, as well as intra- and inter-patient registration evaluation. We established an easily accessible framework for training and validation of 3D registration methods, which enabled the compilation of results of over 65 individual method submissions from more than 20 unique teams. We used a complementary set of metrics, including robustness, accuracy, plausibility, and runtime, enabling unique insight into the current state-of-the-art of medical image registration. This paper describes datasets, tasks, evaluation methods and results of the challenge, as well as results of further analysis of transferability to new datasets, the importance of label supervision, and resulting bias. While no single approach worked best across all tasks, many methodological aspects could be identified that push the performance of medical image registration to new state-of-the-art performance. Furthermore, we demystified the common belief that conventional registration methods have to be much slower than deep-learning-based methods.
Alessa Hering, Lasse Hansen, Tony C. W. Mok, Albert C. S. Chung, Hanna Siebert, Stephanie Häger, Annkristin Lange, Sven Kuckertz, Stefan Heldmann, Wei Shao 0008, Sulaiman Vesal, Mirabela Rusu, Geoffrey A. Sonn, Théo Estienne, Maria Vakalopoulou, Luyi Han, Yunzhi Huang, Pew-Thian Yap, Mikael Brudfors, Yaël Balbastre, Samuel Joutard, Marc Modat, Gal Lifshitz, Dan Raviv, Jinxin Lv, Qiang Li 0018, Vincent Jaouen, Dimitris Visvikis, Constance Fourcade, Mathieu Rubeaux, Wentao Pan 0001, Zhe Xu 0012, Bailiang Jian, Francesca De Benetti, Marek Wodzinski, Niklas Gunnarsson, Jens Sjölund, Daniel Grzech, Huaqi Qiu, Zeju Li, Alexander Thorley, Jinming Duan 0001, Christoph Großbröhmer, Andrew Hoopes, Ingerid Reinertsen, Yiming Xiao 0001, Bennett A. Landman, Yuankai Huo, Keelin Murphy, Nikolas Leßmann, Bram van Ginneken, Adrian V. Dalca, Mattias P. Heinrich
IEEE Trans. Medical Imaging26
2023 Bidirectional Semi-Supervised Dual-Branch CNN for Robust 3D Reconstruction of Stereo Endoscopic Images via Adaptive Cross and Parallel Supervisions
abstract
Semi-supervised learning via teacher-student network can train a model effectively on a few labeled samples. It enables a student model to distill knowledge from the teacher's predictions of extra unlabeled data. However, such knowledge flow is typically unidirectional, having the accuracy vulnerable to the quality of teacher model. In this paper, we seek to robust 3D reconstruction of stereo endoscopic images by proposing a novel fashion of bidirectional learning between two learners, each of which can play both roles of teacher and student concurrently. Specifically, we introduce two self-supervisions, i.e., Adaptive Cross Supervision (ACS) and Adaptive Parallel Supervision (APS), to learn a dual-branch convolutional neural network. The two branches predict two different disparity probability distributions for the same position, and output their expectations as disparity values. The learned knowledge flows across branches along two directions: a cross direction (disparity guides distribution in ACS) and a parallel direction (disparity guides disparity in APS). Moreover, each branch also learns confidences to dynamically refine its provided supervisions. In ACS, the predicted disparity is softened into a unimodal distribution, and the lower the confidence, the smoother the distribution. In APS, the incorrect predictions are suppressed by lowering the weights of those with low confidence. With the adaptive bidirectional learning, the two branches enjoy well-tuned mutual supervisions, and eventually converge on a consistent and more accurate disparity estimation. The experimental results on four public datasets demonstrate our superior accuracy over other state-of-the-arts with a relative decrease of averaged disparity error by at least 9.76%.
Hongkuan Shi, Zhiwei Wang 0002, Dun Li, Xin Yang 0008, Qiang Li 0018
IEEE Trans. Medical Imaging6
2022 Improving the Quality of Sparse-view Cone-Beam Computed Tomography via Reconstruction-Friendly Interpolation Network
Lianying Chao, Wenqi Shan, Haobo Zhang 0003, Zhiwei Wang 0002, Qiang Li 0018
ACCV (6)6
2022 Sparse-view cone beam CT reconstruction using dual CNNs in projection domain and image domain
Lianying Chao, Zhiwei Wang 0002, Haobo Zhang 0003, Wenting Xu, Peng Zhang 0106, Qiang Li 0018
Neurocomputing6
2022 Dual-domain attention-guided convolutional neural network for low-dose cone-beam computed tomography reconstruction
Lianying Chao, Peng Zhang 0106, Zhiwei Wang 0002, Wenting Xu, Qiang Li 0018
Knowl. Based Syst.6
2022 Semi-Supervised Learning for Automatic Atrial Fibrillation Detection in 24-Hour Holter Monitoring
abstract
Paroxysmal atrial fibrillation (AF) is generally diagnosed by long-term dynamic electrocardiogram (ECG) monitoring. Identifying AF episodes from long-term ECG data can place a heavy burden on clinicians. Many machine-learning-based automatic AF detection methods have been proposed to solve this issue. However, these methods require numerous annotated data to train the model, and the annotation of AF in long-term ECG is extremely time-consuming. Reducing the demand for labeled data can effectively improve the clinical practicability of automatic AF detection methods. In this study, we developed a novel semi-supervised learning method that generated modified low-entropy labels of unlabeled samples for training a deep learning model to automatically detect paroxysmal AF in 24 h Holter monitoring data. Our method employed a 1D CNN-LSTM neural network with RR intervals as input and used few labeled training data with numerous unlabeled data for training the neural network. This method was evaluated using a 24 h Holter monitoring dataset collected from 1000 paroxysmal AF patients. Using labeled samples from only 10 patients for model training, our method achieved a sensitivity of 97.8%, specificity of 97.9%, and accuracy of 97.9% in five-fold cross-validation. Compared to the supervised learning method with complete labeled samples, the detection accuracy of our method was only 0.5% lower, while the workload of data annotation was significantly reduced by more than 98%. In general, this is the first study to apply semi-supervised learning techniques for automatic AF detection using ECG. Our method can effectively reduce the demand for AF data annotations and can improve the clinical practicability of automatic AF detection.
Peng Zhang 0106, Fan Lin, Xiaoyun Yang, Qiang Li 0018
IEEE J. Biomed. Health Informatics6
2022 Joint Progressive and Coarse-to-Fine Registration of Brain MRI via Deformation Field Integration and Non-Rigid Feature Fusion
abstract
Registration of brain MRI images requires to solve a deformation field, which is extremely difficult in aligning intricate brain tissues, e.g., subcortical nuclei, etc. Existing efforts resort to decomposing the target deformation field into intermediate sub-fields with either tiny motions, i.e., progressive registration stage by stage, or lower resolutions, i.e., coarse-to-fine estimation of the full-size deformation field. In this paper, we argue that those efforts are not mutually exclusive, and propose a unified framework for robust brain MRI registration in both progressive and coarse-to-fine manners simultaneously. Specifically, building on a dual-encoder U-Net, the fixed-moving MRI pair is encoded and decoded into multi-scale sub-fields from coarse to fine. Each decoding block contains two proposed novel modules: i) in Deformation Field Integration (DFI), a single integrated deformation sub-field is calculated, warping by which is equivalent to warping progressively by sub-fields from all previous decoding blocks, and ii) in Non-rigid Feature Fusion (NFF), features of the fixed-moving pair are aligned by DFI-integrated deformation field, and then fused to predict a finer sub-field. Leveraging both DFI and NFF, the target deformation field is factorized into multi-scale sub-fields, where the coarser fields alleviate the estimate of a finer one and the finer field learns to make up those misalignments insolvable by previous coarser ones. The extensive and comprehensive experimental results on both private and two public datasets demonstrate a superior registration performance of brain MRI images over progressive registration only and coarse-to-fine estimation only, with an increase by at most 8% in the average Dice.
Jinxin Lv, Zhiwei Wang 0002, Hongkuan Shi, Haobo Zhang 0003, Sheng Wang 0013, Yilang Wang, Qiang Li 0018
IEEE Trans. Medical Imaging7
2021 Semi-supervised Learning via Improved Teacher-Student Network for Robust 3D Reconstruction of Stereo Endoscopic Image
abstract
3D reconstruction of stereo endoscope image, as an enabling technique for varied surgical systems, e.g., medical droids, navigations, etc., suffers from severe overfitting problems due to scarce labels. Semi-supervised learning based on Teacher-Student Network (TSN) is a potential solution, which utilizes a supervised teacher model trained on available labeled data to teach a student model on all images via assigning them pseudo labels. However, TSN often faces a dilemma: if given only few labeled endoscope images, the teacher model will be trained to be defective and induce high-noised pseudo labels, degrading the student model significantly. To solve this, we propose an improved TSN for a robust 3D reconstruction of stereo endoscope image. Specifically, two novel modules are introduced: 1) a semi-supervised teacher model based on adversarial learning to produce mostly correct pseudo labels by forcing a consistency in predictions for both labeled and unlabeled data, and 2) a confidence network to further filter out noisy pseudo labels by estimating a confidence for each prediction of the teacher model. By doing so, the student model is able to distill knowledge from more accurate and noiseless pseudo labels, thus achieving improved performance. Experimental results on two public datasets show that our improved TSN achieves a superior performance than the state-of-the-arts by reducing the averaged disparity error by at least 13.5%.
Hongkuan Shi, Zhiwei Wang 0002, Jinxin Lv, Yilang Wang, Peng Zhang 0106, Qiang Li 0018
ACM Multimedia7
2021 Deep learning based data-adaptive descriptor for non-rigid multi-modal medical image registration
Xingxing Zhu, Zhiwen Huang, Mingyue Ding, Qiang Li 0018, Xuming Zhang 0003
Signal Process.5