EDBT 2026 Demo / reviewers in the wild / expert
Kang Li 0004
dblp:l/KangLi
· DBLP profile ↗
67ranked-venue papers
0as first author
53since 2021 · last 2026
0000-0002-8136-9816ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 39 · 33 since 2021Artificial intelligence and machine learning · 19 · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 9 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DICE: Discrete Inversion Enabling Controllable Editing for Masked Generative ModelsabstractRecent advances in discrete diffusion models have demonstrated strong performance in image generation and masked language modeling, yet they remain limited in their capacity for controlled content editing. We propose DICE (Discrete Inversion for Controllable Editing), a novel framework that pioneers precise inversion capabilities for discrete diffusion models, including both masked generative and multinomial diffusion variants. Our key innovation lies in capturing noise sequences and masking patterns during reverse diffusion process, enabling both accurate reconstruction and flexible editing without relying on predefined masks or attention-based manipulations. Through comprehensive experiments across image and text modalities using models such as Paella, VQ-Diffusion, RoBERTa and LLaDA, we demonstrate that DICE successfully maintains high fidelity to the original data while significantly expanding editing capabilities. These results establish new possibilities for fine-grained content manipulation in discrete spaces. Xiaoxiao He, Quan Dao, Ligong Han, Song Wen 0001, Minhao Bai, Di Liu 0003, Han Zhang 0010, Felix Juefei-Xu, Chaowei Tan, Bo Liu 0005, Martin Renqiang Min, Kang Li 0004, Faez Ahmed, Akash Srivastava, Hongdong Li, Junzhou Huang, Dimitris N. Metaxas |
WACV | 12 |
| 2026 | Multi-VLM collaborated adaptive sampling for enhanced data pruning
Changfan Wang, Wei Xu 0046, Huahui Yi, Kang Li 0004, Boyu Wang 0004, Qicheng Lao |
Neurocomputing | 5 |
| 2026 | An adaptive large neighborhood search with dynamic exit reassignment for guided multi-story hospital fire evacuation
Qu Wei, Ruisan Zhang, Hao Yu 0003, Zhaoxia Guo, Kang Li 0004 |
Inf. Sci. | 7 |
| 2026 | PL-Seg: Partially labeled abdominal organ segmentation via classwise orthogonal contrastive learning and progressive self-distillation
Xiangde Luo, Ran Gu, Wenjun Liao, Shichuan Zhang, Kang Li 0004, Guotai Wang, Shaoting Zhang 0001 |
Medical Image Anal. | 7 |
| 2026 | Motion Intention Decoding and Training Effect Evaluation in Robot-assisted Upper Limb Rehabilitation Based on MR-BCIabstractThe interaction between stroke patients and rehabilitation training robots is difficult because the motion intention cannot be executed by limbs due to their weak motor ability. Although brain–computer interaction (BCI) is beneficial for perceiving the motion of patients, it cannot solve the problem independently without receiving enough data about the effectiveness of the robotic assistance. This study proposes a BCI method that integrates rehabilitation training and training effect evaluation in a mixed reality (MR) environment. Three Electroencephalogram (EEG) experimental paradigms were designed, using motor imagery to convert the four-classification problem of motor execution into three binary classification problems. An EEG classification model of a multi-scale convolution residual network based on a multi-head attention mechanism was built, consisting of a multi-head attention mechanism layer, four convolutional layers, and a pooling layer. The highest accuracy and best Kappa coefficient of 12 participants on the testing dataset were recorded. The average classification accuracy of MI, ME1, and ME2 achieves 80.35%, 87.52%, and 86.51%, respectively, proving the proposed classification method has sufficient decoding accuracy. Comparing the hemodynamic response curves of the participants’ upper limbs before and after robotic assistance during rehabilitation training, hemodynamic response curves were analyzed quantitatively. With the paired-sample t -test to compare the hemodynamics of a single participant, the significance of the overall mean difference was calculated to evaluate the influence of robotic assistance on enhancing local tissue blood oxygen, improving tissue function, and accelerating local tissue repair. The study verifies the effectiveness of robotic rehabilitation training and helps develop new types of BCI with better coordination between motion intention perception and motor ability. Dongxian Ye, Yutong Fu, Xiangyun Li, Kang Li 0004 |
ACM Trans. Appl. Percept. | 6 |
| 2026 | A Deep Reinforcement Learning-Based Hyper-Heuristic for Time-Dependent Green Logistics With CrowdsourcingabstractThe rapid growth of e-commerce has increased the complexity of supply chain management, particularly in urban logistics where efficiency and sustainability are critical concerns. In response, this study proposes a selection hyper-heuristic framework for time-dependent green logistics, incorporating key factors such as economic costs, carbon emissions, rider types, real-time traffic conditions, and time-window constraints. To address the complexities introduced by these real-world factors, we design a two-layer distribution model with crowdsourced delivery that covers the flow from city distribution centers to regional hubs and ultimately to end customers. The first layer involves location selection and the delivery process, while the second layer focuses on order allocation and last-mile delivery. For the location selection and order allocation problems, exact optimization models are developed to obtain high-quality solutions. In the delivery process, we integrate Deep Reinforcement Learning to replace the traditional adaptive layer of the Adaptive Large Neighborhood Search algorithm, enabling dynamic and intelligent adjustments during the search. Comparative analyses against existing and traditional methods across various benchmark instances demonstrate the superior efficiency and solution quality of the proposed framework. Simulation experiments based on a real-world road network in China validate the effectiveness of the proposed framework. In addition, the well-trained model can be directly applied to various scenarios, highlighting its strong generalization capability. Chu Tang, Qu Wei, Jingbin He, Guido Perboli, Kang Li 0004 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2026 | Source-Free Active Domain Adaptation via Influential-Points-Guided Progressive Teacher for Medical Image SegmentationabstractDomain adaptation in medical image segmentation enables pre-trained models to generalize to new target domains. Given limited annotated data and privacy constraints, Source-Free Active Domain Adaptation (SFADA) methods provide promising solutions by selecting a few target samples for labeling without accessing source samples. However, in a fully source-free setting, existing works have not fully explored how to select these target samples in a class-balanced manner and how to conduct robust model adaptation using both labeled and unlabeled samples. In this study, we discover that boundary samples with source-like semantics but sharp predictive discrepancies are beneficial for SFADA. We define these samples as the most influential points and propose a slice-wise framework using influential points learning to explore them. Specifically, we detect source-like samples to retain source-specific knowledge. For each target sample, an adaptive K-nearest neighbor algorithm based on local density is introduced to construct neighborhoods of source-like samples for knowledge transfer. We then propose a class-balanced Kullback-Leibler divergence for these neighborhoods, calculating it to obtain an influential score ranking. A diverse subset of the highest-ranked target samples (considered influential points) is manually annotated. Furthermore, we design a progressive teacher model to facilitate SFADA for medical image segmentation. With the guidance of influential points, this model independently generates and utilizes pseudo-labels to mitigate error accumulation. To further suppress noise, curriculum learning is incorporated into the model to progressively leverage reliable supervision signals from pseudo-labels. Experiments on multiple benchmarks demonstrate that our method outperforms state-of-the-art methods even with only 2.5% of the labeling budget. Yong Chen 0024, Xiangde Luo, Renyi Chen, Yiyue Li, Han Zhang 0010, He Lyu, Huan Song, Kang Li 0004 |
IEEE Trans. Medical Imaging | 8 |
| 2026 | Scan-Invariant Mamba With Differentiated Sequence Contrastive Learning in Computational PathologyabstractMultiple instance learning (MIL) is a commonly used paradigm for histopathological analysis due to the ultra-high resolution and coarse-grained labels of Whole Slide Images (WSIs). Recent studies apply Mamba architecture to WSI classification by modeling MIL as long-sequence tasks, but a key discrepancy remains: Mamba's output is sensitive to scanning modes, whereas MIL requires scan-invariant predictions. To address this problem, we propose Scan-invariant Mamba with Differentiated Sequence Contrastive Learning (SMDC-MIL), a novel Mamba-based MIL approach enabling bag-level feature learning independent of input modes. Our method mitigates scanning-mode impacts and adapts Mamba to learn the bag discrimination features that are independent of the input mode via two innovations: 1) a differentiated sequence generation mechanism that employs instance rearrangement, augmentation, and masking to simulate real-world scanning variations by maximizing differences in sequence order, length, and composition from the same WSI; and 2) a differentiated sequence contrastive learning architecture that enforces consistent bag-level representations and predictions across diverse sequences using the same Mamba model, guiding it to prioritize scan-invariant discriminative features. Experimental results on 4 computational pathology tasks and 10 datasets demonstrate that our SMDC-MIL achieves state-of-the-art performance compared to other methods. The corresponding code is available at https://github.com/LianYueZ/SMDCMIL.git. Sheng Huang 0001, Xin Zhang 0131, Bo Liu 0005, Fengtao Zhou, Kang Li 0004, Hao Chen 0011, Meng Wang 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2025 | iDPA: Instance Decoupled Prompt Attention for Incremental Medical Object DetectionabstractExisting prompt-based approaches have demonstrated impressive performance in continual learning, leveraging pre-trained large-scale models for classification tasks; however, the tight coupling between foreground-background information and the coupled attention between prompts and image-text tokens present significant challenges in incremental medical object detection tasks, due to the conceptual gap between medical and natural domains. To overcome these challenges, we introduce the iDPA framework, which comprises two main components: 1) Instance-level Prompt Generation (IPG), which decouples fine-grained instance-level knowledge from images and generates prompts that focus on dense predictions, and 2) Decoupled Prompt Attention (DPA), which decouples the original prompt attention, enabling a more direct and efficient transfer of prompt information while reducing memory usage and mitigating catastrophic forgetting. We collect 13 clinical, cross-modal, multi-organ, and multi-category datasets, referred to as ODinM-13, and experiments demonstrate that iDPA outperforms existing SOTA methods, with FAP improvements of f 5.44%, 4.83%, 12.88%, and 4.59% in full data, 1-shot, 10-shot, and 50-shot settings, respectively. Huahui Yi, Wei Xu 0046, Ziyuan Qin 0001, Xi Chen 0119, Kang Li 0004, Qicheng Lao |
ICML | 6 |
| 2025 | D2MAE: Diffusional Deblurring MAE for Ultrasound Image Pre-training
Qingbo Kang, Hongkai Zhao, Zhu He, Kang Li 0004, Qicheng Lao |
MICCAI (13) | 5 |
| 2025 | DiffOSeg: Omni Medical Image Segmentation via Multi-Expert Collaboration Diffusion Model
Han Zhang 0010, Xiangde Luo, Yong Chen 0024, Kang Li 0004 |
MICCAI (13) | 4 |
| 2025 | Guiding Medical Vision-Language Models with Diverse Visual Prompts: Framework Design and Comprehensive Exploration of Prompt VariationsabstractKangyu Zhu, Ziyuan Qin, Huahui Yi, Zekun Jiang, Qicheng Lao, Shaoting Zhang, Kang Li. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Kangyu Zhu, Ziyuan Qin 0001, Huahui Yi, Zekun Jiang, Qicheng Lao, Shaoting Zhang 0001, Kang Li 0004 |
NAACL (Long Papers) | 7 |
| 2025 | Dynamic graph based weakly supervised deep hashing for whole slide image classification and retrieval
Haochen Jin, Xiaoshuang Shi, Kang Li 0004, Xiaofeng Zhu 0001 |
Medical Image Anal. | 5 |
| 2025 | MedLSAM: Localize and segment anything model for 3D CT images
Wenhui Lei, Wei Xu 0046, Kang Li 0004, Xiaofan Zhang 0002, Shaoting Zhang 0001 |
Medical Image Anal. | 3 |
| 2025 | Interpretable 2.5D network by hierarchical attention and consistency learning for 3D MRI classification
Shuting Pang, Xiaoshuang Shi, Rui Wang 0108, Mingzhe Dai, Xiaofeng Zhu 0001, Bin Song 0002, Kang Li 0004 |
Pattern Recognit. | 8 |
| 2025 | Volume Fusion-Based Self-Supervised Pretraining for 3D Medical Image SegmentationabstractThe performance of deep learning models for medical image segmentation is often limited in scenarios where training data or annotations are limited. Self-Supervised Learning (SSL) is an appealing solution for this dilemma due to its feature learning ability from a large amount of unannotated images. Existing SSL methods have focused on pretraining either an encoder for global feature representation or an encoder-decoder structure for image restoration, where the gap between pretext and downstream tasks limits the usefulness of pretrained decoders in downstream segmentation. In this work, we propose a novel SSL strategy named Volume Fusion (VolF) for pretraining 3D segmentation models. It minimizes the gap between pretext and downstream tasks by introducing a pseudo-segmentation pretext task, where two sub-volumes are fused by a discretized block-wise fusion coefficient map. The model takes the fused result as input and predicts the category of fusion coefficient for each voxel, which can be trained with standard supervised segmentation loss functions without manual annotations. Experiments with an abdominal CT dataset for pretraining and both in-domain and out-domain downstream datasets showed that VolF led to large performance gain from training from scratch with faster convergence speed, and outperformed several state-of-the-art SSL methods. In addition, it is general to different network structures, and the learned features have high generalizability to different body parts and modalities. Guotai Wang, Jianghao Wu 0001, Xiangde Luo, Yubo Zhou, Xinglong Liu, Kang Li 0004, Jingsheng Lin, Baiyong Shen, Shaoting Zhang 0001 |
IEEE Trans. Image Process. | 7 |
| 2025 | Boosting Your Context by Dual Similarity Checkup for In-Context Learning Medical Image SegmentationabstractThe recent advent of in-context learning (ICL) capabilities in large pre-trained models has yielded significant advancements in the generalization of segmentation models. By supplying domain-specific image-mask pairs, the ICL model can be effectively guided to produce optimal segmentation outcomes, eliminating the necessity for model fine-tuning or interactive prompting. However, current existing ICL-based segmentation models exhibit significant limitations when applied to medical segmentation datasets with substantial diversity. To address this issue, we propose a dual similarity checkup approach to guarantee the effectiveness of selected in-context samples so that their guidance can be maximally leveraged during inference. We first employ large pre-trained vision models for extracting strong semantic representations from input images and constructing a feature embedding memory bank for semantic similarity checkup during inference. Assuring the similarity in the input semantic space, we then minimize the discrepancy in the mask appearance distribution between the support set and the estimated mask appearance prior through similarity-weighted sampling and augmentation. We validate our proposed dual similarity checkup approach on eight publicly available medical segmentation datasets, and extensive experimental results demonstrate that our proposed method significantly improves the performance metrics of existing ICL-based segmentation models, particularly when applied to medical image datasets characterized by substantial diversity. Qicheng Lao, Qingbo Kang, Paul Liu 0003, Chenlin Du, Kang Li 0004, Le Zhang 0004 |
IEEE Trans. Medical Imaging | 6 |
| 2024 | Patch Target Guided Dual-Branch Deep Multiple Instance Learning for 3D MRI AnalysisabstractDeep multiple instance learning (MIL) has attracted considerable attention in medical image analysis, since it only requires image-level labels for model training without using fine-grained (or patch) annotations. Unfortunately, MIL-based methods might lose some significant patch features. Although pseudo-label-based methods, which assign a pre-defined label to each patch, can explore more patch-level features, they might bring label noise and make the patch-level features lose diversity, thereby possibly restricting the model performance. To overcome this issue, we propose a novel gradient-based patch target generation (PTG) module to dynamically produce a feature vector for each patch as its target. Additionally, based on the PTG module, we propose a patch-target guided dual-branch deep MIL framework for 3D MRI data analysis, where both the two branches consist of a CNN model to extract patch-level features, an attention module to interpret the significance of patches, and a bag-level classifier, while the second branch also contains the PTG module to generate patch targets of patches. Moreover, the two branches are alternatively updated in our framework, resulting in a bi-level optimization problem, and thus we design a bi-level optimization algorithm to solve our proposed objective function. Extensive experiments demonstrate the superior classification and interpretation performance of the proposed framework over recent state-of-the-art methods. Codes are available at https://github.com/daimz1213/mil2024. Mingzhe Dai, Xiaoshuang Shi, Xiaofeng Zhu 0001, Tingrui Pan, Kang Li 0004 |
BIBM | 5 |
| 2024 | SGX2CT: Self-part Guided 3D CT and Spine Model Reconstruction from Biplanar Lumbar X-raysabstractComputed tomography (CT) imaging, characterized by its high contrast sensitivity and spatial resolution, provides detailed insights into a patient’s internal anatomical structures, particularly the intricate details of the spinal vertebrae. Compared to conventional X-ray imaging, CT scans entail higher radiation exposure and increased costs. The reconstruction of three-dimensional CT images and precise spinal models from two-dimensional X-ray images has garnered clinical interest due to lower radiation risks and improved accessibility. Inspired by self-supervised learning principles, this study introduces a novel self-guided cross-domain learning approach, SGX2CT, aimed at simultaneously reconstructing three-dimensional lumbar CT scans and spinal models from two-dimensional X-rays while facilitating mutual reinforcement. Specifically, a multi-domain shared generator is employed to acquire robust global information, coupled with perceptual loss functions for alignment in image space. Domain-specific group normalization is utilized to disentangle distinct features across domains to enhance outcomes in each domain. Experimental findings demonstrate that SGX2CT achieves state-of-the-art performance in both lumbar CT reconstruction and spinal model generation, underscoring its potential utility in orthopedic practice. Kaiyu Guo, Liang Zhao 0018, Xiandi Wang, Kang Li 0004 |
BIBM | 4 |
| 2024 | One-to-Normal: Anomaly Personalization for Few-shot Anomaly DetectionabstractTraditional Anomaly Detection (AD) methods have predominantly relied on unsupervised learning from extensive normal data. Recent AD methods have evolved with the advent of large pre-trained vision-language models, enhancing few-shot anomaly detection capabilities. However, these latest AD methods still exhibit limitations in accuracy improvement. One contributing factor is their direct comparison of a query image's features with those of few-shot normal images. This direct comparison often leads to a loss of precision and complicates the extension of these techniques to more complex domains—an area that remains underexplored in a more refined and comprehensive manner. To address these limitations, we introduce the anomaly personalization method, which performs a personalized one-to-normal transformation of query images using an anomaly-free customized generation model, ensuring close alignment with the normal manifold. Moreover, to further enhance the stability and robustness of prediction results, we propose a triplet contrastive anomaly inference strategy, which incorporates a comprehensive comparison between the query and generated anomaly-free data pool and prompt information. Extensive evaluations across eleven datasets in three domains demonstrate our model's effectiveness compared to the latest AD methods. Additionally, our method has been proven to transfer flexibly to other AD methods, with the generated image data effectively improving the performance of other AD methods. Yiyue Li, Shaoting Zhang 0001, Kang Li 0004, Qicheng Lao |
NeurIPS | 3 |
| 2024 | Healthcare facilities management: A novel data-driven model for predictive maintenance of computed tomography equipment
Haopeng Zhou, Qilin Liu, Zhenlin Li, Yixuan Zhuo, Kang Li 0004, Changxi Wang |
Artif. Intell. Medicine | 7 |
| 2024 | Prognostic and Health Management of CT Equipment via a Distance Self-Attention Network Using Internet of ThingsabstractAnomalies or failures in medical equipment may lead to severe consequences. Data-driven prognostic and health management (PHM) approaches can improve maintenance efficiency and reduce maintenance costs at hospitals while protecting patients’ lives. However, currently, the research and application of PHM in medical equipment is still rather limited. The development of the Internet of Things (IoT) technology provides new opportunities for PHM, which can safely collect, analyze, and store real-time equipment data in hospitals. The data-driven models used in PHM predict anomalies or failures. However, current data-driven models’ performance may be limited due to lack of consideration for the interaction of similar features and the importance of different time steps. Hence, this article proposes a new deep-learning network called similar feature interaction (SFI) with distance self-attention (SA) for the PHM of medical equipment. First, an SFI module which uses clustering algorithms and causal convolution layers is proposed to consider the interaction of similar features. Second, a distance SA mechanism is proposed to allocate more attention to important time steps. The experiments on millions of computed tomography (CT) equipment operating status instants collected by IoT in the hospital and the public data set show that the proposed model is superior to existing models. The results show that the accuracy, recall, precision, and f1-score of the proposed model on the real CT log data achieve 0.865, 0.682, 0.469, and 0.556, respectively. The proposed PHM model can assist the equipment maintenance team of hospitals in decision making under the IoT framework. Haopeng Zhou, Zhenlin Li, Tong Wu 0026, Changxi Wang, Kang Li 0004 |
IEEE Internet Things J. | 5 |
| 2024 | Deblurring masked image modeling for ultrasound image analysis
Qingbo Kang, Qicheng Lao, Jingyan Liu, Huahui Yi, Buyun Ma, Xiaofan Zhang 0002, Kang Li 0004 |
Medical Image Anal. | 8 |
| 2024 | Diabetic foot ulcers segmentation challenge report: Benchmark and analysisabstractMonitoring the healing progress of diabetic foot ulcers is a challenging process. Accurate segmentation of foot ulcers can help podiatrists to quantitatively measure the size of wound regions to assist prediction of healing status. The main challenge in this field is the lack of publicly available manual delineation, which can be time consuming and laborious. Recently, methods based on deep learning have shown excellent results in automatic segmentation of medical images, however, they require large-scale datasets for training, and there is limited consensus on which methods perform the best. The 2022 Diabetic Foot Ulcers segmentation challenge was held in conjunction with the 2022 International Conference on Medical Image Computing and Computer Assisted Intervention, which sought to address these issues and stimulate progress in this research domain. A training set of 2000 images exhibiting diabetic foot ulcers was released with corresponding segmentation ground truth masks. Of the 72 (approved) requests from 47 countries, 26 teams used this data to develop fully automated systems to predict the true segmentation masks on a test set of 2000 images, with the corresponding ground truth segmentation masks kept private. Predictions from participating teams were scored and ranked according to their average Dice similarity coefficient of the ground truth masks and prediction masks. The winning team achieved a Dice of 0.7287 for diabetic foot ulcer segmentation. This challenge has now entered a live leaderboard stage where it serves as a challenging benchmark for diabetic foot ulcer segmentation. Moi Hoon Yap, Bill Cassidy, Michal Byra, Ting-Yu Liao, Huahui Yi, Adrian Galdran, Yung-Han Chen, Raphael Brüngel, Sven Koitka, Christoph M. Friedrich, Yu-Wen Lo, Ching-Hui Yang, Kang Li 0004, Qicheng Lao, Miguel Ángel González Ballester, Gustavo Carneiro 0001, Yi-Jen Ju, Juinn-Dar Huang, Joseph Pappachan, Neil D. Reeves, Vishnu Chandrabalan, Darren Dancey, Connah Kendrick |
Medical Image Anal. | 13 |
| 2024 | Prognosis prediction of high grade serous adenocarcinoma based on multi-modal convolution neural network
Zongyuan Gan, Kang Li 0004 |
Neural Comput. Appl. | 4 |
| 2024 | FlowX: Towards Explainable Graph Neural Networks via Message FlowsabstractWe investigate the explainability of graph neural networks (GNNs) as a step toward elucidating their working mechanisms. While most current methods focus on explaining graph nodes, edges, or features, we argue that, as the inherent functional mechanism of GNNs, message flows are more natural for performing explainability. To this end, we propose a novel method here, known as FlowX, to explain GNNs by identifying important message flows. To quantify the importance of flows, we propose to follow the philosophy of Shapley values from cooperative game theory. To tackle the complexity of computing all coalitions' marginal contributions, we propose a flow sampling scheme to compute Shapley value approximations as initial assessments of further training. We then propose an information-controlled learning algorithm to train flow scores toward diverse explanation targets: necessary or sufficient explanations. Experimental studies on both synthetic and real-world datasets demonstrate that our proposed FlowX and its variants lead to improved explainability of GNNs. Shurui Gui, Hao Yuan 0001, Jie Wang 0005, Qicheng Lao, Kang Li 0004, Shuiwang Ji |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Hierarchical-Instance Contrastive Learning for Minority Detection on Imbalanced Medical DatasetsabstractDeep learning methods are often hampered by issues such as data imbalance and data-hungry. In medical imaging, malignant or rare diseases are frequently of minority classes in the dataset, featured by diversified distribution. Besides that, insufficient labels and unseen cases also present conundrums for training on the minority classes. To confront the stated problems, we propose a novel Hierarchical-instance Contrastive Learning (HCLe) method for minority detection by only involving data from the majority class in the training stage. To tackle inconsistent intra-class distribution in majority classes, our method introduces two branches, where the first branch employs an auto-encoder network augmented with three constraint functions to effectively extract image-level features, and the second branch designs a novel contrastive learning network by taking into account the consistency of features among hierarchical samples from majority classes. The proposed method is further refined with a diverse mini-batch strategy, enabling the identification of minority classes under multiple conditions. Extensive experiments have been conducted to evaluate the proposed method on three datasets of different diseases and modalities. The experimental results show that the proposed method outperforms the state-of-the-art methods. Yiyue Li, Guangwu Qian, Xiaoshuang Jiang, Zekun Jiang, Shaoting Zhang 0001, Kang Li 0004, Qicheng Lao |
IEEE Trans. Medical Imaging | 7 |
| 2024 | Masked Conditional Variational Autoencoders for Chromosome StraighteningabstractKaryotyping is of importance for detecting chromosomal aberrations in human disease. However, chromosomes easily appear curved in microscopic images, which prevents cytogeneticists from analyzing chromosome types. To address this issue, we propose a framework for chromosome straightening, which comprises a preliminary processing algorithm and a generative model called masked conditional variational autoencoders (MC-VAE). The processing method utilizes patch rearrangement to address the difficulty in erasing low degrees of curvature, providing reasonable preliminary results for the MC-VAE. The MC-VAE further straightens the results by leveraging chromosome patches conditioned on their curvatures to learn the mapping between banding patterns and conditions. During model training, we apply a masking strategy with a high masking ratio to train the MC-VAE with eliminated redundancy. This yields a non-trivial reconstruction task, allowing the model to effectively preserve chromosome banding patterns and structure details in the reconstructed results. Extensive experiments on three public datasets with two stain styles show that our framework surpasses the performance of state-of-the-art methods in retaining banding patterns and structure details. Compared to using real-world bent chromosomes, the use of high-quality straightened chromosomes generated by our proposed method can improve the performance of various deep learning models for chromosome classification by a large margin. Such a straightening approach has the potential to be combined with other karyotyping systems to assist cytogeneticists in chromosome analysis. Jingxiong Li, Sunyi Zheng, Zhongyi Shui, Shichuan Zhang, Linyi Yang, Yuxuan Sun 0002, Honglin Li 0001, Yuanxin Ye, Peter M. A. van Ooijen, Kang Li 0004, Lin Yang 0002 |
IEEE Trans. Medical Imaging | 11 |
| 2024 | FPL+: Filtered Pseudo Label-Based Unsupervised Cross-Modality Adaptation for 3D Medical Image SegmentationabstractAdapting a medical image segmentation model to a new domain is important for improving its cross-domain transferability, and due to the expensive annotation process, Unsupervised Domain Adaptation (UDA) is appealing where only unlabeled images are needed for the adaptation. Existing UDA methods are mainly based on image or feature alignment with adversarial training for regularization, and they are limited by insufficient supervision in the target domain. In this paper, we propose an enhanced Filtered Pseudo Label (FPL+)-based UDA method for 3D medical image segmentation. It first uses cross-domain data augmentation to translate labeled images in the source domain to a dual-domain training set consisting of a pseudo source-domain set and a pseudo target-domain set. To leverage the dual-domain augmented images to train a pseudo label generator, domain-specific batch normalization layers are used to deal with the domain shift while learning the domain-invariant structure features, generating high-quality pseudo labels for target-domain images. We then combine labeled source-domain images and target-domain images with pseudo labels to train a final segmentor, where image-level weighting based on uncertainty estimation and pixel-level weighting based on dual-domain consensus are proposed to mitigate the adverse effect of noisy pseudo labels. Experiments on three public multi-modal datasets for Vestibular Schwannoma, brain tumor and whole heart segmentation show that our method surpassed ten state-of-the-art UDA methods, and it even achieved better results than fully supervised learning in the target domain in some cases. Jianghao Wu 0001, Guotai Wang, Qiang Yue 0005, Huijun Yu, Kang Li 0004, Shaoting Zhang 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2023 | Pre-Trained Tabular Transformer for Real-Time, Efficient, Stable Radiomics Data Processing: A Comprehensive StudyabstractRadiomics is an important research direction in the field of medical image analysis. Although the number of publications is increasing year by year, it has been difficult to translate into clinical practice due to the small size of clinical data. In most cases, Radiomics data can be considered as small tabular data. Deep learning is often less effective than classical machine learning algorithms in processing tabular data. Recently, table representation learning has started to receive more widespread attention which is often an easily overlooked but very important area of research, helping improve the status of tabular deep learning. Here, we first apply a pre-trained Transformer model named Tabular Prior-Data Fitted Network (TabPFN) to the field of Radiomics analysis. We implement extensive experiments on three real-world clinical datasets: (a) Ultrasound Radiomics dataset for the classificatory diagnosis of Kidney tumor, (b) CT Radiomics dataset for the prediction of EGFR gene mutations in non-small cell lung cancer, (c) MRI Radiomics dataset for the prediction of treatment response of brain metastases to gamma knife radiosurgery. By comprehensive analysis, we demonstrate that the pre-trained tabular Transformer can be used as a realtime, efficient, and stable Radiomics data processor with superior performance over other tabular machine learning methods in different clinical tasks. We also simulate an ideal clinical practice scenario for evaluating the clinical translation potential of pretrained models. Finally, we explore the advantages and limitations of pre-trained tabular models for Radiomics analysis. Zekun Jiang, Ruchun Jia, Le Zhang 0004, Kang Li 0004 |
HealthCom | 4 |
| 2023 | MEDICAL IMAGE UNDERSTANDING WITH PRETRAINED VISION LANGUAGE MODELS: A COMPREHENSIVE STUDY
Ziyuan Qin 0001, Huahui Yi, Qicheng Lao, Kang Li 0004 |
ICLR | 4 |
| 2023 | DMCVR: Morphology-Guided Diffusion Model for 3D Cardiac Volume Reconstruction
Xiaoxiao He, Chaowei Tan, Ligong Han, Bo Liu 0005, Leon Axel, Kang Li 0004, Dimitris N. Metaxas |
MICCAI (7) | 6 |
| 2023 | Deblurring Masked Autoencoder Is Better Recipe for Ultrasound Image Recognition
Qingbo Kang, Kang Li 0004, Qicheng Lao |
MICCAI (1) | 3 |
| 2023 | Deep learning-based estimation of whole-body kinematics from multi-view images
Kien X. Nguyen 0001, Liying Zheng, Ashley L. Hawke, Robert E. Carey, Scott P. Breloff, Kang Li 0004, Xi Peng 0005 |
Comput. Vis. Image Underst. | 6 |
| 2023 | MTMVC: Semi-supervised 3D hand pose estimation using multi-task and multi-view consistency
Donghai Xiang, Wei Xu 0046, Bei Peng 0002, Guotai Wang, Kang Li 0004 |
J. Vis. Commun. Image Represent. | 6 |
| 2023 | CDDSA: Contrastive domain disentanglement and style augmentation for generalizable medical image segmentation
Ran Gu, Guotai Wang, Jiangshan Lu, Jingyang Zhang, Wenhui Lei, Wenjun Liao, Shichuan Zhang, Kang Li 0004, Dimitris N. Metaxas, Shaoting Zhang 0001 |
Medical Image Anal. | 9 |
| 2023 | Self-supervised anomaly detection, staging and segmentation for retinal images
Yiyue Li, Qicheng Lao, Qingbo Kang, Zekun Jiang, Shiyi Du, Shaoting Zhang 0001, Kang Li 0004 |
Medical Image Anal. | 7 |
| 2023 | A novel one-to-multiple unsupervised domain adaptation framework for abdominal organ segmentation
Jianghao Wu 0001, Jiangshan Lu, Yuxiang Ye, Yechong Huang, Xin Dou, Kang Li 0004, Guotai Wang, Shaoting Zhang 0001 |
Medical Image Anal. | 8 |
| 2023 | Self-paced resistance learning against overfitting on noisy labels
Xiaoshuang Shi, Zhenhua Guo 0001, Kang Li 0004, Yun Liang 0012, Xiaofeng Zhu 0001 |
Pattern Recognit. | 3 |
| 2023 | Anatomically Guided Cross-Domain Repair and Screening for Ultrasound Fetal BiometryabstractUltrasound based estimation of fetal biometry is extensively used to diagnose prenatal abnormalities and to monitor fetal growth, for which accurate segmentation of the fetal anatomy is a crucial prerequisite. Although deep neural network-based models have achieved encouraging results on this task, inevitable distribution shifts in ultrasound images can still result in severe performance drop in real world deployment scenarios. In this article, we propose a complete ultrasound fetal examination system to deal with this troublesome problem by repairing and screening the anatomically implausible results. Our system consists of three main components: A routine segmentation network, a fetal anatomical key points guided repair network, and a shape-coding based selective screener. Guided by the anatomical key points, our repair network has stronger cross-domain repair capabilities, which can substantially improve the outputs of the segmentation network. By quantifying the distance between an arbitrary segmentation mask to its corresponding anatomical shape class, the proposed shape-coding based selective screener can then effectively reject the entire implausible results that cannot be fully repaired. Extensive experiments demonstrate that our proposed framework has strong anatomical guarantee and outperforms other methods in three different cross-domain scenarios. Qicheng Lao, Paul Liu 0003, Huahui Yi, Qingbo Kang, Zekun Jiang, Kang Li 0004, Yuanyuan Chen 0006, Le Zhang 0004 |
IEEE J. Biomed. Health Informatics | 8 |
| 2023 | Contrastive Semi-Supervised Learning for Domain Adaptive Segmentation Across Similar Anatomical StructuresabstractConvolutional Neural Networks (CNNs) have achieved state-of-the-art performance for medical image segmentation, yet need plenty of manual annotations for training. Semi-Supervised Learning (SSL) methods are promising to reduce the requirement of annotations, but their performance is still limited when the dataset size and the number of annotated images are small. Leveraging existing annotated datasets with similar anatomical structures to assist training has a potential for improving the model's performance. However, it is further challenged by the cross-anatomy domain shift due to the image modalities and even different organs in the target domain. To solve this problem, we propose Contrastive Semi-supervised learning for Cross Anatomy Domain Adaptation (CS-CADA) that adapts a model to segment similar structures in a target domain, which requires only limited annotations in the target domain by leveraging a set of existing annotated images of similar structures in a source domain. We use Domain-Specific Batch Normalization (DSBN) to individually normalize feature maps for the two anatomical domains, and propose a cross-domain contrastive learning strategy to encourage extracting domain invariant features. They are integrated into a Self-Ensembling Mean-Teacher (SE-MT) framework to exploit unlabeled target domain images with a prediction consistency constraint. Extensive experiments show that our CS-CADA is able to solve the challenging cross-anatomy domain shift problem, achieving accurate segmentation of coronary arteries in X-ray images with the help of retinal vessel images and cardiac MR images with the help of fundus images, respectively, given only a small number of annotations in the target domain. Our code is available at https://github.com/HiLab-git/DAG4MIA. Ran Gu, Jingyang Zhang, Guotai Wang, Wenhui Lei, Tao Song 0002, Xiaofan Zhang 0002, Kang Li 0004, Shaoting Zhang 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2023 | PA-Seg: Learning From Point Annotations for 3D Medical Image Segmentation Using Contextual Regularization and Cross Knowledge DistillationabstractThe success of Convolutional Neural Networks (CNNs) in 3D medical image segmentation relies on massive fully annotated 3D volumes for training that are time-consuming and labor-intensive to acquire. In this paper, we propose to annotate a segmentation target with only seven points in 3D medical images, and design a two-stage weakly supervised learning framework PA-Seg. In the first stage, we employ geodesic distance transform to expand the seed points to provide more supervision signal. To further deal with unannotated image regions during training, we propose two contextual regularization strategies, i.e., multi-view Conditional Random Field (mCRF) loss and Variance Minimization (VM) loss, where the first one encourages pixels with similar features to have consistent labels, and the second one minimizes the intensity variance for the segmented foreground and background, respectively. In the second stage, we use predictions obtained by the model pre-trained in the first stage as pseudo labels. To overcome noises in the pseudo labels, we introduce a Self and Cross Monitoring (SCM) strategy, which combines self-training with Cross Knowledge Distillation (CKD) between a primary model and an auxiliary model that learn from soft labels generated by each other. Experiments on public datasets for Vestibular Schwannoma (VS) segmentation and Brain Tumor Segmentation (BraTS) demonstrated that our model trained in the first stage outperformed existing state-of-the-art weakly supervised approaches by a large margin, and after using SCM for additional training, the model's performance was close to its fully supervised counterpart on the BraTS dataset. Shuwei Zhai, Guotai Wang, Xiangde Luo, Qiang Yue 0005, Kang Li 0004, Shaoting Zhang 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2023 | A Robust Region Control Approach for Simultaneous Trajectory Tracking and Compliant Physical Human-Robot InteractionabstractFor the safe and smooth robot-assisted healthcare task execution, real-time motion tracking controls and compliant physical human–robot interactions are concurrently important control objectives. In this work, the uncertainty and disturbance estimator (UDE)-based robust region tracking controller for a robot manipulator is developed. The regional feedback error is derived from the potential function to drive the robot manipulator end-effector converging into the target region, where the safe and compliant physical human–robot interaction can be achieved. Utilizing the back-stepping control approach, the regional feedback error is seamlessly integrated into the UDE-based control framework, where the UDE is employed to estimate and compensate model uncertainties such that only the minimum model information is needed for implementation. The Lyapunov method is used to analyze the stability of the closed-loop control system. Extensive experimental studies including trajectory tracking, human–robot interaction and benchmark comparison are carried out for controller effectiveness validation. Xiangyun Li, Qi Lu 0003, Ning Jiang 0001, Kang Li 0004 |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2022 | Distilling Knowledge from Topological Representations for Pathological Complete Response Prediction
Shiyi Du, Qicheng Lao, Qingbo Kang, Yiyue Li, Zekun Jiang, Kang Li 0004 |
MICCAI (2) | 7 |
| 2022 | Unsupervised Cross-disease Domain Adaptation by Lesion Scale Matching
Qicheng Lao, Qingbo Kang, Paul Liu 0003, Le Zhang 0004, Kang Li 0004 |
MICCAI (8) | 6 |
| 2022 | Thyroid nodule segmentation and classification in ultrasound images through intra- and inter-task consistent learning
Qingbo Kang, Qicheng Lao, Yiyue Li, Zekun Jiang, Shaoting Zhang 0001, Kang Li 0004 |
Medical Image Anal. | 7 |
| 2022 | WORD: A large scale dataset, benchmark and clinical applicable study for abdominal organ segmentation from CT image
Xiangde Luo, Wenjun Liao, Jianghong Xiao, Jieneng Chen, Tao Song 0002, Xiaofan Zhang 0002, Kang Li 0004, Dimitris N. Metaxas, Guotai Wang, Shaoting Zhang 0001 |
Medical Image Anal. | 7 |
| 2022 | SCPM-Net: An anchor-free 3D lung nodule detection network using sphere representation and center points matching
Xiangde Luo, Tao Song 0002, Guotai Wang, Jieneng Chen, Kang Li 0004, Dimitris N. Metaxas, Shaoting Zhang 0001 |
Medical Image Anal. | 6 |
| 2022 | Multiview Video-Based 3-D Pose Estimation of Patients in Computer-Assisted Rehabilitation Environment (CAREN)abstractThe computer-assisted rehabilitation environment (CAREN) system plays an important role in the training of rehabilitation patients, where the capture of the patient's 3-D pose and gait is critical for assessing the patient's requirements for effective training. Vision-based methods are highly effective for this task due to their low cost, high speed, and noninterference. Although various general vision-based pose estimation methods were developed recently, their performance is limited in the CAREN system due to the specific environment. To address these problems, we propose an improved framework for accurate 2-D and 3-D pose estimation for the CAREN system through using multiview videos. First, for 2-D pose estimation, we propose a coarse-to-fine heatmap shrinking (CFHS) strategy that gradually reduces the kernel size of the heatmap of joints during training to improve the performance. Second, to further obtain 3-D pose estimations, we propose a novel spatial-temporal perception network that fuses the 2-D results from multiple views and multiple moments; multiview early fusion uses complementary spatial information from different views, and multimoment late fusion leverages temporal information from the sequential input for higher accuracy. The experimental results, based on CAREN videos of 225 orthopedic patients, showed that the accuracy of 2-D human pose estimations with the CFHS training strategy reached 99.85% [email protected]. For 3-D results, the mean per joint position error was 25.22 mm, and the 3DPCK reached 98.71%, which outperformed existing general video-based methods. The results showed that the proposed system is capable of estimating human poses with high accuracy for clinical applications. Wei Xu 0046, Donghai Xiang, Guotai Wang, Ruisong Liao, Ming Shao, Kang Li 0004 |
IEEE Trans. Hum. Mach. Syst. | 6 |
| 2022 | HMRNet: High and Multi-Resolution Network With Bidirectional Feature Calibration for Brain Structure Segmentation in RadiotherapyabstractAccurate segmentation of Anatomical brain Barriers to Cancer spread (ABCs) plays an important role for automatic delineation of Clinical Target Volume (CTV) of brain tumors in radiotherapy. Despite that variants of U-Net are state-of-the-art segmentation models, they have limited performance when dealing with ABCs structures with various shapes and sizes, especially thin structures (e.g., the falx cerebri) that span only few slices. To deal with this problem, we propose a High and Multi-Resolution Network (HMRNet) that consists of a multi-scale feature learning branch and a high-resolution branch, which can maintain the high-resolution contextual information and extract more robust representations of anatomical structures with various scales. We further design a Bidirectional Feature Calibration (BFC) block to enable the two branches to generate spatial attention maps for mutual feature calibration. Considering the different sizes and positions of ABCs structures, our network was applied after a rough localization of each structure to obtain fine segmentation results. Experiments on the MICCAI 2020 ABCs challenge dataset showed that: 1) Our proposed two-stage segmentation strategy largely outperformed methods segmenting all the structures in just one stage; 2) The proposed HMRNet with two branches can maintain high-resolution representations and is effective to improve the performance on thin structures; 3) The proposed BFC block outperformed existing attention methods using monodirectional feature calibration. Our method won the second place of ABCs 2020 challenge and has a potential for more accurate and reasonable delineation of CTV of brain tumors. Hao Fu 0014, Guotai Wang, Wenhui Lei, Wei Xu 0046, Qianfei Zhao, Shichuan Zhang, Kang Li 0004, Shaoting Zhang 0001 |
IEEE J. Biomed. Health Informatics | 7 |
| 2022 | Learning COVID-19 Pneumonia Lesion Segmentation From Imperfect Annotations via Divergence-Aware Selective TrainingabstractAutomatic segmentation of COVID-19 pneumonia lesions is critical for quantitative measurement for diagnosis and treatment management. For this task, deep learning is the state-of-the-art method while requires a large set of accurately annotated images for training, which is difficult to obtain due to limited access to experts and the time-consuming annotation process. To address this problem, we aim to train the segmentation network from imperfect annotations, where the training set consists of a small clean set of accurately annotated images by experts and a large noisy set of inaccurate annotations by non-experts. To avoid the labels with different qualities corrupting the segmentation model, we propose a new approach to train segmentation networks to deal with noisy labels. We introduce a dual-branch network to separately learn from the accurate and noisy annotations. To fully exploit the imperfect annotations as well as suppressing the noise, we design a Divergence-Aware Selective Training (DAST) strategy, where a divergence-aware noisiness score is used to identify severely noisy annotations and slightly noisy annotations. For severely noisy samples we use an regularization through dual-branch consistency between predictions from the two branches. We also refine slightly noisy samples and use them as supplementary data for the clean branch to avoid overfitting. Experimental results show that our method achieves a higher performance than standard training process for COVID-19 pneumonia lesion segmentation when learning from imperfect labels, and our framework outperforms the state-of-the-art noise-tolerate methods significantly with various clean label percentages. Shuojue Yang, Guotai Wang, Xiangde Luo, Kang Li 0004, Qijun Wang, Shaoting Zhang 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2021 | On Explainability of Graph Neural Networks via Subgraph ExplorationsabstractWe consider the problem of explaining the predictions of graph neural networks (GNNs), which otherwise are considered as black boxes. Existing methods invariably focus on explaining the importance of graph nodes or edges but ignore the substructures of graphs, which are more intuitive and human-intelligible. In this work, we propose a novel method, known as SubgraphX, to explain GNNs by identifying important subgraphs. Given a trained GNN model and an input graph, our SubgraphX explains its predictions by efficiently exploring different subgraphs with Monte Carlo tree search. To make the tree search more effective, we propose to use Shapley values as a measure of subgraph importance, which can also capture the interactions among different subgraphs. To expedite computations, we propose efficient approximation schemes to compute Shapley values for graph data. Our work represents the first attempt to explain GNNs via identifying subgraphs explicitly and directly. Experimental results show that our SubgraphX achieves significantly improved explanations, while keeping computations at a reasonable level. Hao Yuan 0001, Haiyang Yu 0005, Jie Wang 0005, Kang Li 0004, Shuiwang Ji |
ICML | 4 |
| 2021 | Surgical planning of pelvic tumor using multi-view CNN with relation-context representation learningabstractLimb salvage surgery of malignant pelvic tumors is the most challenging procedure in musculoskeletal oncology due to the complex anatomy of the pelvic bones and soft tissues. It is crucial to accurately resect the pelvic tumors with appropriate margins in this procedure. However, there is still a lack of efficient and repetitive image planning methods for tumor identification and segmentation in many hospitals. In this paper, we present a novel deep learning-based method to accurately segment pelvic bone tumors in MRI. Our method uses a multi-view fusion network to extract pseudo-3D information from two scans in different directions and improves the feature representation by learning a relational context. In this way, it can fully utilize spatial information in thick MRI scans and reduce over-fitting when learning from a small dataset. Our proposed method was evaluated on two independent datasets collected from 90 and 15 patients, respectively. The segmentation accuracy of our method was superior to several comparing methods and comparable to the expert annotation, while the average time consumed decreased about 100 times from 1820.3 seconds to 19.2 seconds. In addition, we incorporate our method into an efficient workflow to improve the surgical planning process. Our workflow took only 15 minutes to complete surgical planning in a phantom study, which is a dramatic acceleration compared with the 2-day time span in a traditional workflow. Zhennan Yan, Liang Zhao 0018, Lichi Zhang, Shuaining Xie, Kang Li 0004, Dimitris N. Metaxas, Yongqiang Hao, Kerong Dai, Shaoting Zhang 0001, Xiaofeng Tao 0002, Songtao Ai |
Medical Image Anal. | 8 |
| 2020 | Towards Efficient U-Nets: A Coupled and Quantized ApproachabstractIn this paper, we propose to couple stacked U-Nets for efficient visual landmark localization. The key idea is to globally reuse features of the same semantic meanings across the stacked U-Nets. The feature reuse makes each U-Net light-weighted. Specially, we propose an order- K coupling design to trim off long-distance shortcuts, together with an iterative refinement and memory sharing mechanism. To further improve the efficiency, we quantize the parameters, intermediate features, and gradients of the coupled U-Nets to low bit-width numbers. We validate our approach in two tasks: human pose estimation and facial landmark localization. The results show that our approach achieves state-of-the-art localization accuracy but using ∼ 70% fewer parameters, ∼ 30% less inference time, ∼ 98% less model size, and saving ∼ 75% training memory compared with benchmark localizers. Zhiqiang Tang 0001, Xi Peng 0005, Kang Li 0004, Dimitris N. Metaxas |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | Vertebrae Identification and Localization Utilizing Fully Convolutional Networks and a Hidden Markov ModelabstractAutomated identification and localization of vertebrae in spinal computed tomography (CT) imaging is a complicated hybrid task. This task requires detecting and indexing a long sequence in a 3-D image, and both image feature extraction and sequence modeling are needed to address the problem. In this paper, the powerful fully convolutional neural network (FCN) technique performs both of these tasks simultaneously because FCNs directly encode and decode the spatial interdependence of different components in images. The key module of our proposed framework is a 3-D FCN trained in an end-to-end manner at the spine level to capture the long-range contextual information in CT volumes. The large increase in the calculation due to the full-size image inputs is alleviated by the scale-down of the inputs and the use of an auxiliary FCN to compensate for the loss of details. The composite network pipeline design enables the integration of local image details and global image patterns. Furthermore, explicit spatial and sequential constraints are imposed by the hidden Markov model (HMM) for a higher robustness and a clearer interpretation of network outputs. The proposed framework is quantitatively evaluated on the public dataset from the MICCAI 2014 Computational Challenge on Vertebrae Localization and Identification and demonstrates an identification rate (within 20 mm) of 94.67%, a mean identification rate of 87.97%, and a mean error distance of 2.56 mm on the test set, thus achieving the highest performance reported on this dataset. Yizhi Chen, Yunhe Gao, Kang Li 0004, Liang Zhao 0018, Jun Zhao 0010 |
IEEE Trans. Medical Imaging | 3 |
| 2020 | Weakly Supervised Deep Nuclei Segmentation Using Partial Points Annotation in Histopathology ImagesabstractNuclei segmentation is a fundamental task in histopathology image analysis. Typically, such segmentation tasks require significant effort to manually generate accurate pixel-wise annotations for fully supervised training. To alleviate such tedious and manual effort, in this paper we propose a novel weakly supervised segmentation framework based on partial points annotation, i.e., only a small portion of nuclei locations in each image are labeled. The framework consists of two learning stages. In the first stage, we design a semi-supervised strategy to learn a detection model from partially labeled nuclei locations. Specifically, an extended Gaussian mask is designed to train an initial model with partially labeled data. Then, self-training with background propagation is proposed to make use of the unlabeled regions to boost nuclei detection and suppress false positives. In the second stage, a segmentation model is trained from the detected nuclei locations in a weakly-supervised fashion. Two types of coarse labels with complementary information are derived from the detected points and are then utilized to train a deep neural network. The fully-connected conditional random field loss is utilized in training to further refine the model without introducing extra computational complexity during inference. The proposed method is extensively evaluated on two nuclei segmentation datasets. The experimental results demonstrate that our method can achieve competitive performance compared to the fully supervised counterpart and the state-of-the-art methods while requiring significantly less annotation effort. Pengxiang Wu, Qiaoying Huang, Jingru Yi, Zhennan Yan, Kang Li 0004, Gregory M. Riedlinger, Subhajyoti De, Shaoting Zhang 0001, Dimitris N. Metaxas |
IEEE Trans. Medical Imaging | 6 |
| 2020 | A Noise-Robust Framework for Automatic Segmentation of COVID-19 Pneumonia Lesions From CT ImagesabstractSegmentation of pneumonia lesions from CT scans of COVID-19 patients is important for accurate diagnosis and follow-up. Deep learning has a potential to automate this task but requires a large set of high-quality annotations that are difficult to collect. Learning from noisy training labels that are easier to obtain has a potential to alleviate this problem. To this end, we propose a novel noise-robust framework to learn from noisy labels for the segmentation task. We first introduce a noise-robust Dice loss that is a generalization of Dice loss for segmentation and Mean Absolute Error (MAE) loss for robustness against noise, then propose a novel COVID-19 Pneumonia Lesion segmentation network (COPLE-Net) to better deal with the lesions with various scales and appearances. The noise-robust Dice loss and COPLE-Net are combined with an adaptive self-ensembling framework for training, where an Exponential Moving Average (EMA) of a student model is used as a teacher model that is adaptively updated by suppressing the contribution of the student to EMA when the student has a large training loss. The student model is also adaptive by learning from the teacher only when the teacher outperforms the student. Experimental results showed that: (1) our noise-robust Dice loss outperforms existing noise-robust loss functions, (2) the proposed COPLE-Net achieves higher performance than state-of-the-art image segmentation networks, and (3) our framework with adaptive self-ensembling significantly outperforms a standard training process and surpasses other noise-robust training approaches in the scenario of learning from noisy labels for COVID-19 pneumonia lesion segmentation. Guotai Wang, Xinglong Liu, Chaoping Li, Jiugen Ruan, Kang Li 0004, Shaoting Zhang 0001 |
IEEE Trans. Medical Imaging | 8 |
| 2019 | Collaborative Multi-agent Learning for MR Knee Articular Cartilage Segmentation
Chaowei Tan, Zhennan Yan, Shaoting Zhang 0001, Kang Li 0004, Dimitris N. Metaxas |
MICCAI (2) | 4 |
| 2019 | Predicting 3-D Lower Back Joint Load in Lifting: A Deep Pose Estimation ApproachabstractGoal: Lifting is a common manual material handling task performed in the workplaces. It is considered as one of the main risk factors for work-related musculoskeletal disorders. An important criterion to identify the unsafe lifting task is the values of the net force and moment at L5/S1 joint. These values are mainly calculated in a laboratory environment, which utilizes marker-based sensors to collect three-dimensional (3-D) information and force plates to measure the external forces and moments. However, this method is usually expensive to set up, time-consuming in process, and sensitive to the surrounding environment. In this study, we propose a deep neural network (DNN)-based framework for 3-D pose estimation, which addresses the aforementioned limitations, and we employ the results for L5/S1 moment and force calculation. Methods: At the first step of the proposed framework, full body 3-D pose is captured using a DNN, then at the second step, estimated 3-D body pose along with the subject's anthropometric information is utilized to calculate L5/S1 join's kinetic by a top-down inverse dynamic algorithm. Results: To fully evaluate our approach, we conducted experiments using a lifting dataset consisting of 12 subjects performing various types of lifting tasks. The results are validated against a marker-based motion capture system as a reference. The grand mean ± SD of the total moment/force absolute errors across all the dataset was 9.06 ± 7.60 N·m/4.85 ± 4.85 N. Conclusion: The proposed method provides a reliable tool for assessment of the lower back kinetics during lifting and can be an alternative when the use of marker-based motion capture systems is not possible. Rahil Mehrizi, Xi Peng 0005, Dimitris N. Metaxas, Shaoting Zhang 0001, Kang Li 0004 |
IEEE Trans. Hum. Mach. Syst. | 6 |
| 2018 | Toward Marker-Free 3D Pose Estimation in Lifting: A Deep Multi-View SolutionabstractLifting is a common manual material handling task performed in the workplaces. It is considered as one of the main risk factors for Work-related Musculoskeletal Disorders. To improve work place safety, it is necessary to assess musculoskeletal and biomechanical risk exposures associated with these tasks, which requires very accurate 3D pose. Existing approaches mainly utilize marker-based sensors to collect 3D information. However, these methods are usually expensive to setup, timeconsuming in process, and sensitive to the surrounding environment. In this study, we propose a multi-view based deep perceptron approach to address aforementioned limitations. Our approach consists of two modules: a "view-specific perceptron" network extracts rich information independently from the image of view, which includes both 2D shape and hierarchical texture information; while a "multi-view integration" network synthesizes information from all available views to predict accurate 3D pose. To fully evaluate our approach, we carried out comprehensive experiments to compare different variants of our design. The results prove that our approach achieves comparable performance with former marker-based methods, i.e. an average error of 14:72 ± 2:96 mm on the lifting dataset. The results are also compared with state-of-the-art methods on HumanEva-I dataset [1], which demonstrates the superior performance of our approach. Rahil Mehrizi, Xi Peng 0005, Zhiqiang Tang 0001, Dimitris N. Metaxas, Kang Li 0004 |
FG | 6 |
| 2018 | Skeleton model based behavior recognition for pedestrians and cyclists from vehicle sce ne cameraabstractWith the significant advances in computer vision research, skeleton model based human pose recognition has become more accurate and time-efficient, although most of the applications are limited in laboratory environment or on surveillance videos. This paper proposes a pose tracking and behavior recognition method from in-vehicle scene camera. It will not only detect pedestrians on the road, but also generate their skeleton models describing head, limb, and trunk movements. Based on these more detailed movements of body parts, the proposed method is designed to track poses of pedestrians and cyclists with the potentials to enable automated pedestrian gesture reading and non-verbal interactions between autonomous vehicles and pedestrians. The proposed algorithm has been tested on different databases including TASI 110-car naturalistic driving database and Joint Attention for Autonomous Driving (JAAD) database. Results show that key frames describing different pedestrian and cyclist negotiation gestures are detected from the raw video streams using the proposed method. These results will improve our understanding of pedestrian and cyclist's intentions and can be further used for autonomous vehicle control algorithm development. Qiwen Deng, Renran Tian, Yaobin Chen, Kang Li 0004 |
Intelligent Vehicles Symposium | 4 |
| 2017 | A computationally efficient 3D/2D registration method based on image gradient direction probability density function
Soheil Ghafurian, Ilker Hacihaliloglu, Dimitris N. Metaxas, Virak Tan, Kang Li 0004 |
Neurocomputing | 5 |
| 2017 | Towards large-scale MR thigh image analysis via an integrated quantification framework
Chaowei Tan, Kang Li 0004, Zhennan Yan, Jingru Yi, Pengxiang Wu, Hui Jing Yu, Klaus Engelke, Dimitris N. Metaxas |
Neurocomputing | 2 |
| 2016 | A detection-driven and sparsity-constrained deformable model for fascia lata labeling and thigh inter-muscular adipose quantification
Chaowei Tan, Kang Li 0004, Zhennan Yan, Dong Yang 0005, Shaoting Zhang 0001, Hui Jing Yu, Klaus Engelke, Colin Miller, Dimitris N. Metaxas |
Comput. Vis. Image Underst. | 2 |
| 2014 | Deformable models with sparsity constraints for cardiac motion analysis
Yang Yu 0010, Shaoting Zhang 0001, Kang Li 0004, Dimitris N. Metaxas, Leon Axel |
Medical Image Anal. | 3 |
| 2013 | Mobile user classification and authorization based on gesture usage recognitionabstractIntelligent mobile devices have been widely serving in almost all aspects of everyday life, spanning from communication, web surfing, entertainment, to daily organizer. A large amount of sensitive and private information is stored on the mobile device, leading to severe data security concern. In this work, we propose a novel mobile user classification and authorization scheme based on the recognition of user's gesture. Compared to other security solutions like password, track pattern and finger print etc., our scheme can continuously evolve for better protection during the usage cycle of the mobile device. Besides the regular interactive screen and sensors of modern mobile devices, our scheme does not require any additional hardware supports. Kent W. Nixon, Xiang Chen 0010, Zhi-Hong Mao, Yiran Chen 0001, Kang Li 0004 |
ASP-DAC | 5 |
| 2013 | Collaborative Multi Organ Segmentation by Integrating Deformable and Graphical Models
Mustafa Gökhan Uzunbas, Chao Chen 0012, Shaoting Zhang 0001, Kilian M. Pohl, Kang Li 0004, Dimitris N. Metaxas |
MICCAI (2) | 5 |