VLDB 2026 Research / reviewers in the wild / expert
Kai Ma 0002
dblp:86/7113-2
· DBLP profile ↗
75ranked-venue papers
0as first author
45since 2021 · last 2024
0000-0003-2805-3692ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 49 · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 47 · 22 since 2021Artificial intelligence and machine learning · 21 · 14 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | StableDrag: Stable Dragging for Point-Based Image Editing
Yutao Cui, Shengming Cao, Kai Ma 0002, Limin Wang 0002 |
ECCV (58) | 5 |
| 2024 | VFIMamba: Video Frame Interpolation with State Space ModelsabstractInter-frame modeling is pivotal in generating intermediate frames for video frame interpolation (VFI). Current approaches predominantly rely on convolution or attention-based models, which often either lack sufficient receptive fields or entail significant computational overheads. Recently, Selective State Space Models (S6) have emerged, tailored specifically for long sequence modeling, offering both linear complexity and data-dependent modeling capabilities. In this paper, we propose VFIMamba, a novel frame interpolation method for efficient and dynamic inter-frame modeling by harnessing the S6 model. Our approach introduces the Mixed-SSM Block (MSB), which initially rearranges tokens from adjacent frames in an interleaved fashion and subsequently applies multi-directional S6 modeling. This design facilitates the efficient transmission of information across frames while upholding linear complexity. Furthermore, we introduce a novel curriculum learning strategy that progressively cultivates proficiency in modeling inter-frame dynamics across varying motion magnitudes, fully unleashing the potential of the S6 model. Experimental findings showcase that our method attains state-of-the-art performance across diverse benchmarks, particularly excelling in high-resolution scenarios. In particular, on the X-TEST dataset, VFIMamba demonstrates a noteworthy improvement of 0.80 dB for 4K frames and 0.96 dB for 2K frames. Chunxu Liu, Yutao Cui, Kai Ma 0002, Limin Wang 0002 |
NeurIPS | 5 |
| 2024 | Triplet-branch network with contrastive prior-knowledge embedding for disease grading
Yuexiang Li, Yawen Huang, Jingxin Liu 0005, Yi Lin 0009, Dong Wei 0004, Qirui Zhang 0004, Kai Ma 0002, Guangming Lu 0001, Yefeng Zheng 0001 |
Artif. Intell. Medicine | 9 |
| 2024 | Improving vision transformer for medical image classification via token-wise perturbation
Yuexiang Li, Yawen Huang, Nanjun He, Kai Ma 0002, Yefeng Zheng 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2024 | LENAS: Learning-Based Neural Architecture Search and Ensemble for 3-D Radiotherapy Dose PredictionabstractRadiation therapy treatment planning requires balancing the delivery of the target dose while sparing normal tissues, making it a complex process. To streamline the planning process and enhance its quality, there is a growing demand for knowledge-based planning (KBP). Ensemble learning has shown impressive power in various deep learning tasks, and it has great potential to improve the performance of KBP. However, the effectiveness of ensemble learning heavily depends on the diversity and individual accuracy of the base learners. Moreover, the complexity of model ensembles is a major concern, as it requires maintaining multiple models during inference, leading to increased computational cost and storage overhead. In this study, we propose a novel learning-based ensemble approach named LENAS, which integrates neural architecture search with knowledge distillation for 3-D radiotherapy dose prediction. Our approach starts by exhaustively searching each block from an enormous architecture space to identify multiple architectures that exhibit promising performance and significant diversity. To mitigate the complexity introduced by the model ensemble, we adopt the teacher-student paradigm, leveraging the diverse outputs from multiple learned networks as supervisory signals to guide the training of the student network. Furthermore, to preserve high-level semantic information, we design a hybrid loss to optimize the student network, enabling it to recover the knowledge embedded within the teacher networks. The proposed method has been evaluated on two public datasets: 1) OpenKBP and 2) AIMIS. Extensive experimental results demonstrate the effectiveness of our method and its superior performance to the state-of-the-art methods. Code: github.com/hust-linyi/LENAS. Yi Lin 0009, Hao Chen 0011, Xin Yang 0008, Kai Ma 0002, Yefeng Zheng 0001, Kwang-Ting Cheng |
IEEE Trans. Cybern. | 5 |
| 2024 | Unsupervised Domain Adaptation for Medical Image Segmentation by Disentanglement Learning and Self-TrainingabstractUnsupervised domain adaption (UDA), which aims to enhance the segmentation performance of deep models on unlabeled data, has recently drawn much attention. In this paper, we propose a novel UDA method (namely DLaST) for medical image segmentation via disentanglement learning and self-training. Disentanglement learning factorizes an image into domain-invariant anatomy and domain-specific modality components. To make the best of disentanglement learning, we propose a novel shape constraint to boost the adaptation performance. The self-training strategy further adaptively improves the segmentation performance of the model for the target domain through adversarial learning and pseudo label, which implicitly facilitates feature alignment in the anatomy space. Experimental results demonstrate that the proposed method outperforms the state-of-the-art UDA methods for medical image segmentation on three public datasets, i.e., a cardiac dataset, an abdominal dataset and a brain dataset. The code will be released soon. Qingsong Xie, Yuexiang Li, Nanjun He, Munan Ning, Kai Ma 0002, Guoxing Wang, Yong Lian 0001, Yefeng Zheng 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2024 | Adversarial Medical Image With Hierarchical Feature HidingabstractDeep learning based methods for medical images can be easily compromised by adversarial examples (AEs), posing a great security flaw in clinical decision-making. It has been discovered that conventional adversarial attacks like PGD which optimize the classification logits, are easy to distinguish in the feature space, resulting in accurate reactive defenses. To better understand this phenomenon and reassess the reliability of the reactive defenses for medical AEs, we thoroughly investigate the characteristic of conventional medical AEs. Specifically, we first theoretically prove that conventional adversarial attacks change the outputs by continuously optimizing vulnerable features in a fixed direction, thereby leading to outlier representations in the feature space. Then, a stress test is conducted to reveal the vulnerability of medical images, by comparing with natural images. Interestingly, this vulnerability is a double-edged sword, which can be exploited to hide AEs. We then propose a simple-yet-effective hierarchical feature constraint (HFC), a novel add-on to conventional white-box attacks, which assists to hide the adversarial feature in the target feature distribution. The proposed method is evaluated on three medical datasets, both 2D and 3D, with different modalities. The experimental results demonstrate the superiority of HFC,i.e., it bypasses an array of state-of-the-art adversarial medical AE detectors more efficiently than competing adaptive attacks1, which reveals the deficiencies of medical reactive defense and allows to develop more robust defenses in future. Qingsong Yao, Zecheng He, Yuexiang Li, Yi Lin 0009, Kai Ma 0002, Yefeng Zheng 0001, Shaohua Kevin Zhou |
IEEE Trans. Medical Imaging | 5 |
| 2024 | Relational Experience Replay: Continual Learning by Adaptively Tuning Task-Wise RelationshipabstractContinual learning is a promising machine learning paradigm to learn new tasks while retaining previously learned knowledge over streaming training data. Till now,rehearsal-basedmethods, keeping a small part of data from old tasks as a memory buffer, have shown good performance in mitigating catastrophic forgetting for previously learned knowledge. However, most of these methods typically treat each new task equally, which may not adequately consider the relationship or similarity between old and new tasks. Furthermore, these methods commonly neglect sample importance in the continual training process and result in sub-optimal performance on certain tasks. To address this challenging problem, we propose Relational Experience Replay (RER), a bi-level learning framework, to adaptively tune task-wise relationships and sample importance within each task to achieve a better ‘stability’ and ‘plasticity’ trade-off. As such, the proposed method is capable of accumulating new knowledge while consolidating previously learned old knowledge during continual learning. Extensive experiments conducted on three benchmark image datasets (CIFAR-10, CIFAR-100, and Tiny ImageNet) and two text datasets (20News and DBpedia) show that the proposed method can consistently improve the performance of all baselines and surpass current state-of-the-art methods. Quanziang Wang, Renzhen Wang, Yuexiang Li, Dong Wei 0004, Hong Wang 0021, Kai Ma 0002, Yefeng Zheng 0001, Deyu Meng |
IEEE Trans. Multim. | 6 |
| 2023 | MIL-ViT: A multiple instance vision transformer for fundus image classification
Qi Bi, Xu Sun 0006, Kai Ma 0002, Cheng Bian, Munan Ning, Nanjun He, Yawen Huang, Yuexiang Li, Hanruo Liu, Yefeng Zheng 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2023 | Nuclei segmentation with point annotations from pathology images via self-supervised learning and co-training
Yi Lin 0009, Zhiyong Qu, Hao Chen 0011, Zhongke Gao, Yuexiang Li, Kai Ma 0002, Yefeng Zheng 0001, Kwang-Ting Cheng |
Medical Image Anal. | 7 |
| 2022 | Boost Supervised Pretraining for Visual Transfer Learning: Implications of Self-Supervised Contrastive Representation LearningabstractUnsupervised pretraining based on contrastive learning has made significant progress recently and showed comparable or even superior transfer learning performance to traditional supervised pretraining on various tasks. In this work, we first empirically investigate when and why unsupervised pretraining surpasses supervised counterparts for image classification tasks with a series of control experiments. Besides the commonly used accuracy, we further analyze the results qualitatively with the class activation maps and assess the learned representations quantitatively with the representation entropy and uniformity. Our core finding is that it is the amount of information effectively perceived by the learning model that is crucial to transfer learning, instead of absolute size of the dataset. Based on this finding, we propose Classification Activation Map guided contrastive (CAMtrast) learning which better utilizes the label supervsion to strengthen supervised pretraining, by making the networks perceive more information from the training images. CAMtrast is evaluated with three fundamental visual learning tasks: image recognition, object detection, and semantic segmentation, on various public datasets. Experimental results show that our CAMtrast effectively improves the performance of supervised pretraining, and that its performance is superior to both unsupervised counterparts and a recent related work which similarly attempted improving supervised pretraining. Jinghan Sun, Dong Wei 0004, Kai Ma 0002, Liansheng Wang 0002, Yefeng Zheng 0001 |
AAAI | 3 |
| 2022 | Dense Cross-Query-and-Support Attention Weighted Mask Aggregation for Few-Shot Segmentation
Xinyu Shi 0003, Dong Wei 0004, Yu Zhang 0185, Donghuan Lu, Munan Ning, Jiashun Chen, Kai Ma 0002, Yefeng Zheng 0001 |
ECCV (20) | 7 |
| 2022 | Cross-Domain Gated Learning for Domain Generalization
Dapeng Du, Jiawei Chen 0009, Yuexiang Li, Kai Ma 0002, Gangshan Wu, Yefeng Zheng 0001, Limin Wang 0002 |
Int. J. Comput. Vis. | 4 |
| 2022 | TW-GAN: Topology and width aware GAN for retinal artery/vein classification
Wenting Chen, Kai Ma 0002, Wei Ji 0011, Cheng Bian, Chunyan Chu, LinLin Shen, Yefeng Zheng 0001 |
Medical Image Anal. | 3 |
| 2022 | Mix-and-Interpolate: A Training Strategy to Deal With Source-Biased Medical DataabstractTill March 31st, 2021, the coronavirus disease 2019 (COVID-19) had reportedly infected more than 127 million people and caused over 2.5 million deaths worldwide. Timely diagnosis of COVID-19 is crucial for management of individual patients as well as containment of the highly contagious disease. Having realized the clinical value of non-contrast chest computed tomography (CT) for diagnosis of COVID-19, deep learning (DL) based automated methods have been proposed to aid the radiologists in reading the huge quantities of CT exams as a result of the pandemic. In this work, we address an overlooked problem for training deep convolutional neural networks for COVID-19 classification using real-world multi-source data, namely, the data source bias problem. The data source bias problem refers to the situation in which certain sources of data comprise only a single class of data, and training with such source-biased data may make the DL models learn to distinguish data sources instead of COVID-19. To overcome this problem, we propose MIx-aNd-Interpolate (MINI), a conceptually simple, easy-to-implement, efficient yet effective training strategy. The proposed MINI approach generates volumes of the absent class by combining the samples collected from different hospitals, which enlarges the sample space of the original source-biased dataset. Experimental results on a large collection of real patient data (1,221 COVID-19 and 1,520 negative CT images, and the latter consisting of 786 community acquired pneumonia and 734 non-pneumonia) from eight hospitals and health institutions show that: 1) MINI can improve COVID-19 classification performance upon the baseline (which does not deal with the source bias), and 2) MINI is superior to competing methods in terms of the extent of improvement. Yuexiang Li, Jiawei Chen 0009, Dong Wei 0004, Yanchun Zhu, Junfeng Xiong, Yadong Gang, Tianyi Qian, Kai Ma 0002, Yefeng Zheng 0001 |
IEEE J. Biomed. Health Informatics | 11 |
| 2022 | All-Around Real Label Supervision: Cyclic Prototype Consistency Learning for Semi-Supervised Medical Image SegmentationabstractSemi-supervised learning has substantially advanced medical image segmentation since it alleviates the heavy burden of acquiring the costly expert-examined annotations. Especially, the consistency-based approaches have attracted more attention for their superior performance, wherein the real labels are only utilized to supervise their paired images via supervised loss while the unlabeled images are exploited by enforcing the perturbation-based "unsupervised" consistency without explicit guidance from those real labels. However, intuitively, the expert-examined real labels contain more reliable supervision signals. Observing this, we ask an unexplored but interesting question: can we exploit the unlabeled data via explicit real label supervision for semi-supervised training? To this end, we discard the previous perturbation-based consistency but absorb the essence of non-parametric prototype learning. Based on the prototypical networks, we then propose a novel cyclic prototype consistency learning (CPCL) framework, which is constructed by a labeled-to-unlabeled (L2U) prototypical forward process and an unlabeled-to-labeled (U2L) backward process. Such two processes synergistically enhance the segmentation network by encouraging morediscriminative and compact features. In this way, our framework turns previous "unsupervised" consistency into new "supervised" consistency, obtaining the "all-around real label supervision" property of our method. Extensive experiments on brain tumor segmentation from MRI and kidney segmentation from CT images show that our CPCL can effectively exploit the unlabeled data and outperform other state-of-the-art semi-supervised medical image segmentation methods. Zhe Xu 0012, Yixin Wang 0003, Donghuan Lu, Lequan Yu, Jiangpeng Yan, Jie Luo 0003, Kai Ma 0002, Yefeng Zheng 0001, Raymond Kai-Yu Tong |
IEEE J. Biomed. Health Informatics | 7 |
| 2022 | Domain Adaptation Meets Zero-Shot Learning: An Annotation-Efficient Approach to Multi-Modality Medical Image SegmentationabstractDue to the lack of properly annotated medical data, exploring the generalization capability of the deep model is becoming a public concern. Zero-shot learning (ZSL) has emerged in recent years to equip the deep model with the ability to recognize unseen classes. However, existing studies mainly focus on natural images, which utilize linguistic models to extract auxiliary information for ZSL. It is impractical to apply the natural image ZSL solutions directly to medical images, since the medical terminology is very domain-specific, and it is not easy to acquire linguistic models for the medical terminology. In this work, we propose a new paradigm of ZSL specifically for medical images utilizing cross-modality information. We make three main contributions with the proposed paradigm. First, we extract the prior knowledge about the segmentation targets, called relation prototypes, from the prior model and then propose a cross-modality adaptation module to inherit the prototypes to the zero-shot model. Second, we propose a relation prototype awareness module to make the zero-shot model aware of information contained in the prototypes. Last but not least, we develop an inheritance attention module to recalibrate the relation prototypes to enhance the inheritance process. The proposed framework is evaluated on two public cross-modality datasets including a cardiac dataset and an abdominal dataset. Extensive experiments show that the proposed framework significantly outperforms the state of the arts. Cheng Bian, Chenglang Yuan, Kai Ma 0002, Dong Wei 0004, Yefeng Zheng 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2022 | Beyond Mutual Information: Generative Adversarial Network for Domain Adaptation Using Information Bottleneck ConstraintabstractMedical images from multicentres often suffer from the domain shift problem, which makes the deep learning models trained on one domain usually fail to generalize well to another. One of the potential solutions for the problem is the generative adversarial network (GAN), which has the capacity to translate images between different domains. Nevertheless, the existing GAN-based approaches are prone to fail at preserving image-objects in image-to-image (I2I) translation, which reduces their practicality on domain adaptation tasks. In this regard, a novel GAN (namely IB-GAN) is proposed to preserve image-objects during cross-domain I2I adaptation. Specifically, we integrate the information bottleneck constraint into the typical cycle-consistency-based GAN to discard the superfluous information (e.g., domain information) and maintain the consistency of disentangled content features for image-object preservation. The proposed IB-GAN is evaluated on three tasks-polyp segmentation using colonoscopic images, the segmentation of optic disc and cup in fundus images and the whole heart segmentation using multi-modal volumes. We show that the proposed IB-GAN can generate realistic translated images and remarkably boost the generalization of widely used segmentation networks (e.g., U-Net). Jiawei Chen 0009, Ziqi Zhang 0013, Xinpeng Xie, Yuexiang Li, Tao Xu 0026, Kai Ma 0002, Yefeng Zheng 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2022 | DICDNet: Deep Interpretable Convolutional Dictionary Network for Metal Artifact Reduction in CT ImagesabstractComputed tomography (CT) images are often impaired by unfavorable artifacts caused by metallic implants within patients, which would adversely affect the subsequent clinical diagnosis and treatment. Although the existing deep-learning-based approaches have achieved promising success on metal artifact reduction (MAR) for CT images, most of them treated the task as a general image restoration problem and utilized off-the-shelf network modules for image quality enhancement. Hence, such frameworks always suffer from lack of sufficient model interpretability for the specific task. Besides, the existing MAR techniques largely neglect the intrinsic prior knowledge underlying metal-corrupted CT images which is beneficial for the MAR performance improvement. In this paper, we specifically propose a deep interpretable convolutional dictionary network (DICDNet) for the MAR task. Particularly, we first explore that the metal artifacts always present non-local streaking and star-shape patterns in CT images. Based on such observations, a convolutional dictionary model is deployed to encode the metal artifacts. To solve the model, we propose a novel optimization algorithm based on the proximal gradient technique. With only simple operators, the iterative steps of the proposed algorithm can be easily unfolded into corresponding network modules with specific physical meanings. Comprehensive experiments on synthesized and clinical datasets substantiate the effectiveness of the proposed DICDNet as well as its superior interpretability, compared to current state-of-the-art MAR methods. Code is available at https://github.com/hongwang01/DICDNet. Hong Wang 0021, Yuexiang Li, Nanjun He, Kai Ma 0002, Deyu Meng, Yefeng Zheng 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2022 | Anti-Interference From Noisy Labels: Mean-Teacher-Assisted Confident Learning for Medical Image SegmentationabstractManually segmenting medical images is expertise-demanding, time-consuming and laborious. Acquiring massive high-quality labeled data from experts is often infeasible. Unfortunately, without sufficient high-quality pixel-level labels, the usual data-driven learning-based segmentation methods often struggle with deficient training. As a result, we are often forced to collect additional labeled data from multiple sources with varying label qualities. However, directly introducing additional data with low-quality noisy labels may mislead the network training and undesirably offset the efficacy provided by those high-quality labels. To address this issue, we propose a Mean-Teacher-assisted Confident Learning (MTCL) framework constructed by a teacher-student architecture and a label self-denoising process to robustly learn segmentation from a small set of high-quality labeled data and plentiful low-quality noisy labeled data. Particularly, such a synergistic framework is capable of simultaneously and robustly exploiting (i) the additional dark knowledge inside the images of low-quality labeled set via perturbation-based unsupervised consistency, and (ii) the productive information of their low-quality noisy labels via explicit label refinement. Comprehensive experiments on left atrium segmentation with simulated noisy labels and hepatic and retinal vessel segmentation with real-world noisy labels demonstrate the superior segmentation performance of our approach as well as its effectiveness on label denoising. Zhe Xu 0012, Donghuan Lu, Jie Luo 0003, Yixin Wang 0003, Jiangpeng Yan, Kai Ma 0002, Yefeng Zheng 0001, Raymond Kai-Yu Tong |
IEEE Trans. Medical Imaging | 6 |
| 2022 | Multiattention Adaptation Network for Motor Imagery RecognitionabstractBrain–computer interface (BCI) based on motor imagery electroencephalogram (EEG) has been widely used in various applications. Despite the previous efforts, the remained major challenges are effective feature extraction and the time-consuming calibration procedure. To address these issues, a novel multiattention adaptation network integrating the multiple attention mechanism and transfer learning is proposed to classify the EEG signals. First, the multiattention layer is introduced to automatically capture the dominant brain regions relevant to mental tasks without incorporating any prior knowledge about the physiology. Then, a multiattention convolutional neural network is employed to extract deep representation from raw EEG signals. Especially, a domain discriminator is applied to deep representation to reduce the differences between sessions for target subjects. The extensive experiments are conducted on three public EEG datasets (Dataset IIa and IIb of BCI Competition IV, and High Gamma dataset), achieving the competitive performance with average classification accuracy of 81.48%, 82.54%, and 93.97%, respectively. All the results outperform the state-of-the-art algorithms demonstrate the effectiveness and robustness of the proposed method. Importantly, we confirm that it is easier and more appropriate to transfer the information from local brain regions than from the whole brain. This enhances the transfer ability of deep features and, hence, it improves the performance of BCI systems. Peiyin Chen, Zhongke Gao, Miaomiao Yin, Jialing Wu, Kai Ma 0002, Celso Grebogi |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2021 | Alternative Baselines for Low-Shot 3D Medical Image Segmentation - An Atlas PerspectiveabstractLow-shot (one/few-shot) segmentation has attracted increasing attention as it works well with limited annotation. State-of-the-art low-shot segmentation methods on natural images usually focus on implicit representation learning for each novel class, such as learning prototypes, deriving guidance features via masked average pooling, and segmenting using cosine similarity in feature space. We argue that low-shot segmentation on medical images should step further to explicitly learn dense correspondences between images to utilize the anatomical similarity. The core ideas are inspired by the classical practice of multi-atlas segmentation, where the indispensable parts of atlas-based segmentation, i.e., registration, label propagation, and label fusion are unified into a single framework in our work. Specifically, we propose two alternative baselines, i.e., the Siamese-Baseline and Individual-Difference-Aware Baseline, where the former is targeted at anatomically stable structures (such as brain tissues), and the latter possesses a strong generalization ability to organs suffering large morphological variations (such as abdominal organs). In summary, this work sets up a benchmark for low-shot 3D medical image segmentation and sheds light on further understanding of atlas-based few-shot segmentation. Shilei Cao 0001, Dong Wei 0004, Kai Ma 0002, Liansheng Wang 0002, Deyu Meng, Yefeng Zheng 0001 |
AAAI | 5 |
| 2021 | Alleviating Noisy-label Effects in Image Classification via Probability Transition Matrix
Ziqi Zhang 0013, Yuexiang Li, Hongxin Wei, Kai Ma 0002, Tao Xu 0026, Yefeng Zheng 0001 |
BMVC | 4 |
| 2021 | Calibrated RGB-D Salient Object DetectionabstractComplex backgrounds and similar appearances between objects and their surroundings are generally recognized as challenging scenarios in Salient Object Detection (SOD). This naturally leads to the incorporation of depth information in addition to the conventional RGB image as input, known as RGB-D SOD or depth-aware SOD. Meanwhile, this emerging line of research has been considerably hindered by the noise and ambiguity that prevail in raw depth images. To address the aforementioned issues, we propose a Depth Calibration and Fusion (DCF) framework that contains two novel components: 1) a learning strategy to calibrate the latent bias in the original depth maps towards boosting the SOD performance; 2) a simple yet effective cross reference module to fuse features from both RGB and depth modalities. Extensive empirical experiments demonstrate that the proposed approach achieves superior performance against 27 state-of-the-art methods. Moreover, our depth calibration strategy alone can work as a preprocessing step; empirically it results in noticeable improvements when being applied to existing cutting-edge RGB-D SOD models. Source code is available at https://github.com/jiwei0921/DCF. Wei Ji 0011, Miao Zhang 0004, Yongri Piao, Shunyu Yao 0004, Qi Bi, Kai Ma 0002, Yefeng Zheng 0001, Huchuan Lu, Li Cheng 0001 |
CVPR | 8 |
| 2021 | Learning Calibrated Medical Image Segmentation via Multi-Rater Agreement ModelingabstractIn medical image analysis, it is typical to collect multiple annotations, each from a different clinical expert or rater, in the expectation that possible diagnostic errors could be mitigated. Meanwhile, from the computer vision practitioner viewpoint, it has been a common practice to adopt the ground-truth labels obtained via either the majority-vote or simply one annotation from a preferred rater. This process, however, tends to overlook the rich information of agreement or disagreement ingrained in the raw multi-rater annotations. To address this issue, we propose to explicitly model the multi-rater (dis-)agreement, dubbed MRNet, which has two main contributions. First, an expertise-aware inferring module or EIM is devised to embed the expertise level of individual raters as prior knowledge, to form high-level semantic features. Second, our approach is capable of reconstructing multi-rater gradings from coarse predictions, with the multi-rater (dis-)agreement cues being further exploited to improve the segmentation performance. To our knowledge, our work is the first in producing calibrated predictions under different expertise levels for medical image segmentation. Extensive empirical experiments are conducted across five medical segmentation tasks of diverse imaging modalities. In these experiments, superior performance of our MRNet is observed comparing to the state-of-the-arts, indicating the effectiveness and applicability of our MRNet toward a wide range of medical segmentation tasks. Source code is publicly available. Wei Ji 0011, Kai Ma 0002, Cheng Bian, Qi Bi, Hanruo Liu, Li Cheng 0001, Yefeng Zheng 0001 |
CVPR | 4 |
| 2021 | Multi-Anchor Active Domain Adaptation for Semantic SegmentationabstractUnsupervised domain adaption has proven to be an effective approach for alleviating the intensive workload of manual annotation by aligning the synthetic source-domain data and the real-world target-domain samples. Unfortunately, mapping the target-domain distribution to the source-domain unconditionally may distort the essential structural information of the target-domain data. To this end, we firstly propose to introduce a novel multi-anchor based active learning strategy to assist domain adaptation regarding the semantic segmentation task. By innovatively adopting multiple anchors instead of a single centroid, the source domain can be better characterized as a multimodal distribution, thus more representative and complimentary samples are selected from the target domain. With little workload to manually annotate these active samples, the distortion of the target-domain distribution can be effectively alleviated, resulting in a large performance gain. The multi-anchor strategy is additionally employed to model the target-distribution. By regularizing the latent representation of the target samples compact around multiple anchors through a novel soft alignment loss, more precise segmentation can be achieved. Extensive experiments are conducted on public datasets to demonstrate that the proposed approach outperforms state-of-the-art methods significantly, along with thorough ablation study to verify the effectiveness of each component. The code will be released soon at https://github.com/munanning/MADA. Munan Ning, Donghuan Lu, Dong Wei 0004, Cheng Bian, Chenglang Yuan, Kai Ma 0002, Yefeng Zheng 0001 |
ICCV | 7 |
| 2021 | Stabilized Medical Image Attacks
Gege Qi, Lijun Gong, Yibing Song, Kai Ma 0002, Yefeng Zheng 0001 |
ICLR | 4 |
| 2021 | Local-Global Dual Perception Based Deep Multiple Instance Learning for Retinal Disease Classification
Qi Bi, Wei Ji 0011, Cheng Bian, Lijun Gong, Hanruo Liu, Kai Ma 0002, Yefeng Zheng 0001 |
MICCAI (8) | 7 |
| 2021 | Deep Reinforcement Exemplar Learning for Annotation Refinement
Yuexiang Li, Nanjun He, Sixiang Peng, Kai Ma 0002, Yefeng Zheng 0001 |
MICCAI (8) | 4 |
| 2021 | Triplet-Branch Network with Prior-Knowledge Embedding for Fatigue Fracture Grading
Yuexiang Li, Yi Lin 0009, Dong Wei 0004, Qirui Zhang 0004, Kai Ma 0002, Guangming Lu 0001, Yefeng Zheng 0001 |
MICCAI (5) | 7 |
| 2021 | Seg4Reg+: Consistency Learning Between Spine Segmentation and Cobb Angle Regression
Yi Lin 0009, Luyan Liu, Kai Ma 0002, Yefeng Zheng 0001 |
MICCAI (5) | 3 |
| 2021 | Simultaneous Alignment and Surface Regression Using Hybrid 2D-3D Networks for 3D Coherent Layer Segmentation of Retina OCT Images
Dong Wei 0004, Donghuan Lu, Yuexiang Li, Kai Ma 0002, Liansheng Wang 0002, Yefeng Zheng 0001 |
MICCAI (8) | 5 |
| 2021 | Unsupervised Representation Learning Meets Pseudo-Label Supervised Self-Distillation: A New Approach to Rare Disease Classification
Jinghan Sun, Dong Wei 0004, Kai Ma 0002, Liansheng Wang 0002, Yefeng Zheng 0001 |
MICCAI (5) | 3 |
| 2021 | InDuDoNet: An Interpretable Dual Domain Network for CT Metal Artifact Reduction
Hong Wang 0021, Yuexiang Li, Haimiao Zhang, Jiawei Chen 0009, Kai Ma 0002, Deyu Meng, Yefeng Zheng 0001 |
MICCAI (6) | 5 |
| 2021 | Training Automatic View Planner for Cardiac MR Imaging via Self-supervision by Spatial Relationship Between Views
Dong Wei 0004, Kai Ma 0002, Yefeng Zheng 0001 |
MICCAI (6) | 2 |
| 2021 | Noisy Labels are Treasure: Mean-Teacher-Assisted Confident Learning for Hepatic Vessel Segmentation
Zhe Xu 0012, Donghuan Lu, Yixin Wang 0003, Jie Luo 0003, Jayender Jagadeesan, Kai Ma 0002, Yefeng Zheng 0001, Xiu Li 0001 |
MICCAI (1) | 6 |
| 2021 | A Hierarchical Feature Constraint to Camouflage Medical Adversarial Attacks
Qingsong Yao, Zecheng He, Yi Lin 0009, Kai Ma 0002, Yefeng Zheng 0001, Shaohua Kevin Zhou |
MICCAI (3) | 4 |
| 2021 | MIL-VT: Multiple Instance Learning Enhanced Vision Transformer for Fundus Image Classification
Kai Ma 0002, Qi Bi, Cheng Bian, Munan Ning, Nanjun He, Yuexiang Li, Hanruo Liu, Yefeng Zheng 0001 |
MICCAI (8) | 2 |
| 2021 | GRAND: A large-scale dataset and benchmark for cervical intraepithelial Neoplasia grading with fine-grained lesion description
Yuexiang Li, Zhi-Hua Liu, Peng Xue 0001, Jiawei Chen 0009, Kai Ma 0002, Tianyi Qian, Yefeng Zheng 0001, Youlin Qiao |
Medical Image Anal. | 5 |
| 2021 | S-CUDA: Self-cleansing unsupervised domain adaptation for medical image segmentationabstractMedical image segmentation tasks hitherto have achieved excellent progresses with large-scale datasets, which empowers us to train potent deep convolutional neural networks (DCNNs). However, labeling such large-scale datasets is laborious and error-prone, which leads the noisy (or incorrect) labels to be an ubiquitous problem in the real-world scenarios. In addition, data collected from different sites usually exhibit significant data distribution shift (or domain shift). As a result, noisy label and domain shift become two common problems in medical imaging application scenarios, especially in medical image segmentation, which degrade the performance of deep learning models significantly. In this paper, we identify a novel problem hidden in medical image segmentation, which is unsupervised domain adaptation on noisy labeled data, and propose a novel algorithm named "Self-Cleansing Unsupervised Domain Adaptation" (S-CDUA) to address such issue. S-CUDA sets up a realistic scenario to solve the above problems simultaneously where training data (i.e., source domain) not only shows domain shift w.r.t. unsupervised test data (i.e., target domain) but also contains noisy labels. The key idea of S-CUDA is to learn noise-excluding and domain invariant knowledge from noisy supervised data, which will be applied on the highly corrupted data for label cleansing and further data-recycling, as well as on the test data with domain shift for supervised propagation. To this end, we propose a novel framework leveraging noisy-label learning and domain adaptation techniques to cleanse the noisy labels and learn from trustable clean samples, thus enabling robust adaptation and prediction on the target domain. Specifically, we train two peer adversarial networks to identify high-confidence clean data and exchange them in companions to eliminate the error accumulation problem and narrow the domain gap simultaneously. In the meantime, the high-confidence noisy data are detected and cleansed in order to reuse the contaminated training data. Therefore, our proposed method can not only cleanse the noisy labels in the training set but also take full advantage of the existing noisy data to update the parameters of the network. For evaluation, we conduct experiments on two popular datasets (REFUGE and Drishti-GS) for optic disc (OD) and optic cup (OC) segmentation, and on another public multi-vendor dataset for spinal cord gray matter (SCGM) segmentation. Experimental results show that our proposed method can cleanse noisy labels efficiently and obtain a model with better generalization performance at the same time, which outperforms previous state-of-the-art methods by large margin. Our code can be found at https://github.com/zzdxjtu/S-cuda. Luyan Liu, Shuai Li 0001, Kai Ma 0002, Yefeng Zheng 0001 |
Medical Image Anal. | 4 |
| 2021 | Pairwise learning for medical image segmentation
Renzhen Wang, Shilei Cao 0001, Kai Ma 0002, Yefeng Zheng 0001, Deyu Meng |
Medical Image Anal. | 3 |
| 2021 | A Unified Framework for Generalized Low-Shot Medical Image Segmentation With Scarce DataabstractMedical image segmentation has achieved remarkable advancements using deep neural networks (DNNs). However, DNNs often need big amounts of data and annotations for training, both of which can be difficult and costly to obtain. In this work, we propose a unified framework for generalized low-shot (one- and few-shot) medical image segmentation based on distance metric learning (DML). Unlike most existing methods which only deal with the lack of annotations while assuming abundance of data, our framework works with extreme scarcity of both, which is ideal for rare diseases. Via DML, the framework learns a multimodal mixture representation for each category, and performs dense predictions based on cosine distances between the pixels' deep embeddings and the category representations. The multimodal representations effectively utilize the inter-subject similarities and intraclass variations to overcome overfitting due to extremely limited data. In addition, we propose adaptive mixing coefficients for the multimodal mixture distributions to adaptively emphasize the modes better suited to the current input. The representations are implicitly embedded as weights of the fc layer, such that the cosine distances can be computed efficiently via forward propagation. In our experiments on brain MRI and abdominal CT datasets, the proposed framework achieves superior performances for low-shot segmentation towards standard DNN-based (3D U-Net) and classical registration-based (ANTs) methods, e.g., achieving mean Dice coefficients of 81%/69% for brain tissue/abdominal multi-organ segmentation using a single training sample, as compared to 52%/31% and 72%/35% by the U-Net and ANTs, respectively. Hengji Cui, Dong Wei 0004, Kai Ma 0002, Shi Gu, Yefeng Zheng 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2021 | Anomaly Detection for Medical Images Using Self-Supervised and Translation-Consistent FeaturesabstractAs the labeled anomalous medical images are usually difficult to acquire, especially for rare diseases, the deep learning based methods, which heavily rely on the large amount of labeled data, cannot yield a satisfactory performance. Compared to the anomalous data, the normal images without the need of lesion annotation are much easier to collect. In this paper, we propose an anomaly detection framework, namely [Formula: see text], extracting [Formula: see text]elf-supervised and tr [Formula: see text]ns [Formula: see text]ation-consistent features for [Formula: see text]nomaly [Formula: see text]etection. The proposed SALAD is a reconstruction-based method, which learns the manifold of normal data through an encode-and-reconstruct translation between image and latent spaces. In particular, two constraints (i.e., structure similarity loss and center constraint loss) are proposed to regulate the cross-space (i.e., image and feature) translation, which enforce the model to learn translation-consistent and representative features from the normal data. Furthermore, a self-supervised learning module is engaged into our framework to further boost the anomaly detection accuracy by deeply exploiting useful information from the raw normal data. An anomaly score, as a measure to separate the anomalous data from the healthy ones, is constructed based on the learned self-supervised-and-translation-consistent features. Extensive experiments are conducted on optical coherence tomography (OCT) and chest X-ray datasets. The experimental results demonstrate the effectiveness of our approach. He Zhao 0002, Yuexiang Li, Nanjun He, Kai Ma 0002, Leyuan Fang, Huiqi Li, Yefeng Zheng 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2021 | Deep Representation-Based Domain Adaptation for Nonstationary EEG ClassificationabstractIn the context of motor imagery, electroencephalography (EEG) data vary from subject to subject such that the performance of a classifier trained on data of multiple subjects from a specific domain typically degrades when applied to a different subject. While collecting enough samples from each subject would address this issue, it is often too time-consuming and impractical. To tackle this problem, we propose a novel end-to-end deep domain adaptation method to improve the classification performance on a single subject (target domain) by taking the useful information from multiple subjects (source domain) into consideration. Especially, the proposed method jointly optimizes three modules, including a feature extractor, a classifier, and a domain discriminator. The feature extractor learns the discriminative latent features by mapping the raw EEG signals into a deep representation space. A center loss is further employed to constrain an invariant feature space and reduce the intrasubject nonstationarity. Furthermore, the domain discriminator matches the feature distribution shift between source and target domains by an adversarial learning strategy. Finally, based on the consistent deep features from both domains, the classifier is able to leverage the information from the source domain and accurately predict the label in the target domain at the test time. To evaluate our method, we have conducted extensive experiments on two real public EEG data sets, data set IIa, and data set IIb of brain-computer interface (BCI) Competition IV. The experimental results validate the efficacy of our method. Therefore, our method is promising to reduce the calibration time for the use of BCI and promote the development of BCI. He Zhao 0002, Qingqing Zheng, Kai Ma 0002, Huiqi Li, Yefeng Zheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Classification of EEG Signals on VEP-Based BCI Systems With Broad LearningabstractBrain–computer interface (BCI) systems based on electroencephalography (EEG) signals have been extensively used in medical practice. To enhance the BCI performance, improving the classification accuracy of EEG signals is the key, which has always been the focus of research and development. In this article, a novel method integrating complex network and broad learning system (BLS) is proposed for visual evoked potential (VEP)-based BCI research. First, systematic VEP-based brain experiments are conducted for obtaining EEG signals, including steady-state VEP (SSVEP) and steady-state motion VEP (SSMVEP). Then, limited penetrable visibility graph (LPVG) and its degree sequence are employed to implement the preliminary feature extraction. All these features are finally fed into a BLS to study and classify the SSVEP and SSMVEP signals, respectively. The classification results show that our LPVG-based BLS can effectively classify VEP-based EEG signals, with average classification accuracy 96.22% for SSVEP and 74.54% for SSMVEP. These results are significantly better than other comparison methods as well as traditional CCA-based methods. All these open up new venues for studying EEG-based BCI systems via the fusion of network science and BLS. Zhongke Gao, Wei-Dong Dang, Mingxu Liu, Wei Guo 0026, Kai Ma 0002, Guanrong Chen |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2020 | Generative Adversarial Networks for Video-to-Video Domain AdaptationabstractEndoscopic videos from multicentres often have different imaging conditions, e.g., color and illumination, which make the models trained on one domain usually fail to generalize well to another. Domain adaptation is one of the potential solutions to address the problem. However, few of existing works focused on the translation of video-based data. In this work, we propose a novel generative adversarial network (GAN), namely VideoGAN, to transfer the video-based data across different domains. As the frames of a video may have similar content and imaging conditions, the proposed VideoGAN has an X-shape generator to preserve the intra-video consistency during translation. Furthermore, a loss function, namely color histogram loss, is proposed to tune the color distribution of each translated frame. Two colonoscopic datasets from different centres, i.e., CVC-Clinic and ETIS-Larib, are adopted to evaluate the performance of domain adaptation of our VideoGAN. Experimental results demonstrate that the adapted colonoscopic video generated by our VideoGAN can significantly boost the segmentation accuracy, i.e., an improvement of 5%, of colorectal polyps on multicentre datasets. As our VideoGAN is a general network architecture, we also evaluate its performance with the CamVid driving video dataset on the cloudy-to-sunny translation task. Comprehensive experiments show that the domain gap could be substantially narrowed down by our VideoGAN. Jiawei Chen 0009, Yuexiang Li, Kai Ma 0002, Yefeng Zheng 0001 |
AAAI | 3 |
| 2020 | LT-Net: Label Transfer by Learning Reversible Voxel-Wise Correspondence for One-Shot Medical Image SegmentationabstractWe introduce a one-shot segmentation method to alleviate the burden of manual annotation for medical images. The main idea is to treat one-shot segmentation as a classical atlas-based segmentation problem, where voxel-wise correspondence from the atlas to the unlabelled data is learned. Subsequently, segmentation label of the atlas can be transferred to the unlabelled data with the learned correspondence. However, since ground truth correspondence between images is usually unavailable, the learning system must be well-supervised to avoid mode collapse and convergence failure. To overcome this difficulty, we resort to the forward-backward consistency, which is widely used in correspondence problems, and additionally learn the backward correspondences from the warped atlases back to the original atlas. This cycle-correspondence learning design enables a variety of extra, cycle-consistency-based supervision signals to make the training process stable, while also boost the performance. We demonstrate the superiority of our method over both deep learning-based one-shot segmentation methods and a classical multi-atlas segmentation method via thorough experiments. Shilei Cao 0001, Dong Wei 0004, Renzhen Wang, Kai Ma 0002, Liansheng Wang 0002, Deyu Meng, Yefeng Zheng 0001 |
CVPR | 5 |
| 2020 | Dual Adversarial Network for Deep Active Learning
Shuo Wang 0008, Yuexiang Li, Kai Ma 0002, Ruhui Ma, Haibing Guan, Yefeng Zheng 0001 |
ECCV (24) | 3 |
| 2020 | Self-Supervised CycleGAN for Object-Preserving Image-to-Image Domain Adaptation
Xinpeng Xie, Jiawei Chen 0009, Yuexiang Li, LinLin Shen, Kai Ma 0002, Yefeng Zheng 0001 |
ECCV (20) | 5 |
| 2020 | Deep Image Clustering with Category-Style Representation
Donghuan Lu, Kai Ma 0002, Yu Zhang 0185, Yefeng Zheng 0001 |
ECCV (14) | 3 |
| 2020 | Cross-denoising Network against Corrupted Labels in Medical Image Segmentation with Domain ShiftabstractDeep convolutional neural networks (DCNNs) have contributed many breakthroughs in segmentation tasks, especially in the field of medical imaging. However, domain shift and corrupted annotations, which are two common problems in medical imaging, dramatically degrade the performance of DCNNs in practice. In this paper, we propose a novel robust cross-denoising framework using two peer networks to address domain shift and corrupted label problems with a peer-review strategy. Specifically, each network performs as a mentor, mutually supervised to learn from reliable samples selected by the peer network to combat with corrupted labels. In addition, a noise-tolerant loss is proposed to encourage the network to capture the key location and filter the discrepancy under various noise-contaminant labels. To further reduce the accumulated error, we introduce a class-imbalanced cross learning using most confident predictions at class-level. Experimental results on REFUGE and Drishti-GS datasets for optic disc (OD) and optic cup (OC) segmentation demonstrate the superior performance of our proposed approach to the state-of-the-art methods. Qinming Zhang, Luyan Liu, Kai Ma 0002, Cheng Zhuo, Yefeng Zheng 0001 |
IJCAI | 3 |
| 2020 | TR-GAN: Topology Ranking GAN with Triplet Loss for Retinal Artery/Vein Classification
Wenting Chen, Kai Ma 0002, Cheng Bian, Chunyan Chu, LinLin Shen, Yefeng Zheng 0001 |
MICCAI (5) | 4 |
| 2020 | Distractor-Aware Neuron Intrinsic Learning for Generic 2D Medical Image Classifications
Lijun Gong, Kai Ma 0002, Yefeng Zheng 0001 |
MICCAI (2) | 2 |
| 2020 | Self-Loop Uncertainty: A Novel Pseudo-Label for Semi-supervised Medical Image Segmentation
Yuexiang Li, Jiawei Chen 0009, Xinpeng Xie, Kai Ma 0002, Yefeng Zheng 0001 |
MICCAI (1) | 4 |
| 2020 | Superpixel-Guided Label Softening for Medical Image Segmentation
Dong Wei 0004, Shilei Cao 0001, Kai Ma 0002, Liansheng Wang 0002, Yefeng Zheng 0001 |
MICCAI (4) | 4 |
| 2020 | GREEN: a Graph REsidual rE-ranking Network for Grading Diabetic Retinopathy
Shaoteng Liu, Lijun Gong, Kai Ma 0002, Yefeng Zheng 0001 |
MICCAI (5) | 3 |
| 2020 | Learning Crisp Edge Detector Using Logical Refinement Network
Luyan Liu, Kai Ma 0002, Yefeng Zheng 0001 |
MICCAI (4) | 2 |
| 2020 | A Macro-Micro Weakly-Supervised Framework for AS-OCT Tissue Segmentation
Munan Ning, Cheng Bian, Donghuan Lu, Chenglang Yuan, Yang Guo 0003, Kai Ma 0002, Yefeng Zheng 0001 |
MICCAI (5) | 9 |
| 2020 | Revisiting Rubik's Cube: Self-supervised Learning with Volume-Wise Transformation for 3D Medical Image Segmentation
Xing Tao, Yuexiang Li, Wenhui Zhou 0001, Kai Ma 0002, Yefeng Zheng 0001 |
MICCAI (4) | 4 |
| 2020 | Learning and Exploiting Interclass Visual Correlations for Medical Image Classification
Dong Wei 0004, Shilei Cao 0001, Kai Ma 0002, Yefeng Zheng 0001 |
MICCAI (1) | 3 |
| 2020 | Leveraging Undiagnosed Data for Glaucoma Classification with Teacher-Student Learning
Wenting Chen, Kai Ma 0002, Hanruo Liu, Xiaoguang Di, Yefeng Zheng 0001 |
MICCAI (1) | 4 |
| 2020 | MI2GAN: Generative Adversarial Network for Medical Image Domain Adaptation Using Mutual Information Constraint
Xinpeng Xie, Jiawei Chen 0009, Yuexiang Li, LinLin Shen, Kai Ma 0002, Yefeng Zheng 0001 |
MICCAI (2) | 5 |
| 2020 | Instance-Aware Self-supervised Learning for Nuclei Segmentation
Xinpeng Xie, Jiawei Chen 0009, Yuexiang Li, LinLin Shen, Kai Ma 0002, Yefeng Zheng 0001 |
MICCAI (5) | 5 |
| 2020 | Difficulty-Aware Glaucoma Classification with Multi-rater Consensus Modeling
Kai Ma 0002, Cheng Bian, Chunyan Chu, Hanruo Liu, Yefeng Zheng 0001 |
MICCAI (1) | 3 |
| 2020 | Comparing to Learn: Surpassing ImageNet Pretraining on Radiographs by Comparing Image Representations
Cheng Bian, Kai Ma 0002, Yefeng Zheng 0001 |
MICCAI (1) | 5 |
| 2020 | Uncertainty-aware domain alignment for anatomical structure segmentation
Cheng Bian, Chenglang Yuan, Jiexiang Wang, Meng Li 0090, Xin Yang 0009, Kai Ma 0002, Yefeng Zheng 0001 |
Medical Image Anal. | 7 |
| 2020 | Rubik's Cube+: A self-supervised feature learning framework for 3D medical image analysis
Jiuwen Zhu, Yuexiang Li, Kai Ma 0002, Shaohua Kevin Zhou, Yefeng Zheng 0001 |
Medical Image Anal. | 4 |
| 2020 | Efficient and Effective Training of COVID-19 Classification Networks With Self-Supervised Dual-Track Learning to RankabstractCoronavirus Disease 2019 (COVID-19) has rapidly spread worldwide since first reported. Timely diagnosis of COVID-19 is crucial both for disease control and patient care. Non-contrast thoracic computed tomography (CT) has been identified as an effective tool for the diagnosis, yet the disease outbreak has placed tremendous pressure on radiologists for reading the exams and may potentially lead to fatigue-related mis-diagnosis. Reliable automatic classification algorithms can be really helpful; however, they usually require a considerable number of COVID-19 cases for training, which is difficult to acquire in a timely manner. Meanwhile, how to effectively utilize the existing archive of non-COVID-19 data (the negative samples) in the presence of severe class imbalance is another challenge. In addition, the sudden disease outbreak necessitates fast algorithm development. In this work, we propose a novel approach for effective and efficient training of COVID-19 classification networks using a small number of COVID-19 CT exams and an archive of negative samples. Concretely, a novel self-supervised learning method is proposed to extract features from the COVID-19 and negative samples. Then, two kinds of soft-labels ('difficulty' and 'diversity') are generated for the negative samples by computing the earth mover's distances between the features of the negative and COVID-19 samples, from which data 'values' of the negative samples can be assessed. A pre-set number of negative samples are selected accordingly and fed to the neural network for training. Experimental results show that our approach can achieve superior performance using about half of the negative samples, substantially reducing model training time. Yuexiang Li, Dong Wei 0004, Jiawei Chen 0009, Shilei Cao 0001, Yanchun Zhu, Lan Lan 0002, Tianyi Qian, Kai Ma 0002, Yefeng Zheng 0001 |
IEEE J. Biomed. Health Informatics | 11 |
| 2020 | Computer-Aided Cervical Cancer Diagnosis Using Time-Lapsed Colposcopic ImagesabstractCervical cancer causes the fourth most cancer-related deaths of women worldwide. Early detection of cervical intraepithelial neoplasia (CIN) can significantly increase the survival rate of patients. In this paper, we propose a deep learning framework for the accurate identification of LSIL+ (including CIN and cervical cancer) using time-lapsed colposcopic images. The proposed framework involves two main components, i.e., key-frame feature encoding networks and feature fusion network. The features of the original (pre-acetic-acid) image and the colposcopic images captured at around 60s, 90s, 120s and 150s during the acetic acid test are encoded by the feature encoding networks. Several fusion approaches are compared, all of which outperform the existing automated cervical cancer diagnosis systems using a single time slot. A graph convolutional network with edge features (E-GCN) is found to be the most suitable fusion approach in our study, due to its excellent explainability consistent with the clinical practice. A large-scale dataset, containing time-lapsed colposcopic images from 7,668 patients, is collected from the collaborative hospital to train and validate our deep learning framework. Colposcopists are invited to compete with our computer-aided diagnosis system. The proposed deep learning framework achieves a classification accuracy of 78.33%-comparable to that of an in-service colposcopist-which demonstrates its potential to provide assistance in the realistic clinical scenario. Yuexiang Li, Jiawei Chen 0009, Peng Xue 0001, Jia Chang, Chunyan Chu, Kai Ma 0002, Yefeng Zheng 0001, Youlin Qiao |
IEEE Trans. Medical Imaging | 7 |
| 2020 | Conquering Data Variations in Resolution: A Slice-Aware Multi-Branch Decoder NetworkabstractFully convolutional neural networks have made promising progress in joint liver and liver tumor segmentation. Instead of following the debates over 2D versus 3D networks (for example, pursuing the balance between large-scale 2D pretraining and 3D context), in this paper, we novelly identify the wide variation in the ratio between intra- and inter-slice resolutions as a crucial obstacle to the performance. To tackle the mismatch between the intra- and inter-slice information, we propose a slice-aware 2.5D network that emphasizes extracting discriminative features utilizing not only in-plane semantics but also out-of-plane coherence for each separate slice. Specifically, we present a slice-wise multi-input multi-output architecture to instantiate such a design paradigm, which contains a Multi-Branch Decoder (MD) with a Slice-centric Attention Block (SAB) for learning slice-specific features and a Densely Connected Dice (DCD) loss to regularize the inter-slice predictions to be coherent and continuous. Based on the aforementioned innovations, we achieve state-of-the-art results on the MICCAI 2017 Liver Tumor Segmentation (LiTS) dataset. Besides, we also test our model on the ISBI 2019 Segmentation of THoracic Organs at Risk (SegTHOR) dataset, and the result proves the robustness and generalizability of the proposed method in other segmentation tasks. Shilei Cao 0001, Zhizhong Chai, Dong Wei 0004, Kai Ma 0002, Liansheng Wang 0002, Yefeng Zheng 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2019 | X2CT-GAN: Reconstructing CT From Biplanar X-Rays With Generative Adversarial NetworksabstractComputed tomography (CT) can provide a 3D view of the patient's internal organs, facilitating disease diagnosis, but it incurs more radiation dose to a patient and a CT scanner is much more cost prohibitive than an X-ray machine too. Traditional CT reconstruction methods require hundreds of X-ray projections through a full rotational scan of the body, which cannot be performed on a typical X-ray machine. In this work, we propose to reconstruct CT from two orthogonal X-rays using the generative adversarial network (GAN) framework. A specially designed generator network is exploited to increase data dimension from 2D (X-rays) to 3D (CT), which is not addressed in previous research of GAN. A novel feature fusion method is proposed to combine information from two X-rays. The mean squared error (MSE) loss and adversarial loss are combined to train the generator, resulting in a high-quality CT volume both visually and quantitatively. Extensive experiments on a publicly available chest CT dataset demonstrate the effectiveness of the proposed method. It could be a nice enhancement of a low-cost X-ray machine to provide physicians a CT-like 3D volume in several niche applications. Xingde Ying, Heng Guo 0003, Kai Ma 0002, Zhengxin Weng, Yefeng Zheng 0001 |
CVPR | 3 |
| 2019 | Multi-task Neural Networks with Spatial Activation for Retinal Vessel Segmentation and Artery/Vein Classification
Wenao Ma, Kai Ma 0002, Jiexiang Wang, Xinghao Ding, Yefeng Zheng 0001 |
MICCAI (1) | 3 |
| 2019 | Attentive CT Lesion Detection Using Deep Pyramid Inference with Multi-scale Booster
Qingbin Shao, Lijun Gong, Kai Ma 0002, Hualuo Liu, Yefeng Zheng 0001 |
MICCAI (6) | 3 |
| 2019 | Pairwise Semantic Segmentation via Conjugate Fully Convolutional Network
Renzhen Wang, Shilei Cao 0001, Kai Ma 0002, Deyu Meng, Yefeng Zheng 0001 |
MICCAI (6) | 3 |
| 2019 | Self-supervised Feature Learning for 3D Medical Images by Playing a Rubik's Cube
Xinrui Zhuang, Yuexiang Li, Kai Ma 0002, Yujiu Yang 0001, Yefeng Zheng 0001 |
MICCAI (4) | 4 |