Yulei Qin

dblp:226/3329 · DBLP profile ↗
← Back
25ranked-venue papers
8as first author
18since 2021 · last 2026
0000-0002-0996-3984ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 11 since 2021
YearPublicationVenuePosition
2026 Rethinking Long-tailed Dataset Distillation: A Uni-Level Framework with Unbiased Recovery and Relabeling
abstract
Dataset distillation creates a small distilled set that enables efficient training by capturing key information from the full dataset. While existing dataset distillation methods perform well on balanced datasets, they struggle under long-tailed distributions, where imbalanced class frequencies induce biased model representations and corrupt statistical estimates such as Batch Normalization (BN) statistics. In this paper, we rethink long-tailed dataset distillation by revisiting the limitations of trajectory-based methods, and instead adopt the statistical alignment perspective to jointly mitigate model bias and restore fair supervision. To this end, we introduce three dedicated components that enable unbiased recovery of distilled images and soft relabeling: (1) enhancing expert models (an observer model for recovery and a teacher model for relabeling) to enable reliable statistics estimation and soft-label generation; (2) recalibrating BN statistics via a full forward pass with dynamically adjusted momentum to reduce representation skew; (3) initializing synthetic images by incrementally selecting high-confidence and diverse augmentations via a multi-round mechanism that promotes coverage and diversity. Extensive experiments on four long-tailed benchmarks show consistent improvements over state-of-the-art methods across varying degrees of class imbalance.Notably, our approach improves top-1 accuracy by 15.6% on CIFAR-100-LT and 11.8% on Tiny-ImageNet-LT under IPC=10 and IF=10.
Yulei Qin, Wengang Zhou 0001, Houqiang Li
AAAI2
2025 Multi-Level Optimal Transport for Universal Cross-Tokenizer Knowledge Distillation on Language Models
abstract
Knowledge distillation (KD) has become a prevalent technique for compressing large language models (LLMs). Existing KD methods are constrained by the need for identical tokenizers (i.e., vocabularies) between teacher and student models, limiting their versatility in handling LLMs of different architecture families. In this paper, we introduce the Multi-Level Optimal Transport (MultiLevelOT), a novel approach that advances the optimal transport for universal cross-tokenizer knowledge distillation. Our method aligns the logit distributions of the teacher and the student at both token and sequence levels using diverse cost matrices, eliminating the need for dimensional or token-by-token correspondence. At the token level, MultiLevelOT integrates both global and local information by jointly optimizing all tokens within a sequence to enhance robustness. At the sequence level, we efficiently capture complex distribution structures of logits via the Sinkhorn distance, which approximates the Wasserstein distance for divergence measures. Extensive experiments on tasks such as extractive QA, generative QA, and summarization demonstrate that the MultiLevelOT outperforms state-of-the-art cross-tokenizer KD methods under various settings. Our approach is robust to different student and teacher models across model families, architectures, and parameter sizes.
Mo Zhu, Yulei Qin, Liang Xie 0003, Wengang Zhou 0001, Houqiang Li
AAAI3
2025 OPTICAL: Leveraging Optimal Transport for Contribution Allocation in Dataset Distillation
abstract
The demands for increasingly large-scale datasets pose substantial storage and computation challenges to building deep learning models. Dataset distillation methods, especially those via sample generation techniques, rise in response to condensing large original datasets into small synthetic ones while preserving critical information. Existing subset synthesis methods simply minimize the homogeneous distance where uniform contributions from all real instances are allocated to shaping each synthetic sample. We demonstrate that such equal allocation fails to consider the instance-level relationship between each real-synthetic pair and gives rise to insufficient modeling of geometric structural nuances between the distilled and original sets. In this paper, we propose a novel framework named OPTICAL to reformulate the homogeneous distance minimization into a bi-level optimization problem via matching-and-approximating. In the matching step, we leverage optimal transport matrix to dynamically allocate contributions from real instances. Subsequently, we polish the generated samples in accordance with the established allocation scheme for approximating the real ones. Such a strategy better measures intricate geometric characteristics and handles intra-class variations for high fidelity of data distillation. Extensive experiments across seven datasets and three model architectures demonstrate our method’s versatility and effectiveness. Its plug-and-play characteristic makes it compatible with a wide range of distillation frameworks.
Yulei Qin, Wengang Zhou 0001, Houqiang Li
CVPR2
2025 Optimizing Distributional Geometry Alignment with Optimal Transport for Generative Dataset Distillation
abstract
Dataset distillation seeks to synthesize a compact distilled dataset, enabling models trained on it to achieve performance comparable to models trained on the full dataset. Recent methods for large-scale datasets focus on matching global distributional statistics (e.g., mean and variance), but overlook critical instance-level characteristics and intraclass variations, leading to suboptimal generalization. We address this limitation by reformulating dataset distillation as an Optimal Transport (OT) distance minimization problem, enabling fine-grained alignment at both global and instance levels throughout the pipeline. OT offers a geometrically faithful framework for distribution matching. It effectively preserves local modes, intra-class patterns, and fine-grained variations that characterize the geometry of complex, high-dimensional distributions. Our method comprises three components tailored for preserving distributional geometry: (1) OT-guided diffusion sampling, which aligns latent distributions of real and distilled images; (2) label-image-aligned soft relabeling, which adapts label distributions based on the complexity of distilled image distributions; and (3) OT-based logit matching, which aligns the output of student models with soft-label distributions. Extensive experiments across diverse architectures and large-scale datasets demonstrate that our method consistently outperforms state-of-the-art approaches in an efficient manner, achieving at least 4\% accuracy improvement under IPC=10 settings for each architecture on ImageNet-1K.
Yulei Qin, Wengang Zhou 0001, Houqiang Li
NeurIPS2
2025 MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
abstract
Multimodal Large Language Model (MLLM) relies on the powerful LLM to perform multimodal tasks, showing amazing emergent abilities in recent studies, such as writing poems based on an image. However, it is difficult for these case studies to fully reflect the performance of MLLM, lacking a comprehensive evaluation. In this paper, we fill in this blank, presenting the first comprehensive MLLM Evaluation benchmark MME. It measures both perception and cognition abilities on a total of 14 subtasks. In order to avoid data leakage that may arise from direct use of public datasets for evaluation, the annotations of instruction-answer pairs are all manually designed. The concise instruction design allows us to fairly compare MLLMs, instead of struggling in prompt engineering. Besides, with such an instruction, we can also easily carry out quantitative statistics. A total of 30 advanced MLLMs are comprehensively evaluated on our MME, which not only suggests that existing MLLMs still have a large room for improvement, but also reveals the potential directions for the subsequent model optimization. The data are released at the project page: https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models/tree/Evaluation.
Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Jinrui Yang, Xiawu Zheng, Ke Li 0015, Xing Sun 0001, Yunsheng Wu, Rongrong Ji, Caifeng Shan, Ran He 0001
NeurIPS4
2025 LTD-Bench: Evaluating Large Language Models by Letting Them Draw
abstract
Current evaluation paradigms for large language models (LLMs) represent a critical blind spot in AI research—relying on opaque numerical metrics that conceal fundamental limitations in spatial reasoning while providing no intuitive understanding of model capabilities. This deficiency creates a dangerous disconnect between reported performance and practical abilities, particularly for applications requiring physical world understanding. We introduce LTD-Bench, a breakthrough benchmark that transforms LLM evaluation from abstract scores to directly observable visual outputs by requiring models to generate drawings through dot matrices or executable code. This approach makes spatial reasoning limitations immediately apparent even to non-experts, bridging the fundamental gap between statistical performance and intuitive assessment. LTD-Bench implements a comprehensive methodology with complementary generation tasks (testing spatial imagination) and recognition tasks (assessing spatial perception) across three progressively challenging difficulty levels, methodically evaluating both directions of the critical language-spatial mapping. Our extensive experiments with state-of-the-art models expose an alarming capability gap: even LLMs achieving impressive results on traditional benchmarks demonstrate profound deficiencies in establishing bidirectional mappings between language and spatial concepts—a fundamental limitation that undermines their potential as genuine world models. Furthermore, LTD-Bench's visual outputs enable powerful diagnostic analysis, offering a potential approach to investigate model similarity. Our dataset and codes are available at https://github.com/walktaster/LTD-Bench.
Liuhao Lin, Ke Li 0015, Yulei Qin, Yan Zhang 0109, Xing Sun 0001, Rongrong Ji
NeurIPS5
2025 Incentivizing Reasoning for Advanced Instruction-Following of Large Language Models
abstract
Existing large language models (LLMs) face challenges of following complex instructions, especially when multiple constraints are present and organized in paralleling, chaining, and branching structures. One intuitive solution, namely chain-of-thought (CoT), is expected to universally improve capabilities of LLMs. However, we find that the vanilla CoT exerts a negative impact on performance due to its superficial reasoning pattern of simply paraphrasing the instructions. It fails to peel back the compositions of constraints for identifying their relationship across hierarchies of types and dimensions. To this end, we propose RAIF, a systematic method to boost LLMs in dealing with complex instructions via incentivizing reasoning for test-time compute scaling. First, we stem from the decomposition of complex instructions under existing taxonomies and propose a reproducible data acquisition method. Second, we exploit reinforcement learning (RL) with verifiable rule-centric reward signals to cultivate reasoning specifically for instruction following. We address the shallow, non-essential nature of reasoning under complex instructions via sample-wise contrast for superior CoT enforcement. We also exploit behavior cloning of experts to facilitate steady distribution shift from fast-thinking LLMs to skillful reasoners. Extensive evaluations on seven comprehensive benchmarks confirm the validity of the proposed method, where a 1.5B LLM achieves 11.74% gains with performance comparable to a 8B LLM. Evaluation on OOD constraints also confirms the generalizability of our RAIF.
Yulei Qin, Zongyi Li, Zhekai Lin, Ke Li 0015, Xing Sun 0001
NeurIPS1
2025 SinKD: Sinkhorn Distance Minimization for Knowledge Distillation
abstract
Knowledge distillation (KD) has been widely adopted to compress large language models (LLMs). Existing KD methods investigate various divergence measures including the Kullback-Leibler (KL), reverse KL (RKL), and Jensen-Shannon (JS) divergences. However, due to limitations inherent in their assumptions and definitions, these measures fail to deliver effective supervision when a distribution overlap exists between the teacher and the student. In this article, we show that the aforementioned KL, RKL, and JS divergences, respectively, suffer from issues of mode-averaging, mode-collapsing, and mode-underestimation, which deteriorates logits-based KD for diverse natural language processing (NLP) tasks. We propose the Sinkhorn KD (SinKD) that exploits the Sinkhorn distance to ensure a nuanced and precise assessment of the disparity between distributions of teacher and student models. Besides, thanks to the properties of the Sinkhorn metric, we get rid of sample-wise KD that restricts the perception of divergences inside each teacher-student sample pair. Instead, we propose a batch-wise reformulation to capture the geometric intricacies of distributions across samples in the high-dimensional space. A comprehensive evaluation of GLUE and SuperGLUE, in terms of comparability, validity, and generalizability, highlights our superiority over state-of-the-art (SOTA) methods on all kinds of LLMs with encoder-only, encoder-decoder, and decoder-only architectures. Codes and models are available at https://github.com/2018cx/SinKD.
Yulei Qin, Enwei Zhang, Ke Li 0015, Xing Sun 0001, Wengang Zhou 0001, Houqiang Li
IEEE Trans. Neural Networks Learn. Syst.2
2024 Sinkhorn Distance Minimization for Knowledge Distillation
abstract
Knowledge distillation (KD) has been widely adopted to compress large language models (LLMs). Existing KD methods investigate various divergence measures including the Kullback-Leibler (KL), reverse Kullback-Leibler (RKL), and Jensen-Shannon (JS) divergences. However, due to limitations inherent in their assumptions and definitions, these measures fail to deliver effective supervision when few distribution overlap exists between the teacher and the student. In this paper, we show that the aforementioned KL, RKL, and JS divergences respectively suffer from issues of mode-averaging, mode-collapsing, and mode-underestimation, which deteriorates logits-based KD for diverse NLP tasks. We propose the Sinkhorn Knowledge Distillation (SinKD) that exploits the Sinkhorn distance to ensure a nuanced and precise assessment of the disparity between teacher and student distributions. Besides, profit by properties of the Sinkhorn metric, we can get rid of sample-wise KD that restricts the perception of divergence in each teacher-student sample pair. Instead, we propose a batch-wise reformulation to capture geometric intricacies of distributions across samples in the high-dimensional space. Comprehensive evaluation on GLUE and SuperGLUE, in terms of comparability, validity, and generalizability, highlights our superiority over state-of-the-art methods on all kinds of LLMs with encoder-only, encoder-decoder, and decoder-only architectures.
Yulei Qin, Enwei Zhang, Ke Li 0015, Xing Sun 0001, Wengang Zhou 0001, Houqiang Li
LREC/COLING2
2023 FoPro: Few-Shot Guided Robust Webly-Supervised Prototypical Learning
abstract
Recently, webly supervised learning (WSL) has been studied to leverage numerous and accessible data from the Internet. Most existing methods focus on learning noise-robust models from web images while neglecting the performance drop caused by the differences between web domain and real-world domain. However, only by tackling the performance gap above can we fully exploit the practical value of web datasets. To this end, we propose a Few-shot guided Prototypical (FoPro) representation learning method, which only needs a few labeled examples from reality and can significantly improve the performance in the real-world domain. Specifically, we initialize each class center with few-shot real-world data as the ``realistic" prototype. Then, the intra-class distance between web instances and ``realistic" prototypes is narrowed by contrastive learning. Finally, we measure image-prototype distance with a learnable metric. Prototypes are polished by adjacent high-quality web images and involved in removing distant out-of-distribution samples. In experiments, FoPro is trained on web datasets with a few real-world examples guided and evaluated on real-world datasets. Our method achieves the state-of-the-art performance on three fine-grained datasets and two large-scale datasets. Compared with existing WSL methods under the same few-shot settings, FoPro still excels in real-world generalization. Code is available at https://github.com/yuleiqin/fopro.
Yulei Qin, Chao Chen 0026, Yunhang Shen, Bo Ren 0002, Yun Gu, Jie Yang 0002, Chunhua Shen
AAAI1
2023 CAPro: Webly Supervised Learning with Cross-modality Aligned Prototypes
abstract
Webly supervised learning has attracted increasing attention for its effectiveness in exploring publicly accessible data at scale without manual annotation. However, most existing methods of learning with web datasets are faced with challenges from label noise, and they have limited assumptions on clean samples under various noise. For instance, web images retrieved with queries of ”tiger cat“ (a cat species) and ”drumstick“ (a musical instrument) are almost dominated by images of tigers and chickens, which exacerbates the challenge of fine-grained visual concept learning. In this case, exploiting both web images and their associated texts is a requisite solution to combat real-world noise. In this paper, we propose Cross-modality Aligned Prototypes (CAPro), a unified prototypical contrastive learning framework to learn visual representations with correct semantics. For one thing, we leverage textual prototypes, which stem from the distinct concept definition of classes, to select clean images by text matching and thus disambiguate the formation of visual prototypes. For another, to handle missing and mismatched noisy texts, we resort to the visual feature space to complete and enhance individual texts and thereafter improve text matching. Such semantically aligned visual prototypes are further polished up with high-quality samples, and engaged in both cluster regularization and noise removal. Besides, we propose collective bootstrapping to encourage smoother and wiser label reference from appearance-similar instances in a manner of dictionary look-up. Extensive experiments on WebVision1k and NUS-WIDE (Web) demonstrate that CAPro well handles realistic noise under both single-label and multi-label scenarios. CAPro achieves new state-of-the-art performance and exhibits robustness to open-set recognition. Codes are available at https://github.com/yuleiqin/capro.
Yulei Qin, Yunhang Shen, Chaoyou Fu, Yun Gu, Ke Li 0015, Xing Sun 0001, Rongrong Ji
NeurIPS1
2023 Trustworthy learning with (un)sure annotation for lung nodule diagnosis with CT
Liang Chen 0023, Xiao Gu 0003, Yulei Qin, Zhexin Wang, Yun Gu, Guang-Zhong Yang
Medical Image Anal.5
2023 Multi-site, Multi-domain Airway Tree Modeling
Yangqian Wu, Yulei Qin, Hao Zheng 0008, Wen Tang 0005, Corey W. Arnold, Chenhao Pei, Pengxin Yu, Yang Nan 0002, Guang Yang 0006, Simon Walsh, Dominic C. Marshall, Matthieu Komorowski, Puyang Wang, Dazhou Guo, Dakai Jin, Shuiqing Zhao, Runsheng Chang, Abdul Qayyum 0002, Moona Mazher, Yonghuang Wu, Ying'ao Liu, Jiancheng Yang, Ashkan Pakzad, Bojidar Rangelov, Raúl San José Estépar, Carlos Cano-Espinosa, Jiayuan Sun, Guang-Zhong Yang, Yun Gu
Medical Image Anal.4
2023 KaryoNet: Chromosome Recognition With End-to-End Combinatorial Optimization Network
abstract
Chromosome recognition is a critical way to diagnose various hematological malignancies and genetic diseases, which is however a repetitive and time-consuming process in karyotyping. To explore the relative relation between chromosomes, in this work, we start from a global perspective and learn the contextual interactions and class distribution features between chromosomes within a karyotype. We propose an end-to-end differentiable combinatorial optimization method, KaryoNet, which captures long-range interactions between chromosomes with the proposed Masked Feature Interaction Module (MFIM) and conducts label assignment in a flexible and differentiable way with Deep Assignment Module (DAM). Specially, a Feature Matching Sub-Network is built to predict the mask array for attention computation in MFIM. Lastly, Type and Polarity Prediction Head can predict chromosome type and polarity simultaneously. Extensive experiments on R-band and G-band two clinical datasets demonstrate the merits of the proposed method. For normal karyotypes, the proposed KaryoNet achieves the accuracy of 98.41% on R-band chromosome and 99.58% on G-band chromosome. Owing to the extracted internal relation and class distribution features, KaryoNet can also achieve state-of-the-art performances on karyotypes of patients with different types of numerical abnormalities. The proposed method has been applied to assist clinical karyotype diagnosis. Our code is available at: https://github.com/xiabc612/KaryoNet.
Jiyue Wang, Yulei Qin, Zhaojiang Liu, Lingqian Wu, Yun Gu, Jie Yang 0002
IEEE Trans. Medical Imaging3
2022 An End-to-End Combinatorial Optimization Method for R-band Chromosome Recognition with Grouping Guided Attention
Jiyue Wang, Yulei Qin, Yun Gu, Jie Yang 0002
MICCAI (4)3
2021 Refined Local-imbalance-based Weight for Airway Segmentation in CT
Hao Zheng 0008, Yulei Qin, Yun Gu, Fangfang Xie, Jiayuan Sun, Jie Yang 0002, Guang-Zhong Yang
MICCAI (1)2
2021 Learning Tubule-Sensitive CNNs for Pulmonary Airway and Artery-Vein Segmentation in CT
abstract
Training convolutional neural networks (CNNs) for segmentation of pulmonary airway, artery, and vein is challenging due to sparse supervisory signals caused by the severe class imbalance between tubular targets and background. We present a CNNs-based method for accurate airway and artery-vein segmentation in non-contrast computed tomography. It enjoys superior sensitivity to tenuous peripheral bronchioles, arterioles, and venules. The method first uses a feature recalibration module to make the best use of features learned from the neural networks. Spatial information of features is properly integrated to retain relative priority of activated regions, which benefits the subsequent channel-wise recalibration. Then, attention distillation module is introduced to reinforce representation learning of tubular objects. Fine-grained details in high-resolution attention maps are passing down from one layer to its previous layer recursively to enrich context. Anatomy prior of lung context map and distance transform map is designed and incorporated for better artery-vein differentiation capacity. Extensive experiments demonstrated considerable performance gains brought by these components. Compared with state-of-the-art methods, our method extracted much more branches while maintaining competitive overall segmentation performance. Codes and models are available at http://www.pami.sjtu.edu.cn/News/56.
Yulei Qin, Hao Zheng 0008, Yun Gu, Xiaolin Huang, Jie Yang 0002, Lihui Wang 0002, Yue Min Zhu, Guang-Zhong Yang
IEEE Trans. Medical Imaging1
2021 Alleviating Class-Wise Gradient Imbalance for Pulmonary Airway Segmentation
abstract
Automated airway segmentation is a prerequisite for pre-operative diagnosis and intra-operative navigation for pulmonary intervention. Due to the small size and scattered spatial distribution of peripheral bronchi, this is hampered by a severe class imbalance between foreground and background regions, which makes it challenging for CNN-based methods to parse distal small airways. In this paper, we demonstrate that this problem is arisen by gradient erosion and dilation of the neighborhood voxels. During back-propagation, if the ratio of the foreground gradient to background gradient is small while the class imbalance is local, the foreground gradients can be eroded by their neighborhoods. This process cumulatively increases the noise information included in the gradient flow from top layers to the bottom ones, limiting the learning of small structures in CNNs. To alleviate this problem, we use group supervision and the corresponding WingsNet to provide complementary gradient flows to enhance the training of shallow layers. To further address the intra-class imbalance between large and small airways, we design a General Union loss function that obviates the impact of airway size by distance-based weights and adaptively tunes the gradient ratio based on the learning process. Extensive experiments on public datasets demonstrate that the proposed method can predict the airway structures with higher accuracy and better morphological completeness than the baselines.
Hao Zheng 0008, Yulei Qin, Yun Gu, Fangfang Xie, Jie Yang 0002, Jiayuan Sun, Guang-Zhong Yang
IEEE Trans. Medical Imaging2
2020 Learning Bronchiole-Sensitive Airway Segmentation CNNs by Feature Recalibration and Attention Distillation
Yulei Qin, Hao Zheng 0008, Yun Gu, Xiaolin Huang, Jie Yang 0002, Lihui Wang 0002, Yue Min Zhu
MICCAI (1)1
2020 Learning with Sure Data for Nodule-Level Lung Cancer Prediction
Yun Gu, Yulei Qin, Guang-Zhong Yang
MICCAI (6)3
2020 Weakly Supervised Deep Learning for Breast Cancer Segmentation with Coarse Annotations
Hao Zheng 0008, Zhiguo Zhuang, Yulei Qin, Yun Gu, Jie Yang 0002, Guang-Zhong Yang
MICCAI (4)3
2019 AirwayNet: A Voxel-Connectivity Aware Approach for Accurate Airway Segmentation Using Convolutional Neural Networks
Yulei Qin, Hao Zheng 0008, Yun Gu, Mali Shen, Jie Yang 0002, Xiaolin Huang, Yue Min Zhu, Guang-Zhong Yang
MICCAI (6)1
2019 Varifocal-Net: A Chromosome Classification Approach Using Deep Convolutional Networks
abstract
Chromosome classification is critical for karyotyping in abnormality diagnosis. To expedite the diagnosis, we present a novel method named Varifocal-Net for simultaneous classification of chromosome's type and polarity using deep convolutional networks. The approach consists of one global-scale network (G-Net) and one local-scale network (L-Net). It follows three stages. The first stage is to learn both global and local features. We extract global features and detect finer local regions via the G-Net. By proposing a varifocal mechanism, we zoom into local parts and extract local features via the L-Net. Residual learning and multi-task learning strategies are utilized to promote high-level feature extraction. The detection of discriminative local parts is fulfilled by a localization subnet of the G-Net, whose training process involves both supervised and weakly supervised learning. The second stage is to build two multi-layer perceptron classifiers that exploit features of both two scales to boost classification performance. The third stage is to introduce a dispatch strategy of assigning each chromosome to a type within each patient case, by utilizing the domain knowledge of karyotyping. The evaluation results from 1909 karyotyping cases showed that the proposed Varifocal-Net achieved the highest accuracy per patient case (%) of 99.2 for both type and polarity tasks. It outperformed state-of-the-art methods, demonstrating the effectiveness of our varifocal mechanism, multi-scale feature ensemble, and dispatch strategy. The proposed method has been applied to assist practical karyotype diagnosis.
Yulei Qin, Hao Zheng 0008, Xiaolin Huang, Jie Yang 0002, Yue Min Zhu, Lingqian Wu, Guang-Zhong Yang
IEEE Trans. Medical Imaging1
2018 Simultaneous Accurate Detection of Pulmonary Nodules and False Positive Reduction Using 3D CNNs
abstract
Accurate detection of nodules in CT images is vital for lung cancer diagnosis, which greatly influences the patient's chance for survival. Motivated by successful application of convolutional neural networks (CNNs) on natural images, we propose a computer-aided diagnosis (CAD) system for simultaneous accurate pulmonary nodule detection and false positive reduction. To generate nodule candidates, we build a full 3D CNN model that employs 3D U-Net architecture as the backbone of a region proposal network (RPN). We adopt multi-task residual learning and online hard negative example mining strategy to accelerate the training process and improve the accuracy of nodule detection. Then, a 3D DenseNet-based model is presented to reduce false positive nodules. The densely connected structure reuses nodules' features and boosts feature propagation. Experimental results on LUNA16 datasets demonstrate the superior effectiveness of our approach over state-of-the-art methods.
Yulei Qin, Hao Zheng 0008, Yue Min Zhu, Jie Yang 0002
ICASSP1
2018 Small Lesion Classification in Dynamic Contrast Enhancement MRI for Breast Cancer Early Detection
Hao Zheng 0008, Yun Gu, Yulei Qin, Xiaolin Huang, Jie Yang 0002, Guang-Zhong Yang
MICCAI (2)3