VLDB 2026 Research / reviewers in the wild / expert
Dong Nie
dblp:130/8299
· DBLP profile ↗
58ranked-venue papers
10as first author
31since 2021 · last 2026
0000-0003-0385-8988ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 30 · 3 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 6 first-author · 12 since 2021Artificial intelligence and machine learning · 23 · 7 first-author · 15 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Optimizing LoRA Allocation of MoE with the Alignment of Topic CorrelationabstractMixture of experts (MoE) dynamically routes inputs to specialized expert networks to scale model capacity with low inference overhead. However, the excessive parameter growth in MoE models poses challenges in low-resource settings. To address these issues, MoE with parameter-efficient fine-tuning (PEFT) methods have emerged as a lightweight adaptation paradigm that distributes knowledge among experts via multiple LoRA blocks. Existing MoE-PEFT methods can be broadly categorized into External and Internal PEFT methods. External PEFT methods incorporate lightweight models into existing MoE architectures without modifying their routing, which limits the model’s parameter efficiency. To overcome these issues, Internal PEFT methods integrate MoE architectures into PEFT, enabling minimal parameter overhead. However, they still face two major challenges: (1) lack of expert functional differentiation, resulting in overlapping specialization across modules, and (2) absence of a structured attribution mechanism to guide expert selection based on semantic relevance. To alleviate these challenges, we propose TopicLoRA, a novel three-stage framework that leverages topic knowledge as semantic anchors to guide expert allocation. Specifically, (1) to address expert redundancy, we construct a topic-level prior graph using Graph Neural Network-enhanced representation learning over Big-Bench categories, enforcing structural separation among expert embeddings, and (2) to introduce semantic attribution, we design a dual-loss training mechanism that softly aligns input-query relevance with topic-guided routing distributions via KL divergence. Extensive experiments on representative datasets (e.g., MMLU, GSM8K, Flanv2) demonstrate that TopicLoRA outperforms state-of-the-art PEFT baselines by 2.40% on average in accuracy. Notably, the maximum improvement is 4.21%. Furthermore, ablation studies demonstrate that our framework's robustness to intricate topics and input sequence variations, which stems from the dual-loss training mechanism. Hengyuan Xu, Wenjun Ke 0002, Jiajun Liu 0005, Dong Nie, Peng Wang 0004, Ziyu Shang, Zijie Xu 0003 |
AAAI | 5 |
| 2026 | Large Language Models in Document Intelligence: A Comprehensive Survey, Recent Advances, Challenges, and Future TrendsabstractThe rapid proliferation of documents has made document intelligence increasingly critical across various industries. In recent years, Large Language Models (LLMs) have dramatically transformed the field of document intelligence, allowing for more advanced and accurate document processing solutions. Despite these advancements, most existing surveys have failed to focus on these breakthroughs, instead concentrating on traditional methods and earlier machine learning techniques. This survey seeks to fill that gap by offering an in-depth analysis of approximately 300 papers published between 2021 and mid-2025, thus providing a comprehensive overview of the impact of LLMs in document intelligence. The key topics explored include Retrieval-Augmented Generation (RAG), long-context processing, and fine-tuning LLMs for document comprehension. Furthermore, the survey highlights essential datasets, practical applications, current challenges, and future research directions, offering critical insights for both researchers and industry practitioners looking to advance the field. Wenjun Ke 0002, Hengyuan Xu, Dong Nie, Peng Wang 0004 |
ACM Trans. Inf. Syst. | 5 |
| 2025 | BrainX: A Universal Brain Decoding Framework with Feature Disentanglement and Neuro-Geometric Representation LearningabstractDecoding visual stimuli from human brain activity is a fundamental challenge in cognitive neuroscience and neuroimaging. While recent advances in deep learning have significantly improved the performance of fMRI-to-image decoding, most existing methods overlook the issue of inter-subject variability in fMRI data, which leads to poor generalization across subjects. Current approaches often rely on partially shared model architectures that offer limited generalization and still require subject-specific components, restricting their applicability to unseen subjects. To address this limitation, we propose BrainX, a universal brain decoding framework that constructs a unified fMRI encoder and image generator to achieve subject-agnostic modeling. Specifically, we introduce a feature disentanglement mechanism that extracts subject-shared features from the fMRI embeddings, which are then fed into the image generator to reconstruct visual stimuli. This design eliminates the need for subject-specific models and significantly enhances cross-subject generalization. Additionally, we develop a neuro-geometric fMRI representation learning method that projects 3D cortical structures onto a 2D surface space, effectively mitigating the inaccuracies caused by imprecise geodesic distance estimation in 3D Euclidean space. Extensive experiments on the Natural Scenes Dataset (NSD) demonstrate that BrainX consistently outperforms existing state-of-the-art methods across three decoding settings: within-subject, cross-subject with finetuning, and cross-subject without finetuning. Dong Nie, Pengcheng Xue, Xia Wu 0001, Daoqiang Zhang, Xuyun Wen |
CIKM | 2 |
| 2025 | Generalized Zero-Shot Classification via Semantics-Free Inter-Class Feature GenerationabstractGeneralized Zero-Shot Learning (GZSL) addresses the challenge of classifying unseen classes in the presence of seen classes by leveraging semantic attributes to bridge the gap for unseen classes. However, in image based disease classification, such as glioma sub-typing, distinguishing between classes using image semantic attributes can be challenging. To address this challenge, we introduce a novel GZSL method that eliminates the dependency on semantic information. Specifically, we propose that the primary of most classification in clinic is risk stratification, and classes are inherently ordered rather than purely categorical. Based on this insight, we present an inter-class feature augmentation (IFA) module, where distributions of different classes are ordered by their risk levels in a learned feature space using pre-defined joint conditional Gaussian distribution model. This ordering enables the generation of unseen class features through feature mixing of adjacent seen classes, effectively transforming the zero-shot learning problem into a supervised learning task. Our method eliminates the need for explicit semantic information, avoiding the cross-modal alignment between visual and semantic features. Moreover, the IFA module for GZSL requires no structural modifications to the existing classification models. In the experiment, both in-house and public datasets are used to evaluate our method across different tasks, including glioma subtyping, Alzheimer’s disease (AD) classification and diabetic retinopathy classification. Experimental results demonstrate that our method outperforms the state-of-the-art GZSL methods with statistical significance. Libiao Chen, Dong Nie, JunJun Pan, Zhenyu Tang 0002 |
CVPR | 2 |
| 2025 | μ 2 Tokenizer: Differentiable Multi-Scale Multi-Modal Tokenizer for Radiology Report Generation
Siyou Li, Pengyao Qin, Huanan Wu, Dong Nie, Arun James Thirunavukarasu, Juntao Yu |
MICCAI (5) | 4 |
| 2025 | Iterative Foundation-Dedicated Learning: Optimized Key Frames, Prompts and Memories for Semi-supervised Segmentation
Ziman Yin, Dong Nie, Shuo Li 0001, JunJun Pan, Zhenyu Tang 0002 |
MICCAI (8) | 2 |
| 2025 | Brain-Inspired fMRI-to-Text Decoding via Incremental and Wrap-Up Language ModelingabstractDecoding natural language text from non-invasive brain signals, such as functional magnetic resonance imaging (fMRI), remains a central challenge in brain-computer interface research. While recent advances in large language models (LLMs) have enabled open-vocabulary fMRI-to-text decoding, existing frameworks typically process the entire fMRI sequence in a single step, leading to performance degradation when handling long input sequences due to memory overload and semantic drift. To address this limitation, we propose a brain-inspired sequential fMRI-to-text decoding framework that mimics the human cognitive strategy of segmented and inductive language processing. Specifically, we divide long fMRI time series into consecutive segments aligned with optimal language comprehension length. Each segment is decoded incrementally, followed by a wrap-up mechanism that summarizes the semantic content and incorporates it as prior knowledge into subsequent decoding steps. This sequence-wise approach alleviates memory burden and ensures semantic continuity across segments. In addition, we introduce a text-guided masking strategy integrated with a masked autoencoder (MAE) framework for fMRI representation learning. This method leverages attention distributions over key semantic tokens to selectively mask the corresponding fMRI time points, and employs MAE to guide the model toward focusing on neural activity at semantically salient moments, thereby enhancing the capability of fMRI embeddings to represent textual information. Experimental results on the two datasets demonstrate that our method significantly outperforms state-of-the-art approaches, with performance gains increasing as decoding length grows. Dong Nie, Pengcheng Xue, Piji Li, Daoqiang Zhang, Xuyun Wen |
NeurIPS | 2 |
| 2025 | ROP lesion segmentation via sequence coding and block balancingabstractRetinopathy of prematurity (ROP) is a potentially blinding retinal disease that often affects low birth weight premature infants. Lesion detection and recognition are crucial for ROP diagnosis and clinical treatment. However, this task poses challenges for both ophthalmologists and computer-based systems due to the small size and subtle nature of many ROP lesions. To address these challenges, we present a Sequence encoding and Block balancing-based Segmentation Network (SeBSNet), which incorporates domain knowledge coding, sequence coding learning (SCL), and block-weighted balancing (BWB) techniques into the segmentation of ROP lesions. The experimental results demonstrate that SeBSNet outperforms existing state-of-the-art methods in the segmentation of ROP lesions, with average ROC_AUC, PR_AUC, and Dice scores of 98.84%, 71.90%, and 66.88%, respectively. Furthermore, the integration of the proposed techniques into ROP classification networks as an enhancing module leads to considerable improvements in classification performance. Xiping Jia, Jianying Qiu, Dong Nie |
Medical Image Anal. | 3 |
| 2025 | Structure-Aware Brain Tissue Segmentation for Isointense Infant MRI Data Using Multi-Phase Multi-Scale Assistance NetworkabstractAccurate and automatic brain tissue segmentation is crucial for tracking brain development and diagnosing brain disorders. However, due to inherently ongoing myelination and maturation during the first postnatal year, the intensity distributions of gray matter and white matter in the infant brain MRI at the age of around 6 months old (a.k.a. isointense phase) are highly overlapped, which makes tissue segmentation very challenging, even for experts. To address this issue, in this study, we propose a multi-phase multi-scale assistance segmentation framework, which comprises a structure-preserved generative adversarial network (SPGAN) and a multi-phase multi-scale assisted segmentation network (MASN). SPGAN bi-directionally synthesizes isointense and adult-like data. The synthetic isointense data essentially augment the training dataset, combined with high-quality annotations transferred from its adult-like counterpart. By contrast, the synthetic adult-like data offers clear tissue structures and is concatenated with isointense data to serve as the input of MASN. In particular, MASN is designed with two-branch networks, which simultaneously segment tissues with two phases (isointense and adult-like) and two scales by also preserving their correspondences. We further propose a boundary refinement module to extract maximum gradients from local feature maps to indicate tissue boundaries, prompting MASN to focus more on boundaries where segmentation errors are prone to occur. Extensive experiments on the National Database for Autism Research and Baby Connectome Project datasets quantitatively and qualitatively demonstrate the superiority of our proposed framework compared with seven state-of-the-art methods. Jiameng Liu, Feihong Liu, Dong Nie, Yuning Gu, Dinggang Shen |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | WSSADN: A Weakly Supervised Spherical Age-Disentanglement Network for Detecting Developmental Disorders with Structural MRI
Pengcheng Xue, Dong Nie, Meijiao Zhu, Han Zhang 0002, Daoqiang Zhang, Xuyun Wen |
MICCAI (11) | 2 |
| 2024 | TARDRL: Task-Aware Reconstruction for Dynamic Representation Learning of fMRI
Yunxi Zhao, Dong Nie, Xia Wu 0001, Daoqiang Zhang, Xuyun Wen |
MICCAI (11) | 2 |
| 2024 | Unveiling LoRA Intrinsic Ranks via Salience AnalysisabstractThe immense parameter scale of large language models underscores the necessity for parameter-efficient fine-tuning methods. Methods based on Low-Rank Adaptation (LoRA) assume the low-rank characteristics of the incremental matrix and optimize the matrix obtained from low-rank decomposition. Although effective, these methods are constrained by a fixed and unalterable intrinsic rank, neglecting the variable importance of matrices. Consequently, methods for adaptive rank allocation are proposed, among which AdaLoRA demonstrates excellent fine-tuning performance. AdaLoRA conducts adaptation based on singular value decomposition (SVD), dynamically allocating intrinsic ranks according to importance. However, it still struggles to achieve a balance between fine-tuning effectiveness and efficiency, leading to limited rank allocation space. Additionally, the importance measurement focuses only on parameters with minimal impact on the loss, neglecting the dominant role of singular values in SVD-based matrices and the fluctuations during training. To address these issues, we propose SalientLoRA, which adaptively optimizes intrinsic ranks of LoRA via salience measurement. Firstly, during rank allocation, the salience measurement analyses the variation of singular value magnitudes across multiple time steps and establishes their inter-dependency relationships to assess the matrix importance. This measurement mitigates instability and randomness that may arise during importance assessment. Secondly, to achieve a balance between fine-tuning performance and efficiency, we propose an adaptive adjustment of time-series window, which adaptively controls the size of time-series for significance measurement and rank reduction during training, allowing for rapid rank allocation while maintaining training stability. This mechanism enables matrics to set a higher initial rank, thus expanding the allocation space for ranks. To evaluate the generality of our method across various tasks, we conduct experiments on natural language understanding (NLU), natural language generation (NLG), and large model instruction tuning tasks. Experimental results demonstrate the superiority of SalientLoRA, which outperforms state-of-the-art methods by 0.96\%-3.56\% on multiple datasets. Furthermore, as the rank allocation space expands, our method ensures fine-tuning efficiency, achieving a speed improvement of 94.5\% compared to AdaLoRA. The code is publicly available at https://github.com/Heyest/SalientLoRA. Wenjun Ke 0002, Peng Wang 0004, Jiajun Liu 0005, Dong Nie |
NeurIPS | 5 |
| 2024 | Multimodal Brain Tumor Segmentation Boosted by Monomodal Normal Brain ImagesabstractMany deep learning based methods have been proposed for brain tumor segmentation. Most studies focus on deep network internal structure to improve the segmentation accuracy, while valuable external information, such as normal brain appearance, is often ignored. Inspired by the fact that radiologists often screen lesion regions with normal appearance as reference in mind, in this paper, we propose a novel deep framework for brain tumor segmentation, where normal brain images are adopted as reference to compare with tumor brain images in a learned feature space. In this way, features at tumor regions, i.e., tumor-related features, can be highlighted and enhanced for accurate tumor segmentation. It is known that routine tumor brain images are multimodal, while normal brain images are often monomodal. This causes the feature comparison a big issue, i.e., multimodal vs. monomodal. To this end, we present a new feature alignment module (FAM) to make the feature distribution of monomodal normal brain images consistent/inconsistent with multimodal tumor brain images at normal/tumor regions, making the feature comparison effective. Both public (BraTS2022) and in-house tumor brain image datasets are used to evaluate our framework. Experimental results demonstrate that for both datasets, our framework can effectively improve the segmentation accuracy and outperforms the state-of-the-art segmentation methods. Codes are available at https://github.com/hb-liu/Normal-Brain-Boost-Tumor-Segmentation. Huabing Liu, Zhengze Ni, Dong Nie, Dinggang Shen, Jinda Wang, Zhenyu Tang 0002 |
IEEE Trans. Image Process. | 3 |
| 2024 | A New Multi-Atlas Based Deep Learning Segmentation Framework With Differentiable Atlas Feature WarpingabstractDeep learning based multi-atlas segmentation (DL-MA) has achieved the state-of-the-art performance in many medical image segmentation tasks, e.g., brain parcellation. In DL-MA methods, atlas-target correspondence is the key for accurate segmentation. In most existing DL-MA methods, such correspondence is usually established using traditional or deep learning based registration methods at image level with no further feature level adaption. This could cause possible atlas-target feature inconsistency. As a result, the information from atlases often has limited positive and even counteractive impact on the final segmentation results. To tackle this issue, in this paper, we propose a new DL-MA framework, where a novel differentiable atlas feature warping module with a new smooth regularization term is presented to establish feature level atlas-target correspondence. Comparing with the existing DL-MA methods, in our framework, atlas features containing anatomical prior knowledge are more relevant to the target image feature, leading the final segmentation results to a high accuracy level. We evaluate our framework in the context of brain parcellation using two public MR brain image datasets: LPBA40 and NIREP-NA0. The experimental results demonstrate that our framework outperforms both traditional multi-atlas segmentation (MAS) and state-of-the-art DL-MA methods with statistical significance. Further ablation studies confirm the effectiveness of the proposed differentiable atlas feature warping module. Huabing Liu, Dong Nie, Jian Yang 0009, Jinda Wang, Zhenyu Tang 0002 |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | Multi-Label Clinical Time-Series Generation via Conditional GANabstractIn recent years, deep learning has been successfully adopted in a wide range of applications related to electronic health records (EHRs) such as representation learning and clinical event prediction. However, due to privacy constraints, limited access to EHR becomes a bottleneck for deep learning research. To mitigate these concerns, generative adversarial networks (GANs) have been successfully used for generating EHR data. However, there are still challenges in high-quality EHR generation, including generating time-series EHR data and imbalanced uncommon diseases. In this work, we propose aMulti-labelTime-seriesGAN(MTGAN) to generate EHR and simultaneously improve the quality of uncommon disease generation. The generator of MTGAN uses a gated recurrent unit (GRU) with a smooth conditional matrix to generate sequences and uncommon diseases. The critic gives scores using Wasserstein distance to recognize real samples from synthetic samples by considering both data and temporal features. We also propose a training strategy to calculate temporal features for real data and stabilize GAN training. Furthermore, we design multiple statistical metrics and prediction tasks to evaluate the generated data. Experimental results demonstrate the quality of the synthetic data and the effectiveness of MTGAN in generating realistic sequential EHR data, especially for uncommon diseases. Chang Lu 0004, Chandan K. Reddy, Ping Wang 0024, Dong Nie, Yue Ning 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2023 | The Devil is in the Upsampling: Architectural Decisions Made Simpler for Denoising with Deep Image PriorabstractDeep Image Prior (DIP) shows that some network architectures inherently tend towards generating smooth images while resisting noise, a phenomenon known as spectral bias. Image denoising is a natural application of this property. Although denoising with DIP mitigates the need for large training sets, two often intertwined practical challenges need to be overcome: architectural design and noise fitting. Existing methods either handcraft or search for suitable architectures from a vast design space, due to the limited understanding of how architectural choices affect the denoising outcome. In this study, we demonstrate from a frequency perspective that unlearnt upsampling is the main driving force behind the denoising phenomenon with DIP. This finding leads to straightforward strategies for identifying a suitable architecture for every image without laborious search. Extensive experiments show that the estimated architectures achieve superior denoising results than existing methods with up to 95% fewer parameters. Thanks to this under-parameterization, the resulting architectures are less prone to noise-fitting. Yunkui Pang, Dong Nie, Pew-Thian Yap |
ICCV | 4 |
| 2023 | Multi-Target Domain Adaptation with Prompt Learning for Medical Image Segmentation
Dong Nie, Daoqiang Zhang, Xuyun Wen |
MICCAI (1) | 2 |
| 2023 | TransDose: Transformer-based radiotherapy dose prediction from CT images guided by super-pixel-level GCN classification
Zhengyang Jiao, Xingchen Peng, Yan Wang 0015, Jianghong Xiao, Dong Nie, Xi Wu 0004, Xin Wang 0045, Jiliu Zhou, Dinggang Shen |
Medical Image Anal. | 5 |
| 2023 | Semi-Supervised Standard-Dose PET Image Generation via Region-Adaptive Normalization and Structural Consistency ConstraintabstractPositron Emission Tomography (PET) is an important nuclear medical imaging technique, and has been widely used in clinical applications, e.g., tumor detection and brain disease diagnosis. As PET imaging could put patients at risk of radiation, the acquisition of high-quality PET images with standard-dose tracers should be cautious. However, if dose is reduced in PET acquisition, the imaging quality could become worse and thus may not meet clinical requirement. To safely reduce the tracer dose and also maintain high quality of PET imaging, we propose a novel and effective approach to estimate high-quality Standard-dose PET (SPET) images from Low-dose PET (LPET) images. Specifically, to fully utilize both the rare paired and the abundant unpaired LPET and SPET images, we propose a semi-supervised framework for network training. Meanwhile, based on this framework, we further design a Region-adaptive Normalization (RN) and a structural consistency constraint to track the task-specific challenges. RN performs region-specific normalization in different regions of each PET image to suppress negative impact of large intensity variation across different regions, while the structural consistency constraint maintains structural details during the generation of SPET images from LPET images. Experiments on real human chest-abdomen PET images demonstrate that our proposed approach achieves state-of-the-art performance quantitatively and qualitatively. Caiwen Jiang, Yongsheng Pan, Zhiming Cui 0001, Dong Nie, Dinggang Shen |
IEEE Trans. Medical Imaging | 4 |
| 2023 | Fast Multi-Contrast MRI Acquisition by Optimal Sampling of Information Complementary to Pre-Acquired MRI ContrastabstractRecent studies on multi-contrast MRI reconstruction have demonstrated the potential of further accelerating MRI acquisition by exploiting correlation between contrasts. Most of the state-of-the-art approaches have achieved improvement through the development of network architectures for fixed under-sampling patterns, without considering inter-contrast correlation in the under-sampling pattern design. On the other hand, sampling pattern learning methods have shown better reconstruction performance than those with fixed under-sampling patterns. However, most under-sampling pattern learning algorithms are designed for single contrast MRI without exploiting complementary information between contrasts. To this end, we propose a framework to optimize the under-sampling pattern of a target MRI contrast which complements the acquired fully-sampled reference contrast. Specifically, a novel image synthesis network is introduced to extract the redundant information contained in the reference contrast, which is exploited in the subsequent joint pattern optimization and reconstruction network. We have demonstrated superior performance of our learned under-sampling patterns on both public and in-house datasets, compared to the commonly used under-sampling patterns and state-of-the-art methods that jointly optimize the reconstruction network and the under-sampling patterns, up to 8-fold under-sampling factor. Xiaoxin Li 0001, Feihong Liu, Dong Nie, Pietro Liò, Haikun Qi, Dinggang Shen |
IEEE Trans. Medical Imaging | 4 |
| 2022 | Weakly-supervised Metric Learning with Cross-Module Communications for the Classification of Anterior Chamber Angle ImagesabstractAs the basis for developing glaucoma treatment strategies, Anterior Chamber Angle (ACA) evaluation is usually dependent on experts' Judgements. However, experienced ophthalmologists needed for these Judgements are not widely available. Thus, computer-aided ACA evaluations become a pressing and efficient solution for this issue. In this paper, we propose a novel end-to-end frame-work GCNet for automated Glaucoma Classification based on ACA images or other Glaucoma-related medical images. We first collect and label an ACA image dataset with some pixel-level annotations. Next, we introduce a segmentation module and an embedding module to enhance the performance of classifying ACA images. Within GCNet, we design a Cross-Module Aggregation Net (CMANet) which is a weakly-supervised metric learning network to capture contextual information exchanging across these modules. We conduct experiments on the ACA dataset and two public datasets REFUGE and SIGF. Our experimental results demonstrate that GCNet outperforms several state-of-the-art deep models in the tasks of glaucoma medical image classifications. The source code of GCNet can be found at https://github.com/Jingqi-H/GCNet. Jingqi Huang, Yue Ning 0001, Dong Nie, Linan Guan, Xiping Jia |
CVPR | 3 |
| 2022 | Pyramid Architecture for Multi-Scale Processing in Point Cloud SegmentationabstractSemantic segmentation of point cloud data is a critical task for autonomous driving and other applications. Recent advances of point cloud segmentation are mainly driven by new designs of local aggregation operators and point sampling methods. Unlike image segmentation, few efforts have been made to understand the fundamental issue of scale and how scales should interact and be fused. In this work, we investigate how to efficiently and effectively integrate features at varying scales and varying stages in a point cloud segmentation network. In particular, we open up the commonly used encoder-decoder architecture, and design scale pyramid architectures that allow information to flow more freely and systematically, both laterally and upward/downward in scale. Moreover, a cross-scale attention feature learning block has been designed to enhance the multi-scale feature fusion which occurs everywhere in the network. Such a design of multi-scale processing and fusion gains large improvements in accuracy without adding much additional computation. When built on top of the popular KPConv network, we see consistent improvements on a wide range of datasets, including achieving state-of-the-art performance on NPM3D and S3DIS. Moreover, the pyramid architecture is generic and can be applied to other network designs: we show an example of similar improvements over RandLANet. Dong Nie, Rui Lan, Xiaofeng Ren |
CVPR | 1 |
| 2022 | Doubly-Fused ViT: Fuse Information from Vision Transformer Doubly with Local Representation
Dong Nie, Xiaofeng Ren |
ECCV (23) | 2 |
| 2022 | An Efficient Semi-Supervised Framework with Multi-Task and Curriculum Learning for Medical Image SegmentationabstractA practical problem in supervised deep learning for medical image segmentation is the lack of labeled data which is expensive and time-consuming to acquire. In contrast, there is a considerable amount of unlabeled data available in the clinic. To make better use of the unlabeled data and improve the generalization on limited labeled data, in this paper, a novel semi-supervised segmentation method via multi-task curriculum learning is presented. Here, curriculum learning means that when training the network, simpler knowledge is preferentially learned to assist the learning of more difficult knowledge. Concretely, our framework consists of a main segmentation task and two auxiliary tasks, i.e. the feature regression task and target detection task. The two auxiliary tasks predict some relatively simpler image-level attributes and bounding boxes as the pseudo labels for the main segmentation task, enforcing the pixel-level segmentation result to match the distribution of these pseudo labels. In addition, to solve the problem of class imbalance in the images, a bounding-box-based attention (BBA) module is embedded, enabling the segmentation network to concern more about the target region rather than the background. Furthermore, to alleviate the adverse effects caused by the possible deviation of pseudo labels, error tolerance mechanisms are also adopted in the auxiliary tasks, including inequality constraint and bounding-box amplification. Our method is validated on ACDC2017 and PROMISE12 datasets. Experimental results demonstrate that compared with the full supervision method and state-of-the-art semi-supervised methods, our method yields a much better segmentation performance on a small labeled dataset. Code is available at https://github.com/DeepMedLab/MTCL. Kaiping Wang, Yan Wang 0015, Bo Zhan, Chen Zu, Xi Wu 0004, Jiliu Zhou, Dong Nie, Luping Zhou |
Int. J. Neural Syst. | 8 |
| 2022 | Explainable attention guided adversarial deep network for 3D radiotherapy dose distribution prediction
Huidong Li, Xingchen Peng, Jie Zeng 0003, Jianghong Xiao, Dong Nie, Chen Zu, Xi Wu 0004, Jiliu Zhou, Yan Wang 0015 |
Knowl. Based Syst. | 5 |
| 2022 | Unified medical image segmentation by learning from uncertainty in an end-to-end manner
Pin Tang, Pinli Yang, Dong Nie, Xi Wu 0004, Jiliu Zhou, Yan Wang 0015 |
Knowl. Based Syst. | 3 |
| 2022 | Semantic instance segmentation with discriminative deep supervision for medical images
Sihang Zhou 0001, Dong Nie, Ehsan Adeli-Mosabbeb, Xuhua Ren, Xinwang Liu 0002, En Zhu, Jianping Yin, Qian Wang 0001, Dinggang Shen |
Medical Image Anal. | 2 |
| 2021 | Automated Generation of Accurate & Fluent Medical X-ray ReportsabstractOur paper aims to automate the generation of medical reports from chest X-ray image inputs, a critical yet time-consuming task for radiologists.Existing medical report generation efforts emphasize producing human-readable reports, yet the generated text may not be well aligned to the clinical facts.Our generated medical reports, on the other hand, are fluent and, more importantly, clinically accurate.This is achieved by our fully differentiable and end-to-end paradigm that contains three complementary modules: taking the chest X-ray images and clinical history document of patients as inputs, our classification module produces an internal checklist of disease-related topics, referred to as enriched disease embedding; the embedding representation is then passed to our transformer-based generator, to produce the medical report; meanwhile, our generator also creates a weighted embedding representation, which is fed to our interpreter to ensure consistency with respect to diseaserelated topics.Empirical evaluations demonstrate very promising results achieved by our approach on commonly-used metrics concerning language fluency and clinical accuracy.Moreover, noticeable performance gains are consistently observed when additional input information is available, such as the clinical document and extra scans from different views. Hoang T. N. Nguyen, Dong Nie, Taivanbat Badamdorj, Yingying Zhu 0004, Jason Truong |
EMNLP (1) | 2 |
| 2021 | Edge-preserving MRI image synthesis via adversarial network with iterative multi-scale fusion
Yanmei Luo, Dong Nie, Bo Zhan, Xi Wu 0004, Jiliu Zhou, Yan Wang 0015, Dinggang Shen |
Neurocomputing | 2 |
| 2021 | Cascaded MultiTask 3-D Fully Convolutional Networks for Pancreas SegmentationabstractAutomatic pancreas segmentation is crucial to the diagnostic assessment of diabetes or pancreatic cancer. However, the relatively small size of the pancreas in the upper body, as well as large variations of its location and shape in retroperitoneum, make the segmentation task challenging. To alleviate these challenges, in this article, we propose a cascaded multitask 3-D fully convolution network (FCN) to automatically segment the pancreas. Our cascaded network is composed of two parts. The first part focuses on fast locating the region of the pancreas, and the second part uses a multitask FCN with dense connections to refine the segmentation map for fine voxel-wise segmentation. In particular, our multitask FCN with dense connections is implemented to simultaneously complete tasks of the voxel-wise segmentation and skeleton extraction from the pancreas. These two tasks are complementary, that is, the extracted skeleton provides rich information about the shape and size of the pancreas in retroperitoneum, which can boost the segmentation of pancreas. The multitask FCN is also designed to share the low- and mid-level features across the tasks. A feature consistency module is further introduced to enhance the connection and fusion of different levels of feature maps. Evaluations on two pancreas datasets demonstrate the robustness of our proposed method in correctly segmenting the pancreas in various settings. Our experimental results outperform both baseline and state-of-the-art methods. Moreover, the ablation study shows that our proposed parts/modules are critical for effective multitask learning. Jie Xue 0001, Kelei He, Dong Nie, Ehsan Adeli-Mosabbeb, Zhenshan Shi, Seong-Whan Lee, Yuanjie Zheng, Xiyu Liu 0001, Dengwang Li, Dinggang Shen |
IEEE Trans. Cybern. | 3 |
| 2021 | HF-UNet: Learning Hierarchically Inter-Task Relevance in Multi-Task U-Net for Accurate Prostate Segmentation in CT ImagesabstractAccurate segmentation of the prostate is a key step in external beam radiation therapy treatments. In this paper, we tackle the challenging task of prostate segmentation in CT images by a two-stage network with 1) the first stage to fast localize, and 2) the second stage to accurately segment the prostate. To precisely segment the prostate in the second stage, we formulate prostate segmentation into a multi-task learning framework, which includes a main task to segment the prostate, and an auxiliary task to delineate the prostate boundary. Here, the second task is applied to provide additional guidance of unclear prostate boundary in CT images. Besides, the conventional multi-task deep networks typically share most of the parameters (i.e., feature representations) across all tasks, which may limit their data fitting ability, as the specificity of different tasks are inevitably ignored. By contrast, we solve them by a hierarchically-fused U-Net structure, namely HF-UNet. The HF-UNet has two complementary branches for two tasks, with the novel proposed attention-based task consistency learning block to communicate at each level between the two decoding branches. Therefore, HF-UNet endows the ability to learn hierarchically the shared representations for different tasks, and preserve the specificity of learned representations for different tasks simultaneously. We did extensive evaluations of the proposed method on a large planning CT image dataset and a benchmark prostate zonal dataset. The experimental results show HF-UNet outperforms the conventional multi-task network architectures and the state-of-the-art methods. Kelei He, Chunfeng Lian, Bing Zhang 0012, Xin Zhang 0013, Xiaohuan Cao, Dong Nie, Yang Gao 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 6 |
| 2020 | Hybrid Graph Neural Networks for Crowd CountingabstractCrowd counting is an important yet challenging task due to the large scale and density variation. Recent investigations have shown that distilling rich relations among multi-scale features and exploiting useful information from the auxiliary task, i.e., localization, are vital for this task. Nevertheless, how to comprehensively leverage these relations within a unified network architecture is still a challenging problem. In this paper, we present a novel network structure called Hybrid Graph Neural Network (HyGnn) which targets to relieve the problem by interweaving the multi-scale features for crowd density as well as its auxiliary task (localization) together and performing joint reasoning over a graph. Specifically, HyGnn integrates a hybrid graph to jointly represent the task-specific feature maps of different scales as nodes, and two types of relations as edges: (i) multi-scale relations capturing the feature dependencies across scales and (ii) mutual beneficial relations building bridges for the cooperation between counting and localization. Thus, through message passing, HyGnn can capture and distill richer relations between nodes to obtain more powerful representations, providing robust and accurate results. Our HyGnn performs significantly well on four challenging datasets: ShanghaiTech Part A, ShanghaiTech Part B, UCF_CC_50 and UCF_QNRF, outperforming the state-of-the-art algorithms by a large margin. Ao Luo, Fan Yang 0054, Xin Li 0079, Dong Nie, Zhicheng Jiao, Shangchen Zhou, Hong Cheng 0002 |
AAAI | 4 |
| 2020 | Bidirectional Pyramid Networks for Semantic Segmentation
Dong Nie, Jia Xue, Xiaofeng Ren |
ACCV (1) | 1 |
| 2020 | Adversarial Confidence Learning for Medical Image Segmentation and Synthesis
Dong Nie, Dinggang Shen |
Int. J. Comput. Vis. | 1 |
| 2020 | Task Decomposition and Synchronization for Semantic Biomedical Image SegmentationabstractSemantic segmentation is essentially important to biomedical image analysis. Many recent works mainly focus on integrating the Fully Convolutional Network (FCN) architecture with sophisticated convolution implementation and deep supervision. Such complex networks need large training datasets, a requirement which is challenging for medical image analysis. In this paper, we propose to decompose the single segmentation task into three subsequent sub-tasks, including (1) pixel-wise image semantic segmentation, (2) prediction of the instance class labels of the objects within the image, and (3) classification of the scene the image belonging to. While these three sub-tasks are trained to optimize their individual loss functions at different perceptual levels, we propose to allow their interaction within the task-task context ensemble. Moreover, we propose a novel sync-regularization to penalize the deviation between the outputs of the pixel-wise semantic segmentation and the instance class prediction tasks. These effective regularizations help FCN utilize context information comprehensively and attain accurate segmentation, even though the number of images for training may be limited in many biomedical applications. We have successfully applied our framework to three diverse 2D/3D medical image datasets, including Robotic Scene Segmentation Challenge 18 (ROBOT18), Brain Tumor Segmentation Challenge 18 (BRATS18), and Retinal Fundus Glaucoma Challenge (REFUGE18). We have achieved outperformed or comparable performance in all the three challenges. Our code, typical data and trained models are available athttps://github.com/xuhuaren/TDSNet. Xuhua Ren, Sahar Ahmad, Lichi Zhang, Lei Xiang 0001, Dong Nie, Fan Yang 0054, Qian Wang 0001, Dinggang Shen |
IEEE Trans. Image Process. | 5 |
| 2020 | High-Resolution Encoder-Decoder Networks for Low-Contrast Medical Image SegmentationabstractAutomatic image segmentation is an essential step for many medical image analysis applications, include computer-aided radiation therapy, disease diagnosis, and treatment effect evaluation. One of the major challenges for this task is the blurry nature of medical images (e.g., CT, MR and, microscopic images), which can often result in low-contrast and vanishing boundaries. With the recent advances in convolutional neural networks, vast improvements have been made for image segmentation, mainly based on the skip-connection-linked encoder-decoder deep architectures. However, in many applications (with adjacent targets in blurry images), these models often fail to accurately locate complex boundaries and properly segment tiny isolated parts. In this paper, we aim to provide a method for blurry medical image segmentation and argue that skip connections are not enough to help accurately locate indistinct boundaries. Accordingly, we propose a novel high-resolution multi-scale encoder-decoder network (HMEDN), in which multi-scale dense connections are introduced for the encoder-decoder structure to finely exploit comprehensive semantic information. Besides skip connections, extra deeply-supervised high-resolution pathways (comprised of densely connected dilated convolutions) are integrated to collect high-resolution semantic information for accurate boundary localization. These pathways are paired with a difficulty-guided cross-entropy loss function and a contour regression task to enhance the quality of boundary detection. Extensive experiments on a pelvic CT image dataset, a multi-modal brain tumor dataset, and a cell segmentation dataset show the effectiveness of our method for 2D/3D semantic segmentation and 2D instance segmentation, respectively. Our experimental results also show that besides increasing the network complexity, raising the resolution of semantic feature maps can largely affect the overall model performance. For different tasks, finding a balance between these two factors can further improve the performance of the corresponding network. Sihang Zhou 0001, Dong Nie, Ehsan Adeli-Mosabbeb, Jianping Yin, Jun Lian, Dinggang Shen |
IEEE Trans. Image Process. | 2 |
| 2020 | One-Shot Generative Adversarial Learning for MRI Segmentation of Craniomaxillofacial Bony StructuresabstractCompared to computed tomography (CT), magnetic resonance imaging (MRI) delineation of craniomaxillofacial (CMF) bony structures can avoid harmful radiation exposure. However, bony boundaries are blurry in MRI, and structural information needs to be borrowed from CT during the training. This is challenging since paired MRI-CT data are typically scarce. In this paper, we propose to make full use of unpaired data, which are typically abundant, along with a single paired MRI-CT data to construct a one-shot generative adversarial model for automated MRI segmentation of CMF bony structures. Our model consists of a cross-modality image synthesis sub-network, which learns the mapping between CT and MRI, and an MRI segmentation sub-network. These two sub-networks are trained jointly in an end-to-end manner. Moreover, in the training phase, a neighbor-based anchoring method is proposed to reduce the ambiguity problem inherent in cross-modality synthesis, and a feature-matching-based semantic consistency constraint is proposed to encourage segmentation-oriented MRI synthesis. Experimental results demonstrate the superiority of our method both qualitatively and quantitatively in comparison with the state-of-the-art MRI segmentation methods. Xu Chen 0020, James J. Xia, Dinggang Shen, Chunfeng Lian, Li Wang 0026, Hannah H. Deng, Steve H. Fung, Dong Nie, Kim-Han Thung, Pew-Thian Yap, Jaime Gateno |
IEEE Trans. Medical Imaging | 8 |
| 2020 | Multi-View Spatial Aggregation Framework for Joint Localization and Segmentation of Organs at Risk in Head and Neck CT ImagesabstractAccurate segmentation of organs at risk (OARs) from head and neck (H&N) CT images is crucial for effective H&N cancer radiotherapy. However, the existing deep learning methods are often not trained in an end-to-end fashion, i.e., they independently predetermine the regions of target organs before organ segmentation, causing limited information sharing between related tasks and thus leading to suboptimal segmentation results. Furthermore, when conventional segmentation network is used to segment all the OARs simultaneously, the results often favor big OARs over small OARs. Thus, the existing methods often train a specific model for each OAR, ignoring the correlation between different segmentation tasks. To address these issues, we propose a new multi-view spatial aggregation framework for joint localization and segmentation of multiple OARs using H&N CT images. The core of our framework is a proposed region-of-interest (ROI)-based fine-grained representation convolutional neural network (CNN), which is used to generate multi-OAR probability maps from each 2D view (i.e., axial, coronal, and sagittal view) of CT images. Specifically, our ROI-based fine-grained representation CNN (1) unifies the OARs localization and segmentation tasks and trains them in an end-to-end fashion, and (2) improves the segmentation results of various-sized OARs via a novel ROI-based fine-grained representation. Our multi-view spatial aggregation framework then spatially aggregates and assembles the generated multi-view multi-OAR probability maps to segment all the OARs simultaneously. We evaluate our framework using two sets of H&N CT images and achieve competitive and highly robust segmentation performance for OARs of various sizes. Shujun Liang, Kim-Han Thung, Dong Nie, Yu Zhang 0064, Dinggang Shen |
IEEE Trans. Medical Imaging | 3 |
| 2020 | CT Male Pelvic Organ Segmentation via Hybrid Loss Network With Incomplete AnnotationabstractSufficient data with complete annotation is essential for training deep models to perform automatic and accurate segmentation of CT male pelvic organs, especially when such data is with great challenges such as low contrast and large shape variation. However, manual annotation is expensive in terms of both finance and human effort, which usually results in insufficient completely annotated data in real applications. To this end, we propose a novel deep framework to segment male pelvic organs in CT images with incomplete annotation delineated in a very user-friendly manner. Specifically, we design a hybrid loss network derived from both voxel classification and boundary regression, to jointly improve the organ segmentation performance in an iterative way. Moreover, we introduce a label completion strategy to complete the labels of the rich unannotated voxels and then embed them into the training data to enhance the model capability. To reduce the computation complexity and improve segmentation performance, we locate the pelvic region based on salient bone structures to focus on the candidate segmentation organs. Experimental results on a large planning CT pelvic organ dataset show that our proposed method with incomplete annotation achieves comparable segmentation performance to the state-of-the-art methods with complete annotation. Moreover, our proposed method requires much less effort of manual contouring from medical professionals such that an institutional specific model can be more easily established. Shuai Wang 0003, Dong Nie, Liangqiong Qu, Yeqin Shao, Jun Lian, Qian Wang 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 2 |
| 2019 | Difficulty-Aware Attention Network with Confidence Learning for Medical Image SegmentationabstractMedical image segmentation is a key step for various applications, such as image-guided radiation therapy and diagnosis. Recently, deep neural networks provided promising solutions for automatic image segmentation; however, they often perform good on regular samples (i.e., easy-to-segment samples), since the datasets are dominated by easy and regular samples. For medical images, due to huge inter-subject variations or disease-specific effects on subjects, there exist several difficult-to-segment cases that are often overlooked by the previous works. To address this challenge, we propose a difficulty-aware deep segmentation network with confidence learning for end-to-end segmentation. The proposed framework has two main contributions: 1) Besides the segmentation network, we also propose a fully convolutional adversarial network for confidence learning to provide voxel-wise and region-wise confidence information for the segmentation network. We relax the adversarial learning to confidence learning by decreasing the priority of adversarial learning, so that we can avoid the training imbalance between generator and discriminator. 2) We propose a difficulty-aware attention mechanism to properly handle hard samples or hard regions considering structural information, which may go beyond the shortcomings of focal loss. We further propose a fusion module to selectively fuse the concatenated feature maps in encoder-decoder architectures. Experimental results on clinical and challenge datasets show that our proposed network can achieve state-of-the-art segmentation accuracy. Further analysis also indicates that each individual component of our proposed network contributes to the overall performance improvement. Dong Nie, Li Wang 0026, Lei Xiang 0001, Sihang Zhou 0001, Ehsan Adeli-Mosabbeb, Dinggang Shen |
AAAI | 1 |
| 2019 | RCA-U-Net: Residual Channel Attention U-Net for Fast Tissue Quantification in Magnetic Resonance Fingerprinting
Zhenghan Fang, Yong Chen 0026, Dong Nie, Weili Lin, Dinggang Shen |
MICCAI (3) | 3 |
| 2019 | Automatic brain labeling via multi-atlas guided fully convolutional networks
Longwei Fang, Lichi Zhang, Dong Nie, Xiaohuan Cao, Islem Rekik, Seong-Whan Lee, Huiguang He, Dinggang Shen |
Medical Image Anal. | 3 |
| 2019 | CT male pelvic organ segmentation using fully convolutional networks with boundary sensitive representation
Shuai Wang 0002, Kelei He, Dong Nie, Sihang Zhou 0001, Yaozong Gao, Dinggang Shen |
Medical Image Anal. | 3 |
| 2019 | 3-D Fully Convolutional Networks for Multimodal Isointense Infant Brain Image SegmentationabstractAccurate segmentation of infant brain images into different regions of interest is one of the most important fundamental steps in studying early brain development. In the isointense phase (approximately 6-8 months of age), white matter and gray matter exhibit similar levels of intensities in magnetic resonance (MR) images, due to the ongoing myelination and maturation. This results in extremely low tissue contrast and thus makes tissue segmentation very challenging. Existing methods for tissue segmentation in this isointense phase usually employ patch-based sparse labeling on single modality. To address the challenge, we propose a novel 3-D multimodal fully convolutional network (FCN) architecture for segmentation of isointense phase brain MR images. Specifically, we extend the conventional FCN architectures from 2-D to 3-D, and, rather than directly using FCN, we intuitively integrate coarse (naturally high-resolution) and dense (highly semantic) feature maps to better model tiny tissue regions, in addition, we further propose a transformation module to better connect the aggregating layers; we also propose a fusion module to better serve the fusion of feature maps. We compare the performance of our approach with several baseline and state-of-the-art methods on two sets of isointense phase brain images. The comparison results show that our proposed 3-D multimodal FCN model outperforms all previous methods by a large margin in terms of segmentation accuracy. In addition, the proposed framework also achieves faster segmentation results compared to all other methods. Our experiments further demonstrate that: 1) carefully integrating coarse and dense feature maps can considerably improve the segmentation performance; 2) batch normalization can speed up the convergence of the networks, especially when hierarchical feature aggregations occur; and 3) integrating multimodal information can further boost the segmentation performance. Dong Nie, Li Wang 0026, Ehsan Adeli-Mosabbeb, Cuijin Lao, Weili Lin, Dinggang Shen |
IEEE Trans. Cybern. | 1 |
| 2019 | Pelvic Organ Segmentation Using Distinctive Curve Guided Fully Convolutional NetworksabstractAccurate segmentation of pelvic organs (i.e., prostate, bladder, and rectum) from CT image is crucial for effective prostate cancer radiotherapy. However, it is a challenging task due to: 1) low soft tissue contrast in CT images and 2) large shape and appearance variations of pelvic organs. In this paper, we employ a two-stage deep learning-based method, with a novel distinctive curve-guided fully convolutional network (FCN), to solve the aforementioned challenges. Specifically, the first stage is for fast and robust organ detection in the raw CT images. It is designed as a coarse segmentation network to provide region proposals for three pelvic organs. The second stage is for fine segmentation of each organ, based on the region proposal results. To better identify those indistinguishable pelvic organ boundaries, a novel morphological representation, namely, distinctive curve, is also introduced to help better conduct the precise segmentation. To implement this, in this second stage, a multi-task FCN is initially utilized to learn the distinctive curve and the segmentation map separately and then combine these two tasks to produce accurate segmentation map. The final segmentation results of all three pelvic organs are generated by a weighted max-voting strategy. We have conducted exhaustive experiments on a large and diverse pelvic CT data set for evaluating our proposed method. The experimental results demonstrate that our proposed method is accurate and robust for this challenging segmentation task, by also outperforming the state-of-the-art segmentation methods. Kelei He, Xiaohuan Cao, Yinghuan Shi, Dong Nie, Yang Gao 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 4 |
| 2019 | Benchmark on Automatic Six-Month-Old Infant Brain Segmentation Algorithms: The iSeg-2017 ChallengeabstractAccurate segmentation of infant brain magnetic resonance (MR) images into white matter (WM), gray matter (GM), and cerebrospinal fluid (CSF) is an indispensable foundation for early studying of brain growth patterns and morphological changes in neurodevelopmental disorders. Nevertheless, in the isointense phase (approximately 6-9 months of age), due to inherent myelination and maturation process, WM and GM exhibit similar levels of intensity in both T1-weighted (T1w) and T2-weighted (T2w) MR images, making tissue segmentation very challenging. Despite many efforts were devoted to brain segmentation, only few studies have focused on the segmentation of 6-month infant brain images. With the idea of boosting methodological development in the community, iSeg-2017 challenge (http://iseg2017.web.unc.edu) provides a set of 6-month infant subjects with manual labels for training and testing the participating methods. Among the 21 automatic segmentation methods participating in iSeg-2017, we review the 8 top-ranked teams, in terms of Dice ratio, modified Hausdorff distance and average surface distance, and introduce their pipelines, implementations, as well as source codes. We further discuss limitations and possible future directions. We hope the dataset in iSeg-2017 and this review article could provide insights into methodological development for the community. Li Wang 0026, Dong Nie, Élodie Puybareau, Jose Dolz, Qian Zhang 0066, Fan Wang 0023, Zhengwang Wu, Jiawei Chen 0001, Kim-Han Thung, Toan Duc Bui, Jitae Shin, Guodong Zeng, Guoyan Zheng, Vladimir S. Fonov, Andrew Doyle, Yongchao Xu, Pim Moeskops, Josien P. W. Pluim, Christian Desrosiers, Ismail Ben Ayed, Gerard Sanroma, Oualid M. Benkarim, Adrià Casamitjana, Verónica Vilaplana, Weili Lin, Gang Li 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 2 |
| 2019 | STRAINet: Spatially Varying sTochastic Residual AdversarIal Networks for MRI Pelvic Organ SegmentationabstractAccurate segmentation of pelvic organs is important for prostate radiation therapy. Modern radiation therapy starts to use a magnetic resonance image (MRI) as an alternative to computed tomography image because of its superior soft tissue contrast and also free of risk from radiation exposure. However, segmentation of pelvic organs from MRI is a challenging problem due to inconsistent organ appearance across patients and also large intrapatient anatomical variations across treatment days. To address such challenges, we propose a novel deep network architecture, called "Spatially varying sTochastic Residual AdversarIal Network" (STRAINet), to delineate pelvic organs from MRI in an end-to-end fashion. Compared to the traditional fully convolutional networks (FCN), the proposed architecture has two main contributions: 1) inspired by the recent success of residual learning, we propose an evolutionary version of the residual unit, i.e., stochastic residual unit, and use it to the plain convolutional layers in the FCN. We further propose long-range stochastic residual connections to pass features from shallow layers to deep layers; and 2) we propose to integrate three previously proposed network strategies to form a new network for better medical image segmentation: a) we apply dilated convolution in the smallest resolution feature maps, so that we can gain a larger receptive field without overly losing spatial information; b) we propose a spatially varying convolutional layer that adapts convolutional filters to different regions of interest; and c) an adversarial network is proposed to further correct the segmented organ structures. Finally, STRAINet is used to iteratively refine the segmentation probability maps in an autocontext manner. Experimental results show that our STRAINet achieved the state-of-the-art segmentation accuracy. Further analysis also indicates that our proposed network components contribute most to the performance. Dong Nie, Li Wang 0026, Yaozong Gao, Jun Lian, Dinggang Shen |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2018 | ASDNet: Attention Based Semi-supervised Deep Networks for Medical Image Segmentation
Dong Nie, Yaozong Gao, Li Wang 0026, Dinggang Shen |
MICCAI (4) | 1 |
| 2018 | Volume-Based Analysis of 6-Month-Old Infant Brain MRI for Autism Biomarker Identification and Early Diagnosis
Li Wang 0026, Gang Li 0001, Feng Shi 0001, Xiaohuan Cao, Chunfeng Lian, Dong Nie, Mingxia Liu 0001, Han Zhang 0002, Zhengwang Wu, Weili Lin, Dinggang Shen |
MICCAI (3) | 6 |
| 2018 | Craniomaxillofacial Bony Structures Segmentation from MRI with Deep-Supervision Adversarial Learning
Miaoyun Zhao, Li Wang 0026, Jiawei Chen 0001, Dong Nie, Yulai Cong, Sahar Ahmad, Angela Ho, Peng Yuan 0001, Steve H. Fung, Hannah H. Deng, James J. Xia, Dinggang Shen |
MICCAI (4) | 4 |
| 2018 | Fine-Grained Segmentation Using Hierarchical Dilated Neural Networks
Sihang Zhou 0001, Dong Nie, Ehsan Adeli-Mosabbeb, Yaozong Gao, Li Wang 0026, Jianping Yin, Dinggang Shen |
MICCAI (4) | 2 |
| 2018 | Deep embedding convolutional neural network for synthesizing CT image from T1-Weighted MR image
Lei Xiang 0001, Qian Wang 0001, Dong Nie, Lichi Zhang, Xiyao Jin, Yu Qiao 0001, Dinggang Shen |
Medical Image Anal. | 3 |
| 2018 | Anatomical Landmark Based Deep Feature Representation for MR Images in Brain Disease DiagnosisabstractMost automated techniques for brain disease diagnosis utilize hand-crafted (e.g., voxel-based or region-based) biomarkers from structural magnetic resonance (MR) images as feature representations. However, these hand-crafted features are usually high-dimensional or require regions-of-interest defined by experts. Also, because of possibly heterogeneous property between the hand-crafted features and the subsequent model, existing methods may lead to sub-optimal performances in brain disease diagnosis. In this paper, we propose a landmark-based deep feature learning (LDFL) framework to automatically extract patch-based representation from MRI for automatic diagnosis of Alzheimer's disease. We first identify discriminative anatomical landmarks from MR images in a data-driven manner, and then propose a convolutional neural network for patch-based deep feature learning. We have evaluated the proposed method on subjects from three public datasets, including the Alzheimer's disease neuroimaging initiative (ADNI-1), ADNI-2, and the minimal interval resonance imaging in alzheimer's disease (MIRIAD) dataset. Experimental results of both tasks of brain disease classification and MR image retrieval demonstrate that the proposed LDFL method improves the performance of disease classification and MR image retrieval. Mingxia Liu 0001, Jun Zhang 0018, Dong Nie, Pew-Thian Yap, Dinggang Shen |
IEEE J. Biomed. Health Informatics | 3 |
| 2017 | Deformable Image Registration Based on Similarity-Steered CNN Regression
Xiaohuan Cao, Jianhua Yang 0005, Jun Zhang 0018, Dong Nie, Minjeong Kim 0001, Qian Wang 0001, Dinggang Shen |
MICCAI (1) | 4 |
| 2017 | Medical Image Synthesis with Context-Aware Generative Adversarial Networks
Dong Nie, Roger Trullo, Jun Lian, Caroline Petitjean, Su Ruan, Qian Wang 0001, Dinggang Shen |
MICCAI (3) | 1 |
| 2017 | Deep auto-context convolutional neural networks for standard-dose PET image estimation from low-dose PET/MRI
Lei Xiang 0001, Yu Qiao 0001, Dong Nie, Weili Lin, Qian Wang 0001, Dinggang Shen |
Neurocomputing | 3 |
| 2016 | 3D Deep Learning for Multi-modal Imaging-Guided Survival Time Prediction of Brain Tumor Patients
Dong Nie, Han Zhang 0002, Ehsan Adeli-Mosabbeb, Luyan Liu, Dinggang Shen |
MICCAI (2) | 1 |
| 2013 | Movie Recommendation Using Unrated DataabstractModel based movie recommender systems have been thoroughly investigated in the past few years, and they rely on rating data. In this paper, we take into account unrateddata of genre information to improve the performance of movie recommendation. We propose a novel method to measure users' preference on movie genres, and use Pearson Correlation Coefficient(PCC) to compute the user similarity. A matrix factorization framework is introduced for genre preference regularization. Experimental results on Movie Lens data set demonstrate that the approach performs well. Our method can also be used to increase the genre diversity of recommendations to some extent. Dong Nie, Lingzi Hong, Tingshao Zhu |
ICMLA (1) | 1 |