EDBT 2026 Demo / reviewers in the wild / expert
Xiao Han 0011
dblp:01/2095-11
· DBLP profile ↗
30ranked-venue papers
1as first author
22since 2021 · last 2026
0000-0002-5151-6547ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 9 since 2021Artificial intelligence and machine learning · 12 · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Label Hierarchy Transition: Delving Into Class Hierarchies to Enhance Deep ClassifiersabstractHierarchical classification aims to sort the object into a hierarchical structure of categories. For example, a bird can be categorized according to a three-level hierarchy of order, family, and species. Existing methods commonly address hierarchical classification by decoupling it into a series of multi-class classification tasks. However, such a multi-task learning strategy fails to fully exploit the correlation among various categories across different levels of the hierarchy. In this paper, we propose Label Hierarchy Transition (LHT), a unified probabilistic framework based on deep learning, to address the challenges of hierarchical classification. The LHT framework consists of a transition network and a confusion loss. The transition network focuses on explicitly learning the label hierarchy transition matrices, which has the potential to effectively encode the underlying correlations embedded within class hierarchies. The confusion loss encourages the classification network to learn correlations across different label hierarchies during training. The proposed framework can be readily adapted to any existing deep network with only minor modifications. We experiment with a series of public benchmark datasets for hierarchical classification problems, and the results demonstrate the superiority of our approach beyond current state-of-the-art methods. Furthermore, we extend our proposed LHT framework to the skin lesion diagnosis task and validate its great potential in computer-aided diagnosis. Renzhen Wang, De Cai, Kaiwen Xiao, Xixi Jia, Xiao Han 0011, Deyu Meng |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Boosting Consistency in Story Visualization with Rich-Contextual Conditional Diffusion ModelsabstractRecent research showcases the considerable potential of conditional diffusion models for generating consistent stories. However, current methods, which primarily generate stories in a caption-dependent manner, often overlook the importance of contextual consistency and the relevance of frames during sequential generation. To address this, we propose a novel Rich-contextual Conditional Diffusion Models (RCDMs), a two-stage approach designed to enhance story generation's semantic consistency and temporal consistency. Specifically, in the first stage, the frame-prior transformer diffusion model is presented to predict the frame semantic embedding of the unknown clip by aligning the semantic correlations between the captions and frames of the known clip. The second stage establishes a robust model with rich contextual conditions, including reference images of the known clip, the predicted frame semantic embedding of the unknown clip, and text embeddings of all captions. By jointly injecting these rich contextual conditions at the image and feature levels, RCDMs can generate semantic and temporal consistency stories. Moreover, RCDMs can generate consistent stories with a single forward inference compared to autoregressive models. Our qualitative and quantitative results demonstrate that our proposed RCDMs outperform in challenging scenarios. Fei Shen 0004, Hu Ye, Sibo Liu, Jun Zhang 0018, Cong Wang 0034, Xiao Han 0011 |
AAAI | 6 |
| 2025 | StdGEN: Semantic-Decomposed 3D Character Generation from Single ImagesabstractWe present StdGEN, an innovative pipeline for generating semantically decomposed high-quality 3D characters from single images, enabling broad applications in virtual reality, gaming, and filmmaking, etc. Unlike previous methods which struggle with limited decomposability, unsatisfactory quality, and long optimization times, StdGEN features decomposability, effectiveness and efficiency; i.e., it generates intricately detailed 3D characters with separated semantic components such as the body, clothes, and hair, in three minutes. At the core of StdGEN is our proposed Semantic-aware Large Reconstruction Model (S-LRM), a transformer-based generalizable model that jointly reconstructs geometry, color and semantics from multi-view images in a feed-forward manner. A differentiable multilayer semantic surface extraction scheme is introduced to acquire meshes from hybrid implicit fields reconstructed by our S-LRM. Additionally, a specialized efficient multi-view diffusion model and an iterative multi-layer surface refinement module are integrated into the pipeline to facilitate high-quality, decomposable 3D character generation. Extensive experiments demonstrate our state-of-theart performance in 3D anime character generation, surpassing existing baselines by a significant margin in geometry, texture and decomposability. StdGEN offers ready-touse semantic-decomposed 3D characters and enables flexible customization for a wide range of applications. Project page: https://stdgen.github.io Yanning Zhou 0003, Wang Zhao 0001, Zhongkai Wu, Kaiwen Xiao, Wei Yang 0013, Yong-Jin Liu 0001, Xiao Han 0011 |
CVPR | 8 |
| 2024 | SoftDedup: an Efficient Data Reweighting Method for Speeding Up Language Model Pre-trainingabstractNan He, Weichen Xiong, Hanwen Liu, Yi Liao, Lei Ding, Kai Zhang, Guohua Tang, Xiao Han, Yang Wei. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Weichen Xiong, Guohua Tang, Xiao Han 0011 |
ACL (1) | 8 |
| 2024 | Advancing Pose-Guided Image Synthesis with Progressive Conditional Diffusion ModelsabstractRecent work has showcased the significant potential of diffusion models in pose-guided person image synthesis.
However, owing to the inconsistency in pose between the source and target images, synthesizing an image with a distinct pose, relying exclusively on the source image and target pose information, remains a formidable challenge.
This paper presents Progressive Conditional Diffusion Models (PCDMs) that incrementally bridge the gap between person images under the target and source poses through three stages.
Specifically, in the first stage, we design a simple prior conditional diffusion model that predicts the global features of the target image by mining the global alignment relationship between pose coordinates and image appearance.
Then, the second stage establishes a dense correspondence between the source and target images using the global features from the previous stage, and an inpainting conditional diffusion model is proposed to further align and enhance the contextual features, generating a coarse-grained person image.
In the third stage, we propose a refining conditional diffusion model to utilize the coarsely generated image from the previous stage as a condition, achieving texture restoration and enhancing fine-detail consistency.
The three-stage PCDMs work progressively to generate the final high-quality and high-fidelity synthesized image.
Both qualitative and quantitative results demonstrate the consistency and photorealism of our proposed PCDMs under challenging scenarios.
The code and model will be available at https://github.com/tencent-ailab/PCDMs. Fei Shen 0004, Hu Ye, Jun Zhang 0018, Cong Wang 0034, Xiao Han 0011 |
ICLR | 5 |
| 2024 | VersVideo: Leveraging Enhanced Temporal Diffusion Models for Versatile Video GenerationabstractCreating stable, controllable videos is a complex task due to the need for significant variation in temporal dynamics and cross-frame temporal consistency. To address this, we enhance the spatial-temporal capability and introduce a versatile video generation model, VersVideo, which leverages textual, visual, and stylistic conditions. Current video diffusion models typically extend image diffusion architectures by supplementing 2D operations (such as convolutions and attentions) with temporal operations. While this approach is efficient, it often restricts spatial-temporal performance due to the oversimplification of standard 3D operations. To counter this, we incorporate two key elements: (1) multi-excitation paths for spatial-temporal convolutions with dimension pooling across different axes, and (2) multi-expert spatial-temporal attention blocks. These enhancements boost the model's spatial-temporal performance without significantly escalating training and inference costs. We also tackle the issue of information loss that arises when a variational autoencoder is used to transform pixel space into latent features and then back into pixel frames. To mitigate this, we incorporate temporal modules into the decoder to maintain inter-frame consistency. Lastly, by utilizing the innovative denoising UNet and decoder, we develop a unified ControlNet model suitable for various conditions, including image, Canny, HED, depth, and style. Examples of the videos generated by our model can be found at https://jinxixiang.github.io/versvideo/. Jinxi Xiang, Ricong Huang, Jun Zhang 0018, Guanbin Li, Xiao Han 0011 |
ICLR | 5 |
| 2023 | RLogist: Fast Observation Strategy on Whole-Slide Images with Deep Reinforcement LearningabstractWhole-slide images (WSI) in computational pathology have high resolution with gigapixel size, but are generally with sparse regions of interest, which leads to weak diagnostic relevance and data inefficiency for each area in the slide. Most of the existing methods rely on a multiple instance learning framework that requires densely sampling local patches at high magnification. The limitation is evident in the application stage as the heavy computation for extracting patch-level features is inevitable. In this paper, we develop RLogist, a benchmarking deep reinforcement learning (DRL) method for fast observation strategy on WSIs. Imitating the diagnostic logic of human pathologists, our RL agent learns how to find regions of observation value and obtain representative features across multiple resolution levels, without having to analyze each part of the WSI at the high magnification. We benchmark our method on two whole-slide level classification tasks, including detection of metastases in WSIs of lymph node sections, and subtyping of lung cancer. Experimental results demonstrate that RLogist achieves competitive classification performance compared to typical multiple instance learning algorithms, while having a significantly short observation path. In addition, the observation path given by RLogist provides good decision-making interpretability, and its ability of reading path navigation can potentially be used by pathologists for educational/assistive purposes. Our code is available at: https://github.com/tencent-ailab/RLogist. Boxuan Zhao, Jun Zhang 0018, Deheng Ye, Jian Cao 0001, Xiao Han 0011, Qiang Fu 0016, Wei Yang 0032 |
AAAI | 5 |
| 2023 | Dynamic Low-Rank Instance Adaptation for Universal Neural Image Compression
Yue Lv, Jinxi Xiang, Jun Zhang 0018, Wenming Yang, Xiao Han 0011, Wei Yang 0032 |
ACM Multimedia | 5 |
| 2023 | Towards Real-Time Neural Video Codec for Cross-Platform Application Using Calibration InformationabstractThe state-of-the-art neural video codecs have outperformed the most sophisticated traditional codecs in terms of rate-distortion (RD) performance in certain cases. However, utilizing them for practical applications is still challenging for two major reasons. 1) Cross-platform computational errors resulting from floating point operations can lead to inaccurate decoding of the bitstream. 2) The high computational complexity of the encoding and decoding process poses a challenge in achieving real-time performance. In this paper, we propose a real-time cross-platform neural video codec, which is capable of efficiently decoding (25FPS) of 720P video bitstream from other encoding platforms on a consumer-grade GPU (e.g., NVIDIA RTX 2080). First, to solve the problem of inconsistency of codec caused by the uncertainty of floating point calculations across platforms, we design a calibration transmitting system to guarantee the consistent quantization of entropy parameters between the encoding and decoding stages. The parameters that may have transboundary quantization between encoding and decoding are identified in the encoding stage, and their coordinates will be delivered by auxiliary transmitted bitstream. By doing so, these inconsistent parameters can be processed properly in the decoding stage. Furthermore, to reduce the bitrate of the auxiliary bitstream, we rectify the distribution of entropy parameters using a piecewise Gaussian constraint. Second, to match the computational limitations on the decoding side for real-time video codec, we design a lightweight model. A series of efficiency techniques, such as model pruning, motion downsampling, and arithmetic coding skipping, enable our model to achieve 25 FPS decoding speed on NVIDIA RTX 2080 GPU. Experimental results demonstrate that our model can achieve real-time decoding of 720P videos while encoding on another platform. Furthermore, the real-time model brings up to a maximum of 24.2% BD-rate improvement from the perspective of PSNR with the anchor H.265 (medium). Kuan Tian, Yonghang Guan, Jinxi Xiang, Jun Zhang 0018, Xiao Han 0011, Wei Yang 0032 |
ACM Multimedia | 5 |
| 2023 | CLC-Net: Contextual and local collaborative network for lesion segmentation in diabetic retinopathy images
Yuqi Fang, Sen Yang 0006, Delong Zhu 0001, Jing Zhang 0051, Jun Zhang 0018, Jun Cheng 0003, Raymond Kai-Yu Tong, Xiao Han 0011 |
Neurocomputing | 10 |
| 2023 | RetCCL: Clustering-guided contrastive learning for whole-slide image retrieval
Yuexi Du, Sen Yang 0006, Jun Zhang 0018, Jing Zhang 0051, Wei Yang 0032, Junzhou Huang, Xiao Han 0011 |
Medical Image Anal. | 9 |
| 2023 | A generalizable and robust deep learning algorithm for mitosis detection in multicenter breast histopathological images
Jun Zhang 0018, Sen Yang 0006, Jingxi Xiang, Feng Luo 0003, Jing Zhang 0051, Wei Yang 0032, Junzhou Huang, Xiao Han 0011 |
Medical Image Anal. | 10 |
| 2023 | Merging nucleus datasets by correlation-based cross-training
Jun Zhang 0018, Sen Yang 0006, Junzhou Huang, Wei Yang 0032, Xiao Han 0011 |
Medical Image Anal. | 8 |
| 2022 | Node-aligned Graph Convolutional Network for Whole-slide Image Representation and ClassificationabstractThe large-scale whole-slide images (WSIs) facilitate the learning-based computational pathology methods. However, the gigapixel size of WSIs makes it hard to train a conventional model directly. Current approaches typically adopt multiple-instance learning (MIL) to tackle this problem. Among them, MIL combined with graph convolutional network (GCN) is a significant branch, where the sampled patches are regarded as the graph nodes to further discover their correlations. However, it is difficult to build correspondence across patches from different WSIs. Therefore, most methods have to perform non-ordered node pooling to generate the bag-level representation. Direct non-ordered pooling will lose much structural and contextual information, such as patch distribution and heterogeneous patterns, which is critical for WSI representation. In this paper, we propose a hierarchical global-to-local clustering strategy to build a Node-Aligned GCN (NAGCN) to represent WSI with rich local structural information as well as global distribution. We first deploy a global clustering operation based on the instance features in the dataset to build the correspondence across different WSIs. Then, we perform a local clustering-based sampling strategy to select typical instances belonging to each cluster within the WSI. Finally, we employ the graph convolution to obtain the representation. Since our graph construction strategy ensures the alignment among different WSIs, WSI-level representation can be easily generated and used for the subsequent classification. The experiment results on two cancer subtype classification datasets demonstrate our method achieves better performance compared with the state-of-the-art methods. Yonghang Guan, Jun Zhang 0018, Kuan Tian, Sen Yang 0006, Pei Dong, Jinxi Xiang, Wei Yang 0032, Junzhou Huang, Yuyao Zhang 0005, Xiao Han 0011 |
CVPR | 10 |
| 2022 | SCL-WC: Cross-Slide Contrastive Learning for Weakly-Supervised Whole-Slide Image ClassificationabstractWeakly-supervised whole-slide image (WSI) classification (WSWC) is a challenging task where a large number of unlabeled patches (instances) exist within each WSI (bag) while only a slide label is given. Despite recent progress for the multiple instance learning (MIL)-based WSI analysis, the major limitation is that it usually focuses on the easy-to-distinguish diagnosis-positive regions while ignoring positives that occupy a small ratio in the entire WSI. To obtain more discriminative features, we propose a novel weakly-supervised classification method based on cross-slide contrastive learning (called SCL-WC), which depends on task-agnostic self-supervised feature pre-extraction and task-specific weakly-supervised feature refinement and aggregation for WSI-level prediction. To enable both intra-WSI and inter-WSI information interaction, we propose a positive-negative-aware module (PNM) and a weakly-supervised cross-slide contrastive learning (WSCL) module, respectively. The WSCL aims to pull WSIs with the same disease types closer and push different WSIs away. The PNM aims to facilitate the separation of tumor-like patches and normal ones within each WSI. Extensive experiments demonstrate state-of-the-art performance of our method in three different classification tasks (e.g., over 2% of AUC in Camelyon16, 5% of F1 score in BRACS, and 3% of AUC in DiagSet). Our method also shows superior flexibility and scalability in weakly-supervised localization and semi-supervised classification experiments (e.g., first place in the BRIGHT challenge). Our code will be available at https://github.com/Xiyue-Wang/SCL-WC. Jinxi Xiang, Jun Zhang 0018, Sen Yang 0006, Zhongyi Yang, Jing Zhang 0051, Wei Yang 0032, Junzhou Huang, Xiao Han 0011 |
NeurIPS | 10 |
| 2022 | Transformer-based unsupervised contrastive learning for histopathological image classification
Sen Yang 0006, Jun Zhang 0018, Jing Zhang 0051, Wei Yang 0032, Junzhou Huang, Xiao Han 0011 |
Medical Image Anal. | 8 |
| 2022 | Knowledge-Based Representation Learning for Nucleus Instance Classification From Histopathological ImagesabstractThe classification of nuclei in H&E-stained histopathological images is a fundamental step in the quantitative analysis of digital pathology. Most existing methods employ multi-class classification on the detected nucleus instances, while the annotation scale greatly limits their performance. Moreover, they often downplay the contextual information surrounding nucleus instances that is critical for classification. To explicitly provide contextual information to the classification model, we design a new structured input consisting of a content-rich image patch and a target instance mask. The image patch provides rich contextual information, while the target instance mask indicates the location of the instance to be classified and emphasizes its shape. Benefiting from our structured input format, we propose Structured Triplet for representation learning, a triplet learning framework on unlabelled nucleus instances with customized positive and negative sampling strategies. We pre-train a feature extraction model based on this framework with a large-scale unlabeled dataset, making it possible to train an effective classification model with limited annotated data. We also add two auxiliary branches, namely the attribute learning branch and the conventional self-supervised learning branch, to further improve its performance. As part of this work, we will release a new dataset of H&E-stained pathology images with nucleus instance masks, containing 20,187 patches of size 1024 ×1024 , where each patch comes from a different whole-slide image. The model pre-trained on this dataset with our framework significantly reduces the burden of extensive labeling. We show a substantial improvement in nucleus classification accuracy compared with the state-of-the-art methods. Jun Zhang 0018, Sen Yang 0006, Wei Yang 0032, Junzhou Huang, Xiao Han 0011 |
IEEE Trans. Medical Imaging | 8 |
| 2021 | Diagnose Like A Pathologist: Weakly-Supervised Pathologist-Tree Network for Slide-Level Immunohistochemical ScoringabstractThe immunohistochemistry (IHC) test of biopsy tissue is crucial to develop targeted treatment and evaluate prognosis for cancer patients. The IHC staining slide is usually digitized into the whole-slide image (WSI) with gigapixels for quantitative image analysis. To perform a whole image prediction (e.g., IHC scoring, survival prediction, and cancer grading) from this kind of high-dimensional image, algorithms are often developed based on multi-instance learning (MIL) framework. However, the multi-scale information of WSI and the associations among instances are not well explored in existing MIL based studies. Inspired by the fact that pathologists jointly analyze visual fields at multiple powers of objective for diagnostic predictions, we propose a Pathologist-Tree Network (PTree-Net) to sparsely model the WSI efficiently in multi-scale manner. Specifically, we propose a Focal-Aware Module (FAM) that can approximately estimate diagnosis-related regions with an extractor trained using the thumbnail of WSI. With the initial diagnosis-related regions, we hierarchically model the multi-scale patches in a tree structure, where both the global and local information can be captured. To explore this tree structure in an end-to-end network, we propose a patch Relevance-enhanced Graph Convolutional Network (RGCN) to explicitly model the correlations of adjacent parent-child nodes, accompanied by patch relevance to exploit the implicit contextual information among distant nodes. In addition, tree-based self-supervision is devised to improve representation learning and suppress irrelevant instances adaptively. Extensive experiments are performed on a large-scale IHC HER2 dataset. The ablation study confirms the effectiveness of our design, and our approach outperforms state-of-the-art by a large margin. Zhen Chen 0013, Jun Zhang 0018, Shuanlong Che, Junzhou Huang, Xiao Han 0011, Yixuan Yuan |
AAAI | 5 |
| 2021 | Minimizing Labeling Cost for Nuclei Instance Segmentation and Classification with Cross-domain Images and Weak LabelsabstractNucleus instance segmentation and classification in histopathological images is an essential prerequisite in pathology diagnosis/prognosis. However, nucleus annotations (e.g., segmentation and labeling) require domain experts, and annotating nuclei at pixel-level is time-consuming and labor-intensive. Moreover, nuclei from different cancer types vary in shapes and appearances. These inter-cancer variations require careful annotations for specific cancer types. Therefore, to minimize the labeling cost, we propose a novel application that considers each cancer type as an individual domain and apply domain adaptation techniques to improve the segmentation/classification performance among different cancer types. Unlike the previous studies that focus on unsupervised or weakly-supervised domain adaptation independently, we would like to discover what kinds of labeling can achieve the most cost-effective domain adaptation performance in nucleus instance segmentation and classification. Specifically, we propose a unified framework that is applicable to different level annotations: no annotations, image-level, and point-level annotations. Cyclic adaptation with pseudo labels and adversarial discriminator are utilized for unsupervised domain alignment. Image-level or point-level annotations are additionally adopted to supervise the nucleus classification and refine the pseudo labels. Experiments demonstrate the effectiveness and efficacy of the proposed framework (jointly using unsupervised and weakly supervised learning) on adapting the segmentation and classification model from one cancer type to 18 other cancer types. Siqi Yang 0001, Jun Zhang 0018, Junzhou Huang, Brian C. Lovell, Xiao Han 0011 |
AAAI | 5 |
| 2021 | TransPath: Transformer-Based Self-supervised Learning for Histopathological Image Classification
Sen Yang 0006, Jun Zhang 0018, Jing Zhang 0051, Junzhou Huang, Wei Yang 0032, Xiao Han 0011 |
MICCAI (8) | 8 |
| 2021 | A hybrid network for automatic hepatocellular carcinoma segmentation in H&E-stained whole slide images
Yuqi Fang, Sen Yang 0006, Delong Zhu 0001, Jing Zhang 0051, Raymond Kai-Yu Tong, Xiao Han 0011 |
Medical Image Anal. | 8 |
| 2021 | Joint fully convolutional and graph convolutional networks for weakly-supervised segmentation of pathology images
Jun Zhang 0018, Zhiyuan Hua, Kezhou Yan, Kuan Tian, Jianhua Yao 0001, Eryun Liu, Mingxia Liu 0001, Xiao Han 0011 |
Medical Image Anal. | 8 |
| 2020 | Deep Active Learning for Breast Cancer Segmentation on Immunohistochemistry Images
Haocheng Shen, Kuan Tian, Pei Dong, Jun Zhang 0018, Kezhou Yan, Shannon Che, Jianhua Yao 0001, Pifu Luo, Xiao Han 0011 |
MICCAI (5) | 9 |
| 2020 | Weakly-Supervised Nucleus Segmentation Based on Point Annotations: A Coarse-to-Fine Self-Stimulated Learning Strategy
Kuan Tian, Jun Zhang 0018, Haocheng Shen, Kezhou Yan, Pei Dong, Jianhua Yao 0001, Shannon Che, Pifu Luo, Xiao Han 0011 |
MICCAI (5) | 9 |
| 2019 | Rectified Cross-Entropy and Upper Transition Loss for Weakly Supervised Whole Slide Image Classifier
Hanbo Chen, Xiao Han 0011, Xinjuan Fan, Xiaoying Lou, Hailing Liu, Junzhou Huang, Jianhua Yao 0001 |
MICCAI (1) | 2 |
| 2019 | From Whole Slide Imaging to Microscopy: Deep Microscopy Adaptation Network for Histopathology Cancer Image Classification
Yifan Zhang 0004, Hanbo Chen, Ying Wei 0001, Peilin Zhao, Jiezhang Cao, Xinjuan Fan, Xiaoying Lou, Hailing Liu, Jinlong Hou, Xiao Han 0011, Jianhua Yao 0001, Qingyao Wu, Mingkui Tan, Junzhou Huang |
MICCAI (1) | 10 |
| 2019 | Enhanced Cycle-Consistent Generative Adversarial Network for Color Normalization of H&E Stained Images
Niyun Zhou, De Cai, Xiao Han 0011, Jianhua Yao 0001 |
MICCAI (1) | 3 |
| 2011 | Evaluation of Registration Methods on Thoracic CT: The EMPIRE10 ChallengeabstractEMPIRE10 (Evaluation of Methods for Pulmonary Image REgistration 2010) is a public platform for fair and meaningful comparison of registration algorithms which are applied to a database of intrapatient thoracic CT image pairs. Evaluation of nonrigid registration techniques is a nontrivial task. This is compounded by the fact that researchers typically test only on their own data, which varies widely. For this reason, reliable assessment and comparison of different registration algorithms has been virtually impossible in the past. In this work we present the results of the launch phase of EMPIRE10, which comprised the comprehensive evaluation and comparison of 20 individual algorithms from leading academic and industrial research groups. All algorithms are applied to the same set of 30 thoracic CT pairs. Algorithm settings and parameters are chosen by researchers expert in the configuration of their own method and the evaluation is independent, using the same criteria for all participants. All results are published on the EMPIRE10 website (http://empire10.isi.uu.nl). The challenge remains ongoing and open to new participants. Full results from 24 algorithms have been published at the time of writing. This paper details the organization of the challenge, the data and evaluation methods and the outcome of the initial launch with 20 algorithms. The gain in knowledge and future work are discussed. Keelin Murphy, Bram van Ginneken, Joseph M. Reinhardt, Sven Kabus, Kai Ding 0003, Kunlin Cao, Kaifang Du, Gary E. Christensen, Vincent Garcia, Tom Vercauteren, Nicholas Ayache, Olivier Commowick, Grégoire Malandain, Ben Glocker, Nikos Paragios, Nassir Navab, Vladlena Gorbunova, Jon Sporring, Marleen de Bruijne, Xiao Han 0011, Mattias P. Heinrich, Julia A. Schnabel, Mark Jenkinson, Cristian Lorenz, Marc Modat, Jamie McClelland, Sébastien Ourselin, Sascha E. A. Muenzing, Max A. Viergever, Dante De Nigris, D. Louis Collins, Tal Arbel, Marta Peroni, Rui Li 0053, Gregory C. Sharp, Alexander Schmidt-Richberg, Jan Ehrhardt, René Werner, Dirk Smeets, Dirk Loeckx, Gang Song, Nicholas J. Tustison, Brian B. Avants, James C. Gee, Marius Staring, Stefan Klein 0001, Berend C. Stoel, Martin Urschler, Manuel Werlberger, Jef Vandemeulebroucke, Simon Rit, David Sarrut, Josien P. W. Pluim |
IEEE Trans. Medical Imaging | 21 |
| 2007 | Atlas Renormalization for Improved Brain MR Image Segmentation Across Scanner PlatformsabstractAtlas-based approaches have demonstrated the ability to automatically identify detailed brain structures from 3-D magnetic resonance (MR) brain images. Unfortunately, the accuracy of this type of method often degrades when processing data acquired on a different scanner platform or pulse sequence than the data used for the atlas training. In this paper, we improve the performance of an atlas-based whole brain segmentation method by introducing an intensity renormalization procedure that automatically adjusts the prior atlas intensity model to new input data. Validation using manually labeled test datasets has shown that the new procedure improves the segmentation accuracy (as measured by the Dice coefficient) by 10% or more for several structures including hippocampus, amygdala, caudate, and pallidum. The results verify that this new procedure reduces the sensitivity of the whole brain segmentation method to changes in scanner platforms and improves its accuracy and robustness, which can thus facilitate multicenter or multisite neuroanatomical imaging studies. Xiao Han 0011, Bruce Fischl |
IEEE Trans. Medical Imaging | 1 |
| 2007 | Cortical Surface Shape Analysis Based on Spherical WaveletsabstractIn vivo quantification of neuroanatomical shape variations is possible due to recent advances in medical imaging and has proven useful in the study of neuropathology and neurodevelopment. In this paper, we apply a spherical wavelet transformation to extract shape features of cortical surfaces reconstructed from magnetic resonance images (MRIs) of a set of subjects. The spherical wavelet transformation can characterize the underlying functions in a local fashion in both space and frequency, in contrast to spherical harmonics that have a global basis set. We perform principal component analysis (PCA) on these wavelet shape features to study patterns of shape variation within normal population from coarse to fine resolution. In addition, we study the development of cortical folding in newborns using the Gompertz model in the wavelet domain, which allows us to characterize the order of development of large-scale and finer folding patterns independently. Given a limited amount of training data, we use a regularization framework to estimate the parameters of the Gompertz model to improve the prediction performance on new data. We develop an efficient method to estimate this regularized Gompertz model based on the Broyden-Fletcher-Goldfarb-Shannon (BFGS) approximation. Promising results are presented using both PCA and the folding development model in the wavelet domain. The cortical folding development model provides quantitative anatomic information regarding macroscopic cortical folding development and may be of potential use as a biomarker for early diagnosis of neurologic deficits in newborns. Patricia Ellen Grant, Xiao Han 0011, Florent Ségonne, Rudolph Pienaar, Evelina Busa, Jennifer L. Pacheco, Nikos Makris, Randy L. Buckner, Polina Golland, Bruce Fischl |
IEEE Trans. Medical Imaging | 4 |