Xiang Sean Zhou

dblp:z/XiangSeanZhou · also Xiang Zhou 0009 · DBLP profile ↗
← Back
60ranked-venue papers
15as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 42 · 12 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 26 · 6 since 2021Artificial intelligence and machine learning · 21 · 7 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author
YearPublicationVenuePosition
2026 Hierarchical Contrastive Learning for Precise Whole-Body Anatomical Localization in PET/CT Imaging
abstract
Automatic anatomical localization is critical for radiology report generation. While many studies focus on lesion detection and segmentation, anatomical localization-accurately describing lesion positions in radiology reports-has received less attention. Conventional segmentation-based methods are limited to organ-level localization and often fail in severe disease cases due to low segmentation accuracy. To address these limitations, we reformulate anatomical localization as an image-to-text retrieval task. Specifically, we propose a CLIP-based framework that aligns lesion image patches with anatomically descriptive text embeddings in a shared multimodal space. By projecting lesion features into the semantic space and retrieving the most relevant anatomical descriptions in a coarse-to-fine manner, our method achieves fine-grained lesion localization with high accuracy across the entire body. Our main contributions are as follows: (1) hierarchical anatomical retrieval, which organizes 387 locations into a two-level hierarchy, by retrieving from the first level of 124 coarse categories to narrow down the search space and reduce localization complexity; (2) augmented location descriptions, which integrate domain-specific anatomical knowledge for enhancing semantic representation and improving visual-text alignment; and (3) semi-hard negative sample mining, which improves training stability and discriminative learning by avoiding selecting the overly similar negative samples that may introduce label noise or semantic ambiguity. We validate our method on two whole-body PET/CT datasets, achieving an 84.13% localization accuracy on the internal test set and 80.42% on the external test set, with a per-lesion inference time of 34 ms. The proposed framework also demonstrated superior robustness in complex clinical cases compared to segmentation-based approaches.
Yaozong Gao, Yiran Shu, Mingyang Yu 0009, Yanbo Chen 0003, Jingyu Liu 0002, Shaonan Zhong, Weifang Zhang, Yiqiang Zhan, Xiang Sean Zhou, Xinlu Wang, Meixin Zhao, Dinggang Shen
IEEE Trans. Medical Imaging9
2025 Location-Guided Automated Lesion Captioning in Whole-Body PET/CT Images
Mingyang Yu 0009, Yaozong Gao, Yiran Shu, Yanbo Chen 0003, Jingyu Liu 0002, Caiwen Jiang, Kaicong Sun, Zhiming Cui 0001, Weifang Zhang, Yiqiang Zhan, Xiang Sean Zhou, Shaonan Zhong, Xinlu Wang, Meixin Zhao, Dinggang Shen
MICCAI (5)11
2024 Prompt-Based Segmentation Model of Anatomical Structures and Lesions in CT Images
Xi Ouyang, Dongdong Gu, Qianqian Chen 0002, Yiqiang Zhan, Xiang Sean Zhou, Feng Shi 0001, Zhong Xue, Dinggang Shen
MICCAI (8)7
2024 DTR-Net: Dual-Space 3D Tooth Model Reconstruction From Panoramic X-Ray Images
abstract
In digital dentistry, cone-beam computed tomography (CBCT) can provide complete 3D tooth models, yet suffers from a long concern of requiring excessive radiation dose and higher expense. Therefore, 3D tooth model reconstruction from 2D panoramic X-ray image is more cost-effective, and has attracted great interest in clinical applications. In this paper, we propose a novel dual-space framework, namely DTR-Net, to reconstruct 3D tooth model from 2D panoramic X-ray images in both image and geometric spaces. Specifically, in the image space, we apply a 2D-to-3D generative model to recover intensities of CBCT image, guided by a task-oriented tooth segmentation network in a collaborative training manner. Meanwhile, in the geometric space, we benefit from an implicit function network in the continuous space, learning using points to capture complicated tooth shapes with geometric properties. Experimental results demonstrate that our proposed DTR-Net achieves state-of-the-art performance both quantitatively and qualitatively in 3D tooth model reconstruction, indicating its potential application in dental practice.
Lanzhuju Mei, Yu Fang 0008, Yue Zhao 0012, Xiang Sean Zhou, Zhiming Cui 0001, Dinggang Shen
IEEE Trans. Medical Imaging4
2023 HC-Net: Hybrid Classification Network for Automatic Periodontal Disease Diagnosis
Lanzhuju Mei, Yu Fang 0008, Zhiming Cui 0001, Nizhuan Wang 0001, Xuming He 0001, Yiqiang Zhan, Xiang Sean Zhou, Maurizio Tonetti, Dinggang Shen
MICCAI (6)8
2022 Forecasting Human Trajectory from Scene History
abstract
Predicting the future trajectory of a person remains a challenging problem, due to randomness and subjectivity. However, the moving patterns of human in constrained scenario typically conform to a limited number of regularities to a certain extent, because of the scenario restrictions (\eg, floor plan, roads and obstacles) and person-person or person-object interactivity. Thus, an individual person in this scenario should follow one of the regularities as well. In other words, a person's subsequent trajectory has likely been traveled by others. Based on this hypothesis, we propose to forecast a person's future trajectory by learning from the implicit scene regularities. We call the regularities, inherently derived from the past dynamics of the people and the environment in the scene, \emph{scene history}. We categorize scene history information into two types: historical group trajectories and individual-surroundings interaction. To exploit these information for trajectory prediction, we propose a novel framework Scene History Excavating Network (SHENet), where the scene history is leveraged in a simple yet effective approach. In particular, we design two components, the group trajectory bank module to extract representative group trajectories as the candidate for future path, and the cross-modal interaction module to model the interaction between individual past trajectory and its surroundings for trajectory refinement, respectively. In addition, to mitigate the uncertainty in the evaluation, caused by the aforementioned randomness and subjectivity, we propose to include smoothness into evaluation metrics. We conduct extensive evaluations to validate the efficacy of proposed framework on ETH, UCY, as well as a new, challenging benchmark dataset PAV, demonstrating superior performance compared to state-of-the-art methods.
Mancheng Meng, Ziyan Wu 0001, Terrence Chen, Xiran Cai, Xiang Sean Zhou, Fan Yang 0054, Dinggang Shen
NeurIPS5
2021 Learning Hierarchical Attention for Weakly-Supervised Chest X-Ray Abnormality Localization and Diagnosis
abstract
We consider the problem of abnormality localization for clinical applications. While deep learning has driven much recent progress in medical imaging, many clinical challenges are not fully addressed, limiting its broader usage. While recent methods report high diagnostic accuracies, physicians have concerns trusting these algorithm results for diagnostic decision-making purposes because of a general lack of algorithm decision reasoning and interpretability. One potential way to address this problem is to further train these models to localize abnormalities in addition to just classifying them. However, doing this accurately will require a large amount of disease localization annotations by clinical experts, a task that is prohibitively expensive to accomplish for most applications. In this work, we take a step towards addressing these issues by means of a new attention-driven weakly supervised algorithm comprising a hierarchical attention mining framework that unifies activation- and gradient-based visual attention in a holistic manner. Our key algorithmic innovations include the design of explicit ordinal attention constraints, enabling principled model training in a weakly-supervised fashion, while also facilitating the generation of visual-attention-driven model explanations by means of localization cues. On two large-scale chest X-ray datasets (NIH ChestX-ray14 and CheXpert), we demonstrate significant localization performance improvements over the current state of the art while also achieving competitive classification performance.
Xi Ouyang, Srikrishna Karanam, Ziyan Wu 0001, Terrence Chen, Jiayu Huo, Xiang Sean Zhou, Qian Wang 0001, Jie-Zhi Cheng
IEEE Trans. Medical Imaging6
2019 Novel Iterative Attention Focusing Strategy for Joint Pathology Localization and Prediction of MCI Progression
Xiaodan Xing, Bin Xiao 0010, Quan Huo, Minqing Zhang, Xiang Sean Zhou, Yiqiang Zhan, Zhong Xue, Feng Shi 0001
MICCAI (4)8
2019 Multi-class Gradient Harmonized Dice Loss with Application to Knee MR Image Segmentation
Qin Liu 0004, Xiongfeng Tang, Deming Guo, Yanguo Qin, Yiqiang Zhan, Xiang Sean Zhou, Dijia Wu
MICCAI (6)7
2019 Weakly Supervised Segmentation Framework with Uncertainty: A Study on Pneumothorax Segmentation in Chest X-ray
Xi Ouyang, Zhong Xue, Yiqiang Zhan, Xiang Sean Zhou, Qian Wang 0001, Jie-Zhi Cheng
MICCAI (6)4
2019 Regression-Based Line Detection Network for Delineation of Largely Deformed Brain Midline
Xiangyu Tang, Minqing Zhang, Xiaodan Xing, Xiang Sean Zhou, Zhong Xue, Wenzhen Zhu, Zailiang Chen 0001, Feng Shi 0001
MICCAI (3)6
2019 Dynamic Spectral Graph Convolution Networks with Assistant Task Training for Early MCI Diagnosis
Xiaodan Xing, Minqing Zhang, Yiqiang Zhan, Xiang Sean Zhou, Zhong Xue, Feng Shi 0001
MICCAI (4)6
2018 Towards MR-Only Radiotherapy Treatment Planning: Synthetic CT Generation Using Multi-view Deep Convolutional Neural Networks
Yu Zhao 0007, Shu Liao, Yimo Guo, Liang Zhao 0018, Zhennan Yan, Sungmin Hong, Gerardo Hermosillo, Tianming Liu 0001, Xiang Sean Zhou, Yiqiang Zhan
MICCAI (1)9
2016 Automatic Lumbar Spondylolisthesis Measurement in CT Images
abstract
Lumbar spondylolisthesis is one of the most common spinal diseases. It is caused by the anterior shift of a lumbar vertebrae relative to subjacent vertebrae. In current clinical practices, staging of spondylolisthesis is often conducted in a qualitative way. Although meyerding grading opens the door to stage spondylolisthesis in a more quantitative way, it relies on the manual measurement, which is time consuming and irreproducible. Thus, an automatic measurement algorithm becomes desirable for spondylolisthesis diagnosis and staging. However, there are two challenges. 1) Accurate detection of the most anterior and posterior points on the superior and inferior surfaces of each lumbar vertebrae. Due to the small size of the vertebrae, slight errors of detection may lead to significant measurement errors, hence, wrong disease stages. 2) Automatic localize and label each lumbar vertebrae is required to provide the semantic meaning of the measurement. It is difficult since different lumbar vertebraes have high similarity of both shape and image appearance. To resolve these challenges, a new auto measurement framework is proposed with two major contributions: First, a learning based spine labeling method that integrates both the image appearance and spine geometry information is designed to detect lumbar vertebrae. Second, a hierarchical method using both the population information from atlases and domain-specific information in the target image is proposed for most anterior and posterior points positioning. Validated on 258 CT spondylolisthesis patients, our method shows very similar results to manual measurements by radiologists and significantly increases the measurement efficiency.
Shu Liao, Yiqiang Zhan, Zhongxing Dong, Ruyi Yan, Liyan Gong, Xiang Sean Zhou, Marcos Salganicoff, Jun Fei
IEEE Trans. Medical Imaging6
2016 Multi-Instance Deep Learning: Discover Discriminative Local Anatomies for Bodypart Recognition
abstract
In general image recognition problems, discriminative information often lies in local image patches. For example, most human identity information exists in the image patches containing human faces. The same situation stays in medical images as well. "Bodypart identity" of a transversal slice-which bodypart the slice comes from-is often indicated by local image information, e.g., a cardiac slice and an aorta arch slice are only differentiated by the mediastinum region. In this work, we design a multi-stage deep learning framework for image classification and apply it on bodypart recognition. Specifically, the proposed framework aims at: 1) discover the local regions that are discriminative and non-informative to the image classification problem, and 2) learn a image-level classifier based on these local regions. We achieve these two tasks by the two stages of learning scheme, respectively. In the pre-train stage, a convolutional neural network (CNN) is learned in a multi-instance learning fashion to extract the most discriminative and and non-informative local patches from the training slices. In the boosting stage, the pre-learned CNN is further boosted by these local patches for image classification. The CNN learned by exploiting the discriminative local appearances becomes more accurate than those learned from global image context. The key hallmark of our method is that it automatically discovers the discriminative and non-informative local patches through multi-instance deep learning. Thus, no manual annotation is required. Our method is validated on a synthetic dataset and a large scale CT dataset. It achieves better performances than state-of-the-art approaches, including the standard deep CNN.
Zhennan Yan, Yiqiang Zhan, Zhigang Peng, Shu Liao, Yoshihisa Shinagawa, Shaoting Zhang 0001, Dimitris N. Metaxas, Xiang Sean Zhou
IEEE Trans. Medical Imaging8
2015 A Steering Engine: Learning 3-D Anatomy Orientation Using Regression Forests
Fitsum A. Reda, Yiqiang Zhan, Xiang Sean Zhou
MICCAI (3)3
2013 Automated identification of thoracolumbar vertebrae using orthogonal matching pursuit
Bing Jian, Xiang Sean Zhou
Mach. Vis. Appl.3
2012 Robust MR Spine Detection Using Hierarchical Learning and Local Articulated Model
Yiqiang Zhan, Maneesh Dewan, Martin Harder, Xiang Sean Zhou
MICCAI (1)4
2012 Towards robust and effective shape modeling: Sparse shape composition
Shaoting Zhang 0001, Yiqiang Zhan, Maneesh Dewan, Junzhou Huang, Dimitris N. Metaxas, Xiang Sean Zhou
Medical Image Anal.6
2011 Sparse shape composition: A new framework for shape prior modeling
abstract
Image appearance cues are often used to derive object shapes, which is usually one of the key steps of image understanding tasks. However, when image appearance cues are weak or misleading, shape priors become critical to infer and refine the shape derived by these appearance cues. Effective modeling of shape priors is challenging because: 1) shape variation is complex and cannot always be modeled by a parametric probability distribution; 2) a shape instance derived from image appearance cues (input shape) may have gross errors; and 3) local details of the input shape are difficult to preserve if they are not statistically significant in the training data. In this paper we propose a novel Sparse Shape Composition model (SSC) to deal with these three challenges in a unified framework. In our method, training shapes are adaptively composed to infer/refine an input shape. The a-priori information is thus implicitly incorporated on-the-fly. Our model leverages two sparsity observations of the input shape instance: 1) the input shape can be approximately represented by a sparse linear combination of training shapes; 2) parts of the input shape may contain gross errors but such errors are usually sparse. Using L1 norm relaxation, our model is formulated as a convex optimization problem, which is solved by an efficient alternating minimization framework. Our method is extensively validated on two real world medical applications, 2D lung localization in X-ray images and 3D liver segmentation in low-dose CT scans. Compared to state-of-the-art methods, our model exhibits better performance in both studies.
Shaoting Zhang 0001, Yiqiang Zhan, Maneesh Dewan, Junzhou Huang, Dimitris N. Metaxas, Xiang Sean Zhou
CVPR6
2011 Deformable Segmentation via Sparse Shape Representation
Shaoting Zhang 0001, Yiqiang Zhan, Maneesh Dewan, Junzhou Huang, Dimitris N. Metaxas, Xiang Sean Zhou
MICCAI (2)6
2011 Robust Learning-Based Parsing and Annotation of Medical Radiographs
abstract
In this paper, we propose a learning-based algorithm for automatic medical image annotation based on robust aggregation of learned local appearance cues, achieving high accuracy and robustness against severe diseases, imaging artifacts, occlusion, or missing data. The algorithm starts with a number of landmark detectors to collect local appearance cues throughout the image, which are subsequently verified by a group of learned sparse spatial configuration models. In most cases, a decision could already be made at this stage by simply aggregating the verified detections. For the remaining cases, an additional global appearance filtering step is employed to provide complementary information to make the final decision. This approach is evaluated on a large-scale chest radiograph view identification task, demonstrating a very high accuracy ( > 99.9%) for a posteroanterior/anteroposterior (PA-AP) and lateral view position identification task, compared with the recently reported large-scale result of only 98.2% (Luo, , 2006). Our approach also achieved the best accuracies for a three-class and a multiclass radiograph annotation task, when compared with other state of the art algorithms. Our algorithm was used to enhance advanced image visualization workflows by enabling content-sensitive hanging-protocols and auto-invocation of a computer aided detection algorithm for identified PA-AP chest images. Finally, we show that the same methodology could be utilized for several image parsing applications including anatomy/organ region of interest prediction and optimized image visualization.
Yimo Tao, Zhigang Peng, Arun Krishnan, Xiang Sean Zhou
IEEE Trans. Medical Imaging4
2011 Robust Automatic Knee MR Slice Positioning Through Redundant and Hierarchical Anatomy Detection
abstract
Diagnostic magnetic resonance (MR) image quality is highly dependent on the position and orientation of the slice groups, due to the intrinsic high in-slice and low through-slice resolutions of MR imaging. Hence, the higher speed, accuracy, and reproducibility of automatic slice positioning, make it highly desirable over manual slice positioning. However, imaging artifacts, diseases, joint articulation, variations across ages and demographics as well as the extremely high performance requirements prevent state-of-the-art methods, such as volumetric registration, to be an off-the-shelf solution. In this paper, we address all these issues through an automatic slice positioning framework based on redundant and hierarchical learning. Our method has two hallmarks that are specifically designed to achieve high robustness and accuracy. 1) A redundant set of anatomy detectors are learned to provide local appearance cues. These detections are pruned and assembled according to a distributed anatomy model, which captures group-wise spatial configurations among anatomy primitives. This strategy brings about a high level of robustness and works even if a large portion of the target is distorted, missing, or occluded. 2) The detectors are learned and invoked in a hierarchical fashion, with each local detection scheduled and iterated according to its intrinsic invariance property. This iterative alignment process is shown to dramatically improve alignment accuracy. The proposed system is extensively validated on a large dataset including 744 clinical MR scans. Compared to state-of-the-art methods, our method exhibits superior performance in terms of robustness, accuracy, and reproducibility. The methodology is general and can be applied to other anatomies and other imaging modalities.
Yiqiang Zhan, Maneesh Dewan, Martin Harder, Arun Krishnan, Xiang Sean Zhou
IEEE Trans. Medical Imaging5
2010 Hierarchical Segmentation and Identification of Thoracic Vertebra Using Learning-Based Edge Detection and Coarse-to-Fine Deformable Model
Le Lu 0001, Yiqiang Zhan, Xiang Sean Zhou, Marcos Salganicoff, Arun Krishnan
MICCAI (1)4
2009 Hierarchical learning for tubular structure parsing in medical imaging: A study on coronary arteries using 3D CT Angiography
abstract
Automatic coronary artery centerline extraction from 3D CT Angiography (CTA) has significant clinical importance for diagnosis of atherosclerotic heart disease. The focus of past literature is dominated by segmenting the complete coronary artery system as trees by computer. Though the labeling of different vessel branches (defined by their medical semantics) is much needed clinically, this task has been performed manually. In this paper, we propose a hierarchical machine learning approach to tackle the problem of tubular structure parsing in medical imaging. It has a progressive three-tiered classification process at volumetric voxel level, vessel segment level, and inter-segment level. Generative models are employed to project from low-level, ambiguous data to class-conditional probabilities; and discriminative classifiers are trained on the upper-level structural patterns of probabilities to label and parse the vessel segments. Our method is validated by experiments of detecting and segmenting clinically defined coronary arteries, from the initial noisy vessel segment networks generated by low-level heuristics-based tracing algorithms. The proposed framework is also generically applicable to other tubular structure parsing tasks.
Le Lu 0001, Jinbo Bi, Shipeng Yu, Zhigang Peng, Arun Krishnan, Xiang Sean Zhou
ICCV6
2009 Document Analysis Support for the Manual Auditing of Elections
abstract
Recent developments have resulted in dramatic changes in the way elections are conducted, both in the United States and around the world. Well-publicized flaws in the security of electronic voting systems have led to a push for the use of verifiable paper records in the election process. In this paper, we describe the application of document analysis techniques to facilitate the manual auditing of elections,both to assure the reliability of the final outcome as well as to help reconcile the differences that may arise between repeated scans of the same ballot. We show how techniques developed for document duplicate detection can be applied to this problem, and present experimental results that demonstrate the efficacy of our approach. Related issues concerning machine support for the auditing of elections are also discussed.
Daniel P. Lopresti, Xiang Sean Zhou, Sharon X. Huang, Gang Tan
ICDAR2
2009 Cross Modality Deformable Segmentation Using Hierarchical Clustering and Learning
Yiqiang Zhan, Maneesh Dewan, Xiang Sean Zhou
MICCAI (1)3
2008 Active Scheduling of Organ Detection and Segmentation in Whole-Body Medical Images
Yiqiang Zhan, Xiang Sean Zhou, Zhigang Peng, Arun Krishnan
MICCAI (1)2
2007 Feature selection using principal feature analysis
abstract
Dimensionality reduction of a feature set is a common preprocessing step used for pattern recognition and classification applications. Principal Component Analysis (PCA) is one of the popular methods used, and can be shown to be optimal using different optimality criteria. However, it has the disadvantage that measurements from all the original features are used in the projection to the lower dimensional space. This paper proposes a novel method for dimensionality reduction of a feature set by choosing a subset of the original features that contains most of the essential information, using the same criteria as PCA. We call this method Principal Feature Analysis (PFA). The proposed method is successfully applied for choosing the principal features in face tracking and content-based image retrieval (CBIR) problems. Automated annotation of digital pictures has been a highly challenging problem for computer scientists since the invention of computers. The capability of annotating pictures by computers can lead to breakthroughs in a wide range of applications including Web image search, online picture-sharing communities, and scientific experiments. In our work, by advancing statistical modeling and optimization techniques, we can train computers about hundreds of semantic concepts using example pictures from each concept. The ALIPR (Automatic Linguistic Indexing of Pictures - Real Time) system of fully automatic and high speed annotation for online pictures has been constructed. Thousands of pictures from an Internet photo-sharing site, unrelated to the source of those pictures used in the training process, have been tested. The experimental results show that a single computer processor can suggest annotation terms in real-time and with good accuracy.
Yijuan Lu, Ira Cohen, Xiang Sean Zhou, Qi Tian 0001
ACM Multimedia3
2006 Probabilistic 3D Polyp Detection in CT Images: The Role of Sample Alignment
abstract
Automatic polyp detection is an increasingly important task in medical imaging with virtual colonoscopy [15] being widely used. In this paper, we present a 3D object detection algorithm and show its application on polyp detection from CT images. We make the following contributions: (1) The system adopts Probabilistic Boosting Tree (PBT) to probabilistically detect polyps. Integral volume and 3D Haar filters are introduced to achieve fast feature computation. (2) We give an explicit convergence rate analysis for the AdaBoost algorithm [2] and prove that the error at each step \in t+1. is tightly bounded by the previous error \in t. (3) For a 3D polyp template, a generative model is defined. Given the bound and convergence analysis, we analyze the role of "sample alignment" in the template design and devise a robust and efficient algorithm for polyp detection. The overall system has been tested on 150 volumes and the results obtained are very encouraging.
Zhuowen Tu, Xiang Sean Zhou, Luca Bogoni, Adrian Barbu, Dorin Comaniciu
CVPR (2)2
2006 Simultaneous Registration and Modeling of Deformable Shapes
abstract
Many natural objects vary the shapes as linear combinations of certain bases. The measurement of such deformable shapes is coupling of rigid similarity transformations between the objects and the measuring systems and non-rigid deformations controlled by the linear bases. Thus registration and modeling of deformable shapes are coupled problems, where registration is to compute the rigid transformations and modeling is to construct the linear bases. The previous methods [3, 2] separate the solution into two steps. The first step registers the measurements regarding the shapes as rigid and the deformations as random noise. The second step constructs the linear model using the registered shapes. Since the deformable shapes do not vary randomly but are constrained by the underlying model, such separate steps result in registration biased by nonrigid deformations and shape models involving improper rigid transformations. We for the first time present this bias problem and formulate that, the coupled registration and modeling problems are essentially a single factorization problem and thus require a simultaneous solution. We then propose the Direct Factorization method that extends a structure from motion method [16]. It yields a linear closedform solution that simultaneously registers the deformable shapes at arbitrary dimensions (2D \to 2D, 3D \to 3D, . . .) and constructs the linear bases. The accuracy and robustness of the proposed approach are demonstrated quantitatively on synthetic data and qualitatively on real shapes.
Jing Xiao 0006, Bogdan Georgescu, Xiang Sean Zhou, Dorin Comaniciu, Takeo Kanade
CVPR (2)3
2006 Database-Guided Simultaneous Multi-slice 3D Segmentation for Volumetric Data
Wei Hong 0003, Bogdan Georgescu, Xiang Sean Zhou, Sriram Krishnan, Yi Ma 0001, Dorin Comaniciu
ECCV (4)3
2006 A Learning Based Approach for 3D Segmentation and Colon Detagging
Zhuowen Tu, Xiang Sean Zhou, Dorin Comaniciu, Luca Bogoni
ECCV (3)2
2006 Example Based Non-rigid Shape Detection
Yefeng Zheng 0001, Xiang Sean Zhou, Bogdan Georgescu, Shaohua Kevin Zhou, Dorin Comaniciu
ECCV (4)2
2006 Automatic Hot Spot Detection and Segmentation in Whole Body FDG-PET Images
abstract
We present a system for automatic hot spots detection and segmentation in whole body FDG-PET images. The main contribution of our system is threefold. First, it has a novel body-section labeling module based on spatial hidden-Markov models (HMM); this allows different processing policies to be applied in different body sections. Second, the competition diffusion (CD) segmentation algorithm, which takes into account body-section information, converts the binary thresholding results to probabilistic interpretation and detects hot-spot region candidates. Third, a recursive intensity mode-seeking algorithm finds hot spot centers efficiently, and given these centers, a clinically meaningful protocol is proposed to accurately quantify hot spot volumes. Experimental results show that our system works robustly despite the large variations in clinical PET images.
Haiying Guan, Toshiro Kubota, Sharon X. Huang, Xiang Sean Zhou, Matthew Turk 0001
ICIP4
2006 Total Variation Models for Variable Lighting Face Recognition
abstract
In this paper, we present the logarithmic total variation (LTV) model for face recognition under varying illumination, including natural lighting conditions, where we rarely know the strength, direction, or number of light sources. The proposed LTV model has the ability to factorize a single face image and obtain the illumination invariant facial structure, which is then used for face recognition. Our model is inspired by the SQI model but has better edge-preserving ability and simpler parameter selection. The merit of this model is that neither does it require any lighting assumption nor does it need any training. The LTV model reaches very high recognition rates in the tests using both Yale and CMU PIE face databases as well as a face database containing 765 subjects under outdoor lighting conditions.
Terrence Chen, Wotao Yin, Xiang Sean Zhou, Dorin Comaniciu, Thomas S. Huang
IEEE Trans. Pattern Anal. Mach. Intell.3
2005 Illumination Normalization for Face Recognition and Uneven Background Correction Using Total Variation Based Image Models
abstract
We present a new algorithm for illumination normalization and uneven background correction in images, utilizing the recently proposed TV+L/sup 1/ model: minimizing the total variation of the output cartoon while subject to an L/sup 1/-norm fidelity term. We give intuitive proofs of its main advantages, including the well-known edge preserving capability, minimal signal distortion, and scale-dependent but intensity-independent foreground extraction. We then propose a novel TV-based quotient image model (TVQI) for illumination normalization, an important preprocessing for face recognition under different lighting conditions. Using this model, we achieve 100% face recognition rate on Yale face database B if the reference images are under good lighting condition and 99.45% if not. These results, compared to the average 65% recognition rate of the quotient image model and the average 95% recognition rate of the more recent self quotient image model, show a clear improvement. In addition, this model requires no training data, no assumption on the light source, and no alignment between different images for illumination normalization. We also present the results of the related applications - uneven background correction for cDNA mic roar ray films and digital microscope images. We believe the proposed works can serve important roles in the related fields.
Terrence Chen, Wotao Yin, Xiang Sean Zhou, Dorin Comaniciu, Thomas S. Huang
CVPR (2)3
2005 Database-Guided Segmentation of Anatomical Structures with Complex Appearance
abstract
The segmentation of anatomical structures has been traditionally formulated as a perceptual grouping task, and solved through clustering and variational approaches. However, such strategies require the a priori knowledge to be explicitly defined in the optimization criterion, e.g., "high-gradient border", "smoothness"', or "similar intensity or texture". This approach is limited by the validity of underlying assumptions and cannot capture complex structure appearance. This paper introduces database-guided segmentation as a new data-driven paradigm that directly exploits expert annotation of interest structures in large medical databases. Segmentation is formulated as a two-step learning problem. The first step is structure detection where we learn how to discriminate between the object of interest and background. The resulting classifier based on a boosted cascade of simple features also provides a global rigid transformation of the structure. The second step is shape inference where we use a sample-based representation of the joint distribution of appearance and shape annotations. To learn the association between the complex appearance and shape we propose a feature selection mechanism and the corresponding metric. We show that the selected features are better than using directly the appearance and illustrate the performance of the proposed method on a large set of ultrasound heart images.
Bogdan Georgescu, Xiang Sean Zhou, Dorin Comaniciu, Alok Gupta
CVPR (2)2
2005 Image Based Regression Using Boosting Method
abstract
We present a general algorithm of image based regression that is applicable to many vision problems. The proposed regressor that targets a multiple-output setting is learned using boosting method. We formulate a multiple-output regression problem in such a way that overfitting is decreased and an analytic solution is admitted. Because we represent the image via a set of highly redundant Haar-like features that can be evaluated very quickly and select relevant features through boosting to absorb the knowledge of the training data, during testing we require no storage of the training data and evaluate the regression function almost in no time. We also propose an efficient training algorithm that breaks the computational bottleneck in the greedy feature selection process. We validate the efficiency of the proposed regressor using three challenging tasks of age estimation, tumor detection, and endocardial wall localization and achieve the best performance with a dramatic speed, e.g., more than 1000 times faster than conventional data-driven techniques such as support vector regressor in the experiment of endocardial wall localization.
Shaohua Kevin Zhou, Bogdan Georgescu, Xiang Sean Zhou, Dorin Comaniciu
ICCV3
2005 Background correction for cDNA microarray images using the TV+L1 model
abstract
MOTIVATION: Background correction is an important preprocess in cDNA microarray data analysis. A variety of methods have been used for this purpose. However, many kinds of backgrounds, especially inhomogeneous ones, cannot be estimated correctly using any of the existing methods. In this paper, we propose the use of the TV+L1 model, which minimizes the total variation (TV) of the image subject to an L1-fidelity term, to correct background bias. We demonstrate its advantages over the existing methods by both analytically discussing its properties and numerically comparing it with morphological opening. RESULTS: Experimental results on both synthetic data and real microarray images demonstrate that the TV+L1 model gives the restored intensity that is closer to the true data than morphological opening. As a result, this method can serve an important role in the preprocessing of cDNA microarray data.
Wotao Yin, Terrence Chen, Xiang Sean Zhou, Amit Chakraborty
Bioinform.3
2005 An Information Fusion Framework for Robust Shape Tracking
abstract
Abstract-Existing methods for incorporating subspace model constraints in shape tracking use only partial information from the measurements and model distribution. We propose a unified framework for robust shape tracking, optimally fusing heteroscedastic uncertainties or noise from measurement, system dynamics, and a subspace model. The resulting nonorthogonal subspace projection and fusion are natural extensions of the traditional model constraint using orthogonal projection. We present two motion measurement algorithms and introduce alternative solutions for measurement uncertainty estimation. We build shape models offline from training data and exploit information from the ground truth initialization online through a strong model adaptation. Our framework is applied for tracking in echocardiograms where the motion estimation errors are heteroscedastic in nature, each heart has a distinct shape, and the relative motions of epicardial and endocardial borders reveal crucial diagnostic features. The proposed method significantly outperforms the existing shape-space-constrained tracking algorithm. Due to the complete treatment of heteroscedastic uncertainties, the strong model adaptation, and the coupled tracking of double-contours, robust performance is observed even on the most challenging cases.
Xiang Sean Zhou, Dorin Comaniciu, Alok Gupta
IEEE Trans. Pattern Anal. Mach. Intell.1
2004 A Unified Framework for Uncertainty Propagation in Automatic Shape Tracking
Xiang Sean Zhou, Dorin Comaniciu, Binglong Xie, R. Cruceanu, Alok Gupta
CVPR (1)1
2004 Coupled-Contour Tracking through Non-orthogonal Projections and Fusion for Echocardiography
Xiang Sean Zhou, Dorin Comaniciu, Sriram Krishnan
ECCV (1)1
2004 Real-Time Multi-model Tracking of Myocardium in Echocardiography Using Robust Information Fusion
Bogdan Georgescu, Xiang Sean Zhou, Dorin Comaniciu, R. Bharat Rao
MICCAI (2)2
2004 Robust real-time myocardial border tracking for echocardiography: an information fusion approach
abstract
Ultrasound is a main noninvasive modality for the assessment of the heart function. Wall tracking from ultrasound data is, however, inherently difficult due to weak echoes, clutter, poor signal-to-noise ratio, and signal dropouts. To cope with these artifacts, pretrained shape models can be applied to constrain the tracking. However, existing methods for incorporating subspace shape constraints in myocardial border tracking use only partial information from the model distribution, and do not exploit spatially varying uncertainties from feature tracking. In this paper, we propose a complete fusion formulation in the information space for robust shape tracking, optimally resolving uncertainties from the system dynamics, heteroscedastic measurement noise, and subspace shape model. We also exploit information from the ground truth initialization where this is available. The new framework is applied for tracking of myocardial borders in very noisy echocardiography sequences. Numerous myocardium tracking experiments validate the theory and show the potential of very accurate wall motion measurements. The proposed framework outperforms the traditional shape-space-constrained tracking algorithm by a significant margin. Due to the optimal fusion of different sources of uncertainties, robust performance is observed even for the most challenging cases.
Dorin Comaniciu, Xiang Sean Zhou, Sriram Krishnan
IEEE Trans. Medical Imaging2
2003 Classification Approach towards Banking and Sorting Problems
Shyamsundar Rajaram, Ashutosh Garg 0001, Xiang Sean Zhou, Thomas S. Huang
ECML3
2003 Conditional Feature Sensitivity: A Unifying View on Active Recognition and Feature Selection
abstract
The objective of active recognition is to iteratively collect the next "best" measurements (e.g., camera angles or viewpoints), to maximally reduce ambiguities in recognition. However, existing work largely overlooked feature interaction issues. Feature selection, on the other hand, focuses on the selection of a subset of measurements for a given classification task, but is not context sensitive (i.e., the decision does not depend on the current input). This paper proposes a unified perspective through conditional feature sensitivity analysis, taking into account both current context and feature interactions. Based on different representations of the contextual uncertainties, we present three treatment models and exploit their joint power for dealing with complex feature interactions. Synthetic examples are used to systematically test the validity of the proposed models. A practical application in medical domain is illustrated using an echocardiography database with more than 2000 video segments with both subjective (from experts) and objective validations.
Xiang Sean Zhou, Dorin Comaniciu, Arun Krishnan
ICCV1
2003 Relevance feedback in image retrieval: A comprehensive review
Xiang Sean Zhou, Thomas S. Huang
Multim. Syst.1
2002 Speeding up relevance feedback in image retrieval with triangle-inequality based algorithms
abstract
A content-based image retrieval(CBIR) system has been constructed to integrate relevance feedback with triangle-inequality based algorithms. The system offers typically 20 to 30 times faster retrieving speed with minimum sacrifice of retrieval performance on Corel database consisting of more than 17,000 images. The theoretic framework is built by using triangle-inequality based algorithms at sub-feature level and using relevance feedback techniques at feature level. Results show retrieval performance is clearly improved over the approach with only triangle-inequality based algorithms. A new high level weight updating method for the hierarchical distance model for relevance feedback is proposed.
Ziyou Xiong, Xiang Sean Zhou, William M. Pottenger, Thomas S. Huang
ICASSP2
2002 Optimal temporal sampling of video under channel and buffer constraints
abstract
Over low bit rate channels, we adopt the streaming of nonlinearly sampled video frames (i.e., key-frame slideshow) synchronized with the audio stream. Given the channel and buffer limits, we wish to obtain a set of sampled frames that is not only feasible (i.e., no frame drop) but also optimal in terms of maximal information flow. The contributions of this work include the novel modeling scheme for channel and buffer limits in the video temporal sampling problem; the development of the corresponding efficient algorithms for finding the global optimal solution; and the extension and analysis of these algorithms for practical application scenarios. The proposed algorithms have made possible the automated production of video streaming over low bit rate channels for devices with limited memory and storage capability.
Xiang Sean Zhou, Thomas S. Huang, Shih-Ping Liou
ICME (1)1
2002 Relevance feedback in content-based image retrieval: some recent advances
Xiang Sean Zhou, Thomas S. Huang
Inf. Sci.1
2002 Optimal nonlinear sampling for video streaming at low bit rates
abstract
Over low-bit-rate channels, we adopt the streaming of nonlinearly sampled video frames (i.e., key-frame slideshow) synchronized with the audio stream. Given the channel and buffer limits, we wish to obtain a set of sampled frames that is not only feasible (i.e., streamable), but also optimal in terms of maximal information flow (given that the semantic information contents of each frame can be quantified either automatically or manually). Different application scenarios are considered and modeled in a principle way, for which we propose computationally efficient algorithms for finding the global optimal solution. The contributions of this paper include the novel modeling schemes for channel and buffer limits in the video temporal sampling problem, the analysis and development of the corresponding efficient algorithms for finding the global optimal solution, and the extension and analysis of these algorithms for practical application scenarios. The proposed algorithms have made possible the automated production of the new form of video streaming over low bit rate channels for devices with limited storage capabilities.
Xiang Sean Zhou, Shih-Ping Liou
IEEE Trans. Circuits Syst. Video Technol.1
2001 Small Sample Learning during Multimedia Retrieval using BiasMap
abstract
All positive examples are alike; each negative example is negative in its own way. During interactive multimedia information retrieval, the number of training samples fed-back by the user is usually small; furthermore, they are not representative for the true distributions-especially the negative examples. Adding to the difficulties is the nonlinearity in real-world distributions. Existing solutions fail to address these problems in a principled way. This paper proposes biased discriminant analysis and transforms specifically designed to address the asymmetry between the positive and negative examples, and to trade off generalization for robustness under a small training sample. The kernel version, namely "BiasMap ", is derived to facilitate nonlinear biased discrimination. Extensive experiments are carried out for performance evaluation as compared to the state-of-the-art methods.
Xiang Sean Zhou, Thomas S. Huang
CVPR (1)1
2001 One-class SVM for learning in image retrieval
abstract
Relevance feedback schemes using linear/quadratic estimators have been applied in content-based image retrieval to improve retrieval performance significantly. One major difficulty in relevance feedback is to estimate the support of target images in high dimensional feature space with a relatively small number of training samples. We develop a novel scheme based on one-class SVM, which fits a tight hyper-sphere in the nonlinearly transformed feature space to include most of the target images based on positive examples. The use of a kernel provides us an elegant way to deal with nonlinearity in the distribution of the target images, while the regularization term in SVM provides good generalization ability. To validate the efficacy of the proposed approach, we test it on both synthesized data and real-world images. Promising results are achieved in both cases.
Yunqiang Chen, Xiang Sean Zhou, Thomas S. Huang
ICIP (1)2
2001 Image retrieval with relevance feedback: from heuristic weight adjustment to optimal learning methods
abstract
Various relevance feedback algorithms have been proposed in recent years in the area of content-based image retrieval. This paper gives a brief review and analysis on existing techniques-from early heuristic-based feature weighting schemes to recently proposed optimal learning algorithms. In addition, the kernel-based biased discriminant analysis (KBDA) is proposed to fit the unique nature of relevance feedback as a biased classification problem. As a novel variant of traditional discriminant analysis, the proposed algorithm provides a trade-off between discriminant transform and regression. The kernel form is derived to deal with non-linearity in an elegant way. Experimental results indicate that significant improvement in retrieval performance is achieved by the new scheme.
Xiang Sean Zhou, Thomas S. Huang
ICIP (3)1
2001 ICA-based probabilistic local appearance models
abstract
This paper proposes a novel image modeling scheme for object detection and localization. Object appearance is modeled by the joint distribution of k-tuple salient point feature vectors which are factorized component-wise after an independent component analysis (ICA). Also, we propose a distance-sensitive histograming technique for capturing spatial dependencies. The advantages over existing techniques include the ability to model non-rigid objects (at the expense of modeling accuracy) and the flexibility in modeling spatial relationships. Experiments show that ICA does improve modeling accuracy and detection performance. Experiments in object detection in cluttered scenes have demonstrated promising results.
Xiang Sean Zhou, Baback Moghaddam, Thomas S. Huang
ICIP (1)1
2001 Comparing discriminating transformations and SVM for learning during multimedia retrieval
abstract
On-line learning or techniques for multimedia information retrieval have been explored from many different points of view: from early heuristic-based feature weighting schemes to recently proposed optimal learning algorithms, probabilistic/Bayesian learning algorithms, boosting techniques, discriminant-EM algorithm, support vector machine, and other kernel-based learning machines. Based on a careful examination of the problem and a detailed analysis of the existing solutions, we propose several discriminating transforms as the learning machine during the user interaction. We argue that relevance feedback problem is best represented as a biased classification problem, or a (1+x)-class classification problem. Biased Discriminant Transform (BDT) is shown to outperform all the others. A kernel form is proposed to capture non-linearity in the class distributions.
Xiang Sean Zhou, Thomas S. Huang
ACM Multimedia1
2001 Edge-based structural features for content-based image retrieval
Xiang Sean Zhou, Thomas S. Huang
Pattern Recognit. Lett.1
2000 Image Representation and Retrieval Using Structural Features
abstract
This paper proposes structural features for image representation and retrieval. Among various structural features are the water-filling features, which can be efficiently extracted from edge maps to represent the edge/structural information in the image. The advantages of the feature extraction algorithm include efficiency-it is a linear-time algorithm; and effectiveness-multiple feature components corresponding to different human perceptions can be extracted simultaneously. Since edge maps are usually not scale invariant, also proposed are the multiscale feature extraction and cross-scale matching schemes. Experiments show that the new features can catch salient edge/structure information and cross-scale matching improves the retrieval performance in a real-world image retrieval system.
Xiang Sean Zhou, Thomas S. Huang
ICPR1
1999 Water-Filling: A Novel Way for Image Structural Feature Extraction
abstract
The performance of a content based image retrieval (CBIR) system is inherently constrained by the features adopted to represent the images in the database. In this paper, a new approach is proposed for image feature extraction based on edge maps. The feature vector with multiple feature components is computed through a “water-filling algorithm” applied on the edge map of the original image. The idea of this algorithm is to obtain measures of the edge length and complexity by graph traverse. The new feature is move generally applicable than texture or shape features. We call this structure feature. Experiments show that the new feature is capable of catching salient edge/structure information in the images. An experimental retrieval system utilizing the proposed new features yields better results in retrieving city/building images than some global texture features (wavelet moments). The new feature is ideal for images with clear edge structure. After combining the new features with other features in a relevance feedback framework, satisfactory retrieval results are observed.
Xiang Sean Zhou, Yong Rui, Thomas S. Huang
ICIP (2)1