Pingkun Yan

dblp:y/PingkunYan · DBLP profile ↗
← Back
109ranked-venue papers
13as first author
32since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 51 · 6 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 49 · 7 first-author · 28 since 2021Artificial intelligence and machine learning · 39 · 5 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Facial appearance prediction for orthognathic surgery with diffusion models
Jungwook Lee, Xuanang Xu, Daeseung Kim, Tianshu Kuang, Hannah H. Deng, Xinrui Song, Yasmine Soubra, Michael A. K. Liebschner, Jaime Gateno, Pingkun Yan
Medical Image Anal.10
2025 Facial Appearance Prediction with Conditional Multi-scale Autoregressive Modeling for Orthognathic Surgical Planning
Jungwook Lee, Xuanang Xu, Daeseung Kim, Tianshu Kuang, Hannah H. Deng, Xinrui Song, Yasmine Soubra, Rohan Dharia, Michael A. K. Liebschner, Jaime Gateno, Pingkun Yan
MICCAI (10)11
2025 Phrase-Grounded Fact-Checking for Automatically Generated Chest X-Ray Reports
Razi Mahmood, Diego Machado Reyes, Joy T. Wu, Parisa Kaviani, Ken C. L. Wong, Niharika D'Souza, Mannudeep K. Kalra, Ge Wang 0001, Pingkun Yan, Tanveer F. Syeda-Mahmood
MICCAI (7)9
2025 CXR-LT 2024: A MICCAI challenge on long-tailed, multi-label, and zero-shot disease classification from chest X-ray
Mingquan Lin, Gregory Holste, Song Wang 0026, Yiliang Zhou, Yishu Wei, Imon Banerjee, Pengyi Chen, Tianjie Dai, Yuexi Du, Nicha C. Dvornek, Yuyan Ge, Zuwei Guo, Shohei Hanaoka, Dongkyun Kim, Pablo Messina, Yang Lu 0009, Denis Parra, Donghyun Son, Alvaro Soto, Aisha Urooj Khan, René Vidal, Yosuke Yamagishi, Pingkun Yan, Zefan Yang, Ruichi Zhang, Yang Zhou 0019, Leo A. Celi, Ronald M. Summers, Zhiyong Lu, Hao Chen 0011, Adam E. Flanders, George Shih, Zhangyang Wang, Yifan Peng 0002
Medical Image Anal.23
2025 DINO-Reg: Efficient Multimodal Image Registration With Distilled Features
abstract
Medical image registration is a crucial process for aligning anatomical structures, enabling applications such as atlas mapping, longitudinal analysis, and multimodal data fusion. This paper introduces DINO-Reg, an adaptation-free registration method leveraging the vision foundation model, DINOv2, to extract features for deformable 3D medical image alignment. Although DINOv2 was originally trained on natural images, our study links the vision foundation model with medical image registration and demonstrates that the generic image encoder could readily generalize to medical images with state-of-the-art performance. We further propose DINO-Reg-Eco, a knowledge-distilled version using a UNet-structured 3D convolutional neural network (CNN) for feature extraction. The Eco model reduces encoding time by 99% while maintaining state-of-the-art performance, which is essential for resource-limited settings and significantly lowers the carbon footprint associated with intensive computational demands. Benchmarking across diverse datasets shows that both methods outperform existing supervised and unsupervised approaches without fine-tuning, demonstrating the transformative potential of foundation models in medical image registration. Our code is open-sourced at https://github.com/RPIDIAL/DINO-Reg.
Xinrui Song, Xuanang Xu, Jiajin Zhang, Diego Machado Reyes, Pingkun Yan
IEEE Trans. Medical Imaging5
2025 Chest X-Ray Foundation Model With Global and Local Representations Integration
abstract
Chest X-ray (CXR) is the most frequently ordered imaging test, supporting diverse clinical tasks from thoracic disease detection to postoperative monitoring. However, task-specific classification models are limited in scope, require costly labeled data, and lack generalizability to out-of-distribution datasets. To address these challenges, we introduce CheXFound, a self-supervised vision foundation model that learns robust CXR representations and generalizes effectively across a wide range of downstream tasks. We pretrained CheXFound on a curated CXR-987K dataset, comprising over approximately 987K unique CXRs from 12 publicly available sources. We propose a Global and Local Representations Integration (GLoRI) head for downstream adaptations, by incorporating fine- and coarse-grained disease-specific local features with global image features for enhanced performance in multilabel classification. Our experimental results showed that CheXFound outperformed state-of-the-art models in classifying 40 disease findings across different prevalence levels on the CXR-LT 24 dataset and exhibited superior label efficiency on downstream tasks with limited training data. Additionally, CheXFound achieved significant improvements on downstream tasks with out-of-distribution datasets, including opportunistic cardiovascular disease risk estimation, mortality prediction, malpositioned tube detection, and anatomical structure segmentation. The above results demonstrate CheXFound's strong generalization capabilities, which will enable diverse downstream adaptations with improved label efficiency in future applications. The project source code is publicly available at https://github.com/RPIDIAL/CheXFound.
Zefan Yang, Xuanang Xu, Jiajin Zhang, Ge Wang 0001, Mannudeep K. Kalra, Pingkun Yan
IEEE Trans. Medical Imaging6
2025 Disease-Informed Adaptation of Vision-Language Models
abstract
Expertise scarcity and high cost of data annotation hinder the development of artificial intelligence (AI) foundation models for medical image analysis. Transfer learning provides a way to utilize the off-the-shelf foundation models to address the clinical challenges. However, such models encounter difficulties when adapting to new diseases not presented in their original pre-training datasets. Compounding this challenge is the limited availability of example cases for a new disease, which further leads to the poor performance of the existing transfer learning techniques. This paper proposes a novel method for transfer learning of foundation Vision-Language Models (VLMs) to efficiently adapt them to a new disease with only a few examples. Such an effective adaptation of VLMs hinges on learning the nuanced representation of new disease concepts. By capitalizing on the joint visual-linguistic capabilities of VLMs, we introduce disease-informed contextual prompting in a novel disease prototype learning framework, which enables VLMs to quickly grasp the concept of the new disease, even with limited data. Extensive experiments across multiple pre-trained medical VLMs and multiple tasks showcase the notable enhancements in performance compared to other existing adaptation techniques. The code will be made publicly available at https://github.com/RPIDIAL/Disease-informed-VLM-Adaptation.
Jiajin Zhang, Ge Wang 0001, Mannudeep K. Kalra, Pingkun Yan
IEEE Trans. Medical Imaging4
2024 DINO-Reg: General Purpose Image Encoder for Training-Free Multi-modal Deformable Medical Image Registration
Xinrui Song, Xuanang Xu, Pingkun Yan
MICCAI (2)3
2024 DiRecT: Diagnosis and Reconstruction Transformer for Mandibular Deformity Assessment
Xuanang Xu, Jungwook Lee, Nathan Lampen, Daeseung Kim, Tianshu Kuang, Hannah H. Deng, Michael A. K. Liebschner, Jaime Gateno, Pingkun Yan
MICCAI (3)9
2024 Cardiovascular Disease Detection from Multi-view Chest X-Rays with BI-Mamba
Zefan Yang, Jiajin Zhang, Ge Wang 0001, Mannudeep K. Kalra, Pingkun Yan
MICCAI (5)5
2024 Disease-Informed Adaptation of Vision-Language Models
Jiajin Zhang, Ge Wang 0001, Mannudeep K. Kalra, Pingkun Yan
MICCAI (11)4
2024 Correspondence attention for facial appearance simulation
Xi Fang 0002, Daeseung Kim, Xuanang Xu, Tianshu Kuang, Nathan Lampen, Jungwook Lee, Hannah H. Deng, Michael A. K. Liebschner, James J. Xia, Jaime Gateno, Pingkun Yan
Medical Image Anal.11
2023 When Neural Networks Fail to Generalize? A Model Sensitivity Perspective
abstract
Domain generalization (DG) aims to train a model to perform well in unseen domains under different distributions. This paper considers a more realistic yet more challenging scenario, namely Single Domain Generalization (Single-DG), where only a single source domain is available for training. To tackle this challenge, we first try to understand when neural networks fail to generalize? We empirically ascertain a property of a model that correlates strongly with its generalization that we coin as "model sensitivity". Based on our analysis, we propose a novel strategy of Spectral Adversarial Data Augmentation (SADA) to generate augmented images targeted at the highly sensitive frequencies. Models trained with these hard-to-learn samples can effectively suppress the sensitivity in the frequency space, which leads to improved generalization performance. Extensive experiments on multiple public datasets demonstrate the superiority of our approach, which surpasses the state-of-the-art single-DG methods by up to 2.55%. The source code is available at https://github.com/DIAL-RPI/Spectral-Adversarial-Data-Augmentation.
Jiajin Zhang, Hanqing Chao, Amit Dhurandhar, Ali Tajer, Pingkun Yan
AAAI7
2023 Soft-Tissue Driven Craniomaxillofacial Surgical Planning
Xi Fang 0002, Daeseung Kim, Xuanang Xu, Tianshu Kuang, Nathan Lampen, Jungwook Lee, Hannah H. Deng, Jaime Gateno, Michael A. K. Liebschner, James J. Xia, Pingkun Yan
MICCAI (9)11
2023 Spatiotemporal Incremental Mechanics Modeling of Facial Tissue Change
Nathan Lampen, Daeseung Kim, Xuanang Xu, Xi Fang 0002, Jungwook Lee, Tianshu Kuang, Hannah H. Deng, Michael A. K. Liebschner, James J. Xia, Jaime Gateno, Pingkun Yan
MICCAI (9)11
2023 Spectral Adversarial MixUp for Few-Shot Unsupervised Domain Adaptation
Jiajin Zhang, Hanqing Chao, Amit Dhurandhar, Ali Tajer, Pingkun Yan
MICCAI (1)7
2023 Shape description losses for medical image segmentation
Xi Fang 0002, Xuanang Xu, James J. Xia, Thomas Sanford, Baris Turkbey, Sheng Xu 0001, Bradford J. Wood, Pingkun Yan
Mach. Vis. Appl.8
2023 Toward Adversarial Robustness in Unlabeled Target Domains
abstract
In the past several years, various adversarial training (AT) approaches have been invented to robustify deep learning model against adversarial attacks. However, mainstream AT methods assume the training and testing data are drawn from the same distribution and the training data are annotated. When the two assumptions are violated, existing AT methods fail because either they cannot pass knowledge learnt from a source domain to an unlabeled target domain or they are confused by the adversarial samples in that unlabeled space. In this paper, we first point out this new and challenging problem- adversarial training in unlabeled target domain. We then propose a novel framework named Unsupervised Cross-domain Adversarial Training (UCAT) to address this problem. UCAT effectively leverages the knowledge of the labeled source domain to prevent the adversarial samples from misleading the training process, under the guidance of automatically selected high quality pseudo labels of the unannotated target domain data together with the discriminative and robust anchor representations of the source domain data. The experiments on four public benchmarks show that models trained with UCAT can achieve both high accuracy and strong robustness. The effectiveness of the proposed components is demonstrated through a large set of ablation studies. The source code is publicly available at https://github.com/DIAL-RPI/UCAT.
Jiajin Zhang, Hanqing Chao, Pingkun Yan
IEEE Trans. Image Process.3
2023 Federated Multi-Organ Segmentation With Inconsistent Labels
abstract
Federated learning is an emerging paradigm allowing large-scale decentralized learning without sharing data across different data owners, which helps address the concern of data privacy in medical image analysis. However, the requirement for label consistency across clients by the existing methods largely narrows its application scope. In practice, each clinical site may only annotate certain organs of interest with partial or no overlap with other sites. Incorporating such partially labeled data into a unified federation is an unexplored problem with clinical significance and urgency. This work tackles the challenge by using a novel federated multi-encoding U-Net (Fed-MENU) method for multi-organ segmentation. In our method, a multi-encoding U-Net (MENU-Net) is proposed to extract organ-specific features through different encoding sub-networks. Each sub-network can be seen as an expert of a specific organ and trained for that client. Moreover, to encourage the organ-specific features extracted by different sub-networks to be informative and distinctive, we regularize the training of the MENU-Net by designing an auxiliary generic decoder (AGD). Extensive experiments on six public abdominal CT datasets show that our Fed-MENU method can effectively obtain a federated learning model using the partially labeled datasets with superior performance to other models trained by either localized or centralized learning methods. Source code is publicly available at https://github.com/DIAL-RPI/Fed-MENU.
Xuanang Xu, Hannah H. Deng, Jaime Gateno, Pingkun Yan
IEEE Trans. Medical Imaging4
2022 Regression Metric Loss: Learning a Semantic Representation Space for Medical Images
Hanqing Chao, Jiajin Zhang, Pingkun Yan
MICCAI (8)3
2022 Deep Learning-Based Facial Appearance Simulation Driven by Surgically Planned Craniomaxillofacial Bony Movement
Xi Fang 0002, Daeseung Kim, Xuanang Xu, Tianshu Kuang, Hannah H. Deng, Joshua C. Barber, Nathan Lampen, Jaime Gateno, Michael A. K. Liebschner, James J. Xia, Pingkun Yan
MICCAI (8)11
2022 Overlooked Trustworthiness of Saliency Maps
Jiajin Zhang, Hanqing Chao, Giridhar Dasegowda, Ge Wang 0001, Mannudeep K. Kalra, Pingkun Yan
MICCAI (3)6
2022 OASIS: One-pass aligned atlas set for medical image segmentation
Qikui Zhu, Bo Du 0001, Pingkun Yan
Neurocomputing4
2022 Cross-modal attention for multi-modal image registration
Xinrui Song, Hanqing Chao, Xuanang Xu, Hengtao Guo, Sheng Xu 0001, Baris Turkbey, Bradford J. Wood, Thomas Sanford, Ge Wang 0001, Pingkun Yan
Medical Image Anal.10
2022 Polar transform network for prostate ultrasound segmentation with uncertainty estimation
Xuanang Xu, Thomas Sanford, Baris Turkbey, Sheng Xu 0001, Bradford J. Wood, Pingkun Yan
Medical Image Anal.6
2022 Shadow-Consistent Semi-Supervised Learning for Prostate Ultrasound Segmentation
abstract
Prostate segmentation in transrectal ultrasound (TRUS) image is an essential prerequisite for many prostate-related clinical procedures, which, however, is also a long-standing problem due to the challenges caused by the low image quality and shadow artifacts. In this paper, we propose a Shadow-consistent Semi-supervised Learning (SCO-SSL) method with two novel mechanisms, namely shadow augmentation (Shadow-AUG) and shadow dropout (Shadow-DROP), to tackle this challenging problem. Specifically, Shadow-AUG enriches training samples by adding simulated shadow artifacts to the images to make the network robust to the shadow patterns. Shadow-DROP enforces the segmentation network to infer the prostate boundary using the neighboring shadow-free pixels. Extensive experiments are conducted on two large clinical datasets (a public dataset containing 1,761 TRUS volumes and an in-house dataset containing 662 TRUS volumes). In the fully-supervised setting, a vanilla U-Net equipped with our Shadow-AUG&Shadow-DROP outperforms the state-of-the-arts with statistical significance. In the semi-supervised setting, even with only 20% labeled training data, our SCO-SSL method still achieves highly competitive performance, suggesting great clinical value in relieving the labor of data annotation. Source code is released at https://github.com/DIAL-RPI/SCO-SSL.
Xuanang Xu, Thomas Sanford, Baris Turkbey, Sheng Xu 0001, Bradford J. Wood, Pingkun Yan
IEEE Trans. Medical Imaging6
2021 AnaXNet: Anatomy Aware Multi-label Finding Classification in Chest X-Ray
Nkechinyere Agu, Joy T. Wu, Hanqing Chao, Ismini Lourentzou, Arjun Sharma, Mehdi Moradi, Pingkun Yan, James A. Hendler
MICCAI (5)7
2021 End-to-end Ultrasound Frame to Volume Registration
Hengtao Guo, Xuanang Xu, Sheng Xu 0001, Bradford J. Wood, Pingkun Yan
MICCAI (4)5
2021 Cross-Modal Attention for MRI and Ultrasound Volume Registration
Xinrui Song, Hengtao Guo, Xuanang Xu, Hanqing Chao, Sheng Xu 0001, Baris Turkbey, Bradford J. Wood, Ge Wang 0001, Pingkun Yan
MICCAI (4)9
2021 Task-Oriented Low-Dose CT Image Denoising
Jiajin Zhang, Hanqing Chao, Xuanang Xu, Chuang Niu, Ge Wang 0001, Pingkun Yan
MICCAI (6)6
2021 Integrative analysis for COVID-19 patient outcome prediction
Hanqing Chao, Xi Fang 0002, Jiajin Zhang, Fatemeh Homayounieh, Chiara Daniela Arru, Subba R. Digumarthy, Rosa Babaei, Hadi Karimi Mobin, Iman Mohseni, Luca Saba, Alessandro Carriero, Zeno Falaschi, Alessio Pasche, Ge Wang 0001, Mannudeep K. Kalra, Pingkun Yan
Medical Image Anal.16
2021 Multi-Task Learning for Registering Images With Large Deformation
abstract
Accurate registration of prostate magnetic resonance imaging (MRI) images of the same subject acquired at different time points helps diagnose cancer and monitor the tumor progress. However, it is very challenging especially when one image was acquired with the use of endorectal coil (ERC) but the other was not, which causes significant deformation. Classical iterative image registration methods are also computationally intensive. Deep learning based registration frameworks have recently been developed and demonstrated promising performance. However, the lack of proper constraints often results in unrealistic registration. In this paper, we propose a multi-task learning based registration network with anatomical constraint to address these issues. The proposed approach uses a cycle constraint loss to achieve forward/backward registration and an inverse constraint loss to encourage diffeomorphic registration. In addition, an adaptive anatomical constraint aiming for regularizing the registration network with the use of anatomical labels is introduced through weak supervision. Our experiments on registering prostate MR images of the same subject obtained at different time points with and without ERC show that the proposed method achieves very promising performance under different measures in dealing with the large deformation. Compared with other existing methods, our approach works more efficiently with average running time less than a second and is able to obtain more visually realistic results.
Bo Du 0001, Jiandong Liao, Baris Turkbey, Pingkun Yan
IEEE J. Biomed. Health Informatics4
2020 Unsupervised Domain Adaptation with Dual-Scheme Fusion Network for Medical Image Segmentation
abstract
Domain adaptation aims to alleviate the problem of retraining a pre-trained model when applying it to a different domain, which requires large amount of additional training data of the target domain. Such an objective is usually achieved by establishing connections between the source domain labels and target domain data. However, this imbalanced source-to-target one way pass may not eliminate the domain gap, which limits the performance of the pre-trained model. In this paper, we propose an innovative Dual-Scheme Fusion Network (DSFN) for unsupervised domain adaptation. By building both source-to-target and target-to-source connections, this balanced joint information flow helps reduce the domain gap to further improve the network performance. The mechanism is further applied to the inference stage, where both the original input target image and the generated source images are segmented with the proposed joint network. The results are fused to obtain more robust segmentation. Extensive experiments of unsupervised cross-modality medical image segmentation are conducted on two tasks -- brain tumor segmentation and cardiac structures segmentation. The experimental results show that our method achieved significant performance improvement over other state-of-the-art domain adaptation methods.
Danbing Zou, Qikui Zhu, Pingkun Yan
IJCAI3
2020 Sensorless Freehand 3D Ultrasound Reconstruction via Deep Contextual Learning
Hengtao Guo, Sheng Xu 0001, Bradford J. Wood, Pingkun Yan
MICCAI (3)4
2020 Deep learning in medical image registration: a survey
Grant Haskins, Uwe Krüger 0001, Pingkun Yan
Mach. Vis. Appl.3
2020 Knowledge-Based Analysis for Mortality Prediction From CT Images
abstract
Low-Dose CT (LDCT) can significantly improve the accuracy of lung cancer diagnosis and thus reduce cancer deaths compared to chest X-ray. The lung cancer risk population is also at high risk of other deadly diseases, for instance, cardiovascular diseases. Therefore, predicting the all-cause mortality risks of this population is of great importance. This paper introduces a knowledge-based analytical method using deep convolutional neural network (CNN) for all-cause mortality prediction. The underlying approach combines structural image features extracted from CNNs, based on LDCT volume at different scales, and clinical knowledge obtained from quantitative measurements, to predict the mortality risk of lung cancer screening subjects. The proposed method is referred as Knowledge-based Analysis of Mortality Prediction Network (KAMP-Net). It constitutes a collaborative framework that utilizes both imaging features and anatomical information, instead of completely relying on automatic feature extraction. Our work demonstrates the feasibility of incorporating quantitative clinical measurements to assist CNNs in all-cause mortality prediction from chest LDCT images. The results of this study confirm that radiologist defined features can complement CNNs in performance improvement. The experiments demonstrate that KAMP-Net can achieve a superior performance when compared to other methods.
Hengtao Guo, Uwe Krüger 0001, Ge Wang 0001, Mannudeep K. Kalra, Pingkun Yan
IEEE J. Biomed. Health Informatics5
2020 Multi-Organ Segmentation Over Partially Labeled Datasets With Multi-Scale Feature Abstraction
abstract
Shortage of fully annotated datasets has been a limiting factor in developing deep learning based image segmentation algorithms and the problem becomes more pronounced in multi-organ segmentation. In this paper, we propose a unified training strategy that enables a novel multi-scale deep neural network to be trained on multiple partially labeled datasets for multi-organ segmentation. In addition, a new network architecture for multi-scale feature abstraction is proposed to integrate pyramid input and feature analysis into a U-shape pyramid structure. To bridge the semantic gap caused by directly merging features from different scales, an equal convolutional depth mechanism is introduced. Furthermore, we employ a deep supervision mechanism to refine the outputs in different scales. To fully leverage the segmentation features from all the scales, we design an adaptive weighting layer to fuse the outputs in an automatic fashion. All these mechanisms together are integrated into a Pyramid Input Pyramid Output Feature Abstraction Network (PIPO-FAN). Our proposed method was evaluated on four publicly available datasets, including BTCV, LiTS, KiTS and Spleen, where very promising performance has been achieved. The source code of this work is publicly shared at https://github.com/DIAL-RPI/PIPO-FAN to facilitate others to reproduce the work and build their own models using the introduced mechanisms.
Xi Fang 0002, Pingkun Yan
IEEE Trans. Medical Imaging2
2020 Boundary-Weighted Domain Adaptive Neural Network for Prostate MR Image Segmentation
abstract
Accurate segmentation of the prostate from magnetic resonance (MR) images provides useful information for prostate cancer diagnosis and treatment. However, automated prostate segmentation from 3D MR images faces several challenges. The lack of clear edge between the prostate and other anatomical structures makes it challenging to accurately extract the boundaries. The complex background texture and large variation in size, shape and intensity distribution of the prostate itself make segmentation even further complicated. Recently, as deep learning, especially convolutional neural networks (CNNs), emerging as the best performed methods for medical image segmentation, the difficulty in obtaining large number of annotated medical images for training CNNs has become much more pronounced than ever. Since large-scale dataset is one of the critical components for the success of deep learning, lack of sufficient training data makes it difficult to fully train complex CNNs. To tackle the above challenges, in this paper, we propose a boundary-weighted domain adaptive neural network (BOWDA-Net). To make the network more sensitive to the boundaries during segmentation, a boundary-weighted segmentation loss is proposed. Furthermore, an advanced boundary-weighted transfer leaning approach is introduced to address the problem of small medical imaging datasets. We evaluate our proposed model on three different MR prostate datasets. The experimental results demonstrate that the proposed model is more sensitive to object boundaries and outperformed other state-of-the-art methods.
Qikui Zhu, Bo Du 0001, Pingkun Yan
IEEE Trans. Medical Imaging3
2019 PASiam: Predicting Attention Inspired Siamese Network, for Space-Borne Satellite Video Tracking
abstract
Tracking a moving target of interests from a space-borne satellite video is really challenging. The difficulty lies in that the target usually occupies only several pixels, so that its features are very difficult to obtain. Besides, appearance features of the target would be unobvious when it is occluded, suffers from illumination variation influence or moves to similarity surroundings. In this paper, we propose a PREDICTING ATTENTION Inspired SIAMESE NETWORK (PASiam) for space-borne satellite video tracking, which constructs a fully convolutional Siamese network with shallow-layer features to obtain fine-grained appearance features. Moreover, a predicting attention is proposed to deal with occlusion and obscure. It employs Gaussian mixture models (GMM) to detect the target's motion status, and Kalman filter to predict and correct the target's location. Quantitative evaluations are performed on three real satellite video datasets. The results show our approach outperforms the state-of-the-art tracking methods while running at 54.83 FPS.
Jia Shao, Bo Du 0001, Chen Wu 0003, Pingkun Yan
ICME4
2019 MR Image Super-Resolution via Wide Residual Networks With Fixed Skip Connection
abstract
Spatial resolution is a critical imaging parameter in magnetic resonance imaging. The image super-resolution (SR) is an effective and cost efficient alternative technique to improve the spatial resolution of MR images. Over the past several years, the convolutional neural networks (CNN)-based SR methods have achieved state-of-the-art performance. However, CNNs with very deep network structures usually suffer from the problems of degradation and diminishing feature reuse, which add difficulty to network training and degenerate the transmission capability of details for SR. To address these problems, in this work, a progressive wide residual network with a fixed skip connection (named FSCWRN) based SR algorithm is proposed to reconstruct MR images, which combines the global residual learning and the shallow network based local residual learning. The strategy of progressive wide networks is adopted to replace deeper networks, which can partially relax the above-mentioned problems, while a fixed skip connection helps provide rich local details at high frequencies from a fixed shallow layer network to subsequent networks. The experimental results on one simulated MR image database and three real MR image databases show the effectiveness of the proposed FSCWRN SR algorithm, which achieves improved reconstruction performance compared with other algorithms.
Jun Shi 0004, Shihui Ying, Chaofeng Wang 0003, Qingping Liu, Qi Zhang 0003, Pingkun Yan
IEEE J. Biomed. Health Informatics7
2018 A Deep Learning Health Data Analysis Approach: Automatic 3D Prostate MR Segmentation with Densely-Connected Volumetric ConvNets
abstract
Automated prostate segmentation in 3D medical images play an important role in many clinical applications, such as diagnosis of prostatitis, prostate cancer and enlarged prostate. However, it is still a challenging task due to the complex background, lacking of clear boundary and various shape and texture between the slices. In this paper, we propose a novel 3D convolutional neural network with densely-connected layers to automatically segment the prostate from Magnetic Resonance(MR) images. Compared with other methods, our method has three compelling advantages. First, our model can effectively detect the prostate region in a volume-to-volume manner by utilizing the 3D convolution rather than the 3D convolution, which can fully exploit both spatial and region information. Second, the proposed network architecture alleviates the vanishing-gradient problem, strengthens the information propagation between layers, overcomes the problem of over-fitting and makes the network deeper by adopting a densely-connected manner. Third, besides the densely-connected manner inside each block, we also adopt the long connections strategy between blocks. We evaluate our proposed model on prostate dataset. The experimental results show that our model achieved significant segmentation results and outperformed other state-of-arts methods.
Qikui Zhu, Bo Du 0001, Jia Wu 0001, Pingkun Yan
IJCNN4
2018 Learning from Noisy Label Statistics: Detecting High Grade Prostate Cancer in Ultrasound Guided Biopsy
Shekoofeh Azizi, Pingkun Yan, Amir M. Tahmasebi, Peter A. Pinto, Bradford J. Wood, Jin Tae Kwak, Sheng Xu 0001, Baris Turkbey, Peter L. Choyke, Parvin Mousavi, Purang Abolmaesumi
MICCAI (4)2
2018 Shape prior constrained PSO model for bladder wall MRI segmentation
Qikui Zhu, Bo Du 0001, Pingkun Yan, Hongbing Lu, Liangpei Zhang 0001
Neurocomputing3
2018 Correlation-Based Tracking of Multiple Targets With Hierarchical Layered Structure
abstract
Visual target tracking is one of the most important research areas in the field of computer vision. Within this realm, multiple targets tracking (MTT) under complicated scene stands out for its great availability in real life applications, such as urban traffic surveillance and sports video analysis. However, in MTT, main difficulties arise from large variation in target saliency and significant motion heterogeneity, which may result in the failure of tracking weak targets. To tackle this challenge, a novel hierarchical layered tracking structure is proposed to perform tracking sequentially layer-by-layer. Upon this layered structure, we establish an intertarget mutual assistance mechanism on basis of intertarget correlation exploited among targets. The tracking results of a subset of targets can be utilized as additional prior information for tracking other targets. Specifically, a nonlinear motion model as well as a target interaction model basing on the intertarget correlation are proposed to effectively estimate the possible target region-of-interest to facilitate the prediction-based tracking. Moreover, the concept of motion entropy is introduced to quantitatively measure the degree of motion heterogeneity within the tracking scene for layer construction. Compared to other existing methods, extensive experiments demonstrated that the proposed method is capable of achieving higher tracking performance in complicated scenes, where targets are characterized with great heterogeneity.
Xianbin Cao 0001, Pingkun Yan
IEEE Trans. Cybern.4
2018 Deep Recurrent Neural Networks for Prostate Cancer Detection: Analysis of Temporal Enhanced Ultrasound
abstract
Temporal enhanced ultrasound (TeUS), comprising the analysis of variations in backscattered signals from a tissue over a sequence of ultrasound frames, has been previously proposed as a new paradigm for tissue characterization. In this paper, we propose to use deep recurrent neural networks (RNN) to explicitly model the temporal information in TeUS. By investigating several RNN models, we demonstrate that long short-term memory (LSTM) networks achieve the highest accuracy in separating cancer from benign tissue in the prostate. We also present algorithms for in-depth analysis of LSTM networks. Our in vivo study includes data from 255 prostate biopsy cores of 157 patients. We achieve area under the curve, sensitivity, specificity, and accuracy of 0.96, 0.76, 0.98, and 0.93, respectively. Our result suggests that temporal modeling of TeUS using RNN can significantly improve cancer detection accuracy over previously presented works.
Shekoofeh Azizi, Sharareh Bayat, Pingkun Yan, Amir M. Tahmasebi, Jin Tae Kwak, Sheng Xu 0001, Baris Turkbey, Peter L. Choyke, Peter A. Pinto, Bradford J. Wood, Parvin Mousavi, Purang Abolmaesumi
IEEE Trans. Medical Imaging3
2018 Low-Dose CT Image Denoising Using a Generative Adversarial Network With Wasserstein Distance and Perceptual Loss
abstract
The continuous development and extensive use of computed tomography (CT) in medical practice has raised a public concern over the associated radiation dose to the patient. Reducing the radiation dose may lead to increased noise and artifacts, which can adversely affect the radiologists' judgment and confidence. Hence, advanced image reconstruction from low-dose CT data is needed to improve the diagnostic performance, which is a challenging problem due to its ill-posed nature. Over the past years, various low-dose CT methods have produced impressive results. However, most of the algorithms developed for this application, including the recently popularized deep learning techniques, aim for minimizing the mean-squared error (MSE) between a denoised CT image and the ground truth under generic penalties. Although the peak signal-to-noise ratio is improved, MSE- or weighted-MSE-based methods can compromise the visibility of important structural details after aggressive denoising. This paper introduces a new CT image denoising method based on the generative adversarial network (GAN) with Wasserstein distance and perceptual similarity. The Wasserstein distance is a key concept of the optimal transport theory and promises to improve the performance of GAN. The perceptual loss suppresses noise by comparing the perceptual features of a denoised output against those of the ground truth in an established feature space, while the GAN focuses more on migrating the data noise distribution from strong to weak statistically. Therefore, our proposed method transfers our knowledge of visual perception to the image denoising task and is capable of not only reducing the image noise level but also trying to keep the critical information at the same time. Promising results have been obtained in our experiments with clinical CT images.
Qingsong Yang, Pingkun Yan, Hengyong Yu, Yongyi Shi, Xuanqin Mou, Mannudeep K. Kalra, Yi Zhang 0018, Ling Sun 0006, Ge Wang 0001
IEEE Trans. Medical Imaging2
2017 Deeply-supervised CNN for prostate segmentation
abstract
Prostate segmentation from Magnetic Resonance (MR) images plays an important role in image guided intervention. However, the lack of clear boundary specifically at the apex and base, and huge variation of shape and texture between the images from different patients make the task very challenging. To overcome these problems, in this paper, we propose a deeply supervised convolutional neural network (CNN) utilizing the convolutional information to accurately segment the prostate from MR images. The proposed model can effectively detect the prostate region with additional deeply supervised layers compared with other approaches. Since some information will be abandoned after convolution, it is necessary to pass the features extracted from early stages to later stages. The experimental results show that significant segmentation accuracy improvement has been achieved by our proposed method compared to other reported approaches.
Qikui Zhu, Bo Du 0001, Baris Turkbey, Peter L. Choyke, Pingkun Yan
IJCNN5
2016 Classifying Cancer Grades Using Temporal Ultrasound for Transrectal Prostate Biopsy
Shekoofeh Azizi, Farhad Imani, Jin Tae Kwak, Amir M. Tahmasebi, Sheng Xu 0001, Pingkun Yan, Jochen Kruecker, Baris Turkbey, Peter L. Choyke, Peter A. Pinto, Bradford J. Wood, Parvin Mousavi, Purang Abolmaesumi
MICCAI (1)6
2015 Feature competition and partial sparse shape modeling for cardiac image sequences segmentation
Xianjing Qin, Pingkun Yan
Neurocomputing3
2015 Partial sparse shape constrained sector-driven bladder wall segmentation
Xianjing Qin, Hongbing Lu, Pingkun Yan
Mach. Vis. Appl.4
2015 Label Image Constrained Multiatlas Selection
abstract
Multiatlas based method is commonly used in medical image segmentation. In multiatlas based image segmentation, atlas selection and combination are considered as two key factors affecting the performance. Recently, manifold learning based atlas selection methods have emerged as very promising methods. However, due to the complexity of prostate structures in raw images, it is difficult to get accurate atlas selection results by only measuring the distance between raw images on the manifolds. Although the distance between the regions to be segmented across images can be readily obtained by the label images, it is infeasible to directly compute the distance between the test image (gray) and the label images (binary). This paper tries to address this problem by proposing a label image constrained atlas selection method, which exploits the label images to constrain the manifold projection of raw images. Analyzing the data point distribution of the selected atlases in the manifold subspace, a novel weight computation method for atlas combination is proposed. Compared with other related existing methods, the experimental results on prostate segmentation from T2w MRI showed that the selected atlases are closer to the target structure and more accurate segmentation were obtained by using our proposed method.
Pingkun Yan, Yihui Cao, Yuan Yuan 0001, Baris Turkbey, Peter L. Choyke
IEEE Trans. Cybern.1
2014 Hierarchical incorporation of shape and shape dynamics for flying bird detection
Jun Zhang 0007, Qunyu Xu, Xianbin Cao 0001, Pingkun Yan, Xuelong Li 0001
Neurocomputing4
2014 Ego motion guided particle filter for vehicle tracking in airborne videos
Xianbin Cao 0001, Changcheng Gao, Jinhe Lan, Yuan Yuan 0001, Pingkun Yan
Neurocomputing5
2014 Guest Editorial: Special issue on advanced computing for image-guided intervention
Fei Zuo, Jungong Han, Pingkun Yan, Hans C. van Assen, Kenji Suzuki 0001
Neurocomputing3
2014 Alternatively Constrained Dictionary Learning For Image Superresolution
abstract
Dictionaries are crucial in sparse coding-based algorithm for image superresolution. Sparse coding is a typical unsupervised learning method to study the relationship between the patches of high-and low-resolution images. However, most of the sparse coding methods for image superresolution fail to simultaneously consider the geometrical structure of the dictionary and the corresponding coefficients, which may result in noticeable superresolution reconstruction artifacts. In other words, when a low-resolution image and its corresponding high-resolution image are represented in their feature spaces, the two sets of dictionaries and the obtained coefficients have intrinsic links, which has not yet been well studied. Motivated by the development on nonlocal self-similarity and manifold learning, a novel sparse coding method is reported to preserve the geometrical structure of the dictionary and the sparse coefficients of the data. Moreover, the proposed method can preserve the incoherence of dictionary entries and provide the sparse coefficients and learned dictionary from a new perspective, which have both reconstruction and discrimination properties to enhance the learning performance. Furthermore, to utilize the model of the proposed method more effectively for single-image superresolution, this paper also proposes a novel dictionary-pair learning method, which is named as two-stage dictionary training. Extensive experiments are carried out on a large set of images comparing with other popular algorithms for the same purpose, and the results clearly demonstrate the effectiveness of the proposed sparse representation model and the corresponding dictionary learning algorithm.
Xiaoqiang Lu, Yuan Yuan 0001, Pingkun Yan
IEEE Trans. Cybern.3
2014 Adaptive Shape Prior Constrained Level Sets for Bladder MR Image Segmentation
abstract
Three-dimensional bladder wall segmentation for thickness measuring can be very useful for bladder magnetic resonance (MR) image analysis, since thickening of the bladder wall can indicate abnormality. However, it is a challenging task due to the artifacts inside bladder lumen, weak boundaries in the apex and base areas, and complicated outside intensity distributions. To deal with these difficulties, in this paper, an adaptive shape prior constrained directional level set model is proposed to segment the inner and outer boundaries of the bladder wall. In addition, a coupled directional level set model is presented to refine the segmentation by exploiting the prior knowledge of region information and minimum thickness. With our proposed method, the influence of the artifacts in the bladder lumen and the complicated outside tissues surrounding the bladder can be appreciably reduced. Furthermore, the leakage on the weak boundaries can be avoided. Compared with other related methods, better results were obtained on 11 patients' 3-D bladder MR images by using the proposed method.
Xianjing Qin, Xuelong Li 0001, Yang Liu 0093, Hongbing Lu, Pingkun Yan
IEEE J. Biomed. Health Informatics5
2013 Global structure constrained local shape prior estimation for medical image segmentation
Pingkun Yan, Wuxia Zhang, Baris Turkbey, Peter L. Choyke, Xuelong Li 0001
Comput. Vis. Image Underst.1
2013 Tracking vehicles as groups in airborne videos
Xianbin Cao 0001, Zhengrong Shi, Pingkun Yan, Xuelong Li 0001
Neurocomputing3
2013 Pedestrian detection in unseen scenes by dynamically updating visual words
Xianbin Cao 0001, Bo Ning 0003, Yuan Yuan 0001, Pingkun Yan
Neurocomputing5
2013 Transfer learning for pedestrian detection
Xianbin Cao 0001, Zhong Wang 0008, Pingkun Yan, Xuelong Li 0001
Neurocomputing3
2013 Sparse coding for image denoising using spike and slab prior
Xiaoqiang Lu, Yuan Yuan 0001, Pingkun Yan
Neurocomputing3
2013 Image registration by normalized mapping
Qi Wang 0009, Cuiming Zou, Yuan Yuan 0001, Hongbing Lu, Pingkun Yan
Neurocomputing5
2013 Special issue: Behaviours in video
Huiyu Zhou 0001, Yuan Yuan 0001, Yingzi Du, Pingkun Yan
Neurocomputing4
2013 SIFT on manifold: An intrinsic description
Guokang Zhu, Qi Wang 0009, Yuan Yuan 0001, Pingkun Yan
Neurocomputing4
2013 Greedy regression in sparse coding space for single-image super-resolution
Yi Tang 0003, Yuan Yuan 0001, Pingkun Yan, Xuelong Li 0001
J. Vis. Commun. Image Represent.3
2013 Machine learning in medical imaging
Pingkun Yan, Kenji Suzuki 0001, Fei Wang 0002, Dinggang Shen
Mach. Vis. Appl.1
2013 Robust visual tracking with discriminative sparse learning
Xiaoqiang Lu, Yuan Yuan 0001, Pingkun Yan
Pattern Recognit.3
2013 Multi-spectral saliency detection
Qi Wang 0009, Pingkun Yan, Yuan Yuan 0001, Xuelong Li 0001
Pattern Recognit. Lett.2
2013 Image Super-Resolution Via Double Sparsity Regularized Manifold Learning
abstract
Over the past few years, high resolutions have been desirable or essential, e.g., in online video systems, and therefore, much has been done to achieve an image of higher resolution from the corresponding low-resolution ones. This procedure of recovering/rebuilding is called single-image super-resolution (SR). Performance of image SR has been significantly improved via methods of sparse coding. That is to say, the image frame patch can be sparse linear combinations of basis elements. However, most of these existing methods fail to consider the local geometrical structure in the space of the training data. To take this crucial issue into account, this paper proposes a method named double sparsity regularized manifold learning (DSRML). DSRML can preserve the properties of the aforementioned local geometrical structure by employing manifold learning, e.g., locally linear embedding. Based on a large amount of experimental results, DSRML is demonstrated to be more robust and more effective than previous efforts in the task of single-image SR.
Xiaoqiang Lu, Yuan Yuan 0001, Pingkun Yan
IEEE Trans. Circuits Syst. Video Technol.3
2013 Visual Saliency by Selective Contrast
abstract
Automatic detection of salient objects in visual media (e.g., videos and images) has been attracting much attention. The detected salient objects can be utilized for segmentation, recognition, and retrieval. However, the accuracy of saliency detection remains a challenge. The reason behind this challenge is mainly due to the lack of a well-defined model for interpreting saliency formulation. To tackle this problem, this letter proposes to detect salient objects based on selective contrast. Selective contrast intrinsically explores the most distinguishable component information in color, texture, and location. A large number of experiments are thereafter carried out upon a benchmark dataset, and the results are compared with those of 12 other popular state-of-the-art algorithms. In addition, the advantage of the proposed algorithm is also demonstrated in a retargeting application.
Qi Wang 0009, Yuan Yuan 0001, Pingkun Yan
IEEE Trans. Circuits Syst. Video Technol.3
2013 Saliency Detection by Multiple-Instance Learning
abstract
Saliency detection has been a hot topic in recent years. Its popularity is mainly because of its theoretical meaning for explaining human attention and applicable aims in segmentation, recognition, etc. Nevertheless, traditional algorithms are mostly based on unsupervised techniques, which have limited learning ability. The obtained saliency map is also inconsistent with many properties of human behavior. In order to overcome the challenges of inability and inconsistency, this paper presents a framework based on multiple-instance learning. Low-, mid-, and high-level features are incorporated in the detection procedure, and the learning ability enables it robust to noise. Experiments on a data set containing 1000 images demonstrate the effectiveness of the proposed framework. Its applicability is shown in the context of a seam carving application.
Qi Wang 0009, Yuan Yuan 0001, Pingkun Yan, Xuelong Li 0001
IEEE Trans. Cybern.3
2013 Learning Saliency by MRF and Differential Threshold
abstract
Saliency detection has been an attractive topic in recent years. The reliable detection of saliency can help a lot of useful processing without prior knowledge about the scene, such as content-aware image compression, segmentation, etc. Although many efforts have been spent in this subject, the feature expression and model construction are far from perfect. The obtained saliency maps are therefore not satisfying enough. In order to overcome these challenges, this paper presents a new psychologic visual feature based on differential threshold and applies it in a supervised Markov-random-field framework. Experiments on two public data sets and an image retargeting application demonstrate the effectiveness, robustness, and practicability of the proposed method.
Guokang Zhu, Qi Wang 0009, Yuan Yuan 0001, Pingkun Yan
IEEE Trans. Cybern.4
2013 Manifold Regularized Sparse NMF for Hyperspectral Unmixing
abstract
Hyperspectral unmixing is one of the most important techniques in analyzing hyperspectral images, which decomposes a mixed pixel into a collection of constituent materials weighted by their proportions. Recently, many sparse nonnegative matrix factorization (NMF) algorithms have achieved advanced performance for hyperspectral unmixing because they overcome the difficulty of absence of pure pixels and sufficiently utilize the sparse characteristic of the data. However, most existing sparse NMF algorithms for hyperspectral unmixing only consider the Euclidean structure of the hyperspectral data space. In fact, hyperspectral data are more likely to lie on a low-dimensional submanifold embedded in the high-dimensional ambient space. Thus, it is necessary to consider the intrinsic manifold structure for hyperspectral unmixing. In order to exploit the latent manifold structure of the data during the decomposition, manifold regularization is incorporated into sparsity-constrained NMF for unmixing in this paper. Since the additional manifold regularization term can keep the close link between the original image and the material abundance maps, the proposed approach leads to a more desired unmixing performance. The experimental results on synthetic and real hyperspectral data both illustrate the superiority of the proposed method compared with other state-of-the-art approaches.
Xiaoqiang Lu, Hao Wu 0098, Yuan Yuan 0001, Pingkun Yan, Xuelong Li 0001
IEEE Trans. Geosci. Remote. Sens.4
2012 Geometry constrained sparse coding for single image super-resolution
abstract
The choice of the over-complete dictionary that sparsely represents data is of prime importance for sparse coding-based image super-resolution. Sparse coding is a typical unsupervised learning method to generate an over-complete dictionary. However, most of the sparse coding methods for image super-resolution fail to simultaneously consider the geometrical structure of the dictionary and corresponding coefficients, which may result in noticeable super-resolution reconstruction artifacts. In this paper, a novel sparse coding method is proposed to preserve the geometrical structure of the dictionary and the sparse coefficients of the data. Moreover, the proposed method can preserve the incoherence of dictionary entries, which is critical for sparse representation. Inspired by the development on non-local self-similarity and manifold learning, the proposed sparse coding method can provide the sparse coefficients and learned dictionary from a new perspective, which have both reconstruction and discrimination properties to enhance the learning performance. Extensive experimental results on image super-resolution have demonstrated the effectiveness of the proposed method.
Xiaoqiang Lu, Pingkun Yan, Yuan Yuan 0001, Xuelong Li 0001
CVPR3
2012 Multi-atlas Based Image Selection with Label Image Constraint
abstract
Atlas selection plays an important role in multiatlas based image segmentation. In atlas selection methods, manifold learning based techniques have recently emerged as very promisingly. However, due to the complexity of anatomical structures in raw images, it is difficult to get accurate atlas selection results by measuring only the distance between raw images on the manifolds. In this paper, we tackle this problem by proposing a label image constrained atlas selection (LICAS) method to exploit the shape and size information of the regions to be segmented from the label images. Constrained by the label images, a new manifold projection method is developed to help uncover the intrinsic similarity between the regions of interest across images. Compared with other existing methods, the experimental results of segmentation on 60 Magnetic Resonance (MR) images showed that the selected atlases are closer to the target structure and more accurate segmentation can be obtained by using the proposed method.
Yihui Cao, Xuelong Li 0001, Pingkun Yan
ICMLA (1)3
2012 Coupled Directional Level Set for MR Image Segmentation
abstract
Segmenting bladder wall for thickness measuring is a fundamental operation in bladder magnetic resonance (MR) image analysis since thickening of the bladder wall may indicate abnormality. Active contours have been used for bladder wall segmentation, which can be broadly divided into gradient-based and region-based methods, according to the used image features. However, the artifacts in MR images and the complex background outside the bladder lead to significant challenges for segmentation. In this paper, a coupled directional level set model is proposed to segment the outer and inner boundaries simultaneously by exploiting the directional gradient, region information and thickness prior of the bladder wall. With our proposed method, the influence of the artifacts in the bladder lumen and the complicated intensity distribution of soft tissues surrounding the bladder can be appreciably reduced. Promising results on 119 bladder MR images have demonstrated the performance of the presented method.
Xianjing Qin, Yang Liu 0093, Hongbing Lu, Xuelong Li 0001, Pingkun Yan
ICMLA (1)5
2012 Vehicle detection and tracking in airborne videos by multi-motion layer analysis
Xianbin Cao 0001, Jinhe Lan, Pingkun Yan, Xuelong Li 0001
Mach. Vis. Appl.3
2012 Visual Attention Accelerated Vehicle Detection in Low-Altitude Airborne Video of Urban Environment
abstract
One of the primary goals of the airborne vehicle detection system is to reduce the risks of incident collisions and to relieve traffic jam caused by the increasing number of vehicles. Different from the stationary systems, which are usually fixed on buildings, the airborne systems in unmanned aircrafts or satellites take the advantages of wider view angle and higher mobility. However, detecting vehicles in airborne videos is a challenging task because of the scene complexity and platform movement. The direct application of the traditional image processing techniques to the problem may result in low detection rate or cannot meet the requirements of real-time applications. To address these problems, a new and efficient method composed by two stages, attention focus extraction and vehicle classification is proposed in this paper. Our work makes two key contributions. The first is the introduction of a new attention focus extraction algorithm, which can quickly detect the candidate vehicle regions to make the algorithm focus on much smaller regions for faster computation. The second contribution is a simple and efficient classification process, which is built using the AdaBoost learning algorithm. The classification process, which is a hierarchical structure, is designed to obtain a lower false alarm rate by looking for vehicles in the candidate regions. Experimental results demonstrate that, compared with other representative algorithms, our method can obtain better performance in terms of higher detection rate and lower false positive rate, while meeting the needs of real-time application.
Xianbin Cao 0001, Renjun Lin, Pingkun Yan, Xuelong Li 0001
IEEE Trans. Circuits Syst. Video Technol.3
2012 Selecting Key Poses on Manifold for Pairwise Action Recognition
abstract
In action recognition, bag of visual words based approaches have been shown to be successful, for which the quality of codebook is critical. In a large vocabulary of poses (visual words), some key poses play a more decisive role than others in the codebook. This paper proposes a novel approach for key poses selection, which models the descriptor space utilizing a manifold learning technique to recover the geometric structure of the descriptors on a lower dimensional manifold. A PageRank-based centrality measure is developed to select key poses according to the recovered geometric structure. In each step, a key pose is selected from the manifold and the remaining model is modified to maximize the discriminative power of selected codebook. With the obtained codebook, each action can be represented with a histogram of the key poses. To solve the ambiguity between some action classes, a pairwise subdivision is executed to select discriminative codebooks for further recognition. Experiments on benchmark datasets showed that our method is able to obtain better performance compared with other state-of-the-art methods.
Xianbin Cao 0001, Bo Ning 0003, Pingkun Yan, Xuelong Li 0001
IEEE Trans. Ind. Informatics3
2012 Robust Alternative Minimization for Matrix Completion
abstract
Recently, much attention has been drawn to the problem of matrix completion, which arises in a number of fields, including computer vision, pattern recognition, sensor network, and recommendation systems. This paper proposes a novel algorithm, named robust alternative minimization (RAM), which is based on the constraint of low rank to complete an unknown matrix. The proposed RAM algorithm can effectively reduce the relative reconstruction error of the recovered matrix. It is numerically easier to minimize the objective function and more stable for large-scale matrix completion compared with other existing methods. It is robust and efficient for low-rank matrix completion, and the convergence of the RAM algorithm is also established. Numerical results showed that both the recovery accuracy and running time of the RAM algorithm are competitive with other reported methods. Moreover, the applications of the RAM algorithm to low-rank image recovery demonstrated that it achieves satisfactory performance.
Xiaoqiang Lu, Tieliang Gong, Pingkun Yan, Yuan Yuan 0001, Xuelong Li 0001
IEEE Trans. Syst. Man Cybern. Part B3
2011 Image Denoising via Improved Sparse Coding
abstract
This paper presents a novel dictionary learning method for image denoising, which removes zero-mean independent identically distributed additive noise from a given image. Choosing noisy image itself to train an over-complete dictionary, the dictionary trained by traditional sparse coding methods contains noise information. Through mathematical derivation of equation, we found that a lower bound of dictionary is related with the level of noise in dictionary learning. The proposed idea is to take advantage of the noise information for designing a sparse coding algorithm called improved sparse coding (ISC), which effectively suppresses the noise influence for training a dictionary. This denoising framework utilizes the effective \nmethod, which is based on sparse representations over trained dictionaries. Acquiring an over-complete dictionary by ISC mainly includes three stages. Firstly, we utilize \nK-means method to group the noisy image patches. Secondly, each dictionary is trained by ISC in corresponding class. Finally, an over-complete dictionary is merged \nby these dictionaries. Theory analysis and experimental results both demonstrate that the proposed method yields excellent performance.
Xiaoqiang Lu, Pingkun Yan, Luoqing Li, Xuelong Li 0001
BMVC3
2011 Medical Image Segmentation Using Descriptive Image Features
abstract
Segmentation of medical images is an important component for diagnosis and treatment of diseases using medical imaging technologies. However, automated accurate medical image segmentation is still a challenge due to the difficulties in finding a robust feature descriptor to describe the object boundaries in medical images. In this paper, a new normal vector feature profile (NVFP) is proposed to describe the local image information of a contour point by concatenating a series of local region descriptors along the normal direction at that point. To avoid trapping by false boundaries caused by nonboundary image features, a modified scale invariant feature transform (SIFT) descriptor is developed. The number and locations of sample points for building NVFP are determined for each contour point, which are constrained by the neighboring anatomical structures and the statistical consistency of the training features. NVFP is incorporated into a model based method for image segmentation. The performance of our proposed method was demonstrated by segmenting prostate MR images. The segmentation results indicated that our method can achieve better performance compared with other existing methods.
Meijuan Yang, Yuan Yuan 0001, Xuelong Li 0001, Pingkun Yan
BMVC4
2011 Accelerating Vehicle Detection in Low-Altitude Airborne Urban Video
abstract
The limitation of the existing methods of traffic data collection is that they rely on techniques that are strictly local in nature. The airborne system in unmanned aircrafts provides the advantages of wider view angle and higher mobility. However, detecting vehicles in airborne videos is a challenging task because of the scene complexity and platform movement. Most of the techniques used in stationary platforms cannot perform well in this situation. A new and efficient method based on Bayes model is proposed in this paper. This method can be divided into two stages, attention focus extraction and vehicle classification. Experimental results demonstrated that, compared with other representative algorithms, our method obtained better performance with higher detection rate, lower false positive rate and faster detection speed.
Xianbin Cao 0001, Renjun Lin, Pingkun Yan, Xuelong Li 0001
ICIG3
2011 KLT Feature Based Vehicle Detection and Tracking in Airborne Videos
abstract
Airborne vehicle detection and tracking systems equipped on unmanned aerial vehicles (UAVs) are difficult to develop because of factors like UAV motion, scene complexity and so on. In this paper, we propose a new framework of multi-motion layer analysis to detect and track moving vehicles in airborne platform. Moving vehicles are firstly detected by registration and temporal differencing to establish motion layers. After motion layers are constructed, they are maintained over time for tracking vehicles. All vehicles are tracked by maintaining their corresponding motion layers. Our experimental results showed that compared with other previous algorithms, our method can achieve better results in terms of detection and tracking performance.
Xianbin Cao 0001, Jinhe Lan, Pingkun Yan, Xuelong Li 0001
ICIG3
2011 Single-Image Super-Resolution via Sparse Coding Regression
abstract
In this paper, it has been shown that the sparse coding algorithm for single-image super-resolution is equivalent to a linear regression algorithm in the sparse coding space. Following the idea, the sparse coding algorithm are generalized by a novel L2-Boosting-based single-resolution super-resolution algorithm which focuses on the relationship between sparse codings corresponding to the low- and high-resolution image patches. The experimental results demonstrate the effectiveness of the proposed algorithm by comparing with other state-of-the-art algorithms.
Yi Tang 0003, Yuan Yuan 0001, Pingkun Yan, Xuelong Li 0001
ICIG3
2011 Linear SVM classification using boosting HOG features for vehicle detection in low-altitude airborne videos
abstract
Visual surveillance from low-altitude airborne platforms has been widely addressed in recent years. Moving vehicle detection is an important component of such a system, which is a very challenging task due to illumination variance and scene complexity. Therefore, a boosting Histogram Orientation Gradients (boosting HOG) feature is proposed in this paper. This feature is not sensitive to illumination change and shows better performance in characterizing object shape and appearance. Each of the boosting HOG feature is an output of an adaboost classifier, which is trained using all bins upon a cell in traditional HOG features. All boosting HOG features are combined to establish the final feature vector to train a linear SVM classifier for vehicle classification. Compared with classical approaches, the proposed method achieved better performance in higher detection rate, lower false positive rate and faster detection speed.
Xianbin Cao 0001, Changxia Wu, Pingkun Yan, Xuelong Li 0001
ICIP3
2011 Putting images on a manifold for atlas-based image segmentation
abstract
In medical image analysis, atlas-based segmentation has become a popular approach. Given a target image, how to select the atlases with the similar shape of anatomical structure to the input image is one of the most critical factors affecting the segmentation accuracy. In this paper, we propose a novel strategy by putting the images on a manifold to analyze the intrinsic similarity between the images. A subset of atlases can be selected and the optimal fusion weights are computed in a low-dimensional manifold space. Finally, it combines the selected atlases by using the corresponding weights for image segmentation. The experimental results demonstrated that our proposed method is robust and accurate especially when a large number of training samples are available.
Yihui Cao, Yuan Yuan 0001, Xuelong Li 0001, Pingkun Yan
ICIP4
2011 A novel alternative algorithm for limited angle tomography
abstract
This paper studies incomplete data problems of circular cone-beam computed tomography, which occur frequently in medical imaging and industrial imaging. The incomplete data problems in which projection data are only available in an angular range can be attributed to the limited angle tomography. Limited angle tomography is a severely ill-posed inverse problem. In recent years, image reconstruction based on total variation (TV) was employed to reduce the problem and gave better performance on edge-preserving reconstruction. However, the artificial parameter can only be determined through considerable experimentation. In this paper, an alternating minimization method based on TV is proposed to reduce the data insufficiency in tomographic imaging. This novel alternating minimization method provides a robust and effective reconstruction without any artificial parameter in the iterative processes, by using the TV as a multiplicative constraint. The results demonstrate that this new reconstruction method brings satisfactory performance.
Xiaoqiang Lu, Yuan Yuan 0001, Pingkun Yan, Xuelong Li 0001
ICIP3
2011 Robust color correction in stereo vision
abstract
The phenomenon of color discrepancy between image pairs happens frequently in stereo vision systems. This inconsistence in color domain may cause difficulties when identifying point correspondence to reconstruct the scene depth. In this paper, we propose a robust algorithm to correct the color discrepancy between images. The proposed algorithm neither requires a color calibration chart/object which is a tedious procedure, nor explicitly compensates for the image as a whole, which possibly give bad correction results in local areas of an image. Instead, we correct the image region by region. Experiments show that the presented color correction algorithm is effective and efficient.
Qi Wang 0009, Pingkun Yan, Yuan Yuan 0001, Xuelong Li 0001
ICIP2
2011 Learning shape statistics for hierarchical 3D medical image segmentation
abstract
Accurate image segmentation is important for many medical imaging applications, whereas it remains challenging due to the complexity in medical images, such as the complex shapes and varied neighbor structures. This paper proposes a new hierarchical 3D image segmentation method based on patient-specific shape prior and surface patch shape statistics (SURPASS) model. In the segmentation process, a coarse-to-fine, two-stage strategy is designed, which contains global segmentation and local segmentation. In the global segmentation stage, patient-specific shape prior is estimated by using manifold learning techniques to achieve the overall segmentation. In the second stage, SURPASS is computed to solve the problem of poor segmentation at certain surface patches. The effectiveness of the proposed 3D image segmentation method has been demonstrated by the experiments on segmenting the prostate from a series of MR images.
Wuxia Zhang, Yuan Yuan 0001, Xuelong Li 0001, Pingkun Yan
ICIP4
2011 Segmenting Images by Combining Selected Atlases on Manifold
Yihui Cao, Yuan Yuan 0001, Xuelong Li 0001, Baris Turkbey, Peter L. Choyke, Pingkun Yan
MICCAI (3)6
2011 Local learning-based image super-resolution
abstract
Local learning algorithm has been widely used in single-frame super-resolution reconstruction algorithm, such as neighbor embedding algorithm [1] and locality preserving constraints algorithm [2]. Neighbor embedding algorithm is based on manifold assumption, which defines that the embedded neighbor patches are contained in a single manifold. While manifold assumption does not always hold. In this paper, we present a novel local learning-based image single-frame SR reconstruction algorithm with kernel ridge regression (KRR). Firstly, Gabor filter is adopted to extract texture information from low-resolution patches as the feature. Secondly, each input low-resolution feature patch utilizes K nearest neighbor algorithm to generate a local structure. Finally, KRR is employed to learn a map from input low-resolution (LR) feature patches to high-resolution (HR) feature patches in the corresponding local structure. Experimental results show the effectiveness of our method.
Xiaoqiang Lu, Yuan Yuan 0001, Pingkun Yan, Luoqing Li, Xuelong Li 0001
MMSP4
2011 Local semi-supervised regression for single-image super-resolution
abstract
In this paper, we propose a local semi-supervised learning-based algorithm for single-image super-resolution. Different from most of example-based algorithms, the information of test patches is considered during learning local regression functions which map a low-resolution patch to a high-resolution patch. Localization strategy is generally adopted in single-image super-resolution with nearest neighbor-based algorithms. However, the poor generalization of the nearest neighbor estimation decreases the performance of such algorithms. Though the problem can be fixed by local regression algorithms, the sizes of local training sets are always too small to improve the performance of nearest neighbor-based algorithms significantly. To overcome the difficulty, the semi-supervised regression algorithm is used here. Unlike supervised regression, the information about test samples is considered in semi-supervised regression algorithms, which makes the semi-supervised regression more powerful. Noticing that numerous test patches exist, the performance of nearest neighbor-based algorithms can be further improved by employing a semi-supervised regression algorithm. Experiments verify the effectiveness of the proposed algorithm.
Yi Tang 0003, Xiaoli Pan, Yuan Yuan 0001, Pingkun Yan, Luoqing Li, Xuelong Li 0001
MMSP4
2011 Rapid pedestrian detection in unseen scenes
Xianbin Cao 0001, Zhong Wang 0008, Pingkun Yan, Xuelong Li 0001
Neurocomputing3
2011 Vehicle Detection and Motion Analysis in Low-Altitude Airborne Video Under Urban Environment
abstract
Visual surveillance from low-altitude airborne platforms plays a key role in urban traffic surveillance. Moving vehicle detection and motion analysis are very important for such a system. However, illumination variance, scene complexity, and platform motion make the tasks very challenging. In addition, the used algorithms have to be computationally efficient in order to be used on a real-time platform. To deal with these problems, a new framework for vehicle detection and motion analysis from low-altitude airborne videos is proposed. Our paper has two major contributions. First, to speed up feature extraction and to retain additional global features in different scales for higher classification accuracy, a boosting light and pyramid sampling histogram of oriented gradients feature extraction method is proposed. Second, to efficiently correlate vehicles across different frames for vehicle motion trajectories computation, a spatio-temporal appearance-related similarity measure is proposed. Compared to other representative existing methods, our experimental results showed that the proposed method is able to achieve better performance with higher detection rate, lower false positive rate, and faster detection speed.
Xianbin Cao 0001, Changxia Wu, Jinhe Lan, Pingkun Yan, Xuelong Li 0001
IEEE Trans. Circuits Syst. Video Technol.4
2010 Incremental Shape Statistics Learning for Prostate Tracking in TRUS
Pingkun Yan, Jochen Kruecker
MICCAI (2)1
2009 Modeling Interaction for Segmentation of Neighboring Structures
abstract
This paper presents a new method for segmenting medical images by modeling interaction between neighboring structures. Compared to previously reported methods, the proposed approach enables simultaneous segmentation of multiple neighboring structures for improved robustness. During the segmentation process, the object contour evolution and shape prior estimates are influenced by the interactions between neighboring shapes consisting of attraction, repulsion, and competition. Instead of estimating the a priori shape of each structure independently, an interactive maximum a posteriori shape estimation method is used for estimating the shape priors by considering shape prior distribution, neighboring shapes, and image features. Energy functionals are then formulated to model the interaction and segmentation. With the proposed method, neighboring structures with similar intensities and/or textures, and blurred boundaries can be extracted simultaneously. Experimental results obtained on both synthetic data and medical images demonstrate that the introduced interaction between neighboring structures improves segmentation performance compared with other existing approaches.
Pingkun Yan, Ashraf A. Kassim, Weijia Shen, Mubarak Shah
IEEE Trans. Inf. Technol. Biomed.1
2008 Learning 4D action feature models for arbitrary view action recognition
abstract
In this paper we present a novel approach using a 4D (x,y,z,t) action feature model (4D-AFM) for recognizing actions from arbitrary views. The 4D-AFM elegantly encodes shape and motion of actors observed from multiple views. The modeling process starts with reconstructing 3D visual hulls of actors at each time instant. Spatiotemporal action features are then computed in each view by analyzing the differential geometric properties of spatio-temporal volumes (3D STVs) generated by concatenating the actor’s silhouette over the course of the action (x, y, t). These features are mapped to the sequence of 3D visual hulls over time (4D) to build the initial 4D-AFM. Actions are recognized based on the scores of matching action features from the input videos to the model points of 4D-AFMs by exploiting pairwise interactions of features. Promising recognition results have been demonstrated on the multi-view IXMAS dataset using both single and multi-view input videos.
Pingkun Yan, Saad M. Khan, Mubarak Shah
CVPR1
2008 Action recognition using spatio-temporal regularity based features
abstract
In this paper, a novel feature for capturing information in a spatio-temporal volume based on regularity flow is presented for action recognition. The regularity flow describes the direction of least intensity change within a spatio-temporal volume. Our feature consists of weighted histograms of the computed regularity flow around selected interest points. We then apply this new feature to recognizing actions with experiments on known benchmark dataset. A more discriminating representation of spatio-temporal volume is obtained by using the feature descriptors with the bag of words model. Action recognition is performed by using this new representation with a trained support vector machine. We show that by utilizing regularity flow based features, recognition can be performed with better performance than the best known features. Additionally, results suggest that our descriptor captures information otherwise not harnessed by existing methods.
Taylor Goodhart, Pingkun Yan, Mubarak Shah
ICASSP2
2008 Automatic Segmentation of High-Throughput RNAi Fluorescent Cellular Images
abstract
High-throughput genome-wide RNA interference (RNAi) screening is emerging as an essential tool to assist biologists in understanding complex cellular processes. The large number of images produced in each study make manual analysis intractable; hence, automatic cellular image analysis becomes an urgent need, where segmentation is the first and one of the most important steps. In this paper, a fully automatic method for segmentation of cells from genome-wide RNAi screening images is proposed. Nuclei are first extracted from the DNA channel by using a modified watershed algorithm. Cells are then extracted by modeling the interaction between them as well as combining both gradient and region information in the Actin and Rac channels. A new energy functional is formulated based on a novel interaction model for segmenting tightly clustered cells with significant intensity variance and specific phenotypes. The energy functional is minimized by using a multiphase level set method, which leads to a highly effective cell segmentation method. Promising experimental results demonstrate that automatic segmentation of high-throughput genome-wide multichannel screening can be achieved by using the proposed method, which may also be extended to other multichannel image segmentation problems.
Pingkun Yan, Xiaobo Zhou 0001, Mubarak Shah, Stephen T. C. Wong
IEEE Trans. Inf. Technol. Biomed.1
2007 A Homographic Framework for the Fusion of Multi-view Silhouettes
abstract
This paper presents a purely image-based approach to fusing foreground silhouette information from multiple arbitrary views. Our approach does not require 3D constructs like camera calibration to carve out 3D voxels or project visual cones in 3D space. Using planar homographies and foreground likelihood information from a set of arbitrary views, we show that visual hull intersection can be performed in the image plane without requiring to go in 3D space. This process delivers a 2D grid of object occupancy likelihoods representing a cross-sectional slice of the object. Subsequent slices of the object are obtained by extending the process to planes parallel to a reference plane in a direction along the body of the object. We show that homographies of these new planes between views can be computed in the framework of plane to plane homologies using the homography induced by a reference plane and the vanishing point of the reference direction. Occupancy grids are stacked on top of each other, creating a three dimensional data structure that encapsulates the object shape and location. Object structure is finally segmented out by minimizing an energy functional over the surface of the object in a level sets formulation. We show the application of our method on complicated object shapes as well as cluttered environments containing multiple objects.
Saad M. Khan, Pingkun Yan, Mubarak Shah
ICCV2
2007 3D Model based Object Class Detection in An Arbitrary View
abstract
In this paper, a novel object class detection method based on 3D object modeling is presented. Instead of using a complicated mechanism for relating multiple 2D training views, the proposed method establishes spatial connections between these views by mapping them directly to the surface of 3D model. The 3D shape of an object is reconstructed by using a homographic framework from a set of model views around the object and is represented by a volume consisting of binary slices. Features are computed in each 2D model view and mapped to the 3D shape model using the same homographic framework. To generalize the model for object class detection, features from supplemental views are also considered. A codebook is constructed from all of these features and then a 3D feature model is built. Given a 2D test image, correspondences between the 3D feature model and the testing view are identified by matching the detected features. Based on the 3D locations of the corresponding features, several hypotheses of viewing planes can be made. The one with the highest confidence is then used to detect the object using feature location matching. Performance of the proposed method has been evaluated by using the PASCAL VOC challenge dataset and promising results are demonstrated.
Pingkun Yan, Saad M. Khan, Mubarak Shah
ICCV1
2007 Spatio-Temporal Regularity Flow (SPREF): Its Estimation and Applications
abstract
Feature selection and extraction is a key operation in video analysis for achieving a higher level of abstraction. In this paper, we introduce a general framework to extract a new spatio-temporal feature that represents the directions in which a video is regular, i.e., the pixel appearances change the least. We propose to model the directions of regular variations with a 3-D vector field, which is referred to as spatio-temporal regularity flow (SPREF). SPREF vectors are designed to have three cross-sectional parallel components Fx, Fy, and Ftfor convenient use in different applications. They are estimated using all the frames simultaneously by minimizing an energy functional formulated according to its definition. In this paper, we first introduce translational SPREF (T-SPREF) and then extend our framework to affine SPREF (A-SPREF). The successful use of SPREF in a few applications, including object removal, video inpainting, and video compression, is also demonstrated
Orkun Alatas, Pingkun Yan, Mubarak Shah
IEEE Trans. Circuits Syst. Video Technol.2
2006 Segmentation of volumetric MRA images by using capillary active contour
Pingkun Yan, Ashraf A. Kassim
Medical Image Anal.1
2006 Medical Image Segmentation Using Minimal Path Deformable Models With Implicit Shape Priors
abstract
This paper presents a new method for segmentation of medical images by extracting organ contours, using minimal path deformable models incorporated with statistical shape priors. In our approach, boundaries of structures are considered as minimal paths, i.e., paths associated with the minimal energy, on weighted graphs. Starting from the theory of minimal path deformable models, an intelligent "worm" algorithm is proposed for segmentation, which is used to evaluate the paths and finally find the minimal path. Prior shape knowledge is incorporated into the segmentation process to achieve more robust segmentation. The shape priors are implicitly represented and the estimated shapes of the structures can be conveniently obtained. The worm evolves under the joint influence of the image features, its internal energy, and the shape priors. The contour of the structure is then extracted as the worm trail. The proposed segmentation framework overcomes the short-comings of existing deformable models and has been successfully applied to segmenting various medical images.
Pingkun Yan, Ashraf A. Kassim
IEEE Trans. Inf. Technol. Biomed.1
2005 MRA Image Segmentation with Capillary Active Contour
Pingkun Yan, Ashraf A. Kassim
MICCAI1
2005 Segmentation of Neighboring Organs in Medical Image with Model Competition
Pingkun Yan, Weijia Shen, Ashraf A. Kassim, Mubarak Shah
MICCAI1
2005 Motion compensated lossy-to-lossless compression of 4-D medical images using integer wavelet transforms
abstract
This paper proposes a method for progressive lossy-to-lossless compression of four-dimensional (4-D) medical images (sequences of volumetric images over time) by using a combination of three-dimensional (3-D) integer wavelet transform (IWT) and 3-D motion compensation. A 3-D extension of the set-partitioning in hierarchical trees (SPIHT) algorithm is employed for coding the wavelet coefficients. To effectively exploit the redundancy between consecutive 3-D images, the concepts of key and residual frames from video coding is used. A fast 3-D cube matching algorithm is employed to do motion estimation. The key and the residual volumes are then coded using 3-D IWT and the modified 3-D SPIHT. The experimental results presented in this paper show that our proposed compression scheme achieves better lossy and lossless compression performance on 4-D medical images when compared with JPEG-2000 and volumetric compression based on 3-D SPIHT.
Ashraf A. Kassim, Pingkun Yan, Wei Siong Lee, Kuntal Sengupta
IEEE Trans. Inf. Technol. Biomed.2
2004 Medical image segmentation with minimal path deformable models
Pingkun Yan, Ashraf A. Kassim
ICIP1