David Dagan Feng

dblp:f/DavidDaganFeng · also Dagan Feng 0001, David D. Feng, David Feng 0003 · DBLP profile ↗
← Back
321ranked-venue papers
10as first author
44since 2021 · last 2026
0000-0002-3381-214XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 167 · 1 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 107 · 9 first-author · 22 since 2021Artificial intelligence and machine learning · 75 · 9 since 2021Human-computer interaction and ubiquitous computing · 13 · 1 since 2021Computer networks · 5Databases, data management, data science and information retrieval · 2Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 Dynamic Traceback Learning for Medical Report Generation
abstract
Automated medical report generation has demonstrated the potential to significantly reduce the workload associated with time-consuming medical reporting. Recent generative representation learning methods have shown promise in integrating vision and language modalities for medical report generation. However, when trained end-to-end and applied directly to medical image-to-text generation, they face two significant challenges: i) difficulty in accurately capturing subtle yet crucial pathological details, and ii) reliance on both visual and textual inputs during inference, leading to performance degradation in zero-shot inference when only images are available. To address these challenges, this study proposes a novel multimodal dynamic traceback learning framework (DTrace)1. Specifically, we introduce a traceback mechanism to supervise the semantic validity of generated content and a dynamic learning strategy to adapt to various proportions of image and text input, enabling text generation without strong reliance on the input from both modalities during inference. The learning of cross-modal knowledge is enhanced by supervising the model to recover masked semantic information from a complementary counterpart. Extensive experiments conducted on two benchmark datasets, IU-Xray and MIMIC-CXR, demonstrate that the proposedDTraceframework outperforms state-of-the-art methods for medical report generation.
Shuchang Ye, Mingyuan Meng, Mingjian Li, David Dagan Feng, Usman Naseem, Jinman Kim
IEEE Trans. Multim.4
2025 A Generative Adversarial Network for Upsampling of Direct Volume Rendering Images
abstract
Abstract Direct volume rendering (DVR) is an important tool for scientific and medical imaging visualization. Modern GPU acceleration has made DVR more accessible; however, the production of high‐quality rendered images with high frame rates is computationally expensive. We propose a deep learning method with a reduced computational demand. We leveraged a conditional generative adversarial network (cGAN) to upsample DVR images (a rendered scene), with a reduced sampling rate to obtain similar visual quality to that of a fully sampled method. Our dvrGAN is combined with a colour‐based loss function that is optimized for DVR images where different structures such as skin, bone, etc. are distinguished by assigning them distinct colours. The loss function highlights the structural differences between images, by examining pixel‐level colour, and thus helps identify, for instance, small bones in the limbs that may not be evident with reduced sampling rates. We evaluated our method in DVR of human computed tomography (CT) and CT angiography (CTA) volumes. Our method retained image quality and reduced computation time when compared to fully sampled methods and outperformed existing state‐of‐the‐art upsampling methods.
Ge Jin 0001, Younhyun Jung, Michael J. Fulham, David Dagan Feng, Jinman Kim
Comput. Graph. Forum4
2025 AutoFuse: Automatic fusion networks for deformable medical image registration
Mingyuan Meng, Michael J. Fulham, David Dagan Feng, Lei Bi 0001, Jinman Kim
Pattern Recognit.3
2025 MHFNet: A Multimodal Hybrid-Embedding Fusion Network for Automatic Sleep Staging
abstract
Scoring sleep stages is essential for evaluating the status of sleep continuity and comprehending its structure. Despite previous attempts, automating sleep scoring remains challenging. First, most existing works did not fuse local and global temporal information. Second, the correlation for special waves in different signals is rarely used in sleep staging modeling. Third, the logic of scoring rules based on adjacent epochs is not considered in developing sleep staging models. This paper introduces a multimodal hybrid-embedding fusion network (MHFNet), which aims to tackle these challenges in automating sleep stage scoring. MHFNet comprises multi-stream Xception blocks to extract wave characteristics, a hybrid time-embedding module to combine local and global temporal information, a dual-path gate transformer to fuse and enhance attention features, and a refined output header to reconstruct sleep scoring. We perform experiments using three publicly available datasets (SleepEDF-ST, SleepEDF-SC, and SHHS). Experimental results indicate the superiority of MHFNet over baseline approaches in cross-validation. Moreover, at the individual level, MHFNet yielded an average $R^{2}$ score improvement of 9$\%$ in the testing dataset compared to state-of-the-art models, paving the way for its applications in real-world sleep medicine.
Ruhan Liu, Jiajia Li 0004, Bin Sheng 0001, David Dagan Feng, Ping Zhang 0016
IEEE J. Biomed. Health Informatics6
2025 Federated Pseudo Modality Generation for Incomplete Multi-Modal MRI Reconstruction
abstract
While multi-modal learning has been widely used for MRI reconstruction, it relies on paired multi-modal data, which is difficult to acquire in real clinical scenarios. Especially in the federated setting, there is a common issue that several medical institutions suffer from missing modalities or even only have single-modal data. Therefore, it is infeasible to deploy a standard federated learning framework in such conditions. In this paper, we propose a novel communication-efficient federated learning framework (namely Fed-PMG) to address the missing modality challenge in federated multi-modal MRI reconstruction. Specifically, we utilize a pseudo modality generation mechanism to recover the missing modality for each single-modal client by sharing the distribution information of the amplitude spectrum in frequency space. However, the step of sharing the original amplitude spectrum leads to heavy communication costs. To reduce the communication cost, we introduce a clustering scheme to project the set of amplitude spectrum into a finite number of cluster centroids and share them among the clients. With such an elaborate design, our approach can effectively complete the missing modality within an acceptable communication cost. Extensive experimental results demonstrate that our proposed method can outperform state-of-the-art methods and reach a performance similar to the ideal scenario (i.e., all clients have the full set of modalities).
Yunlu Yan, Chun-Mei Feng 0001, Yuexiang Li, Ping Li 0016, Rick Siow Mong Goh, Bai Ying Lei, Weiming Wang 0002, David Dagan Feng, Lei Zhu 0003
IEEE J. Biomed. Health Informatics8
2025 Explicit Abnormality Extraction for Unsupervised Motion Artifact Reduction in Magnetic Resonance Imaging
abstract
Motion artifacts compromise the quality of magnetic resonance imaging (MRI) and pose challenges to achieving diagnostic outcomes and image-guided therapies. In recent years, supervised deep learning approaches have emerged as successful solutions for motion artifact reduction (MAR). One disadvantage of these methods is their dependency on acquiring paired sets of motion artifact-corrupted (MA-corrupted) and motion artifact-free (MA-free) MR images for training purposes. Obtaining such image pairs is difficult and therefore limits the application of supervised training. In this paper, we propose a novel UNsupervised Abnormality Extraction Network (UNAEN) to alleviate this problem. Our network is capable of working with unpaired MA-corrupted and MA-free images. It converts the MA-corrupted images to MA-reduced images by extracting abnormalities from the MA-corrupted images using a proposed artifact extractor, which intercepts the residual artifact maps from the MA-corrupted MR images explicitly, and a reconstructor to restore the original input from the MA-reduced images. The performance of UNAEN was assessed by experimenting with various publicly available MRI datasets and comparing them with state-of-the-art methods. The quantitative evaluation demonstrates the superiority of UNAEN over alternative MAR methods and visually exhibits fewer residual artifacts. Our results substantiate the potential of UNAEN as a promising solution applicable in real-world clinical environments, with the capability to enhance diagnostic accuracy and facilitate image-guided therapies.
Hao Li 0034, Zhengmin Kong, Tao Huang 0008, Euijoon Ahn, Zhihan Lyu, Jinman Kim, David Dagan Feng
IEEE J. Biomed. Health Informatics9
2025 Enhancing Medical Vision-Language Contrastive Learning via Inter-Matching Relation Modeling
abstract
Medical image representations can be learned through medical vision-language contrastive learning (mVLCL) where medical imaging reports are used as weak supervision through image-text alignment. These learned image representations can be transferred to and benefit various downstream medical vision tasks such as disease classification and segmentation. Recent mVLCL methods attempt to align image sub-regions and the report keywords as local-matchings. However, these methods aggregate all local-matchings via simple pooling operations while ignoring the inherent relations between them. These methods therefore fail to reason between local-matchings that are semantically related, e.g., local-matchings that correspond to the disease word and the location word (semantic-relations), and also fail to differentiate such clinically important local-matchings from others that correspond to less meaningful words, e.g., conjunction words (importance-relations). Hence, we propose a mVLCL method that models the inter-matching relations between local-matchings via a relation-enhanced contrastive learning framework (RECLF). In RECLF, we introduce a semantic-relation reasoning module (SRM) and an importance-relation reasoning module (IRM) to enable more fine-grained report supervision for image representation learning. We evaluated our method using six public benchmark datasets on four downstream tasks, including segmentation, zero-shot classification, linear classification, and cross-modal retrieval. Our results demonstrated the superiority of our RECLF over the state-of-the-art mVLCL methods with consistent improvements across single-modal and cross-modal tasks. These results suggest that our RECLF, by modeling the inter-matching relations, can learn improved medical image representations with better generalization capabilities.
Mingjian Li, Mingyuan Meng, Michael J. Fulham, David Dagan Feng, Lei Bi 0001, Jinman Kim
IEEE Trans. Medical Imaging4
2025 Continuous Bijection Supervised Pyramid Diffeomorphic Deformation for Learning Tooth Meshes From CBCT Images
abstract
Accurate and high-quality tooth mesh generation from cone-beam computerized tomography (CBCT) is an essential computer-aided technology for digital dentistry. However, existing segmentation-based methods require complicated post-processing and significant manual correction to generate regular tooth meshes. In this paper, we propose a method of continuous bijection supervised pyramid diffeomorphic deformation (PDD) for learning tooth meshes, which could be used to directly generate high-quality tooth meshes from CBCT Images. Overall, we adopt a classic two-stage framework. In the first stage, we devise an enhanced detector to accurately locate and crop every tooth. In the second stage, a PDD network is designed to deform a sphere mesh from low resolution to high one according to pyramid flows based on diffeomorphic mesh deformations, so that the generated mesh approximates the ground truth infinitely and efficiently. To achieve that, a novel continuous bijection distance loss on the diffeomorphic sphere is also designed to supervise the deformation learning, which overcomes the shortcoming of loss based on nearest-neighbour mapping and improves the fitting precision. Experiments show that our method outperforms the state-of-the-art methods in terms of both different evaluation metrics and the geometry quality of reconstructed tooth surfaces.
Zechu Zhang, Weilong Peng, Jinyu Wen, Keke Tang, Meie Fang, David Dagan Feng, Ping Li 0016
IEEE Trans. Multim.6
2025 Deep contour attention learning for scleral deformation from OCT images
Hao Chen 0011, Yupeng Xu, Huating Li, Yuan Xie 0006, David Dagan Feng, Jinman Kim, Lei Bi 0001, Xiangui He, Bin Sheng 0001
Vis. Comput.7
2024 Semi-Mamba: Improving Medical Image Segmentation via Semi-Automatic Mamba Network
Lei Bi 0001, Yige Peng, David Dagan Feng, Jinman Kim
CGI (3)3
2024 Correlation-aware Coarse-to-fine MLPs for Deformable Medical Image Registration
abstract
Deformable image registration is a fundamental step for medical image analysis. Recently, transformers have been used for registration and outperformed Convolutional Neural Networks (CNNs). Transformers can capture long-range dependence among image features, which have been shown beneficial for registration. However, due to the high computation/memory loads of self-attention, transformers are typically used at downsampled feature resolutions and cannot capture fine-grained long-range dependence at the full image resolution. This limits deformable registration as it necessitates precise dense correspondence between each image pixel. Multi-layer Perceptrons (MLPs) without self-attention are efficient in computation/memory usage, enabling the feasibility of capturing fine-grained long-range dependence at full resolution. Nevertheless, MLPs have not been extensively explored for image registration and are lacking the consideration of inductive bias crucial for medical registration tasks. In this study, we propose the first correlation-aware MLP-based registration network (CorrMLP) for deformable medical image registration. Our CorrMLP introduces a correlation-aware multi-window MLP block in a novel coarse-to-fine registration architecture, which captures fine-grained multi-range dependence to perform correlation-aware coarse-to-fine registration. Extensive experiments with seven public medical datasets show that our CorrMLP outperforms state-of-the-art deformable registration methods.
Mingyuan Meng, David Dagan Feng, Lei Bi 0001, Jinman Kim
CVPR2
2024 3DPX: Progressive 2D-to-3D Oral Image Reconstruction with Hybrid MLP-CNN Networks
Xiaoshuang Li, Mingyuan Meng, Zimo Huang, Lei Bi 0001, Eduardo Delamare, David Dagan Feng, Bin Sheng 0001, Jinman Kim
MICCAI (7)6
2024 Enabling Text-Free Inference in Language-Guided Segmentation of Chest X-Rays via Self-guidance
Shuchang Ye, Mingyuan Meng, Mingjian Li, David Dagan Feng, Jinman Kim
MICCAI (8)4
2024 A Transformer-Assisted Cascade Learning Network for Choroidal Vessel Segmentation
Lei Bi 0001, Wuzhen Shi, Yupeng Xu, Wenming Cao 0001, David Dagan Feng
J. Comput. Sci. Technol.9
2024 Multi-Label Chest X-Ray Image Classification With Single Positive Labels
abstract
Deep learning approaches for multi-label Chest X-ray (CXR) images classification usually require large-scale datasets. However, acquiring such datasets with full annotations is costly, time-consuming, and prone to noisy labels. Therefore, we introduce a weakly supervised learning problem called Single Positive Multi-label Learning (SPML) into CXR images classification (abbreviated as SPML-CXR), in which only one positive label is annotated per image. A simple solution to SPML-CXR problem is to assume that all the unannotated pathological labels are negative, however, it might introduce false negative labels and decrease the model performance. To this end, we present a Multi-level Pseudo-label Consistency (MPC) framework for SPML-CXR. First, inspired by the pseudo-labeling and consistency regularization in semi-supervised learning, we construct a weak-to-strong consistency framework, where the model prediction on weakly-augmented image is treated as the pseudo label for supervising the model prediction on a strongly-augmented version of the same image, and define an Image-level Perturbation-based Consistency (IPC) regularization to recover the potential mislabeled positive labels. Besides, we incorporate Random Elastic Deformation (RED) as an additional strong augmentation to enhance the perturbation. Second, aiming to expand the perturbation space, we design a perturbation stream to the consistency framework at the feature-level and introduce a Feature-level Perturbation-based Consistency (FPC) regularization as a supplement. Third, we design a Transformer-based encoder module to explore the sample relationship within each mini-batch by a Batch-level Transformer-based Correlation (BTC) regularization. Extensive experiments on the CheXpert and MIMIC-CXR datasets have shown the effectiveness of our MPC framework for solving the SPML-CXR problem.
Jiayin Xiao, Si Li 0005, Tongxu Lin, Jian Zhu 0001, Xiaochen Yuan, David Dagan Feng, Bin Sheng 0001
IEEE Trans. Medical Imaging6
2023 Merging-Diverging Hybrid Transformer Networks for Survival Prediction in Head and Neck Cancer
Mingyuan Meng, Lei Bi 0001, Michael J. Fulham, David Dagan Feng, Jinman Kim
MICCAI (6)4
2023 Non-iterative Coarse-to-Fine Transformer Networks for Joint Affine and Deformable Image Registration
Mingyuan Meng, Lei Bi 0001, Michael J. Fulham, David Dagan Feng, Jinman Kim
MICCAI (10)4
2023 Unsupervised Landmark Detection-Based Spatiotemporal Motion Estimation for 4-D Dynamic Medical Images
abstract
Motion estimation is a fundamental step in dynamic medical image processing for the assessment of target organ anatomy and function. However, existing image-based motion estimation methods, which optimize the motion field by evaluating the local image similarity, are prone to produce implausible estimation, especially in the presence of large motion. In addition, the correct anatomical topology is difficult to be preserved as the image global context is not well incorporated into motion estimation. In this study, we provide a novel motion estimation framework of dense-sparse-dense (DSD), which comprises two stages. In the first stage, we process the raw dense image to extract sparse landmarks to represent the target organ's anatomical topology, and discard the redundant information that is unnecessary for motion estimation. For this purpose, we introduce an unsupervised 3-D landmark detection network to extract spatially sparse but representative landmarks for the target organ's motion estimation. In the second stage, we derive the sparse motion displacement from the extracted sparse landmarks of two images of different time points. Then, we present a motion reconstruction network to construct the motion field by projecting the sparse landmarks' displacement back into the dense image domain. Furthermore, we employ the estimated motion field from our two-stage DSD framework as initialization and boost the motion estimation quality in light-weight yet effective iterative optimization. We evaluate our method on two dynamic medical imaging tasks to model cardiac motion and lung respiratory motion, respectively. Our method has produced superior motion estimation accuracy compared to the existing comparative methods. Besides, the extensive experimental results demonstrate that our solution can extract well-representative anatomical landmarks without any requirement of manual annotation. Our code is publicly available online: https://github.com/yyguo-sjtu/DSD-3D-Unsupervised-Landmark-Detection-Based-Motion-Estimation.
Yuyu Guo 0002, Lei Bi 0001, Dongming Wei, Liyun Chen, Zhengbin Zhu, David Dagan Feng, Ruiyan Zhang, Qian Wang 0001, Jinman Kim
IEEE Trans. Cybern.6
2023 Multi-Level Adversarial Spatio-Temporal Learning for Footstep Pressure Based FoG Detection
abstract
Freezing of gait (FoG) is one of the most common symptoms of Parkinson's disease, which is a neurodegenerative disorder of the central nervous system impacting millions of people around the world. To address the pressing need to improve the quality of treatment for FoG, devising a computer-aided detection and quantification tool for FoG has been increasingly important. As a non-invasive technique for collecting motion patterns, the footstep pressure sequences obtained from pressure sensitive gait mats provide a great opportunity for evaluating FoG in the clinic and potentially in the home environment. In this study, FoG detection is formulated as a sequential modelling task and a novel deep learning architecture, namely Adversarial Spatio-temporal Network (ASTN), is proposed to learn FoG patterns across multiple levels. ASTN introduces a novel adversarial training scheme with a multi-level subject discriminator to obtain subject-independent FoG representations, which helps to reduce the over-fitting risk due to the high inter-subject variance. As a result, robust FoG detection can be achieved for unseen subjects. The proposed scheme also sheds light on improving subject-level clinical studies from other scenarios as it can be integrated with many existing deep architectures. To the best of our knowledge, this is one of the first studies of footstep pressure-based FoG detection and the approach of utilizing ASTN is the first deep neural network architecture in pursuit of subject-independent representations. In our experiments on 393 trials collected from 21 subjects, the proposed ASTN achieved an AUC 0.85, clearly outperforming conventional learning methods.
Kun Hu 0008, Shaohui Mei, Wei Wang 0115, Kaylena A. Ehgoetz Martens, Liang Wang 0001, Simon J. G. Lewis, David Dagan Feng, Zhiyong Wang 0001
IEEE J. Biomed. Health Informatics7
2023 A Shortened Model for Logan Reference Plot Implemented via the Self-Supervised Neural Network for Parametric PET Imaging
abstract
Dynamic PET imaging provides superior physiological information than conventional static PET imaging. However, the dynamic information is gained at the cost of a long scanning protocol; this limits the clinical application of dynamic PET imaging. We developed a modified Logan reference plot model to shorten the acquisition procedure in dynamic PET imaging by omitting the early-time information necessary for the conventional reference Logan model. The proposed model is accurate theoretically, but the straightforward approach raises the sampling problem in implementation and results in noisy parametric images. We then designed a self-supervised convolutional neural network to increase the noise performance of parametric imaging, with dynamic images of only a single subject for training. The proposed method was validated via simulated and real dynamic [Formula: see text]-fallypride PET data. Results showed that it accurately estimated the distribution volume ratio (DVR) in dynamic PET with a shortened scanning protocol, e.g., 20 minutes, where the estimations were comparable with those obtained from a standard dynamic PET study of 120 minutes of acquisition. Further comparisons illustrated that our method outperformed the shortened Logan model implemented with Gaussian filtering, regularization, BM4D and the 4D deep image prior methods in terms of the trade-off between bias and variance. Since the proposed method uses data acquired in a short period of time upon the equilibrium, it has the potential to add clinical values by providing both DVR and Standard Uptake Value (SUV) simultaneously. It thus promotes clinical applications of dynamic PET studies when neuronal receptor functions are studied.
Wenxiang Ding, Qiaoqiao Ding, Kewei Chen 0001, Miao Zhang 0038, David Dagan Feng, Lei Bi 0001, Jinman Kim, Qiu Huang
IEEE Trans. Medical Imaging6
2023 Individualized Statistical Modeling of Lesions in Fundus Images for Anomaly Detection
abstract
Anomaly detection in fundus images remains challenging due to the fact that fundus images often contain diverse types of lesions with various properties in locations, sizes, shapes, and colors. Current methods achieve anomaly detection mainly through reconstructing or separating the fundus image background from a fundus image under the guidance of a set of normal fundus images. The reconstruction methods, however, ignore the constraint from lesions. The separation methods primarily model the diverse lesions with pixel-based independent and identical distributed (i.i.d.) properties, neglecting the individualized variations of different types of lesions and their structural properties. And hence, these methods may have difficulty to well distinguish lesions from fundus image backgrounds especially with the normal personalized variations (NPV). To address these challenges, we propose a patch-based non-i.i.d. mixture of Gaussian (MoG) to model diverse lesions for adapting to their statistical distribution variations in different fundus images and their patch-like structural properties. Further, we particularly introduce the weighted Schatten p-norm as the metric of low-rank decomposition for enhancing the accuracy of the learned fundus image backgrounds and reducing false-positives caused by NPV. With the individualized modeling of the diverse lesions and the background learning, fundus image backgrounds and NPV are finely learned and subsequently distinguished from diverse lesions, to ultimately improve the anomaly detection. The proposed method is evaluated on two real-world databases and one artificial database, outperforming the state-of-the-art methods.
Yuchen Du, Lisheng Wang, Deyu Meng, Benzhi Chen, Chengyang An, Hao Liu 0120, Yupeng Xu, David Dagan Feng, Xiuying Wang 0001
IEEE Trans. Medical Imaging10
2023 EAPT: Efficient Attention Pyramid Transformer for Image Processing
abstract
Recent transformer-based models, especially patch-based methods, have shown huge potentiality in vision tasks. However, the split fixed-size patches divide the input features into the same size patches, which ignores the fact that vision elements are often various and thus may destroy the semantic information. Also, the vanilla patch-based transformer cannot guarantee the information communication between patches, which will prevent the extraction of attention information with a global view. To circumvent those problems, we propose an Efficient Attention Pyramid Transformer (EAPT). Specifically, we first propose the Deformable Attention, which learns an offset for each position in patches. Thus, even with split fixed-size patches, our method can still obtain non-fixed attention information that can cover various vision elements. Then, we design the Encode-Decode Communication module (En-DeC module), which can obtain communication information among all patches to get more complete global attention information. Finally, we propose a position encoding specifically for vision transformers, which can be used for patches of any dimension and any length. Extensive experiments on the vision tasks of image classification, object detection, and semantic segmentation demonstrate the effectiveness of our proposed model. Furthermore, we also conduct rigorous ablation studies to evaluate the key components of the proposed structure.
Xiao Lin 0012, Shuzhou Sun, Bin Sheng 0001, Ping Li 0016, David Dagan Feng
IEEE Trans. Multim.6
2023 SThy-Net: a feature fusion-enhanced dense-branched modules network for small thyroid nodule classification from ultrasound images
Abdulrhman H. Al-Jebrni, Saba Ghazanfar Ali, Huating Li, Xiao Lin 0012, Ping Li 0016, Younhyun Jung, Jinman Kim, David Dagan Feng, Bin Sheng 0001, Lixin Jiang
Vis. Comput.8
2022 Non-iterative Coarse-to-Fine Registration Based on Single-Pass Deep Cumulative Learning
Mingyuan Meng, Lei Bi 0001, David Dagan Feng, Jinman Kim
MICCAI (6)3
2022 Deep multi-scale resemblance network for the sub-class differentiation of adrenal masses on computed tomography images
Lei Bi 0001, Jinman Kim, Tingwei Su, Michael J. Fulham, David Dagan Feng, Guang Ning
Artif. Intell. Medicine5
2022 Deep Cognitive Gate: Resembling Human Cognition for Saliency Detection
abstract
Saliency detection by human refers to the ability to identify pertinent information using our perceptive and cognitive capabilities. While human perception is attracted by visual stimuli, our cognitive capability is derived from the inspiration of constructing concepts of reasoning. Saliency detection has gained intensive interest with the aim of resembling human 'perceptual' system. However, saliency related to human 'cognition', particularly the analysis of complex salient regions ('cogitating' process), is yet to be fully exploited. We propose to resemble human cognition, coupled with human perception, to improve saliency detection. We recognize saliency in three phases ('Seeing' - 'Perceiving' - 'Cogitating), mimicking human's perceptive and cognitive thinking of an image. In our method, 'Seeing' phase is related to human perception, and we formulate the 'Perceiving' and 'Cogitating' phases related to the human cognition systems via deep neural networks (DNNs) to construct a new module (Cognitive Gate) that enhances the DNN features for saliency detection. To the best of our knowledge, this is the first work that established DNNs to resemble human cognition for saliency detection. In our experiments, our approach outperformed 17 benchmarking DNN methods on six well-recognized datasets, demonstrating that resembling human cognition improves saliency detection.
Ke Yan 0005, Xiuying Wang 0001, Jinman Kim, Wangmeng Zuo, David Dagan Feng
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Affective Audio Annotation of Public Speeches with Convolutional Clustering Neural Network
abstract
Public speaking is a critical skill in daily communication. While more practicing such as rehearsal is helpful to improve such a skill, lack of personalized feedback limits the effectiveness of practicing. Therefore, we formulate the task of personalized feedback as an affective audio annotation problem by learning knowledge from online public speech videos. Considering the great success of deep learning techniques such as convolutional neural networks in a wide range of applications including speech recognition and object recognition, we propose a novel convolutional clustering neural network (CCNN) to solve this multi-label classification problem. Instead of aggregating the features of different channels through pooling, we introduce a novel clustering layer to derive intermediate representation for improved annotation performance. In order to evaluate the performance of our proposed method, we purposely built an affective audio annotation dataset by collecting more than 2,000 video clips from the TED website. Experimental results on this dataset demonstrate that our proposed method outperforms traditional CNN-based approaches with a lower hamming loss for affective annotation.
Jiahao Xu 0002, Zhiyong Wang 0001, Yang Wang 0002, Fang Chen 0001, Junbin Gao, David Dagan Feng
IEEE Trans. Affect. Comput.7
2022 Automatic Detection and Classification System of Domestic Waste via Multimodel Cascaded Convolutional Neural Network
abstract
Domestic waste classification was incorporated into legal provisions recently in China. However, relying on manpower to detect and classify domestic waste is highly inefficient. To that end, in this article, we propose a multimodel cascaded convolutional neural network (MCCNN) for domestic waste image detection and classification. MCCNN combined three subnetworks (DSSD, YOLOv4, and Faster-RCNN) to obtain the detections. Moreover, to suppress the false-positive predicts, we utilized a classification model cascaded with the detection part to judge whether the detection results are correct. To train and evaluate MCCNN, we designed a large-scale waste image dataset (LSWID), containing 30 000 domestic waste multilabeled images with 52 categories. To the best of our knowledge, the LSWID is the largest dataset on domestic waste images. Furthermore, a smart trash can is designed and applied to a Shanghai community, which helped to make waste recycling more efficient. Experimental results showed a state-of-the-art performance, with an average improvement of 10% in detection precision.
Jiajia Li 0004, Jie Chen 0097, Bin Sheng 0001, Ping Li 0016, Po Yang 0001, David Dagan Feng, Jun Qi 0001
IEEE Trans. Ind. Informatics6
2022 Graph Convolutional Dictionary Selection With L₂, ₚ Norm for Video Summarization
abstract
Video Summarization (VS) has become one of the most effective solutions for quickly understanding a large volume of video data. Dictionary selection with self representation and sparse regularization has demonstrated its promise for VS by formulating the VS problem as a sparse selection task on video frames. However, existing dictionary selection models are generally designed only for data reconstruction, which results in the neglect of the inherent structured information among video frames. In addition, the sparsity commonly constrained by$L_{2,1}$norm is not strong enough, which causes the redundancy of keyframes, i.e., similar keyframes are selected. Therefore, to address these two issues, in this paper we propose a general framework called graph convolutional dictionary selection with$L_{2,p}$($0< p\leq 1$) norm (GCDS$_{2,p}$) for both keyframe selection and skimming based summarization. Firstly, we incorporate graph embedding into dictionary selection to generate the graph embedding dictionary, which can take the structured information depicted in videos into account. Secondly, we propose to use$L_{2,p}$($0< p\leq 1$) norm constrained row sparsity, in which$p$can be flexibly set for two forms of video summarization. For keyframe selection,$0< p< 1$can be utilized to select diverse and representative keyframes; and for skimming,$p=1$can be utilized to select key shots. In addition, an efficient iterative algorithm is devised to optimize the proposed model, and the convergence is theoretically proved. Experimental results including both keyframe selection and skimming based summarization on four benchmark datasets demonstrate the effectiveness and superiority of the proposed method.
Mingyang Ma 0004, Shaohui Mei, Shuai Wan, Zhiyong Wang 0001, Xian-Sheng Hua 0001, David Dagan Feng
IEEE Trans. Image Process.6
2022 ECSU-Net: An Embedded Clustering Sliced U-Net Coupled With Fusing Strategy for Efficient Intervertebral Disc Segmentation and Classification
abstract
Automatic vertebra segmentation from computed tomography (CT) image is the very first and a decisive stage in vertebra analysis for computer-based spinal diagnosis and therapy support system. However, automatic segmentation of vertebra remains challenging due to several reasons, including anatomic complexity of spine, unclear boundaries of the vertebrae associated with spongy and soft bones. Based on 2D U-Net, we have proposed an Embedded Clustering Sliced U-Net (ECSU-Net). ECSU-Net comprises of three modules named segmentation, intervertebral disc extraction (IDE) and fusion. The segmentation module follows an instance embedding clustering approach, where our three sliced sub-nets use axis of CT images to generate a coarse 2D segmentation along with embedding space with the same size of the input slices. Our IDE module is designed to classify vertebra and find the inter-space between two slices of segmented spine. Our fusion module takes the coarse segmentation (2D) and outputs the refined 3D results of vertebra. A novel adaptive discriminative loss (ADL) function is introduced to train the embedding space for clustering. In the fusion strategy, three modules are integrated via a learnable weight control component, which adaptively sets their contribution. We have evaluated classical and deep learning methods on Spineweb dataset-2. ECSU-Net has provided comparable performance to previous neural network based algorithms achieving the best segmentation dice score of 95.60% and classification accuracy of 96.20%, while taking less time and computation resources.
Anam Nazir, Muhammad Nadeem Cheema, Bin Sheng 0001, Ping Li 0016, Huating Li, Guangtao Xue, Harry Qin, Jinman Kim, David Dagan Feng
IEEE Trans. Image Process.9
2022 DeepMTS: Deep Multi-Task Learning for Survival Prediction in Patients With Advanced Nasopharyngeal Carcinoma Using Pretreatment PET/CT
abstract
Nasopharyngeal Carcinoma (NPC) is a malignant epithelial cancer arising from the nasopharynx. Survival prediction is a major concern for NPC patients, as it provides early prognostic information to plan treatments. Recently, deep survival models based on deep learning have demonstrated the potential to outperform traditional radiomics-based survival prediction models. Deep survival models usually use image patches covering the whole target regions (e.g., nasopharynx for NPC) or containing only segmented tumor regions as the input. However, the models using the whole target regions will also include non-relevant background information, while the models using segmented tumor regions will disregard potentially prognostic information existing out of primary tumors (e.g., local lymph node metastasis and adjacent tissue invasion). In this study, we propose a 3D end-to-end Deep Multi-Task Survival model (DeepMTS) for joint survival prediction and tumor segmentation in advanced NPC from pretreatment PET/CT. Our novelty is the introduction of a hard-sharing segmentation backbone to guide the extraction of local features related to the primary tumors, which reduces the interference from non-relevant background information. In addition, we also introduce a cascaded survival network to capture the prognostic information existing out of primary tumors and further leverage the global tumor information (e.g., tumor size, shape, and locations) derived from the segmentation backbone. Our experiments with two clinical datasets demonstrate that our DeepMTS can consistently outperform traditional radiomics-based survival prediction models and existing deep survival models.
Mingyuan Meng, Bingxin Gu, Lei Bi 0001, Shaoli Song, David Dagan Feng, Jinman Kim
IEEE J. Biomed. Health Informatics5
2022 Improving Breast Tumor Segmentation in PET via Attentive Transformation Based Normalization
abstract
Positron Emission Tomography (PET) has become a preferred imaging modality for cancer diagnosis, radiotherapy planning, and treatment responses monitoring. Accurate and automatic tumor segmentation is the fundamental requirement for these clinical applications. Deep convolutional neural networks have become the state-of-the-art in PET tumor segmentation. The normalization process is one of the key components for accelerating network training and improving the performance of the network. However, existing normalization methods either introduce batch noise into the instance PET image by calculating statistics on batch level or introduce background noise into every single pixel by sharing the same learnable parameters spatially. In this paper, we proposed an attentive transformation (AT)-based normalization method for PET tumor segmentation. We exploit the distinguishability of breast tumor in PET images and dynamically generate dedicated and pixel-dependent learnable parameters in normalization via the transformation on a combination of channel-wise and spatial-wise attentive responses. The attentive learnable parameters allow to re-calibrate features pixel-by-pixel to focus on the high-uptake area while attenuating the background noise of PET images. Our experimental results on two real clinical datasets show that the AT-based normalization method improves breast tumor segmentation performance when compared with the existing normalization methods.
Xiaoya Qiao, Chunjuan Jiang, Panli Li, Yuan Yuan 0022, Qinglong Zeng, Lei Bi 0001, Shaoli Song, Jinman Kim, David Dagan Feng, Qiu Huang
IEEE J. Biomed. Health Informatics9
2022 Machine Learning-Based Noninvasive Quantification of Single-Imaging Session Dual-Tracer 18F-FDG and 68Ga-DOTATATE Dynamic PET-CT in Oncology
abstract
68Ga-DOTATATE PET-CT is routinely used for imaging neuroendocrine tumor (NET) somatostatin receptor subtype 2 (SSTR2) density in patients, and is complementary to FDG PET-CT for improving the accuracy of NET detection, characterization, grading, staging, and predicting/monitoring NET responses to treatment. Performing sequential18F-FDG and68Ga-DOTATATE PET scans would require 2 or more days and can delay patient care. To align temporal and spatial measurements of18F-FDG and68Ga-DOTATATE PET, and to reduce scan time and CT radiation exposure to patients, we propose a single-imaging session dual-tracer dynamic PET acquisition protocol in the study. A recurrent extreme gradient boosting (rXGBoost) machine learning algorithm was proposed to separate the mixed18F-FDG and68Ga-DOTATATE time activity curves (TACs) for the region of interest (ROI) based quantification with tracer kinetic modeling. A conventional parallel multi-tracer compartment modeling method was also implemented for reference. Single-scan dual-tracer dynamic PET was simulated from 12 NET patient studies with18F-FDG and68Ga-DOTATATE 45-min dynamic PET scans separately obtained within 2 days. Our experimental results suggested an18F-FDG injection first followed by68Ga-DOTATATE with a minimum 5 min delayed injection protocol for the separation of mixed18F-FDG and68Ga-DOTATATE TACs using rXGBoost algorithm followed by tracer kinetic modeling is highly feasible.
Wenxiang Ding, Jiangyuan Yu, Chaojie Zheng, Peng Fu 0003, Qiu Huang, David Dagan Feng, Richard L. Wahl, Yun Zhou 0006
IEEE Trans. Medical Imaging6
2021 A Spatial Guided Self-supervised Clustering Network for Medical Image Segmentation
Euijoon Ahn, David Dagan Feng, Jinman Kim
MICCAI (1)2
2021 Unsupervised brain tumor segmentation using a symmetric-driven adversarial network
Xinheng Wu, Lei Bi 0001, Michael J. Fulham, David Dagan Feng, Luping Zhou, Jinman Kim
Neurocomputing4
2021 Similarity Based Block Sparse Subset Selection for Video Summarization
abstract
Video summarization (VS) is generally formulated as a subset selection problem where a set of representative keyframes or key segments is selected from an entire video frame set. Though many sparse subset selection based VS algorithms have been proposed in the past decade, most of them adopt linear sparse formulation in the explicit feature vector space of video frames, and don’t consider the local or global relationships among frames. In this paper, we first extend the conventional sparse subset selection for VS into kernel block sparse subset selection (KBS3) to utilize the advantage of kernel sparse coding and introduce a local inter-frame relationship through packing of frame blocks. Going a step further, we propose a similarity based block sparse subset selection (SB2S3) model by applying a specially designed transformation matrix on the KBS3 model in order to introduce a kind of global inter-frame relationship through the similarity. Finally, a greedy pursuit based algorithm is devised for the proposed NP-hard model optimization. The proposed SB2S3 has the following advantages: 1) through the similarity between each frame and any other frame, the global relationship among all frames can be considered; 2) through block sparse coding, the local relationship of adjacent frames is further considered; and 3) it has a wider application, since features can derive similarity, but not vice versa. It is believed that the effect of modeling such global and local relationships among frames in this paper, is similar to that of modeling the long-range and short-range dependencies among frames in deep learning based methods. Experimental results on three benchmark datasets have demonstrated that the proposed approach is superior to not only other sparse subset selection based VS methods but also most unsupervised deep-learning based VS methods.
Mingyang Ma 0004, Shaohui Mei, Shuai Wan, Zhiyong Wang 0001, David Dagan Feng, Mohammed Bennamoun
IEEE Trans. Circuits Syst. Video Technol.5
2021 A New Aggregation of DNN Sparse and Dense Labeling for Saliency Detection
abstract
As a fundamental requirement to many computer vision systems, saliency detection has experienced substantial progress in recent years based on deep neural networks (DNNs). Most DNN-based methods rely on either sparse or dense labeling, and thus they are subject to the inherent limitations of the chosen labeling schemes. DNN dense labeling captures salient objects mainly from global features, which are often hampered by other visually distinctive regions. On the other hand, DNN sparse labeling is usually impeded by inaccurate presegmentation of the images that it depends on. To address these limitations, we propose a new framework consisting of two pathways and an Aggregator to progressively integrate the DNN sparse and DNN dense labeling schemes to derive the final saliency map. In our "zipper" type aggregation, we propose a multiscale kernels approach to extract optimal criteria for saliency detection where we suppress nonsalient regions in the sparse labeling while guiding the dense labeling to recognize more complete extent of the saliency. We demonstrate that our method outperforms in saliency detection compared to other 11 state-of-the-art methods across six well-recognized benchmarking datasets.
Ke Yan 0005, Xiuying Wang 0001, Jinman Kim, David Dagan Feng
IEEE Trans. Cybern.4
2021 Kernelized Mahalanobis Distance for Fuzzy Clustering
abstract
Data samples of complicated geometry and nonlinear separability are considered as common challenges to clustering algorithms. In this article, we first construct Mahalanobis distance in the kernel space and then propose a novel fuzzy clustering model with a kernelized Mahalanobis distance, namely KMD-FC. The key contributions of KMD-FC include: first, the construction of KMD matrix is innovatively transformed from the Euclidean distance kernel matrix, which is able to effectively avoid the problem of “curse of dimensionality” posed by explicitly calculating the sample covariance matrix in the kernel space; second, for the first time, the kernelized Gustafson–Kessel (GK) fuzzy C-means algorithm is achieved, which is critically important to extend the applications of the GK algorithm to the nonlinear classification tasks; finally, taking account of the overall distribution of samples in the kernel space after kernel mapping to improve the generalizability of the proposed KMD-FC clustering method. Comprehensive experiments conducted on a wide range of datasets, including synthetic datasets and machine learning repository (UCI) datasets, have validated that the proposed clustering algorithm outperformed the state-of-the-art methods in comparison.
Shan Zeng, Xiuying Wang 0001, Xiangjun Duan, Sen Zeng, Zuyin Xiao, David Dagan Feng
IEEE Trans. Fuzzy Syst.6
2021 Modified GAN-CAED to Minimize Risk of Unintentional Liver Major Vessels Cutting by Controlled Segmentation Using CTA/SPET-CT
abstract
This article substantially advances upon state-of-the-art to enhance liver vessels segmentation accuracy by leveraging advantages of synthetic PET-CT (SPET-CT) images in addition to computed tomography angiography (CTA) volumes. Our setup makes a hybrid solution of modified generative adversarial network-convolutional autoencoder (GAN-cAED) combining synthetic ability of GAN to deliver SPET-CT images with generative ability of cAED network in terms of latent learning to more refined segmentation of major liver vessels. We improve time complexity through a novel concept of controlled segmentation by introducing a threshold metric to stop segmentation up to a desired level. The innovative concept of controlled vessel segmentation with a stopping criterion via variant threshold levels will help surgeons to avoid unintentional major blood vessels cutting, reducing the risk of excessive blood loss. Clinically, such solutions offer computer-aided liver surgeries and drug treatment evaluation in a CTA-only environment, shorten the requirement of radioactive and expensive fused PET-CT images.
Muhammad Nadeem Cheema, Anam Nazir, Po Yang 0001, Bin Sheng 0001, Ping Li 0016, Huating Li, Xiaoer Wei, Harry Qin, Jinman Kim, David Dagan Feng
IEEE Trans. Ind. Informatics10
2021 Keyframe Extraction From Laparoscopic Videos via Diverse and Weighted Dictionary Selection
abstract
Laparoscopic videos have been increasingly acquired for various purposes including surgical training and quality assurance, due to the wide adoption of laparoscopy in minimally invasive surgeries. However, it is very time consuming to view a large amount of laparoscopic videos, which prevents the values of laparoscopic video archives from being well exploited. In this paper, a dictionary selection based video summarization method is proposed to effectively extract keyframes for fast access of laparoscopic videos. Firstly, unlike the low-level feature used in most existing summarization methods, deep features are extracted from a convolutional neural network to effectively represent video frames. Secondly, based on such a deep representation, laparoscopic video summarization is formulated as a diverse and weighted dictionary selection model, in which image quality is taken into account to select high quality keyframes, and a diversity regularization term is added to reduce redundancy among the selected keyframes. Finally, an iterative algorithm with a rapid convergence rate is designed for model optimization, and the convergence of the proposed method is also analyzed. Experimental results on a recently released laparoscopic dataset demonstrate the clear superiority of the proposed methods. The proposed method can facilitate the access of key information in surgeries, training of junior clinicians, explanations to patients, and archive of case files.
Mingyang Ma 0004, Shaohui Mei, Shuai Wan, Zhiyong Wang 0001, ZongYuan Ge, Vincent Lam, David Dagan Feng
IEEE J. Biomed. Health Informatics7
2021 Short-Term Lesion Change Detection for Melanoma Screening With Novel Siamese Neural Network
abstract
Short-term monitoring of lesion changes has been a widely accepted clinical guideline for melanoma screening. When there is a significant change of a melanocytic lesion at three months, the lesion will be excised to exclude melanoma. However, the decision on change or no-change heavily depends on the experience and bias of individual clinicians, which is subjective. For the first time, a novel deep learning based method is developed in this paper for automatically detecting short-term lesion changes in melanoma screening. The lesion change detection is formulated as a task measuring the similarity between two dermoscopy images taken for a lesion in a short time-frame, and a novel Siamese structure based deep network is proposed to produce the decision: changed (i.e. not similar) or unchanged (i.e. similar enough). Under the Siamese framework, a novel structure, namely Tensorial Regression Process, is proposed to extract the global features of lesion images, in addition to deep convolutional features. In order to mimic the decision-making process of clinicians who often focus more on regions with specific patterns when comparing a pair of lesion images, a segmentation loss (SegLoss) is further devised and incorporated into the proposed network as a regularization term. To evaluate the proposed method, an in-house dataset with 1,000 pairs of lesion images taken in a short time-frame at a clinical melanoma centre was established. Experimental results on this first-of-a-kind large dataset indicate that the proposed model is promising in detecting the short-term lesion change for objective melanoma screening.
Zhiyong Wang 0001, Junbin Gao, Chantal Rutjes, Kaitlin Nufer, Dacheng Tao, David Dagan Feng, Scott W. Menzies
IEEE Trans. Medical Imaging7
2021 Hybrid Refinement-Correction Heatmaps for Human Pose Estimation
abstract
In this paper, we present a method (Hybrid-Pose) to improve human pose estimation in images. We adopt Stacked Hourglass Networks to design two convolutional neural network models, RNet for pose refinement and CNet for pose correction. The CNet (Correction Network) guides the pose refinement RNet (Refinement Network) to correct the joint location before generating the final pose. Each of the two models is composed of four hourglasses, and each hourglass generates a group of detection heatmaps for the joints. The RNet model hourglasses have the same structure. However, the CNet model is designed with hourglasses of different structures for pose guidance. Since the pose estimation in RGB images is very sensitive to the image scene, our proposed approach generates multiple outputs of detection heatmaps to broaden the searching scope for the correct joints locations. We use the RNet model to refine the joints locations in each hourglass stage horizontally, then the heatmaps of each stage are fused with the heatmaps of all the CNet model hourglasses vertically in a hybrid manner. Our method shows competitive results with the existing state-of-the-art approaches on MPII and FLIC benchmark datasets. Although our proposed method focuses on improving single-person pose estimation, we also show the influence of this improvement on multi-person pose estimation by detecting multiple people using SSD detector, then estimating the pose of each person individually.
Aouaidjia Kamel, Bin Sheng 0001, Ping Li 0016, Jinman Kim, David Dagan Feng
IEEE Trans. Multim.5
2021 Patch Based Video Summarization With Block Sparse Representation
abstract
In recent years, sparse representation has been successfully utilized for video summarization (VS). However, most of the sparse representation based VS methods characterize each video frame with global features. As a result, some important local details could be neglected by global features, which may compromise the performance of summarization. In this paper, we propose to partition each video frame into a number of patches and characterize each patch with global features. Instead of concatenating the features of each patch and utilizing conventional sparse representation, we formulate the VS problem with such video frame representation as block sparse representation by considering each video frame as a block containing a number of patches. By taking the reconstruction constraint into account, we devise a simultaneous version of block-based OMP (Orthogonal Matching Pursuit) algorithm, namely SBOMP, to solve the proposed model. The proposed model is further extended to a neighborhood based model which considers temporally adjacent frames as a super block. This is one of the first sparse representation based VS methods taking both spatial and temporal contexts into account with blocks. Experimental results on two widely used VS datasets have demonstrated that our proposed methods present clear superiority over existing sparse representation based VS methods and are highly comparable to some deep learning ones requiring supervision information for extra model training.
Shaohui Mei, Mingyang Ma 0004, Shuai Wan, Junhui Hou, Zhiyong Wang 0001, David Dagan Feng
IEEE Trans. Multim.6
2021 Efficient Body Motion Quantification and Similarity Evaluation Using 3-D Joints Skeleton Coordinates
abstract
Evaluating whole-body motion is challenging because of the articulated nature of the skeleton structure. Each joint moves in an unpredictable way with uncountable possibilities of movements direction under the influence of one or many of its parent joints. This paper presents a method for human motion quantification via three-dimensional (3-D) body joints coordinates. We calculate a set of metrics that influence the joints movement considering the motion of its parent joints without requiring prior knowledge of the motion parameters. Only the raw joints coordinates data of a motion sequence are needed to automatically estimate the transformation matrix of the joints between frames. We also consider the angles between limbs as a fundamental factor to follow the joints directions. We classify the joints motion as global motion and local motion. The global motion represents the joint movement according to a fixed joint, and the local motion represents the joint movement according to its first parent joint. In order to evaluate the performance of the proposed method, we also propose a comparison algorithm between two skeletons motions based on the quantified metrics. We measured the comparative similarity between the 3-D joints coordinates on Microsoft Kinect V2 and UTD-MHAD dataset. User studies were conducted to evaluate the performance under different factors. Various results and comparisons have shown that our method effectively quantifies and evaluates the motion similarity.
Aouaidjia Kamel, Bin Sheng 0001, Ping Li 0016, Jinman Kim, David Dagan Feng
IEEE Trans. Syst. Man Cybern. Syst.5
2020 A Spatiotemporal Volumetric Interpolation Network for 4D Dynamic Medical Image
abstract
Dynamic medical images are often limited in its application due to the large radiation doses and longer image scanning and reconstruction times. Existing methods attempt to reduce the volume samples in the dynamic sequence by interpolating the volumes between the acquired samples. However, these methods are limited to either 2D images and/or are unable to support large but periodic variations in the functional motion between the image volume samples. In this paper, we present a spatiotemporal volumetric interpolation network (SVIN) designed for 4D dynamic medical images. SVIN introduces dual networks: the first is the spatiotemporal motion network that leverages the 3D convolutional neural network (CNN) for unsupervised parametric volumetric registration to derive spatiotemporal motion field from a pair of image volumes; the second is the sequential volumetric interpolation network, which uses the derived motion field to interpolate image volumes, together with a new regression-based module to characterize the periodic motion cycles in functional organ structures. We also introduce an adaptive multi-scale architecture to capture the volumetric large anatomy motions. Experimental results demonstrated that our SVIN outperformed state-of-the-art temporal medical interpolation methods and natural video interpolation method that has been extended to support volumetric images. Code is available at [1].
Yuyu Guo 0002, Lei Bi 0001, Euijoon Ahn, David Dagan Feng, Qian Wang 0001, Jinman Kim
CVPR4
2020 Malocclusion Treatment Planning via PointNet Based Spatial Transformation Network
Xiaoshuang Li, Lei Bi 0001, Jinman Kim, Tingyao Li, Peng Li 0079, Bin Sheng 0001, David Dagan Feng
MICCAI (3)8
2020 Multi-modality Information Fusion for Radiomics-Based Neural Architecture Search
Yige Peng, Lei Bi 0001, Michael J. Fulham, David Dagan Feng, Jinman Kim
MICCAI (7)4
2020 Video summarization via block sparse dictionary selection
Mingyang Ma 0004, Shaohui Mei, Shuai Wan, Junhui Hou, Zhiyong Wang 0001, David Dagan Feng
Neurocomputing6
2020 SPST-CNN: Spatial pyramid based searching and tagging of liver's intraoperative live views via CNN for minimal invasive surgery
Anam Nazir, Muhammad Nadeem Cheema, Bin Sheng 0001, Ping Li 0016, Huating Li, Po Yang 0001, Younhyun Jung, Harry Qin, David Dagan Feng
J. Biomed. Informatics9
2020 Multi-Label classification of multi-modality skin lesion via hyper-connected convolutional neural network
Lei Bi 0001, David Dagan Feng, Michael J. Fulham, Jinman Kim
Pattern Recognit.2
2020 Abnormality detection in retinal image by individualized background learning
Benzhi Chen, Lisheng Wang, Xiuying Wang 0001, Jian Sun 0009, David Dagan Feng, Zongben Xu
Pattern Recognit.6
2020 Learning visual relationship and context-aware attention for image captioning
Junbo Wang 0003, Wei Wang 0115, Liang Wang 0001, Zhiyong Wang 0001, David Dagan Feng, Tieniu Tan
Pattern Recognit.5
2020 Real-time hand posture recognition using hand geometric features and Fisher Vector
Linpu Fang, Ningxin Liang, Wenxiong Kang, Zhiyong Wang 0001, David Dagan Feng
Signal Process. Image Commun.5
2020 Illumination-Invariant Video Cut-Out Using Octagon Sensitive Optimization
abstract
This paper presents an effective video cut-out approach, which can be utilized to segment the moving object in video shots. We first introduce the Octagon-Sensitive-Filtering (OSF) and its illumination invariant feature (IIF), which is computed on each pixel of the image via adding contributions from neighboring pixels. We integrate our IIF into the variational model and obtain the seeds during preprocessing to help address large displacement and illumination changes. An effective seed update method based on tracking-then-refinement based on IIF is presented to compensate for location ambiguities, and the strategy is effective to deal with illumination variances and objects deformation. Furthermore, we apply the IIF-based graph-cut to deal with fuzzy boundaries. Multiple experiments on quantitative challenging datasets have shown the robustness, high-quality video cut-out and efficiency of our approach to acute variances of illumination and complex motion.
Jingye Wang, Bin Sheng 0001, Ping Li 0016, David Dagan Feng
IEEE Trans. Circuits Syst. Video Technol.5
2020 Automated Decision Support System for Lung Cancer Detection and Classification via Enhanced RFCN With Multilayer Fusion RPN
abstract
Detection of lung cancer at early stages is critical, in most of the cases radiologists read computed tomography (CT) images to prescribe follow-up treatment. The conventional method for detecting nodule presence in CT images is tedious. In this article, we propose an enhanced multidimensional region-based fully convolutional network (mRFCN) based automated decision support system for lung nodule detection and classification. The mRFCN is used as an image classifier backbone for feature extraction along with the novel multilayer fusion region proposal network (mLRPN) with position-sensitive score maps being explored. We applied a median intensity projection to leverage three-dimensional information from CT scans and introduced deconvolutional layer to adopt proposed mLRPN in our architecture to automatically select the potential region of interest. Our system has been trained and evaluated using LIDC dataset, and the experimental results showed promising detection performance in comparison to the state-of-the-art nodule detection/classification methods, achieving a sensitivity of 98.1% and classification accuracy of 97.91%.
Anum Masood, Bin Sheng 0001, Po Yang 0001, Ping Li 0016, Huating Li, Jinman Kim, David Dagan Feng
IEEE Trans. Ind. Informatics7
2020 Graph Sequence Recurrent Neural Network for Vision-Based Freezing of Gait Detection
abstract
Freezing of gait (FoG) is one of the most common symptoms of Parkinson's disease (PD), a neurodegenerative disorder which impacts millions of people around the world. Accurate assessment of FoG is critical for the management of PD and to evaluate the efficacy of treatments. Currently, the assessment of FoG requires well-trained experts to perform time-consuming annotations via vision-based observations. Thus, automatic FoG detection algorithms are needed. In this study, we formulate vision-based FoG detection, as a fine-grained graph sequence modelling task, by representing the anatomic joints in each temporal segment with a directed graph, since FoG events can be observed through the motion patterns of joints. A novel deep learning method is proposed, namely graph sequence recurrent neural network (GS-RNN), to characterize the FoG patterns by devising graph recurrent cells, which take graph sequences of dynamic structures as inputs. For the cases of which prior edge annotations are not available, a data-driven based adjacency estimation method is further proposed. To the best of our knowledge, this is one of the first studies on vision-based FoG detection using deep neural networks designed for graph sequences of dynamic structures. Experimental results on more than 150 videos collected from 45 patients demonstrated promising performance of the proposed GS-RNN for FoG detection with an AUC value of 0.90.
Kun Hu 0008, Zhiyong Wang 0001, Wei Wang 0115, Kaylena A. Ehgoetz Martens, Liang Wang 0001, Tieniu Tan, Simon J. G. Lewis, David Dagan Feng
IEEE Trans. Image Process.8
2020 OFF-eNET: An Optimally Fused Fully End-to-End Network for Automatic Dense Volumetric 3D Intracranial Blood Vessels Segmentation
abstract
Intracranial blood vessels segmentation from computed tomography angiography (CTA) volumes is a promising biomarker for diagnosis and therapeutic treatment in cerebrovascular diseases. These segmentation outputs are a fundamental requirement in the development of automated decision support systems for preoperative assessment or intraoperative guidance in neuropathology. The state-of-the-art in medical image segmentation methods are reliant on deep learning architectures based on convolutional neural networks. However, despite their popularity, there is a research gap in the current deep learning architectures optimized to address the technical challenges in blood vessel segmentation. These challenges include: (i) the extraction of concrete brain vessels close to the skull; and (ii) the precise marking of the vessel locations. We propose an Optimally Fused Fully end-to-end Network (OFF-eNET) for automatic segmentation of the volumetric 3D intracranial vascular structures. OFF-eNET comprises of three modules. In the first module, we exploit the up-skip connections to enhance information flow, and dilated convolution for detailed preservation of spatial feature map that are designed for thin blood vessels. In the second module, we employ residual mapping along with inception module for speedy network convergence and richer visual representation. For the third module, we make use of the transferred knowledge in the form of cascaded training strategy to gradually optimize the three segmentation stages (basic, complete, and enhanced) to segment thin vessels located close to the skull. All these modules are designed to be computationally efficient. Our OFF-eNET, evaluated using 70 CTA image volumes, resulted in 90.75% performance in the segmentation of intracranial blood vessels and outperformed the state-of-the-art counterparts.
Anam Nazir, Muhammad Nadeem Cheema, Bin Sheng 0001, Huating Li, Ping Li 0016, Po Yang 0001, Younhyun Jung, Harry Qin, Jinman Kim, David Dagan Feng
IEEE Trans. Image Process.10
2020 Vision-Based Freezing of Gait Detection With Anatomic Directed Graph Representation
abstract
Parkinson's disease significantly impacts the life quality of millions of people around the world. While freezing of gait (FoG) is one of the most common symptoms of the disease, it is time consuming and subjective to assess FoG for well-trained experts. Therefore, it is highly desirable to devise computer-aided FoG detection methods for the purpose of objective and time-efficient assessment. In this paper, in line with the gold standard of FoG clinical assessment, which requires video or direct observation, we propose one of the first vision-based methods for automatic FoG detection. To better characterize FoG patterns, instead of learning an overall representation of a video, we propose a novel architecture of graph convolution neural network and represent each video as a directed graph where FoG related candidate regions are the vertices. A weakly-supervised learning strategy and a weighted adjacency matrix estimation layer are proposed to eliminate the resource expensive data annotation required for fully supervised learning. As a result, the interference of visual information irrelevant to FoG, such as gait motion of supporting staff involved in clinical assessments, has been reduced to improve FoG detection performance by identifying the vertices contributing to FoG events. To further improve the performance, the global context of a clinical video is also considered and several fusion strategies with graph predictions are investigated. Experimental results on more than 100 videos collected from 45 patients during a clinical assessment demonstrated promising performance of our proposed method with an AUC of 0.887.
Kun Hu 0008, Zhiyong Wang 0001, Shaohui Mei, Kaylena A. Ehgoetz Martens, Simon J. G. Lewis, David Dagan Feng
IEEE J. Biomed. Health Informatics7
2020 A Residual Based Attention Model for EEG Based Sleep Staging
abstract
Sleep staging is to score the sleep state of a subject into different sleep stages such as Wake and Rapid Eye Movement (REM). It plays an indispensable role in the diagnosis and treatment of sleep disorders. As manual sleep staging through well-trained sleep experts is time consuming, tedious, and subjective, many automatic methods have been developed for accurate, efficient, and objective sleep staging. Recently, deep learning based methods have been successfully proposed for electroencephalogram (EEG) based sleep staging with promising results. However, most of these methods directly take EEG raw signals as input of convolutional neural networks (CNNs) without considering the domain knowledge of EEG staging. Apart from that, to capture temporal information, most of the existing methods utilize recurrent neural networks such as LSTM (Long Short Term Memory) which are not effective for modelling global temporal context and difficult to train. Therefore, inspired by the clinical guidelines of sleep staging such as AASM (American Academy of Sleep Medicine) rules where different stages are generally characterized by EEG waveforms of various frequencies, we propose a multi-scale deep architecture by decomposing an EEG signal into different frequency bands as input to CNNs. To model global temporal context, we utilize the multi-head self-attention module of the transformer model to not only improve performance, but also shorten the training time. In addition, we choose residual based architecture which makes training end-to-end. Experimental results on two widely used sleep staging datasets, Montreal Archive of Sleep Studies (MASS) and sleep-EDF datasets, demonstrate the effectiveness and significant efficiency (up to 12 times less training time) of our proposed method over the state-of-the-art.
Zhiyong Wang 0001, Hong Hong 0001, Zheru Chi, David Dagan Feng, Ronald R. Grunstein, Christopher James Gordon
IEEE J. Biomed. Health Informatics5
2020 Non-Contact Sleep Stage Detection Using Canonical Correlation Analysis of Respiratory Sound
abstract
Respiratory sound is able to differentiate sleep stages and provide a non-contact and cost-effective solution for the diagnosis and treatment monitoring of sleep-related diseases. While most of the existing respiratory sound-based methods focus on a limited number of sleep stages such as sleep/wake and wake/rapid eye movement (REM)/non-REM, it is essential to detect sleep stages at a finer level for sleep quality evaluation. In this paper, we for the first time study a sleep stage detection method aiming at classifying sleep states into four sleep stages: wake, REM, light sleep, and deep sleep from the respiratory sound. In addition to extracting time-domain features, frequency-domain features of respiratory sound, non-linear features of snoring sound are devised to better characterize snoring-related signals of respiratory sound. To effectively fuse the three sets of features, a novel feature fusion technique combining the generalized canonical correlation analysis with the ReliefF algorithm is proposed for discriminative feature selection. Final stage detection is achieved with popular classifiers including decision tree, support vector machines, K-nearest neighbor, and the ensemble classifier. To evaluate our proposed method, we built an in-house dataset, which is comprised of 13 nights of sleep audio data from a sleep laboratory. Experimental results indicate that our proposed method outperforms the existing related ones and is promising for large-scale non-contact sleep monitoring.
Biao Xue, Boya Deng, Hong Hong 0001, Zhiyong Wang 0001, Xiaohua Zhu 0001, David Dagan Feng
IEEE J. Biomed. Health Informatics6
2020 Unsupervised Domain Adaptation to Classify Medical Images Using Zero-Bias Convolutional Auto-Encoders and Context-Based Feature Augmentation
abstract
The accuracy and robustness of image classification with supervised deep learning are dependent on the availability of large-scale labelled training data. In medical imaging, these large labelled datasets are sparse, mainly related to the complexity in manual annotation. Deep convolutional neural networks (CNNs), with transferable knowledge, have been employed as a solution to limited annotated data through: 1) fine-tuning generic knowledge with a relatively smaller amount of labelled medical imaging data, and 2) learning image representation that is invariant to different domains. These approaches, however, are still reliant on labelled medical image data. Our aim is to use a new hierarchical unsupervised feature extractor to reduce reliance on annotated training data. Our unsupervised approach uses a multi-layer zero-bias convolutional auto-encoder that constrains the transformation of generic features from a pre-trained CNN (for natural images) to non-redundant and locally relevant features for the medical image data. We also propose a context-based feature augmentation scheme to improve the discriminative power of the feature representation. We evaluated our approach on 3 public medical image datasets and compared it to other state-of-the-art supervised CNNs. Our unsupervised approach achieved better accuracy when compared to other conventional unsupervised methods and baseline fine-tuned CNNs.
Euijoon Ahn, Ashnil Kumar, Michael J. Fulham, David Dagan Feng, Jinman Kim
IEEE Trans. Medical Imaging4
2020 Co-Learning Feature Fusion Maps From PET-CT Images of Lung Cancer
abstract
The analysis of multi-modality positron emission tomography and computed tomography (PET-CT) images for computer aided diagnosis applications (e.g., detection and segmentation) requires combining the sensitivity of PET to detect abnormal regions with anatomical localization from CT. Current methods for PET-CT image analysis either process the modalities separately or fuse information from each modality based on knowledge about the image analysis task. These methods generally do not consider the spatially varying visual characteristics that encode different information across the different modalities, which have different priorities at different locations. For example, a high abnormal PET uptake in the lungs is more meaningful for tumor detection than physiological PET uptake in the heart. Our aim is to improve fusion of the complementary information in multi-modality PET-CT with a new supervised convolutional neural network (CNN) that learns to fuse complementary information for multi-modality medical image analysis. Our CNN first encodes modality-specific features and then uses them to derive a spatially varying fusion map that quantifies the relative importance of each modality's features across different spatial locations. These fusion maps are then multiplied with the modality-specific feature maps to obtain a representation of the complementary multi-modality information at different locations, which can then be used for image analysis. We evaluated the ability of our CNN to detect and segment multiple regions (lungs, mediastinum, tumors) with different fusion requirements using a dataset of PET-CT images of lung cancer. We compared our method to baseline techniques for multi-modality image fusion (fused inputs (FS), multi-branch (MB) techniques, and multichannel (MC) techniques) and segmentation. Our findings show that our CNN had a significantly higher foreground detection accuracy (99.29%, p < 0:05) than the fusion baselines (FS: 99.00%, MB: 99.08%, TC: 98.92%) and a significantly higher Dice score (63.85%) than recent PET-CT tumor segmentation methods.
Ashnil Kumar, Michael J. Fulham, David Dagan Feng, Jinman Kim
IEEE Trans. Medical Imaging3
2019 Deep Local-Global Refinement Network for Stent Analysis in IVOCT Images
Yuyu Guo 0002, Lei Bi 0001, Ashnil Kumar, Yue Gao 0002, Ruiyan Zhang, David Dagan Feng, Qian Wang 0001, Jinman Kim
MICCAI (5)6
2019 Stacked Memory Network for Video Summarization
abstract
In recent years, supervised video summarization has achieved promising progress with various recurrent neural networks (RNNs) based methods, which treats video summarization as a sequence-to-sequence learning problem to exploit temporal dependency among video frames across variable ranges. However, RNN has limitations in modelling the long-term temporal dependency for summarizing videos with thousands of frames due to the restricted memory storage unit. Therefore, in this paper we propose a stacked memory network called SMN to explicitly model the long dependency among video frames so that redundancy could be minimized in the video summaries produced. Our proposed SMN consists of two key components: Long Short-Term Memory (LSTM) layer and memory layer, where each LSTM layer is augmented with an external memory layer. In particular, we stack multiple LSTM layers and memory layers hierarchically to integrate the learned representation from prior layers. By combining the hidden states of the LSTM layers and the read representations of the memory layers, our SMN is able to derive more accurate video summaries for individual video frames. Compared with the existing RNN based methods, our SMN is particularly good at capturing long temporal dependency among frames with few additional training parameters. Experimental results on two widely used public benchmark datasets: SumMe and TVsum, demonstrate that our proposed model is able to clearly outperform a number of state-of-the-art ones under various settings.
Junbo Wang 0003, Wei Wang 0115, Zhiyong Wang 0001, Liang Wang 0001, David Dagan Feng, Tieniu Tan
ACM Multimedia5
2019 IntersectGAN: Learning Domain Intersection for Generating Images with Multiple Attributes
abstract
Generative adversarial networks (GANs) have demonstrated great success in generating various visual content. However, images generated by existing GANs are often of attributes (e.g., smiling expression) learned from one image domain. As a result, generating images of multiple attributes requires many real samples possessing multiple attributes which are very resource expensive to be collected. In this paper, we propose a novel GAN, namely IntersectGAN, to learn multiple attributes from different image domains through an intersecting architecture. For example, given two image domains $X_1$ and $X_2$ with certain attributes, the intersection $X_1 \cap X_2$ denotes a new domain where images possess the attributes from both $X_1$ and $X_2$ domains. The proposed IntersectGAN consists of two discriminators $D_1$ and $D_2$ to distinguish between generated and real samples of different domains, and three generators where the intersection generator is trained against both discriminators. And an overall adversarial loss function is defined over three generators. As a result, our proposed IntersectGAN can be trained on multiple domains of which each presents one specific attribute, and eventually eliminates the need of real sample images simultaneously possessing multiple attributes. By using the CelebFaces Attributes dataset, our proposed IntersectGAN is able to produce high quality face images possessing multiple attributes (e.g., a face with black hair and a smiling expression). Both qualitative and quantitative evaluations are conducted to compare our proposed IntersectGAN with other baseline methods. Besides, several different applications of IntersectGAN have been explored with promising results.
Zehui Yao, Zhiyong Wang 0001, Wanli Ouyang, Dong Xu 0001, David Dagan Feng
ACM Multimedia6
2019 A study on multi-kernel intuitionistic fuzzy C-means clustering with multiple attributes
Shan Zeng, Zhiyong Wang 0001, Rui Huang 0001, David Dagan Feng
Neurocomputing5
2019 Convolutional sparse kernel network for unsupervised medical image analysis
Euijoon Ahn, Ashnil Kumar, Michael J. Fulham, David Dagan Feng, Jinman Kim
Medical Image Anal.4
2019 Robust video summarization using collaborative representation of adjacent frames
Mingyang Ma 0004, Shaohui Mei, Shuai Wan, Zhiyong Wang 0001, David Dagan Feng
Multim. Tools Appl.5
2019 Feature covariance matrix-based dynamic hand gesture recognition
Linpu Fang, Guile Wu, Wenxiong Kang, Qiuxia Wu, Zhiyong Wang 0001, David Dagan Feng
Neural Comput. Appl.6
2019 Step-wise integration of deep class-specific learning for dermoscopic image segmentation
Lei Bi 0001, Jinman Kim, Euijoon Ahn, Ashnil Kumar, David Dagan Feng, Michael J. Fulham
Pattern Recognit.5
2019 Retinal Vessel Segmentation Using Minimum Spanning Superpixel Tree Detector
abstract
The retinal vessel is one of the determining factors in an ophthalmic examination. Automatic extraction of retinal vessels from low-quality retinal images still remains a challenging problem. In this paper, we propose a robust and effective approach that qualitatively improves the detection of low-contrast and narrow vessels. Rather than using the pixel grid, we use a superpixel as the elementary unit of our vessel segmentation scheme. We regularize this scheme by combining the geometrical structure, texture, color, and space information in the superpixel graph. And the segmentation results are then refined by employing the efficient minimum spanning superpixel tree to detect and capture both global and local structure of the retinal images. Such an effective and structure-aware tree detector significantly improves the detection around the pathologic area. Experimental results have shown that the proposed technique achieves advantageous connectivity-area-length (CAL) scores of 80.92% and 69.06% on two public datasets, namely, DRIVE and STARE, thereby outperforming state-of-the-art segmentation methods. In addition, the tests on the challenging retinal image database have further demonstrated the effectiveness of our method. Our approach achieves satisfactory segmentation performance in comparison with state-of-the-art methods. Our technique provides an automated method for effectively extracting the vessel from fundus images.
Bin Sheng 0001, Ping Li 0016, Shuangjia Mo, Huating Li, Xuhong Hou, Harry Qin, Ruogu Fang, David Dagan Feng
IEEE Trans. Cybern.9
2019 Fast and Accurate Retinal Identification System: Using Retinal Blood Vasculature Landmarks
abstract
The expansion of automation techniques and increased risk of identity theft have led emphasis on the tremendous need of automated identification system. Due to the high recognition accuracy and robustness to changes in human physiology, retinal biometric identification system has drawn much attention in this research field. In this paper, we aim to propose an automatic fast and accurate retinal identification system for the multisample dataset. The proposed approach uses a hybrid segmentation technique to segment out both thick/thin vessels for effectively balancing the difference of wavelet response between thick/thin blood vessels. As a result, recognition accuracy is improved. A Principle Component Analysis-based feature processing approach is proposed for efficiently reducing the dimensionality of a large number of vessels features. It significantly reduces computation time and accelerates the matching process in the retinal identification system. The proposed technique is validated on DRIVE, STARE, VARIA, RIDB, HRF, Messidor, DIARETDB0, and a large multisample per subject database created by authors using the images provided by Dr. Chen (Shanghai Jiao Tong University Affiliated Sixth People Hospital). Experimental results demonstrated that the proposed approach outperforms other existing techniques. Segmentation achieves an overall accuracy of 99.65% with the recognition rate of 99.40% on all these databases.
Sidra Aleem, Bin Sheng 0001, Ping Li 0016, Po Yang 0001, David Dagan Feng
IEEE Trans. Ind. Informatics5
2019 Illumination-Guided Video Composition via Gradient Consistency Optimization
abstract
Video composition aims at cloning a patch from the source video into the target scene to create a seamless and harmonious blending frame sequence. Previous work in video composition usually suffer from artifacts around the blending region and spatial-temporal consistency when illumination intensity varies in the input source and target video. We propose an illumination-guided video composition method via a unified spatial and temporal optimization framework. Our method can produce globally consistent composition results and maintain the temporal coherency. We first compute a spatial-temporal blending boundary iteratively. For each frame, the gradient field of the target and source frames are mixed adaptively based on gradients and inter-frame color difference. The temporal consistency is further obtained by optimizing luminance gradients throughout all the composition frames. Moreover, we extend the mean-value cloning by smoothing discrepancies between the source and target frames, then eliminate the color distribution overflow exponentially to reduce falsely blending pixels. Various experiments have shown the effectiveness and high-quality performance of our illumination-guided composition.
Jingye Wang, Bin Sheng 0001, Ping Li 0016, Yuxi Jin, David Dagan Feng
IEEE Trans. Image Process.5
2019 Deep Color Guided Coarse-to-Fine Convolutional Network Cascade for Depth Image Super-Resolution
abstract
Depth image super-resolution is a significant yet challenging task. In this paper, we introduce a novel deep color guided coarse-to-fine convolutional neural network (CNN) framework to address this problem. First, we present a datadriven filter method to approximate the ideal filter for depth image super-resolution instead of hand-designed filters. Based on large data samples, the filter learned is more accurate and stable for upsampling depth image. Second, we introduce a coarse-to-fine CNN to learn different sizes of filter kernels. In coarse stage, larger filter kernels are learned by CNN to achieve crude high-resolution depth image. As to fine stage, the crude high-resolution depth image is used as the input so that smaller filter kernels are learned to gain more accurate results. Benefit from this network, we can progressively recover the high frequency details. Third, we construct a color guidance strategy that fuses color difference and spatial distance for depth image upsampling. We revise the interpolated high-resolution depth image according to the corresponding pixels in highresolution color maps. Guided by color information, the depth of high-resolution image obtained can alleviate texture copying artifacts and preserve edge details effectively. Quantitative and qualitative experimental results demonstrate our state-of-the-art performance for depth map super-resolution.
Bin Sheng 0001, Ping Li 0016, Weiyao Lin, David Dagan Feng
IEEE Trans. Image Process.5
2019 Unsupervised Two-Path Neural Network for Cell Event Detection and Classification Using Spatiotemporal Patterns
abstract
Automatic event detection in cell videos is essential for monitoring cell populations in biomedicine. Deep learning methods have advantages over traditional approaches for cell event detection due to their ability to capture more discriminative features of cellular processes. Supervised deep learning methods, however, are inherently limited due to the scarcity of annotated data. Unsupervised deep learning methods have shown promise in general (non-cell) videos because they can learn the visual appearance and motion of regularly occurring events. Cell videos, however, can have rapid, irregular changes in cell appearance and motion, such as during cell division and death, which are often the events of most interest. We propose a novel unsupervised two-path input neural network architecture to capture these irregular events with three key elements: 1) a visual encoding path to capture regular spatiotemporal patterns of observed objects with convolutional long short-term memory units; 2) an event detection path to extract information related to irregular events with max-pooling layers; and 3) integration of the hidden states of the two paths to provide a comprehensive representation of the video that is used to simultaneously locate and classify cell events. We evaluated our network in detecting cell division in densely packed stem cells in phase-contrast microscopy videos. Our unsupervised method achieved higher or comparable accuracy to standard and state-of-the-art supervised methods.
Ha Tran Hong Phan, Ashnil Kumar, David Dagan Feng, Michael J. Fulham, Jinman Kim
IEEE Trans. Medical Imaging3
2019 Knowledge-based Collaborative Deep Learning for Benign-Malignant Lung Nodule Classification on Chest CT
abstract
The accurate identification of malignant lung nodules on chest CT is critical for the early detection of lung cancer, which also offers patients the best chance of cure. Deep learning methods have recently been successfully introduced to computer vision problems, although substantial challenges remain in the detection of malignant nodules due to the lack of large training data sets. In this paper, we propose a multi-view knowledge-based collaborative (MV-KBC) deep model to separate malignant from benign nodules using limited chest CT data. Our model learns 3-D lung nodule characteristics by decomposing a 3-D nodule into nine fixed views. For each view, we construct a knowledge-based collaborative (KBC) submodel, where three types of image patches are designed to fine-tune three pre-trained ResNet-50 networks that characterize the nodules' overall appearance, voxel, and shape heterogeneity, respectively. We jointly use the nine KBC submodels to classify lung nodules with an adaptive weighting scheme learned during the error back propagation, which enables the MV-KBC model to be trained in an end-to-end manner. The penalty loss function is used for better reduction of the false negative rate with a minimal effect on the overall performance of the MV-KBC model. We tested our method on the benchmark LIDC-IDRI data set and compared it to the five state-of-the-art classification approaches. Our results show that the MV-KBC model achieved an accuracy of 91.60% for lung nodule classification with an AUC of 95.70%. These results are markedly superior to the state-of-the-art approaches.
Yutong Xie 0001, Yong Xia 0001, Yang Song 0001, David Dagan Feng, Michael J. Fulham, Tom Weidong Cai
IEEE Trans. Medical Imaging5
2019 Deep Convolutional Neural Networks for Human Action Recognition Using Depth Maps and Postures
abstract
In this paper, we present a method (Action-Fusion) for human action recognition from depth maps and posture data using convolutional neural networks (CNNs). Two input descriptors are used for action representation. The first input is a depth motion image that accumulates consecutive depth maps of a human action, whilst the second input is a proposed moving joints descriptor which represents the motion of body joints over time. In order to maximize feature extraction for accurate action classification, three CNN channels are trained with different inputs. The first channel is trained with depth motion images (DMIs), the second channel is trained with both DMIs and moving joint descriptors together, and the third channel is trained with moving joint descriptors only. The action predictions generated from the three CNN channels are fused together for the final action classification. We propose several fusion score operations to maximize the score of the right action. The experiments show that the results of fusing the output of three channels are better than using one channel or fusing two channels only. Our proposed method was evaluated on three public datasets: 1) Microsoft action 3-D dataset (MSRAction3D); 2) University of Texas at Dallas-multimodal human action dataset; and 3) multimodal action dataset (MAD) dataset. The testing results indicate that the proposed approach outperforms most of existing state-of-the-art methods, such as histogram of oriented 4-D normals and Actionlet on MSRAction3D. Although MAD dataset contains a high number of actions (35 actions) compared to existing action RGB-D datasets, this paper surpasses a state-of-the-art method on the dataset by 6.84%.
Aouaidjia Kamel, Bin Sheng 0001, Po Yang 0001, Ping Li 0016, Ruimin Shen, David Dagan Feng
IEEE Trans. Syst. Man Cybern. Syst.6
2018 Voxelized Facial Reconstruction Using Deep Neural Network
abstract
This paper presents an approach to predicting variation tendency of human faces with regard to cranium changes based on deep learning. Our work focuses on generating individual customized facial models with high plausibility. Inspired by the performance of encoder-decoder convolutional neural network, the core trainable predicting engine of our learning network is designed for three-dimension voxelized data representation as the encoder-decoder structure and the encoder part is similar to the 7 layers of VGG16 network. To take full consideration of the cranium changes and features of original human face, a novel formation of channeled volumetric data structure is presented, and also the corresponding sub and up-sampling strategies for volume data. Our encoder-decoder neural network consumes discrete 3-channel volume data and generates 1-channel volume data as predicted post-variation human face. This framework is quantified with clinical dataset and it shows that its' performance improves in comparison with the state-of-the-art technologies.
Xiaoshuang Li, Bin Sheng 0001, Ping Li 0016, Jinman Kim, David Dagan Feng
CGI5
2018 Prior Knowledge Driven Energy for Saliency Detection
abstract
Saliency detection on images has experienced substantial progress in recent years on the basis of deep neural network (DNN). However, there may exist secondary saliency in the background that distracts DNN learning and mistakes the secondary salient regions as saliency. To address this issue, we propose a dual-term energy to improve the inference of saliency on top of DNN estimation, where dense term smoothens salient regions in pixel scale and sparse term extracts prior knowledge to differentiate saliency and non-saliency superpixels. Our prior knowledge including extra- and intra-region priors, contributes to improving overall saliency detection. The extra-region prior knowledge estimates the saliency probabilities for different pre-partitioned regions to eliminate the secondary saliency. The intra-region prior knowledge helps to group the salient regions that otherwise could be ignored by DNN predictor, and thus to provide more complete saliency definition. We evaluated our model on 8,465 images from four well-recognized saliency detection benchmarking datasets, and compared our model to six state-of-the-art comparative methods. Experimental results demonstrated that our model outperformed the state-of-the-art counterpart with improvements of up to 2.51% in terms of F-measure.
Ke Yan 0005, Chaojie Zheng, Qiu Huang, Jinman Kim, David Dagan Feng, Xiuying Wang 0001
ICARCV5
2018 Densely Connected Large Kernel Convolutional Network for Semantic Membrane Segmentation in Microscopy Images
abstract
Structural analysis of neurons can provide valuable insights of brain function. Semantic segmentation of neurons thus becomes an important technique in bioinformatics. Deep learning approaches have shown promising performance in various semantic segmentation problems. However, segmentation of neurons in Electron Microscopy (EM) images has some differences compared with typical segmentation tasks due to the image noise and the disturbance of the intracellular structures. In our work, we propose a network with a ResNet encoder and densely connected decoder with large kernels, and then refinement with simple morphological post-possessing. Two main advantages of our method are: 1) the network can prevent the loss of high-resolution information and enlarge the reception field; 2) the post-processing method is simple and can be directly applied to the probability map from the network to enhance the unconfident area. Evaluated on the ISBI2012 EM membrane segmentation challenge, the proposed method achieves competitive performance.
Dongnan Liu, Donghao Zhang 0004, Siqi Liu 0001, Yang Song 0001, Haozhe Jia, David Dagan Feng, Yong Xia 0001, Tom Weidong Cai
ICIP6
2018 Video Summarization via Weighted Neighborhood Based Representation
abstract
The recent explosive growth of multimedia data has posed a new set of challenges in computer vision, and video summarization (VS) techniques are increasingly important to automatically summarize a large amount of multimedia data in an effective and efficient manner. Recent years have witnessed the rise and developments of sparse representation based approaches for VS. While the existing methods select keyframes according to the information contained in the single frame, and such a selection based solely on single-frame information may not be robust. Therefore, in this paper, the information of the single frame's neighborhood is taken into consideration, and different weights are assigned to these neighbouring frames. We formulate the VS problem as a weighted neighborhood based representation model, and design a greedy pursuit algorithm to extract keyframes. Experimental results on a benchmark dataset demonstrate that the proposed method can outperform the state of the arts.
Mingyang Ma 0004, Shaohui Mei, Shuai Wan, Zhiyong Wang 0001, Ah Chung Tsoi, David Dagan Feng
ICIP6
2018 Feature of Interest-Based Direct Volume Rendering Using Contextual Saliency-Driven Ray Profile Analysis
abstract
Abstract Direct volume rendering (DVR) visualization helps interpretation because it allows users to focus attention on the subset of volumetric data that is of most interest to them. The ideal visualization of the features of interest (FOIs) in a volume, however, is still a major challenge. The clear depiction of FOIs depends on accurate identification of the FOIs and appropriate specification of the optical parameters via transfer function (TF) design and it is typically a repetitive trial‐and‐error process. We address this challenge by introducing a new method that uses contextual saliency information to group the voxels along a viewing ray into distinct FOIs where ‘contextual saliency’ is a biologically inspired attribute that aids the identification of features that the human visual system considers important. The saliency information is also used to automatically define the optical parameters that emphasize the visual depiction of the FOIs in DVR. We demonstrate the capabilities of our method by its application to a variety of volumetric data sets and highlight its advantages by comparison to current state‐of‐the‐art ray profile analysis methods.
Younhyun Jung, Jinman Kim, Ashnil Kumar, David Dagan Feng, Michael J. Fulham
Comput. Graph. Forum4
2018 Atlas registration and ensemble deep convolutional neural network-based prostate segmentation using magnetic resonance imaging
Haozhe Jia, Yong Xia 0001, Yang Song 0001, Tom Weidong Cai, Michael J. Fulham, David Dagan Feng
Neurocomputing6
2018 Computer-Assisted Decision Support System in Pulmonary Cancer detection and stage classification on CT images
Anum Masood, Bin Sheng 0001, Ping Li 0016, Xuhong Hou, Xiaoer Wei, Harry Qin, David Dagan Feng
J. Biomed. Informatics7
2018 Exploiting spatial-temporal context for trajectory based action video retrieval
Lelin Zhang, Zhiyong Wang 0001, Shin'ichi Staoh, Tao Mei 0001, David Dagan Feng
Multim. Tools Appl.6
2018 End-User Development for Interactive Data Analytics: Uncertainty, Correlation and User Confidence
abstract
This paper investigates End-User Development (EUD) for interactive data-analytic interfaces-building upon the ideas of making machine learning transparent. The research is carried out in a business operation environment (water pipe failure prediction in our case) motivated to integrate advanced analytics into decision-making processes of an urban Internet of Things (IoT) concept. We explore effects of revealing uncertainty and correlation on user confidence in a data-driven decision making scenario. It was found that user confidence varied significantly amongst various user groups when different machine learning models were displayed with/without supplementary information. Galvanic Skin Response (GSR) signals were analyzed and shown as reasonable indices for predicting user confidence levels. Supplementary data visualizations (of inherent uncertainty and correlation in data) contributed to explicability principles while GSR indexing added towards correctibility principles. We recommend transparent machine learning as the key to effective EUD for interactive data analytics.
Jianlong Zhou, Syed Arshad, Xiuying Wang 0001, Zhidong Li, David Dagan Feng, Fang Chen 0001
IEEE Trans. Affect. Comput.5
2018 Dense and Sparse Labeling With Multidimensional Features for Saliency Detection
abstract
Conventional low-level feature-based saliency detection methods tend to use nonrobust prior knowledge and do not perform well in complex or low-contrast images. In this paper, to address these issues in existing methods, we propose a novel deep neural network (DNN)-based dense and sparse labeling (DSL) framework for saliency detection. DSL consists of three major steps, namely, dense labeling (DL), sparse labeling (SL), and deep convolutional (DC) network. The DL and SL steps conduct initial saliency estimations with macro object contours and low-level image features, respectively, which effectively approximate the location of the salient object and generate accurate guidance channels for the DC step; the DC step, on the other hand, takes in the results of DL and SL, establishes a six-channeled input data structure (including local superpixel information), and conducts accurate final saliency classification. Our DSL framework exploits the saliency estimation guidance from both macro object contours and local low-level features, as well as utilizing the DNN for high-level saliency feature extraction. Extensive experiments are conducted on six well-recognized public data sets against 16 state-of-the-art saliency detection methods, including ten conventional feature-based methods and six learning-based methods. The results demonstrate the superior performance of DSL on various challenging cases in terms of both accuracy and robustness.
Yuchen Yuan, ChangYang Li, Jinman Kim, Tom Weidong Cai, David Dagan Feng
IEEE Trans. Circuits Syst. Video Technol.5
2018 A Unified Collaborative Multikernel Fuzzy Clustering for Multiview Data
abstract
Clustering is increasingly important for multiview data analytics and current algorithms are either based on the collaborative learning of local partitions or directly derived global clustering from multikernel learning. In this paper, we innovate a clustering model that unifies the local partitions and global clustering in a collaborative learning framework. We first construct a common multikernel space from a set of basis kernels to better reflect clustering information of each individual view. Then, considering that joint local partitions would conform to the global clustering, we fuse the local partitions and global clustering guidance as a single objective function in accordance with fuzzy clustering form. The collaborative learning strategy enables the mutual and interactive clustering from local partitions and global clustering. The validation was performed over two synthetic and four public databases and the clustering accuracy was measured by normalized mutual information and rand index. The experimental results demonstrated that the proposed algorithm outperformed the related state-of-the-art algorithms in comparison, which included multitask, multikernel, and multiview clustering approaches.
Shan Zeng, Xiuying Wang 0001, Hui Cui 0002, Chaojie Zheng, David Dagan Feng
IEEE Trans. Fuzzy Syst.5
2018 Reversion Correction and Regularized Random Walk Ranking for Saliency Detection
abstract
In recent saliency detection research, many graph-based algorithms have applied boundary priors as background queries, which may generate completely "reversed" saliency maps if the salient objects are on the image boundaries. Moreover, these algorithms usually depend heavily on pre-processed superpixel segmentation, which may lead to notable degradation in image detail features. In this paper, a novel saliency detection method is proposed to overcome the above issues. First, we propose a saliency reversion correction process, which locates and removes the boundary-adjacent foreground superpixels, and thereby increases the accuracy and robustness of the boundary prior-based saliency estimations. Second, we propose a regularized random walk ranking model, which introduces prior saliency estimation to every pixel in the image by taking both region and pixel image features into account, thus leading to pixel-detailed and superpixel-independent saliency maps. Experiments are conducted on four well-recognized data sets; the results indicate the superiority of our proposed method against 14 state-of-the-art methods, and demonstrate its general extensibility as a saliency optimization algorithm. We further evaluate our method on a new data set comprised of images that we define as boundary adjacent object saliency, on which our method performs better than the comparison methods.
Yuchen Yuan, ChangYang Li, Jinman Kim, Tom Weidong Cai, David Dagan Feng
IEEE Trans. Image Process.5
2018 Classification of Medical Images in the Biomedical Literature by Jointly Using Deep and Handcrafted Visual Features
abstract
The classification of medical images and illustrations from the biomedical literature is important for automated literature review, retrieval, and mining. Although deep learning is effective for large-scale image classification, it may not be the optimal choice for this task as there is only a small training dataset. We propose a combined deep and handcrafted visual feature (CDHVF) based algorithm that uses features learned by three fine-tuned and pretrained deep convolutional neural networks (DCNNs) and two handcrafted descriptors in a joint approach. We evaluated the CDHVF algorithm on the ImageCLEF 2016 Subfigure Classification dataset and it achieved an accuracy of 85.47%, which is higher than the best performance of other purely visual approaches listed in the challenge leaderboard. Our results indicate that handcrafted features complement the image representation learned by DCNNs on small training datasets and improve accuracy in certain medical image classification problems.
Yong Xia 0001, Yutong Xie 0001, Michael J. Fulham, David Dagan Feng
IEEE J. Biomed. Health Informatics5
2018 Multi-Pass Fast Watershed for Accurate Segmentation of Overlapping Cervical Cells
abstract
The task of segmenting cell nuclei and cytoplasm in pap smear images is one of the most challenging tasks in automated cervix cytological analysis due to specifically the presence of overlapping cells. This paper introduces a multi-pass fast watershed-based method (MPFW) to segment both nucleus and cytoplasm from large cell masses of overlapping cervical cells in three watershed passes. The first pass locates the nuclei with barrier-based watershed on the gradient-based edge map of a pre-processed image. The next pass segments the isolated, touching, and partially overlapping cells with a watershed transform adapted to the cell shape and location. The final pass introduces mutual iterative watersheds separately applied to each nucleus in the largely overlapping clusters to estimate the cell shape. In MPFW, the line-shaped contours of the watershed cells are deformed with ellipse fitting and contour adjustment to give a better representation of cell shapes. The performance of the proposed method has been evaluated using synthetic, real extended depth-of-field, and multi-layers cervical cytology images provided by the first and second overlapping cervical cytology image segmentation challenges in ISBI 2014 and ISBI 2015. The experimental results demonstrate superior performance of the proposed MPFW in terms of segmentation accuracy, detection rate, and time complexity, compared with recent peer methods.
Afaf Tareef, Yang Song 0001, Heng Huang 0001, David Dagan Feng, Yue Joseph Wang, Tom Weidong Cai
IEEE Trans. Medical Imaging4
2018 Real-Time Long-Term Tracking With Prediction-Detection-Correction
abstract
Real-time long-term visual tracking is one of the most challenging problems in computer vision due to various factors such as occlusion and motion ambiguity. To achieve robust long-term tracking, most state-of-the-art methods typically construct an online detector in each frame. However, they fail to achieve real-time performance due to high computational complexity. In this paper, we propose a novel real-time long-term tracking algorithm by exploiting a joint Prediction-Detection-Correction Tracking framework (PDCT). We utilize a superpixel optical flow to construct a predictor to estimate the target motion and internal scale variation. To locate the target at a finer level, we develop an improved kernelized correlation detector with an adaptive online learning rate and translation-scale parameters from the predictor. To refine the tracking result and redetect the target in the case of a tracking failure, we devise a corrector utilizing dual online SVMs with dense sampling and reliable history samples. The SVMs are trained with passive-aggressive learning and online retraining strategies. In addition, we employ a selection mechanism for the correlation responses to maintain reliable samples effectively. As a result, our proposed tracker is able to refine tracking results via the corrector and detector and maintains reliable tracking results for subsequent tracking. Extensive experiments on the widely used object tracking benchmark show that the proposed tracker is superior to state-of-the-art trackers in terms of both effectiveness and efficiency, and the integration of each component is effective under the PDCT framework.
Ningxin Liang, Guile Wu, Wenxiong Kang, Zhiyong Wang 0001, David Dagan Feng
IEEE Trans. Multim.5
2018 Dual-Path Adversarial Learning for Fully Convolutional Network (FCN)-Based Medical Image Segmentation
Lei Bi 0001, David Dagan Feng, Jinman Kim
Vis. Comput.2
2017 Exploring the influence of feature representation for dictionary selection based video summarization
abstract
Dictionary selection based video summarization (VS) algorithms, in which keyframes are considered as a dictionary to reconstruct all the video frames, have been demonstrated to be effective and efficient for video summarization. It has been noticed that the feature representation of video plays a great impact of the performance of VS. In this paper, the influence of feature representation of video frames on the performance of dictionary selection-based VS is for the first time investigated. In addition to the traditional hand-crafted features used in VS, such as color histogram, the deep features learned through deep neural networks are firstly used to represent video frames for dictionary selection-based VS. The impact of dimensionality reduction to the high-dimensional deep learning features on VS is further discussed. Experimental results on a benchmark video dataset demonstrate that deep learning features are able to achieve better performance than traditional hand-crafted features for dictionary selection-based VS. Moreover, the dimensionality of deep learning features can be reduced to decrease the computational cost without the degradation of VS performance.
Mingyang Ma 0004, Shaohui Mei, Jingyu Ji, Shuai Wan, Zhiyong Wang 0001, David Dagan Feng
ICIP6
2017 Nonlinear kernel sparse dictionary selection for video summarization
abstract
Sparse dictionary selection (SDS) has demonstrated to be an effective solution for keyframe based video summarization (VS), which generally assumes a linear relation among similar video frames. However, such a linear assumption is not always true for videos. In this paper, the nonlinearity among frames is taken into consideration and a nonlinear SDS model is formulated for VS, in which the nonlinearity is transformed to linearity by projecting a video to a high dimensional feature space induced by a kernel function. Moreover, a kernel simultaneous orthogonal matching pursuit (KSOMP) is proposed to solve the problem. In order to achieve an intuitive and flexible configuration of the VS process, an adaptive criterion is devised to produce video summaries with different lengths for different video content. Experimental results on benchmark video datasets demonstrate that the proposed algorithm outperforms several state-of-the-art VS algorithms.
Mingyang Ma 0004, Shaohui Mei, Junhui Hou, Shuai Wan, Zhiyong Wang 0001, David Dagan Feng
ICME6
2017 Neural net-based and safety-oriented visual analytics for time-spatial data
abstract
Safety-oriented visualization is one of significant approaches to gain insights from time-spatial data while neural net currently serves as a decent way to perform machine learning in data mining industry. This paper proposes a visual analytics pipeline for trajectory data enabling better understanding movements pattern of people using Neural Network as back-end and other visualization techniques as front-end for gaining information of preferences of attractions, similarities of groups, popularities of attractions and pattern of movement flow. Such understandings help to address the management issue by extracting the outstanding features to detect abnormal pattern such as detection of crime and predicting overall movements, and so on. Successfully dealing with those issues would have significant improvements of entire management of public facility such as parks and transportation.
Jianlong Zhou, Xiuying Wang 0001, Jeremy Swanson, Fang Chen 0001, David Dagan Feng
IJCNN6
2017 Transferable Multi-model Ensemble for Benign-Malignant Lung Nodule Classification on Chest CT
Yutong Xie 0001, Yong Xia 0001, David Dagan Feng, Michael J. Fulham, Tom Weidong Cai
MICCAI (3)4
2017 Visual tracking utilizing robust complementary learner and adaptive refiner
Guile Wu, Wenxiong Kang, Zhiyong Wang 0001, David Dagan Feng
Neurocomputing5
2017 Automatic segmentation of overlapping cervical smear cells based on local distinctive features and guided shape deformation
Afaf Tareef, Yang Song 0001, Tom Weidong Cai, Heng Huang 0001, Hang Chang, Yue Joseph Wang, Michael J. Fulham, David Dagan Feng
Neurocomputing8
2017 Optimizing the cervix cytological examination based on deep learning and dynamic shape modeling
Afaf Tareef, Yang Song 0001, Heng Huang 0001, Yue Joseph Wang, David Dagan Feng, Tom Weidong Cai
Neurocomputing5
2017 Dual discriminative local coding for tissue aging analysis
Yang Song 0001, Qing Li 0012, Fan Zhang 0013, Heng Huang 0001, David Dagan Feng, Yue Joseph Wang, Tom Weidong Cai
Medical Image Anal.5
2017 Learning universal multiview dictionary for human action recognition
Zhiyong Wang 0001, Zhao Xie, Jun Gao 0006, David Dagan Feng
Pattern Recognit.5
2017 Saliency-Based Lesion Segmentation Via Background Detection in Dermoscopic Images
abstract
The segmentation of skin lesions in dermoscopic images is a fundamental step in automated computer-aided diagnosis of melanoma. Conventional segmentation methods, however, have difficulties when the lesion borders are indistinct and when contrast between the lesion and the surrounding skin is low. They also perform poorly when there is a heterogeneous background or a lesion that touches the image boundaries; this then results in under- and oversegmentation of the skin lesion. We suggest that saliency detection using the reconstruction errors derived from a sparse representation model coupled with a novel background detection can more accurately discriminate the lesion from surrounding regions. We further propose a Bayesian framework that better delineates the shape and boundaries of the lesion. We also evaluated our approach on two public datasets comprising 1100 dermoscopic images and compared it to other conventional and state-of-the-art unsupervised (i.e., no training required) lesion segmentation methods, as well as the state-of-the-art unsupervised saliency detection methods. Our results show that our approach is more accurate and robust in segmenting lesions compared to other methods. We also discuss the general extension of our framework as a saliency optimization algorithm for lesion segmentation.
Euijoon Ahn, Jinman Kim, Lei Bi 0001, Ashnil Kumar, ChangYang Li, Michael J. Fulham, David Dagan Feng
IEEE J. Biomed. Health Informatics7
2017 Occlusion and Slice-Based Volume Rendering Augmentation for PET-CT
abstract
Dual-modality positron emission tomography and computed tomography (PET-CT) depicts pathophysiological function with PET in an anatomical context provided by CT. Three-dimensional volume rendering approaches enable visualization of a two-dimensional slice of interest (SOI) from PET combined with direct volume rendering (DVR) from CT. However, because DVR depicts the whole volume, it may occlude a region of interest, such as a tumor in the SOI. Volume clipping can eliminate this occlusion by cutting away parts of the volume, but it requires intensive user involvement in deciding on the appropriate depth to clip. Transfer functions that are currently available can make the regions of interest visible, but this often requires complex parameter tuning and coupled preprocessing of the data to define the regions. Hence, we propose a new visualization algorithm where an SOI from PET is augmented by volumetric contextual information from a DVR of the counterpart CT so that the obtrusiveness from the CT in the SOI is minimized. Our approach automatically calculates an augmentation depth parameter by considering the occlusion information derived from the voxels of the CT in front of the PET SOI. The depth parameter is then used to generate an opacity weight function that controls the amount of contextual information visible from the DVR. We outline the improvements with our visualization approach compared to other slice-based and our previous approaches. We present the preliminary clinical evaluation of our visualization in a series of PET-CT studies from patients with nonsmall cell lung cancer.
Younhyun Jung, Jinman Kim, David Dagan Feng, Michael J. Fulham
IEEE J. Biomed. Health Informatics3
2017 An Ensemble of Fine-Tuned Convolutional Neural Networks for Medical Image Classification
abstract
The availability of medical imaging data from clinical archives, research literature, and clinical manuals, coupled with recent advances in computer vision offer the opportunity for image-based diagnosis, teaching, and biomedical research. However, the content and semantics of an image can vary depending on its modality and as such the identification of image modality is an important preliminary step. The key challenge for automatically classifying the modality of a medical image is due to the visual characteristics of different modalities: some are visually distinct while others may have only subtle differences. This challenge is compounded by variations in the appearance of images based on the diseases depicted and a lack of sufficient training data for some modalities. In this paper, we introduce a new method for classifying medical images that uses an ensemble of different convolutional neural network (CNN) architectures. CNNs are a state-of-the-art image classification technique that learns the optimal image features for a given classification task. We hypothesise that different CNN architectures learn different levels of semantic image representation and thus an ensemble of CNNs will enable higher quality features to be extracted. Our method develops a new feature extractor by fine-tuning CNNs that have been initialized on a large dataset of natural images. The fine-tuning process leverages the generic image features from natural images that are fundamental for all images and optimizes them for the variety of medical imaging modalities. These features are used to train numerous multiclass classifiers whose posterior probabilities are fused to predict the modalities of unseen images. Our experiments on the ImageCLEF 2016 medical image public dataset (30 modalities; 6776 training images, and 4166 test images) show that our ensemble of fine-tuned CNNs achieves a higher accuracy than established CNNs. Our ensemble also achieves a higher accuracy than methods in the literature evaluated on the same benchmark dataset and is only overtaken by those methods that source additional training data.
Ashnil Kumar, Jinman Kim, David Lyndon, Michael J. Fulham, David Dagan Feng
IEEE J. Biomed. Health Informatics5
2017 Automatic Measurement of Thalamic Diameter in 2-D Fetal Ultrasound Brain Images Using Shape Prior Constrained Regularized Level Sets
abstract
We derived an automated algorithm for accurately measuring the thalamic diameter from 2-D fetal ultrasound (US) brain images. The algorithm overcomes the inherent limitations of the US image modality: nonuniform density; missing boundaries; and strong speckle noise. We introduced a "guitar" structure that represents the negative space surrounding the thalamic regions. The guitar acts as a landmark for deriving the widest points of the thalamus even when its boundaries are not identifiable. We augmented a generalized level-set framework with a shape prior and constraints derived from statistical shape models of the guitars; this framework was used to segment US images and measure the thalamic diameter. Our segmentation method achieved a higher mean Dice similarity coefficient, Hausdorff distance, specificity, and reduced contour leakage when compared to other well-established methods. The automatic thalamic diameter measurement had an interobserver variability of -0.56 ± 2.29 mm compared to manual measurement by an expert sonographer. Our method was capable of automatically estimating the thalamic diameter, with the measurement accuracy on par with clinical assessment. Our method can be used as part of computer-assisted screening tools that automatically measure the biometrics of the fetal thalamus; these biometrics are linked to neurodevelopmental outcomes.
Pradeeba Sridar, Ashnil Kumar, ChangYang Li, Joyce Woo, Ann Quinton, Ron Benzie, Michael J. Peek, David Dagan Feng, R. Krishna Kumar, Ralph Nanan, Jinman Kim
IEEE J. Biomed. Health Informatics8
2017 Low Dimensional Representation of Fisher Vectors for Microscopy Image Classification
abstract
Microscopy image classification is important in various biomedical applications, such as cancer subtype identification, and protein localization for high content screening. To achieve automated and effective microscopy image classification, the representative and discriminative capability of image feature descriptors is essential. To this end, in this paper, we propose a new feature representation algorithm to facilitate automated microscopy image classification. In particular, we incorporate Fisher vector (FV) encoding with multiple types of local features that are handcrafted or learned, and we design a separation-guided dimension reduction method to reduce the descriptor dimension while increasing its discriminative capability. Our method is evaluated on four publicly available microscopy image data sets of different imaging types and applications, including the UCSB breast cancer data set, MICCAI 2015 CBTC challenge data set, and IICBU malignant lymphoma, and RNAi data sets. Our experimental results demonstrate the advantage of the proposed low-dimensional FV representation, showing consistent performance improvement over the existing state of the art and the commonly used dimension reduction techniques.
Yang Song 0001, Qing Li 0012, Heng Huang 0001, David Dagan Feng, Tom Weidong Cai
IEEE Trans. Medical Imaging4
2017 Stacked fully convolutional networks with multi-channel learning: application to medical image segmentation
Lei Bi 0001, Jinman Kim, Ashnil Kumar, Michael J. Fulham, David Dagan Feng
Vis. Comput.5
2016 Atmospheric turbulence mitigation based on turbulence extraction
abstract
A video taken under the influence of atmospheric turbulence suffers from serious distortion caused by the variation of optical refractive index. In order to reduce geometric distortion and time-space-varying blur, and recover both coarse structure and fine details, a novel turbulence extraction based approach for recovering a latent image from an atmospheric turbulence degraded imagery sequence is proposed. Firstly, a non-rigid image registration method is applied as a preprocessing to reduce geometric deformation. Secondly, the registered image sequence is decomposed into a low-rank background scene component and a sparse turbulent component via matrix decomposition. Different from other approaches, which intend to remove turbulence directly, we manage to extract information of distortion position from the sparse turbulent component to indicate the sharpest turbulence patches. The selected sharpest turbulence patches are then enhanced and fused to generate an enhanced detail layer. Finally, the output image is generated by fusing the deblurred background scene layer and the enhanced detail layer together. Experiments indicate that our approach is capable of significantly alleviating atmospheric turbulence blur and geometric distortion.
Zhiyong Wang 0001, Yangyu Fan, David Dagan Feng
ICASSP4
2016 Multilevel affinity graph for unsupervised image segmentation
abstract
Unsupervised segmentation and contour detection remains a challenging task. In graph-based unsupervised segmentation, the formulation of the affinity graph is pivotal to segmentation performance. Conventional graph-based approaches often only define pixels as graph nodes, and may overlook important regional information. In this paper, we propose a novel scheme for affinity graph construction, where the affinity weight matrix unifies the association across pixel-wise nodes and multilevel region-wise nodes of different scales. Integrating the multilevel regional information, which is formulated using superpixels, into the affinity graph contributes to better capture of image intensity and color cues. Experimental evaluation of our approach on the BSDS500 dataset showed that our proposed method achieved the second best performance compared to other nine unsupervised state-of-art methods commonly used for comparison.
Xiuying Wang 0001, Ke Yan 0005, ChangYang Li, David Dagan Feng
ICIP5
2016 Adaptive background search and foreground estimation for saliency detection via comprehensive autoencoder
abstract
In saliency object detection, inappropriate boundary-background priors is known to degrade performance in challenging image datasets, and even may lead to `inverse' results when saliency regions are attached to the image boundaries. This is an active field where many works have proposed various techniques to lessen such degradation by inappropriate boundary-background priors. Although the use of boundary-background priors has shown to be capable of improving the detection, inherently, these techniques confront serious challenges in background suppression. To overcome this limitation, we propose an adaptive background extractor to search background seeds without the need of boundary-background priors. With the adaptive background seeds, the saliency objects can be then extracted via our proposed hierarchical foreground estimation model. We evaluate our adaptive Background Search and Foreground Estimation (BSFE) algorithm in comparison with six state-of-the-art methods on four well-recognized public datasets. The experimental results demonstrate that our BSFE algorithm outperforms compared methods in majority of the datasets and in particular achieves double-winners in terms of F-measure and mean absolute error on two challenging datasets.
Ke Yan 0005, ChangYang Li, Xiuying Wang 0001, Yuchen Yuan, Jinman Kim, David Dagan Feng
ICIP7
2016 Bioimage classification with subcategory discriminant transform of high dimensional visual descriptors
abstract
BACKGROUND: Bioimage classification is a fundamental problem for many important biological studies that require accurate cell phenotype recognition, subcellular localization, and histopathological classification. In this paper, we present a new bioimage classification method that can be generally applicable to a wide variety of classification problems. We propose to use a high-dimensional multi-modal descriptor that combines multiple texture features. We also design a novel subcategory discriminant transform (SDT) algorithm to further enhance the discriminative power of descriptors by learning convolution kernels to reduce the within-class variation and increase the between-class difference. RESULTS: We evaluate our method on eight different bioimage classification tasks using the publicly available IICBU 2008 database. Each task comprises a separate dataset, and the collection represents typical subcellular, cellular, and tissue level classification problems. Our method demonstrates improved classification accuracy (0.9 to 9%) on six tasks when compared to state-of-the-art approaches. We also find that SDT outperforms the well-known dimension reduction techniques, with for example 0.2 to 13% improvement over linear discriminant analysis. CONCLUSIONS: We present a general bioimage classification method, which comprises a highly descriptive visual feature representation and a learning-based discriminative feature transformation algorithm. Our evaluation on the IICBU 2008 database demonstrates improved performance over the state-of-the-art for six different classification tasks.
Yang Song 0001, Tom Weidong Cai, Heng Huang 0001, David Dagan Feng, Yue Joseph Wang
BMC Bioinform.4
2016 DeepGene: an advanced cancer type classifier based on deep learning and somatic point mutations
abstract
BACKGROUND: With the developments of DNA sequencing technology, large amounts of sequencing data have become available in recent years and provide unprecedented opportunities for advanced association studies between somatic point mutations and cancer types/subtypes, which may contribute to more accurate somatic point mutation based cancer classification (SMCC). However in existing SMCC methods, issues like high data sparsity, small volume of sample size, and the application of simple linear classifiers, are major obstacles in improving the classification performance. RESULTS: To address the obstacles in existing SMCC studies, we propose DeepGene, an advanced deep neural network (DNN) based classifier, that consists of three steps: firstly, the clustered gene filtering (CGF) concentrates the gene data by mutation occurrence frequency, filtering out the majority of irrelevant genes; secondly, the indexed sparsity reduction (ISR) converts the gene data into indexes of its non-zero elements, thereby significantly suppressing the impact of data sparsity; finally, the data after CGF and ISR is fed into a DNN classifier, which extracts high-level features for accurate classification. Experimental results on our curated TCGA-DeepGene dataset, which is a reformulated subset of the TCGA dataset containing 12 selected types of cancer, show that CGF, ISR and DNN all contribute in improving the overall classification performance. We further compare DeepGene with three widely adopted classifiers and demonstrate that DeepGene has at least 24% performance improvement in terms of testing accuracy. CONCLUSIONS: Based on deep learning and somatic point mutation data, we devise DeepGene, an advanced cancer type classifier, which addresses the obstacles in existing SMCC studies. Experiments indicate that DeepGene outperforms three widely adopted existing classifiers, which is mainly attributed to its deep learning module that is able to extract the high level features between combinatorial somatic point mutations and cancer types.
Yuchen Yuan, Yi Shi 0007, ChangYang Li, Jinman Kim, Tom Weidong Cai, Zeguang Han, David Dagan Feng
BMC Bioinform.7
2016 Topology-aware illumination design for volume rendering
abstract
BACKGROUND: Direct volume rendering is one of flexible and effective approaches to inspect large volumetric data such as medical and biological images. In conventional volume rendering, it is often time consuming to set up a meaningful illumination environment. Moreover, conventional illumination approaches usually assign same values of variables of an illumination model to different structures manually and thus neglect the important illumination variations due to structure differences. RESULTS: We introduce a novel illumination design paradigm for volume rendering on the basis of topology to automate illumination parameter definitions meaningfully. The topological features are extracted from the contour tree of an input volumetric data. The automation of illumination design is achieved based on four aspects of attenuation, distance, saliency, and contrast perception. To better distinguish structures and maximize illuminance perception differences of structures, a two-phase topology-aware illuminance perception contrast model is proposed based on the psychological concept of Just-Noticeable-Difference. CONCLUSIONS: The proposed approach allows meaningful and efficient automatic generations of illumination in volume rendering. Our results showed that our approach is more effective in depth and shape depiction, as well as providing higher perceptual differences between structures.
Jianlong Zhou, Xiuying Wang 0001, Hui Cui 0002, Xianglin Miao, Yalin Miao, Chun Xiao, Fang Chen 0001, David Dagan Feng
BMC Bioinform.9
2016 Texture image classification with discriminative neural networks
abstract
Texture provides an important cue for many computer vision applications, and texture image classification has been an active research area over the past years. Recently, deep learning techniques using convolutional neural networks (CNN) have emerged as the state-of-the-art: CNN-based features provide a significant performance improvement over previous handcrafted features. In this study, we demonstrate that we can further improve the discriminative power of CNN-based features and achieve more accurate classification of texture images. In particular, we have designed a discriminative neural network-based feature transformation (NFT) method, with which the CNN-based features are transformed to lower dimensionality descriptors based on an ensemble of neural networks optimized for the classification objective. For evaluation, we used three standard benchmark datasets (KTH-TIPS2, FMD, and DTD) for texture image classification. Our experimental results show enhanced classification performance over the state-of-the-art.
Yang Song 0001, Qing Li 0012, David Dagan Feng, Ju Jia Zou, Tom Weidong Cai
Comput. Vis. Media3
2016 Dictionary pruning with visual word significance for medical image retrieval
Fan Zhang 0013, Yang Song 0001, Tom Weidong Cai, Alex Hauptmann 0001, Sidong Liu, Sonia Pujol, Ron Kikinis, Michael J. Fulham, David Dagan Feng
Neurocomputing9
2016 Robust foreground object segmentation from handheld camera videos with occlusion map
Hao Xiong 0001, Zhiyong Wang 0001, David Dagan Feng
Multim. Tools Appl.4
2016 Investigating the impact of frame rate towards robust human action recognition
Fredro Harjanto, Zhiyong Wang 0001, Shiyang Lu, Ah Chung Tsoi, David Dagan Feng
Signal Process.5
2016 Foreground Detection With Simultaneous Dictionary Learning and Historical Pixel Maintenance
abstract
Foreground detection is fundamental in surveillance video analysis and meaningful toward object tracking and higher level tasks, such as anomaly detection and activity analysis. Nevertheless, existing methods are still limited in accurately detecting the foreground due to the complex scene settings. To robustly handle the diverse background variations and foreground challenges, this paper proposes a Background REpresentation approach With Dictionary Learning and Historical Pixel Maintenance (BREW-DLHPM). Specifically, a dictionary learning problem is formulated at the frame level to adaptively represent the background signals with the varied structure information captured, while a pixel-level maintenance is exploited to grasp the dynamic nature of historical information under the help of the learned background. The simultaneous utilization of dictionary learning and historical pixel maintenance facilitates the accurate description of the background and thus guides a wise foreground detection decision. The proposed BREW-DLHPM has been evaluated on the prestigious change detection challenge data set against 11 state-of-the-art foreground detection approaches and encouraging performances have been achieved by our method.
Pei Dong, Shanshan Wang 0002, Yong Xia 0001, Dong Liang 0001, David Dagan Feng
IEEE Trans. Image Process.5
2016 A Scalable Approach for Content-Based Image Retrieval in Peer-to-Peer Networks
abstract
Peer-to-peer networking offers a scalable solution for sharing multimedia data across the network. With a large amount of visual data distributed among different nodes, it is an important but challenging issue to perform content-based retrieval in peer-to-peer networks. While most of the existing methods focus on indexing high dimensional visual features and have limitations of scalability, in this paper we propose a scalable approach for content-based image retrieval in peer-to-peer networks by employing the bag-of-visual-words model. Compared with centralized environments, the key challenge is to efficiently obtain a global codebook, as images are distributed across the whole peer-to-peer network. In addition, a peer-to-peer network often evolves dynamically, which makes a static codebook less effective for retrieval tasks. Therefore, we propose a dynamic codebook updating method by optimizing the mutual information between the resultant codebook and relevance information, and the workload balance among nodes that manage different codewords. In order to further improve retrieval performance and reduce network cost, indexing pruning techniques are developed. Our comprehensive experimental results indicate that the proposed approach is scalable in evolving and distributed peer-to-peer networks, while achieving improved retrieval accuracy.
Lelin Zhang, Zhiyong Wang 0001, Tao Mei 0001, David Dagan Feng
IEEE Trans. Knowl. Data Eng.4
2015 Robust saliency detection via regularized random walks ranking
abstract
In the field of saliency detection, many graph-based algorithms heavily depend on the accuracy of the pre-processed superpixel segmentation, which leads to significant sacrifice of detail information from the input image. In this paper, we propose a novel bottom-up saliency detection approach that takes advantage of both region-based features and image details. To provide more accurate saliency estimations, we first optimize the image boundary selection by the proposed erroneous boundary removal. By taking the image details and region-based estimations into account, we then propose the regularized random walks ranking to formulate pixel-wised saliency maps from the superpixel-based background and foreground saliency estimations. Experiment results on two public datasets indicate the significantly improved accuracy and robustness of the proposed algorithm in comparison with 12 state-of-the-art saliency detection approaches.
ChangYang Li, Yuchen Yuan, Tom Weidong Cai, Yong Xia 0001, David Dagan Feng
CVPR5
2015 Fusing subcategory probabilities for texture classification
abstract
Texture, as a fundamental characteristic of objects, has attracted much attention in computer vision research. Performance of texture classification is however still lacking for some challenging cases, largely due to the high intra-class variation and low inter-class distinction. To tackle these issues, in this paper, we propose a sub-categorization model for texture classification. By clustering each class into subcategories, classification probabilities at the subcategory-level are computed based on between-subcategory distinctiveness and within-subcategory representativeness. These subcategory probabilities are then fused based on their contribution levels and cluster qualities. This fused probability is added to the multiclass classification probability to obtain the final class label. Our method was applied to texture classification on three challenging datasets - KTH-TIPS2, FMD and DTD, and has shown excellent performance in comparison with the state-of-the-art approaches.
Yang Song 0001, Tom Weidong Cai, Qing Li 0012, Fan Zhang 0013, David Dagan Feng, Heng Huang 0001
CVPR5
2015 Subject-centered multi-view feature fusion for neuroimaging retrieval and classification
abstract
Multi-View neuroimaging retrieval and classification play an important role in computer-aided-diagnosis of brain disorders, as multi-view features could provide more insights of the disease pathology and potentially lead to more accurate diagnosis than single-view features. The large inter-feature and inter-subject variations make the multi-view neuroimaging analysis a challenging task. Many multi-view or multi-modal feature fusion methods have been proposed to reduce the impact of inter-feature variations in neuroimaging data. However, there is not much in-depth work focusing on the inter-subject variations. In this study, we propose a subject-centered multi-view feature fusion method for neuroimaging retrieval and classification based on the propagation graph fusion (PGF) algorithm. Two main advantages of the proposed method are: 1) it evaluates the query online and adaptively reshapes the connections between subjects according to the query; 2) it measures the affinity of the query to the subjects using the subject-centered affinity matrices, which can be easily combined and efficiently solved. Evaluated using a public accessible neuroimaging database, our algorithm outperforms the state-of-the-art methods in retrieval and achieves comparable performance in classification.
Sidong Liu, Tom Weidong Cai, Siqi Liu 0001, Sonia Pujol, Ron Kikinis, David Dagan Feng
ICIP6
2015 Beating cilia identification in fluorescence microscope images for accurate CBF measurement
abstract
Ciliary beating frequency (CBF) is a regulated quantitative measurement to describe ciliary beating properties. It is widely used for diagnosis of defective mucociliary clearance diseases. Image-based methods can be effective for CBF estimation but also affected by the moving objects such as ciliated cells and debris. In this work, we propose a CBF estimation method by removing these unfavorable objects, which we refer to as foreground, so that we can focus on observing the beating cilia only. We firstly design a graph-based method to divide the cilia image into different regions. Next, the foreground regions are extracted and removed from the region division result. The beating cilia are then recognized from the background and used to compute the CBF. Our method conducts the CBF estimation by incorporating the cilia regions only and thus can provide a more accurate description of ciliary beating properties. Preliminary experimental results on cilia images showed the proposed method's potentials for accurate CBF measurement.
Fan Zhang 0013, Tom Weidong Cai, Yang Song 0001, Paul M. Young, Daniela Traini, Lucy Morgan, Hui-Xin Ong, Lachlan Buddle, David Dagan Feng
ICIP9
2015 Learning Shape-Driven Segmentation Based on Neural Network and Sparse Reconstruction Toward Automated Cell Analysis of Cervical Smears
Afaf Tareef, Yang Song 0001, Tom Weidong Cai, Heng Huang 0001, Yue Joseph Wang, David Dagan Feng
ICONIP (1)6
2015 Motion Representation of Ciliated Cell Images with Contour-Alignment for Automated CBF Estimation
Fan Zhang 0013, Yang Song 0001, Siqi Liu 0001, Paul M. Young, Daniela Traini, Lucy Morgan, Hui-Xin Ong, Lachlan Buddle, Sidong Liu, David Dagan Feng, Tom Weidong Cai
MICCAI (3)10
2015 Spatial-temporal correlation for trajectory based action video retrieval
abstract
The bag-of-visual-words model has been widely utilized for content based image and video retrieval due to its scalability. In this paper, we extend this model for human action video retrieval. We adopt dense trajectory features which are able to achieve the state-of-the-art performance on action recognition, while most of the existing video retrieval methods utilize descriptors of local interest points. In order to improve similarity measurement between bag-of-visual-words model based representation, we propose to discover and incorporate spatial-temporal correlation (STC) among the trajectories in a given query video. The spatial-temporal correlation consists of spatial proximity and temporal consistence among trajectories, which is capable of strengthening discriminative power among visual words. Note that such query focused spatial-temporal correlation makes our method dynamic for different queries and is able to improve retrieval performance without significantly increasing the size of a visual vocabulary. The experimental results on an action video dataset demonstrate that our proposed method outperforms other similar methods.
Lelin Zhang, Zhiyong Wang 0001, David Dagan Feng
MMSP4
2015 Resource restricted on-line Video Summarization with Minimum Sparse Reconstruction
abstract
Video Summarization (VS) techniques have been widely utilized to produce a concise video content representation, such that the video content can be quickly explored and the complexity of video based analysis and retrieval applications can be highly reduced. However, little attention has been paid for on-line applications, especially for resource restricted applications, such as onboard VS. In this paper, our previous on-line Minimum Sparse Reconstruction (OnMSR) based VS algorithm is improved for resources restricted applications by confining the size of keyframes for reconstruction. Specially, an on-line reconstruction keyframe set update strategy is designed to meet the requirement of real-time resource restricted situation. Experimental results on various types of videos demonstrate the performance of OnMSR does not vary much by imposing resource constraint in the proposed resource restricted OnMSR (RR-onMSR) algorithm. As a result, the proposed RR-onMSR is very effective for real-time onboard VS applications.
Shaohui Mei, Zhiyong Wang 0001, Mingyi He, David Dagan Feng
PCS4
2015 Locality-constrained Subcluster Representation Ensemble for lung image classification
Yang Song 0001, Tom Weidong Cai, Heng Huang 0001, Yun Zhou 0006, Yue Joseph Wang, David Dagan Feng
Medical Image Anal.6
2015 An iteratively reweighting algorithm for dynamic video summarization
Pei Dong, Yong Xia 0001, Shanshan Wang 0002, Li Zhuo 0001, David Dagan Feng
Multim. Tools Appl.5
2015 Video summarization via minimum sparse reconstruction
Shaohui Mei, Genliang Guan, Zhiyong Wang 0001, Shuai Wan, Mingyi He, David Dagan Feng
Pattern Recognit.6
2015 Guest-Editorial - Telehealth Systems and Applications
abstract
In the 21st century, the convergence of healthcare and information and communications technologies (ICT) offers an opportunity to give patients greater liberty from their health problems. Telehealth systems and applications, supported by advances in ICT, are fostering a diversity of cost-effective and efficient healthcare solutions. These solutions are becoming embedded in all aspects of clinical care and are enhancing the quality, equality, and accessibility of care, while playing a pivotal role in decreasing the rising costs from the growth in aging population. The emergence of affordable health sensors and accessible mobile computing devices, such as smartphones, wearables and tablets, offers opportunities to revolutionize healthcare solutions.
David Dagan Feng, Jinman Kim, Mohamed Khadra, Donna L. Hudson, Christian Roux
IEEE J. Biomed. Health Informatics1
2015 Large Margin Local Estimate With Applications to Medical Image Classification
abstract
Medical images usually exhibit large intra-class variation and inter-class ambiguity in the feature space, which could affect classification accuracy. To tackle this issue, we propose a new Large Margin Local Estimate (LMLE) classification model with sub-categorization based sparse representation. We first sub-categorize the reference sets of different classes into multiple clusters, to reduce feature variation within each subcategory compared to the entire reference set. Local estimates are generated for the test image using sparse representation with reference subcategories as the dictionaries. The similarity between the test image and each class is then computed by fusing the distances with the local estimates in a learning-based large margin aggregation construct to alleviate the problem of inter-class ambiguity. The derived similarities are finally used to determine the class label. We demonstrate that our LMLE model is generally applicable to different imaging modalities, and applied it to three tasks: interstitial lung disease (ILD) classification on high-resolution computed tomography (HRCT) images, phenotype binary classification and continuous regression on brain magnetic resonance (MR) imaging. Our experimental results show statistically significant performance improvements over existing popular classifiers.
Yang Song 0001, Tom Weidong Cai, Heng Huang 0001, Yun Zhou 0006, David Dagan Feng, Yue Joseph Wang, Michael J. Fulham
IEEE Trans. Medical Imaging5
2014 Medical image classification with convolutional neural network
abstract
Image patch classification is an important task in many different medical imaging applications. In this work, we have designed a customized Convolutional Neural Networks (CNN) with shallow convolution layer to classify lung image patches with interstitial lung disease (ILD). While many feature descriptors have been proposed over the past years, they can be quite complicated and domain-specific. Our customized CNN framework can, on the other hand, automatically and efficiently learn the intrinsic image features from lung image patches that are most suitable for the classification purpose. The same architecture can be generalized to perform other medical image or texture classification tasks.
Qing Li 0012, Tom Weidong Cai, Xiaogang Wang 0001, Yun Zhou 0006, David Dagan Feng
ICARCV5
2014 A new statistical and Dirichlet integral framework applied to liver segmentation from volumetric CT images
abstract
Accurate liver segmentation from computed tomography (CT) images is problematic due to non-uniform density, weak boundaries and because there may be multiple liver tumors that have heterogeneous intensities in region(s) of interest (ROIs). So we propose a generalized energy framework that harnesses the statistical intensity approximation with image data on graphs. Our statistical energy term takes advantage of the mixture-of-mixtures Gaussian model to approximate the probability density distribution of the liver and background to better differentiate between the two. The probability density estimation can be combined with the spatial cohesion of the graph-based Dirichlet integral by using graph calculus. Matrix decomposition and differentiation are used to minimize our proposed energy functional. We tested our approach on 20 public high-contrast CT images with single and multiple liver tumors. Our method had an average dice similarity coefficient (DSC) of 93.75±1.29%, an average false positive (FP) rate of 9.43±3.52% and an average false negative (FN) rate of 3.48±1.48%. Our method outperformed the benchmark graph-based Random Walker algorithm (average DSC=81.97±4.09%, average FP rate 34.10±10.53%, and average FN rate 7.10±4.35%).
ChangYang Li, Xiuying Wang 0001, David Dagan Feng, Stefan Eberl, Michael J. Fulham
ICARCV4
2014 Propagation graph fusion for multi-modal medical content-based retrieval
abstract
Medical content-based retrieval (MCBR) plays an important role in computer aided diagnosis and clinical decision support. Multi-modal imaging data have been increasingly used in MCBR, as they could provide more insights of the diseases and complement the deficiencies of single-modal data. However, it is very challenging to fuse data in different modalities since they have different physical fundamentals and large value range variations. In this study, we propose a novel Propagation Graph Fusion (PGF) framework for multi-modal medical data retrieval. PGF models the subjects' relationships in single modalities using the directed propagation graphs, and then fuses the graphs into a single graph by summing up the edge weights. Our proposed PGF method could reduce the large inter-modality and inter-subject variations, and can be solved efficiently using the PageRank algorithm. We test the proposed method on a public medical database with 331 subjects using features extracted from two imaging modalities, PET and MRI. The preliminary results show that our PGF method could enhance multi-modal retrieval and modestly outperform the state-of-the-art single-modal and multi-modal retrieval methods.
Sidong Liu, Siqi Liu 0001, Sonia Pujol, Ron Kikinis, David Dagan Feng, Tom Weidong Cai
ICARCV5
2014 Automated three-stage nucleus and cytoplasm segmentation of overlapping cells
abstract
Developing segmentation techniques for overlapping cells has become a major hurdle for automated analysis of cervical cells. In this paper, an automated three-stage segmentation approach to segment the nucleus and cytoplasm of each overlapping cell is described. First, superpixel clustering is conducted to segment the image into small coherent clusters that are used to generate a refined superpixel map. The refined superpixel map is passed to an adaptive thresholding step to initially segment the image into cellular clumps and background. Second, a linear classifier with superpixel-based features is designed to finalize the separation between nuclei and cytoplasm. Finally, edge and region based cell segmentation are performed based on edge enhancement process, gradient thresholding, morphological operations, and region properties evaluation on all detected nuclei and cytoplasm pairs. The proposed framework has been evaluated using the ISBI 2014 challenge dataset. The dataset consists of 45 synthetic cell images, yielding 270 cells in total. Compared with the state-of-the-art approaches, our approach provides more accurate nuclei boundaries, as well as successfully segments most of overlapping cells.
Afaf Tareef, Yang Song 0001, Tom Weidong Cai, David Dagan Feng
ICARCV4
2014 Non-sparse infinite-kernel learning for automated identification of Alzheimer's disease using PET imaging
abstract
Multi-kernel learning machine (MKLM) has recently been introduced to the research of computer-aided dementia identification and pathology progress tracking. Despite its good performance especially in case of using heterogeneous data, such learning schema and its variants usually utilize a L-l norm constraint that promotes sparse solutions, which may cause loss of potentially important information. In this paper, we propose the non-sparse infinite-kernel learning machine (NS-IKLM) for automated identification of Alzheimer cases from normal controls. In our approach, a modified constraint is utilized to promotes non-sparse solutions and kernel parameters are automatically tuned during the learning process. The proposed algorithm has been evaluated on a set of FDG-PET images selected from the Alzheimer's disease neuroimaing initiative (ADNI) cohort. Our results demonstrate that the proposed non-sparse NS-IKLM is able to achieve satisfying dementia identification at a relatively low computational cost.
Yong Xia 0001, Shen Lu, Wei Wei 0008, David Dagan Feng, Yanning Zhang 0001
ICARCV4
2014 Importance-aware lighting design in volume visualization
abstract
Lighting design plays critical roles in depicting structural details in volume rendering. Insufficient and excessive illumination can both affect effectiveness of presenting structural details in visualization. This paper introduces topological importance into the lighting design and proposes the importance-aware lighting. In the proposed approach, the lighting in volume rendering is enhanced based on topological importance. As a result, importance of structures can be depicted from the lighting perspective. The contour tree, one of topological data structures, is used to represent topology in this paper. Topological importance such as persistence derived from the contour tree is used to modulate lighting coefficients. The experimental results demonstrate that the importance-aware lighting not only helps to depict structural details more clearly but also reveal topological importance of structures in rendering. The importance-aware lighting is more meaningful to users but not a random selection without physical meanings based on preferences.
Jianlong Zhou, Xiuying Wang 0001, David Dagan Feng
ICARCV3
2014 Image noise level estimation based on a new adaptive superpixel classification
abstract
Accurate estimation of noise level in images plays an important role in different image processing applications. The current algorithms can precisely estimate noise with smooth images, but it is still the challenge to approximate noise level from richly textured images. In this paper, we proposed a new adaptive superpixel classification algorithm for noise estimation in complicated textured images. Firstly, our new superpixel algorithm adapts the finite Gaussian clustering approach, which can better approximate homogeneous patches in noisy images. Then noise information is obtained locally from each superpixel patch. Finally, the best estimation of noise level is calculated with a statistical approach. Experimental results with various kinds of images demonstrate that our method is more accurate and robust compared to the five existing common used algorithms.
Peng Fu 0003, ChangYang Li, Quan-Sen Sun, Tom Weidong Cai, David Dagan Feng
ICIP5
2014 Iterative keyframe selection by orthogonal subspace projection
abstract
Recent developments on sparse dictionary selection have demonstrated promising results for Video Summarization (VS). However, the convex relaxation based solution cannot ensure the sparsity of the dictionary directly. In this paper, a selection matrix is proposed to model the VS problem, according to which the L0norm of this selection matrix is imposed to ensure sparsity directly. As a result, a computational efficient Orthogonal Subspace Projection (OSP) based Iterative Keyframe Selection (IKS) algorithm is proposed for VS. In addition, a Percentage Of Reconstruction (POR) criterion is proposed to provide an intuitive and flexible control of the length of final video summaries even without prior knowledge of a given video. Experimental results on a popular benchmark dataset demonstrate that our proposed algorithm outperforms the state-of-the-art methods.
Shaohui Mei, Genliang Guan, Zhiyong Wang 0001, Mingyi He, Shuai Wan, David Dagan Feng
ICIP6
2014 L2, 0 constrained sparse dictionary selection for video summarization
abstract
The ever increasing volume of video content has created profound challenges for developing efficient video summarization (VS) techniques to access the data. Recent developments on sparse dictionary selection have demonstrated promising results for VS, however, the convex relaxation based solution cannot ensure the sparsity of the dictionary directly and it selects keyframes in a local point of view. In this paper, an L2,0constrained sparse dictionary selection model is proposed to reformulate the problem of VS. In addition, a simultaneous orthogonal matching pursuit (SOMP) based method is proposed to obtain an approximate solution for the proposed model without smoothing the penalty function, and thus selects keyframes in a global point of view. In order to allow for intuitive and flexible configuration of VS process, a percentage of residuals (POR) criterion is also developed to produce video summaries in different lengths. Experimental results demonstrate that our proposed method outperforms the state-of-the-art.
Shaohui Mei, Genliang Guan, Zhiyong Wang 0001, Mingyi He, Xian-Sheng Hua 0001, David Dagan Feng
ICME6
2014 Multi-stage Thresholded Region Classification for Whole-Body PET-CT Lymphoma Studies
Lei Bi 0001, Jinman Kim, David Dagan Feng, Michael J. Fulham
MICCAI (1)3
2014 Large Margin Aggregation of Local Estimates for Medical Image Classification
Yang Song 0001, Tom Weidong Cai, Heng Huang 0001, Yun Zhou 0006, David Dagan Feng
MICCAI (2)5
2014 A clonal selection based approach to statistical brain voxel classification in magnetic resonance images
Tong Zhang 0017, Yong Xia 0001, David Dagan Feng
Neurocomputing3
2014 Adaptive scale fuzzy local Gaussian mixture model for brain MR image segmentation
Zexuan Ji, Yong Xia 0001, Quan-Sen Sun, Qiang Chen 0004, David Dagan Feng
Neurocomputing5
2014 Spectral embedding based facial expression recognition with multiple features
Kaimin Yu, Zhiyong Wang 0001, Markus Hagenbuchner, David Dagan Feng
Neurocomputing4
2014 A graph-based approach for the retrieval of multi-modality medical images
Ashnil Kumar, Jinman Kim, Lingfeng Wen, Michael J. Fulham, David Dagan Feng
Medical Image Anal.5
2014 Detecting and tracking dim small targets in infrared image sequences under complex backgrounds
Ying Li 0017, Bendu Bai, David Dagan Feng
Multim. Tools Appl.4
2014 Quality assessment of perceptual color video based on a top-down framework and quaternion
Xiuying Wang 0001, David Dagan Feng
Multim. Tools Appl.5
2014 Relative Saliency Model over Multiple Images with an Application to Yarn Surface Evaluation
abstract
Saliency models have been developed and widely demonstrated to benefit applications in computer vision and image understanding. In most of existing models, saliency is evaluated within an individual image. That is, saliency value of an item (object/region/pixel) represents the conspicuity of it as compared with the remaining items in the same image. We call this saliency as absolute saliency, which is uncomparable among images. However, saliency should be determined in the context of multiple images for some visual inspection tasks. For example, in yarn surface evaluation, saliency of a yarn image should be measured with regard to a set of graded standard images. We call this saliency the relative saliency, which is comparable among images. In this paper, a study of visual attention model for comparison of multiple images is explored, and a relative saliency model of multiple images is proposed based on a combination of bottom-up and top-down mechanisms, to enable relative saliency evaluation for the cases where other image contents are involved. To fully characterize the differences among multiple images, a structural feature extraction strategy is proposed, where two levels of feature (high-level, low-level) and three types of feature (global, local-local, local-global) are extracted. Mapping functions between features and saliency values are constructed and their outputs reflect relative saliency for multiimage contents instead of single image content. The performance of the proposed relative saliency model is well demonstrated in a yarn surface evaluation. Furthermore, the eye tracking technique is employed to verify the proposed concept of relative saliency for multiple images.
Bingang Xu, Zheru Chi, David Dagan Feng
IEEE Trans. Cybern.4
2014 Postreconstruction Nonlocal Means Filtering of Whole-Body PET With an Anatomical Prior
abstract
Positron emission tomography (PET) images usually suffer from poor signal-to-noise ratio (SNR) due to the high level of noise and low spatial resolution, which adversely affect its performance for lesion detection and quantification. The complementary information present in high-resolution anatomical images from multi-modality imaging systems could potentially be used to improve the ability to detect and/or quantify lesions. However, previous methods that use anatomical priors usually require matched organ/lesion boundaries. In this study, we investigated the use of anatomical information to suppress noise in PET images while preserving both quantitative accuracy and the amplitude of prominent signals that do not have corresponding boundaries on computerized tomography (CT). The proposed approach was realized through a postreconstruction filter based on the nonlocal means (NLM) filter, which reduces noise by computing the weighted average of voxels based on the similarity measurement between patches of voxels within the image. Anatomical knowledge obtained from CT was incorporated to constrain the similarity measurement within a subset of voxels. In contrast to other methods that use anatomical priors, the actual number of neighboring voxels and weights used for smoothing were determined from a robust measurement on PET images within the subset. Thus, the proposed approach can be robust to signal mismatches between PET and CT. A 3-D search scheme was also investigated for the volumetric PET/CT data. The proposed anatomically guided median nonlocal means filter (AMNLM) was first evaluated using a computer phantom and a physical phantom to simulate realistic but challenging situations where small lesions are located in homogeneous regions, which can be detected on PET but not on CT. The proposed method was further assessed with a clinical study of a patient with lung lesions. The performance of the proposed method was compared to Gaussian, edge-preserving bilateral and NLM filters, as well as median nonlocal means (MNLM) filtering without an anatomical prior. The proposed AMNLM method yielded improved lesion contrast and SNR compared with other methods even with imperfect anatomical knowledge, such as missing lesion boundaries and mismatched organ boundaries.
Chung Chan, Roger R. Fulton, Robert Barnett, David Dagan Feng, Steven R. Meikle
IEEE Trans. Medical Imaging4
2014 Lesion Detection and Characterization With Context Driven Approximation in Thoracic FDG PET-CT Images of NSCLC Studies
abstract
We present a lesion detection and characterization method for (18)F-fluorodeoxyglucose positron emission tomography-computed tomography (FDG PET-CT) images of the thorax in the evaluation of patients with primary nonsmall cell lung cancer (NSCLC) with regional nodal disease. Lesion detection can be difficult due to low contrast between lesions and normal anatomical structures. Lesion characterization is also challenging due to similar spatial characteristics between the lung tumors and abnormal lymph nodes. To tackle these problems, we propose a context driven approximation (CDA) method. There are two main components of our method. First, a sparse representation technique with region-level contexts was designed for lesion detection. To discriminate low-contrast data with sparse representation, we propose a reference consistency constraint and a spatial consistent constraint. Second, a multi-atlas technique with image-level contexts was designed to represent the spatial characteristics for lesion characterization. To accommodate inter-subject variation in a multi-atlas model, we propose an appearance constraint and a similarity constraint. The CDA method is effective with a simple feature set, and does not require parametric modeling of feature space separation. The experiments on a clinical FDG PET-CT dataset show promising performance improvement over the state-of-the-art.
Yang Song 0001, Tom Weidong Cai, Heng Huang 0001, Xiaogang Wang 0001, Yun Zhou 0006, Michael J. Fulham, David Dagan Feng
IEEE Trans. Medical Imaging7
2014 A Bag-of-Importance Model With Locality-Constrained Coding Based Feature Learning for Video Summarization
abstract
Video summarization helps users obtain quick comprehension of video content. Recently, some studies have utilized local features to represent each video frame and formulate video summarization as a coverage problem of local features. However, the importance of individual local features has not been exploited. In this paper, we propose a novel Bag-of-Importance (BoI) model for static video summarization by identifying the frames with important local features as keyframes, which is one of the first studies formulating video summarization at local feature level, instead of at global feature level. That is, by representing each frame with local features, a video is characterized with a bag of local features weighted with individual importance scores and the frames with more important local features are more representative, where the representativeness of each frame is the aggregation of the weighted importance of the local features contained in the frame. In addition, we propose to learn a transformation from a raw local feature to a more powerful sparse nonlinear representation for deriving the importance score of each local feature, rather than directly utilize the hand-crafted visual features like most of the existing approaches. Specifically, we first employ locality-constrained linear coding (LCC) to project each local feature into a sparse transformed space. LCC is able to take advantage of the manifold geometric structure of the high dimensional feature space and form the manifold of the low dimensional transformed space with the coordinates of a set of anchor points. Then we calculate the l2 norm of each anchor point as the importance score of each local feature which is projected to the anchor point. Finally, the distribution of the importance scores of all the local features in a video is obtained as the BoI representation of the video. We further differentiate the importance of local features with a spatial weighting template by taking the perceptual difference among spatial regions of a frame into account. As a result, our proposed video summarization approach is able to exploit both the inter-frame and intra-frame properties of feature representations and identify keyframes capturing both the dominant content and discriminative details within a video. Experimental results on three video datasets across various genres demonstrate that the proposed approach clearly outperforms several state-of-the-art methods.
Shiyang Lu, Zhiyong Wang 0001, Tao Mei 0001, Genliang Guan, David Dagan Feng
IEEE Trans. Multim.5
2014 A Top-Down Approach for Video Summarization
abstract
While most existing video summarization approaches aim to identify important frames of a video from either a global or local perspective, we propose a top-down approach consisting of scene identification and scene summarization. For scene identification, we represent each frame with global features and utilize a scalable clustering method. We then formulate scene summarization as choosing those frames that best cover a set of local descriptors with minimal redundancy. In addition, we develop a visual word-based approach to make our approach more computationally scalable. Experimental results on two benchmark datasets demonstrate that our proposed approach clearly outperforms the state-of-the-art.
Genliang Guan, Zhiyong Wang 0001, Shaohui Mei, Maximilian Ott, Mingyi He, David Dagan Feng
ACM Trans. Multim. Comput. Commun. Appl.6
2014 Editorial
Jinman Kim, Daniel Thalmann, Kun Zhou 0001, David Dagan Feng, Holly E. Rushmeier
Vis. Comput.4
2013 Graph-based retrieval of PET-CT images using vector space embedding
abstract
Graph-based content-based image retrieval (CBIR) techniques, which use graphs to represent image features and calculate image similarity using the graph edit distance, achieve high retrieval accuracy. However, such techniques suffer from high computational complexity. In this paper, we present a graph-based CBIR algorithm that achieves improved retrieval efficiency. We compute a vector space embedding for every graph, using their distances from a set of prototype graphs, so that each vector component represents a distortion from a prototype. This process is performed offline. We compare images by computing the Euclidean distance of the vector embeddings, which is a faster process than calculating the graph edit distance. We evaluated our work using 50 combined positron emission tomography and computed tomography (PET-CT) volumes of patients with lung tumours. Our results show that our method is at least 21 times faster than the graph edit distance with a mean average precision difference of less than 4%.
Ashnil Kumar, Jinman Kim, David Dagan Feng, Michael J. Fulham
CBMS3
2013 A web-based medical multimedia visualisation interface for personal health records
abstract
The healthcare industry has begun to utilise web-based systems and cloud computing infrastructure to develop an increasing array of online personal health record (PHR) systems. Although these systems provide the technical capacity to store and retrieve medical data in various multimedia formats, including images, videos, voice, and text, individual patient use remains limited by the lack of intuitive data representation and visualisation techniques. As such, further research is necessary to better visualise and present these records, in ways that make the complex medical data more intuitive. In this study, we present a web-based PHR visualisation system, called the 3D medical graphical avatar (MGA), which was designed to explore web-based delivery of a wide array of medical data types including multi-dimensional medical images; medical videos; text-based data; and spatial annotations. Mapping information was extracted from each of the data types and was used to embed spatial and textual annotations, such as regions of interest (ROIs) and time-based video annotations. Our MGA itself is built from clinical patient imaging studies, when available. We have taken advantage of the emerging web technologies of HTML5 and WebGL to make our application available to a wider base of users and devices. We analysed the performance of our proof-of-concept prototype system on mobile and desktop consumer devices. Our initial experiments indicate that our system can render the medical data in a fashion that enables interactive navigation of the MGA.
Michael de Ridder, Liviu Constantinescu, Lei Bi 0001, Younhyun Jung, Ashnil Kumar, Jinman Kim, David Dagan Feng, Michael J. Fulham
CBMS7
2013 A supervised multiview spectral embedding method for neuroimaging classification
abstract
The multi-view/multi-modal features are commonly used in neuroimaging classification because they could provide complementary information to each other and thus result in better classification performance than single-view features. However, it is very challenging to effectively integrate such rich features, since straightforward concatenation or singleview spectral embedding methods rarely leads to physically meaningful integration. In this paper, we present a supervised multi-view/multi-modal spectral embedding method (SMSE) for neuroimaging classification. This method embeds the high dimensional multi-view features derived from multi-modal neuroimaging data into a low dimensional feature space and preserves the optimal local embeddings among different views. The proposed SMSE algorithm, validated using three groups of neuroimaging data, is able to achieve significant classification improvement over the state-of-the-art multi-view spectral embedding methods.
Sidong Liu, Lelin Zhang, Tom Weidong Cai, Yang Song 0001, Zhiyong Wang 0001, Lingfeng Wen, David Dagan Feng
ICIP7
2013 Graph cuts based relevance feedback in image retrieval
abstract
Relevance feedback (RF) allows users to be actively involved in the information retrieval process and has been widely used in various information retrieval tasks. While most existing RF methods in content-based image retrieval (CBIR) focus on visual features of individual images only, in this paper we formulate the relevance feedback process as an energy minimization problem. The energy function takes into account both the feature aspect of each image and the manifold structure among individual images. The solution of labelling images as relevant or irrelevant is obtained with the graph cuts method. As a result, our method enables flexibly partitioning the feature space and labelling of images and is capable of handling challenging scenarios (or queries). Experimental results demonstrate that our proposed method outperforms the popular RF methods.
Lelin Zhang, Sidong Liu, Zhiyong Wang 0001, Tom Weidong Cai, Yang Song 0001, David Dagan Feng
ICIP6
2013 Multifold Bayesian Kernelization in Alzheimer's Diagnosis
Sidong Liu, Yang Song 0001, Tom Weidong Cai, Sonia Pujol, Ron Kikinis, Xiaogang Wang 0001, David Dagan Feng
MICCAI (2)7
2013 Discriminative Data Transform for Image Feature Extraction and Classification
Yang Song 0001, Tom Weidong Cai, Seungil Huh, Takeo Kanade, Yun Zhou 0006, David Dagan Feng
MICCAI (2)7
2013 Similarity Guided Feature Labeling for Lesion Detection
Yang Song 0001, Tom Weidong Cai, Heng Huang 0001, Xiaogang Wang 0001, Stefan Eberl, Michael J. Fulham, David Dagan Feng
MICCAI (1)7
2013 Region-based progressive localization of cell nuclei in microscopic images with data adaptive modeling
abstract
BACKGROUND: Segmenting cell nuclei in microscopic images has become one of the most important routines in modern biological applications. With the vast amount of data, automatic localization, i.e. detection and segmentation, of cell nuclei is highly desirable compared to time-consuming manual processes. However, automated segmentation is challenging due to large intensity inhomogeneities in the cell nuclei and the background. RESULTS: We present a new method for automated progressive localization of cell nuclei using data-adaptive models that can better handle the inhomogeneity problem. We perform localization in a three-stage approach: first identify all interest regions with contrast-enhanced salient region detection, then process the clusters to identify true cell nuclei with probability estimation via feature-distance profiles of reference regions, and finally refine the contours of detected regions with regional contrast-based graphical model. The proposed region-based progressive localization (RPL) method is evaluated on three different datasets, with the first two containing grayscale images, and the third one comprising of color images with cytoplasm in addition to cell nuclei. We demonstrate performance improvement over the state-of-the-art. For example, compared to the second best approach, on the first dataset, our method achieves 2.8 and 3.7 reduction in Hausdorff distance and false negatives; on the second dataset that has larger intensity inhomogeneity, our method achieves 5% increase in Dice coefficient and Rand index; on the third dataset, our method achieves 4% increase in object-level accuracy. CONCLUSIONS: To tackle the intensity inhomogeneities in cell nuclei and background, a region-based progressive localization method is proposed for cell nuclei localization in fluorescence microscopy images. The RPL method is demonstrated highly effective on three different public datasets, with on average 3.5% and 7% improvement of region- and contour-based segmentation performance over the state-of-the-art.
Yang Song 0001, Tom Weidong Cai, Heng Huang 0001, Yue Joseph Wang, David Dagan Feng
BMC Bioinform.5
2013 Layered-based exposure fusion algorithm
abstract
Owing to the limitation of dynamic range, a single still image is usually insufficient to describe a high contrast scene. Fusing multi‐exposure images of the same scene can produce a resulting image with details both in the bright and the dark regions. However, they may be sensitive to the exposure parameters of the input images. In this study, a global layer is introduced to improve the robustness of the fusion method. The global layer is employed to preserve the overall luminance of a real scene and avoid possible luminance reversion artefacts. Then, details are recovered in the gradient domain by a Poisson solver. Experimental results show the superior performance of our approach in terms of robustness and details preservation.
Fenghui Li, Li Zhuo 0001, David Dagan Feng
IET Image Process.4
2013 Multi-instance multi-label image classification: A neural approach
Zenghai Chen, Zheru Chi, Hong Fu, David Dagan Feng
Neurocomputing4
2013 Multi-pose 3D face recognition based on 2D sparse representation
Yanning Zhang 0001, Yong Xia 0001, Zenggang Lin, Yangyu Fan, David Dagan Feng
J. Vis. Commun. Image Represent.6
2013 Fast human action classification and VOI localization with enhanced sparse coding
Shiyang Lu, Jian Zhang 0002, Zhiyong Wang 0001, David Dagan Feng
J. Vis. Commun. Image Represent.4
2013 Discriminative two-level feature selection for realistic human action recognition
Qiuxia Wu, Zhiyong Wang 0001, Feiqi Deng, Yong Xia 0001, Wenxiong Kang, David Dagan Feng
J. Vis. Commun. Image Represent.6
2013 Semantic context based refinement for news video annotation
Zhiyong Wang 0001, Genliang Guan, Li Zhuo 0001, David Dagan Feng
Multim. Tools Appl.5
2013 Learning realistic facial expressions from web images
Kaimin Yu, Zhiyong Wang 0001, Li Zhuo 0001, Zheru Chi, David Dagan Feng
Pattern Recognit.6
2013 Dictionary learning based impulse noise removal via L1-L1 minimization
Shanshan Wang 0002, Qiegen Liu, Yong Xia 0001, Pei Dong, Jianhua Luo, Qiu Huang, David Dagan Feng
Signal Process.7
2013 Keypoint-Based Keyframe Selection
abstract
Keyframe selection has been crucial for effective and efficient video content analysis. While most of the existing approaches represent individual frames with global features, we, for the first time, propose a keypoint-based framework to address the keyframe selection problem so that local features can be employed in selecting keyframes. In general, the selected keyframes should both be representative of video content and containing minimum redundancy. Therefore, we introduce two criteria, coverage and redundancy, based on keypoint matching in the selection process. Comprehensive experiments demonstrate that our approach outperforms the state of the art.
Genliang Guan, Zhiyong Wang 0001, Shiyang Lu, Jeremiah D. Deng, David Dagan Feng
IEEE Trans. Circuits Syst. Video Technol.5
2013 Robust Model for Segmenting Images With/Without Intensity Inhomogeneities
abstract
Intensity inhomogeneities and different types/levels of image noise are the two major obstacles to accurate image segmentation by region-based level set models. To provide a more general solution to these challenges, we propose a novel segmentation model that considers global and local image statistics to eliminate the influence of image noise and to compensate for intensity inhomogeneities. In our model, the global energy derived from a Gaussian model estimates the intensity distribution of the target object and background; the local energy derived from the mutual influences of neighboring pixels can eliminate the impact of image noise and intensity inhomogeneities. The robustness of our method is validated on segmenting synthetic images with/without intensity inhomogeneities, and with different types/levels of noise, including Gaussian noise, speckle noise, and salt and pepper noise, as well as images from different medical imaging modalities. Quantitative experimental comparisons demonstrate that our method is more robust and more accurate in segmenting the images with intensity inhomogeneities than the local binary fitting technique and its more recent systematic model. Our technique also outperformed the region-based Chan–Vese model when dealing with images without intensity inhomogeneities and produce better segmentation results than the graph-based algorithms including graph-cuts and random walker when segmenting noisy images.
ChangYang Li, Xiuying Wang 0001, Stefan Eberl, Michael J. Fulham, David Dagan Feng
IEEE Trans. Image Process.5
2013 A New Energy Framework With Distribution Descriptors for Image Segmentation
abstract
Segmentation of the target object(s) from images that have multiple complicated regions, mixture intensity distributions or are corrupted by noise poses a challenge for the level set models. In addition, the conventional piecewise smooth level set models normally require prior knowledge about the number of image segments. To address these problems, we propose a novel segmentation energy function with two distribution descriptors to model the background and the target. The single background descriptor models the heterogeneous background with multiple regions. Then, the target descriptor takes into account the intensity distribution and incorporates local spatial constraint. Our descriptors, which have more complete distribution information, construct the unique energy function to differentiate the target from the background and are more tolerant of image noise. We compare our approach to three other level set models: 1) the Chan-Vese; 2) the multiphase level set; and 3) the geodesic level set. This comparison using 260 synthetic images with varying levels and types of image noise and medical images with more complicated backgrounds showed that our method outperforms these models for accuracy and immunity to noise. On an additional set of 300 synthetic images, our model is also less sensitive to the contour initialization as well as to different types and levels of noise.
ChangYang Li, Xiuying Wang 0001, Stefan Eberl, Michael J. Fulham, David Dagan Feng
IEEE Trans. Image Process.5
2013 Corrections to "Robust Model for Segmenting Images With/Without Intensity Inhomogeneities" [August 13 3296-3309]
abstract
Equation (16) in the above paper (ibid., vol. 22, no. 8, pp. 3296-3309, Aug. 2013) contained an error in the numerator. Equation (17) in the same paper contained errors in both the numerator and the denominator. The corrected versions of both equations are presented here.
ChangYang Li, Xiuying Wang 0001, Stefan Eberl, Michael J. Fulham, David Dagan Feng
IEEE Trans. Image Process.5
2013 Fenchel Duality Based Dictionary Learning for Restoration of Noisy Images
abstract
Dictionary learning based sparse modeling has been increasingly recognized as providing high performance in the restoration of noisy images. Although a number of dictionary learning algorithms have been developed, most of them attack this learning problem in its primal form, with little effort being devoted to exploring the advantage of solving this problem in a dual space. In this paper, a novel Fenchel duality based dictionary learning (FD-DL) algorithm has been proposed for the restoration of noise-corrupted images. With the restricted attention to the additive white Gaussian noise, the sparse image representation is formulated as an 2-1 minimization problem, whose dual formulation is constructed using a generalization of Fenchel’s duality theorem and solved under the augmented Lagrangian framework. The proposed algorithm has been compared with four state-of-the-art algorithms, including the local pixel grouping-principal component analysis, method of optimal directions, K-singular value decomposition, and beta process factor analysis, on grayscale natural images. Our results demonstrate that the FD-DL algorithm can effectively improve the image quality and its noisy image restoration ability is comparable or even superior to the abilities of the other four widely-used algorithms.
Shanshan Wang 0002, Yong Xia 0001, Qiegen Liu, Pei Dong, David Dagan Feng, Jianhua Luo
IEEE Trans. Image Process.5
2013 Joint Probabilistic Model of Shape and Intensity for Multiple Abdominal Organ Segmentation From Volumetric CT Images
abstract
We propose a novel joint probabilistic model that correlates a new probabilistic shape model with the corresponding global intensity distribution to segment multiple abdominal organs simultaneously. Our probabilistic shape model estimates the probability of an individual voxel belonging to the estimated shape of the object. The probability density of the estimated shape is derived from a combination of the shape variations of target class and the observed shape information. To better capture the shape variations, we used probabilistic principle component analysis optimized by expectation maximization to capture the shape variations and reduce computational complexity. The maximum a posteriori estimation was optimized by the iterated conditional mode-expectation maximization. We used 72 training datasets including low- and high-contrast CT images to construct the shape models for the liver, spleen and both kidneys. We evaluated our algorithm on 40 test datasets that were grouped into normal (34 normal cases) and pathologic (6 datasets) classes. The testing datasets were from different databases and manual segmentation was performed by different clinicians. We measured the volumetric overlap percentage error, relative volume difference, average square symmetric surface distance, false positive rate and false negative rate and our method achieved accurate and robust segmentation for multiple abdominal organs simultaneously.
ChangYang Li, Xiuying Wang 0001, Stefan Eberl, Michael J. Fulham, David Dagan Feng
IEEE J. Biomed. Health Informatics7
2013 Feature-Based Image Patch Approximation for Lung Tissue Classification
abstract
In this paper, we propose a new classification method for five categories of lung tissues in high-resolution computed tomography (HRCT) images, with feature-based image patch approximation. We design two new feature descriptors for higher feature descriptiveness, namely the rotation-invariant Gabor-local binary patterns (RGLBP) texture descriptor and multi-coordinate histogram of oriented gradients (MCHOG) gradient descriptor. Together with intensity features, each image patch is then labeled based on its feature approximation from reference image patches. And a new patch-adaptive sparse approximation (PASA) method is designed with the following main components: minimum discrepancy criteria for sparse-based classification, patch-specific adaptation for discriminative approximation, and feature-space weighting for distance computation. The patch-wise labelings are then accumulated as probabilistic estimations for region-level classification. The proposed method is evaluated on a publicly available ILD database, showing encouraging performance improvements over the state-of-the-arts.
Yang Song 0001, Tom Weidong Cai, Yun Zhou 0006, David Dagan Feng
IEEE Trans. Medical Imaging4
2013 Realistic Human Action Recognition With Multimodal Feature Selection and Fusion
abstract
Although promising results have been achieved for human action recognition under well-controlled conditions, it is very challenging to recognize human actions in realistic scenarios due to increased difficulties such as dynamic backgrounds. In this paper, we propose to take multimodal (i.e., audiovisual) characteristics of realistic human action videos into account in human action recognition for the first time, since, in realistic scenarios, audio signals accompanying an action generally provide a cue to the nature of the action, such as phone ringing to answering the phone . In order to cope with diverse audio cues of an action in realistic scenarios, we propose to identify effective features from a large number of audio features with the generalized multiple kernel learning algorithm. The widely used space-time interest point descriptors are utilized as visual features, and a support vector machine is employed for both audio- and video-based classifications. At the final stage, fuzzy integral is utilized to fuse recognition results of both audio and visual modalities. Experimental results on the challenging Hollywood-2 Human Action data set demonstrate that the proposed approach is able to achieve better recognition performance improvement than that of integrating scene context. It is also discovered how audio context influences realistic action recognition from our comprehensive experiments.
Qiuxia Wu, Zhiyong Wang 0001, Feiqi Deng, Zheru Chi, David Dagan Feng
IEEE Trans. Syst. Man Cybern. Syst.5
2013 Visibility-driven PET-CT visualisation with region of interest (ROI) segmentation
Younhyun Jung, Jinman Kim, Stefan Eberl, Michael J. Fulham, David Dagan Feng
Vis. Comput.5
2012 Graph-based retrieval of multi-modality medical images: A comparison of representations using simulated images
abstract
Content-based image retrieval (CBIR) is an image search technique that utilises visual features as search criteria; it has potential clinical applications in evidence-based diagnosis, physician training, and biomedical research. Graph-based CBIR techniques have high accuracy when retrieving images by the similarity of the spatial arrangement of their constituent objects but these techniques were initially designed for single-modality images and have limited retrieval capabilities when multi-modality images, such as combined positron emission tomography and computed tomography (PET-CT), are considered. In this paper, we present a graph-based CBIR approach for multimodality images that integrates modality-specific features on graph vertices and adapts a well-established graph similarity scheme to account for varying vertex feature sets. Furthermore, we propose a graph pruning method that removes redundant edges using the spatial proximity of image regions. We evaluated our work using two simulated data sets, consisting of 2D liver shapes and 3D whole-body lymphoma images. In our experiments we achieved a higher level of retrieval precision using our graph method when compared to conventional graph-based retrieval, demonstrating that our proposed method enabled new capabilities and improved multi-modality CBIR.
Ashnil Kumar, Jinman Kim, David Dagan Feng, Michael J. Fulham
CBMS3
2012 Multiscale and multiorientation feature extraction with degenerative patterns for 3D neuroimaging retrieval
abstract
Accurate neuroimaging feature extraction is essential for effective content-based management of the large neuroimaging databases, as well as achieving improved diagnosis. In this paper, we presented a multiscale and multi-orientation neuroimaging feature extraction algorithm with degenerative patterns for content-based 3D neuroimaging analysis and retrieval, based on the localized 3D Gabor wavelets. Our proposed approach was evaluated with 209 3D clinical neurological imaging studies and compared with the 3D discrete curvelet transform based method and the 3D spatial grey level co-occurrence matrices based method. The preliminary results suggested that our algorithm could support more reliable 3D neuroimaging retrieval.
Sidong Liu, Tom Weidong Cai, Lingfeng Wen, David Dagan Feng
ICIP4
2012 Unsupervised Spectral Mixture Analysis with Hopfield Neural Network for hyperspectral images
abstract
Spectral Mixture Analysis (SMA) has been widely utilized to address the mixed-pixel problem in the quantitative analysis of hyperspectral remote sensing images. Recently Nonnegative Matrix Factorization (NMF) has been successfully utilized to simultaneously perform endmember extraction (EE) and abundance estimation (AE). In this paper, we formulate the solution of NMF by performing EE and AE iteratively. Based on our previous Hopfield Neural Network (HNN) based AE algorithm, an HNN is also constructed for EE to solve the multiplicative updating problem of NMF for SMA. As a result, SMA is conducted in an unsupervised manner and our algorithm is able to extract virtual endmembers without assuming the presence of spectrally pure constituents in hyperspectral scenes. We further extend such strategy to solve the constrained NMF (cNMF) models for SMA, where extra constraints are imposed to better model the mixed-pixel problem. Experimental results on both synthetic and real hyperspectral images demonstrate the effectiveness of our proposed HNN based unsupervised SMA algorithms.
Shaohui Mei, Mingyi He, Zhiyong Wang 0001, David Dagan Feng
ICIP4
2012 Deformable registration model with local rigidity preservation for radiation therapy of lung tumor
abstract
Deformable registration has an important application in radiation therapy and evaluation of patient response to the treatment. However, the conventional deformable registration may introduce excessive deformation on tumor shape and volume, which thereby may lead to less accurate dose delivery on the targeted tumor and possible lesion on the surrounding healthy organs. In this paper, we proposed a deformable registration model to preserve accurate tumor information while retaining the global alignment of the temporal or multimodal biomedical images. Our experimental results on 12 clinical datasets with 15 tumors demonstrated that our deformable model was able to maintain the unique pathological features of the tumor and provide appropriate alignment and correspondence.
Chaojie Zheng, Xiuying Wang 0001, Jinhu Chen, David Dagan Feng
ICIP5
2012 Real-Time Storyboard Generation for H.264/AVC Compressed Videos
abstract
Video summarization enables convenient and efficient management of large volume of visual data. However, most existing summarization approaches are based on either the pixel domain information or conventional video compression standards. As the most recent and popular international video coding standard, H.264/AVC adopts a number of advanced techniques and brings not only opportunities but also challenges to video summarization. In this paper, we propose a real-time image storyboard generation algorithm for H.264/AVC compressed videos by using both compressed domain and pixel domain information jointly and adaptively. This algorithm extracts compressed domain information for visual content representation, video structuring and candidate representative frame selection. By fusing both compressed domain and pixel domain information, the redundancy in the candidate representative frames is further reduced. Our experimental results show that the proposed algorithm can efficiently produce image storyboards conforming to human interpretation of the essential content in generic videos.
Pei Dong, Yong Xia 0001, David Dagan Feng
ICME3
2012 Thoracic Abnormality Detection with Data Adaptive Structure Estimation
Yang Song 0001, Tom Weidong Cai, Yun Zhou 0006, David Dagan Feng
MICCAI (1)4
2012 Browse-to-search
abstract
This demonstration presents a novel interactive online shopping application based on visual search technologies. When users want to buy something on a shopping site, they usually have the requirement of looking for related information from other web sites. Therefore users need to switch between the web page being browsed and other websites that provide search results. The proposed application enables users to naturally search products of interest when they browse a web page, and make their even causal purchase intent easily satisfied. The interactive shopping experience is characterized by: 1) in session---it allows users to specify the purchase intent in the browsing session, instead of leaving the current page and navigating to other websites; 2) in context---the browsed web page provides implicit context information which helps infer user purchase preferences; 3) in focus---users easily specify their search interest using gesture on touch devices and do not need to formulate queries in search box; 4) natural-gesture inputs and visual-based search provides users a natural shopping experience. The system is evaluated against a data set consisting of several millions commercial product images.
Shiyang Lu, Tao Mei 0001, Jingdong Wang 0001, Jian Zhang 0002, Zhiyong Wang 0001, David Dagan Feng, Jian-Tao Sun, Shipeng Li 0001
ACM Multimedia6
2012 What is happening: annotating images with verbs
abstract
Image annotation has been widely investigated to discover the semantics of an image. However, most of the existing algorithms focus on noun tags (e.g. concepts and objects). Since an image is a snapshot of the real world event, annotating images with verbs will enable richer understanding of an image. In this paper, we propose a data-driven approach to verb oriented image annotation. At first, we obtain verb candidates by generating search queries for a given image with initial noun tags and establishing a sentence corpus from those queries. We utilize visualness to filter tags which are not visually presentable (e.g. pain) and differentiate tags into two categories (i.e. scene based and object based) to impose linguistic rules in verb extraction. Then we further re-rank the candidate verbs with the tag context discovered from the images which are both semantically and visually similar to the given image in the MIRFlickr dataset. Our experimental results from user study demonstrate that our proposed approach is promising.
Gang Tian, Genliang Guan, Zhiyong Wang 0001, David Dagan Feng
ACM Multimedia4
2012 Intelligent characterization and evaluation of yarn surface appearance using saliency map analysis, wavelet transform and fuzzy ARTMAP neural network
Bingang Xu, Zheru Chi, David Dagan Feng
Expert Syst. Appl.4
2012 Gabor feature based nonlocal means filter for textured image denoising
Shanshan Wang 0002, Yong Xia 0001, Qiegen Liu, Jianhua Luo, Yue Min Zhu, David Dagan Feng
J. Vis. Commun. Image Represent.6
2012 Salient object detection using content-sensitive hypergraph representation and partitioning
Zheru Chi, Hong Fu, David Dagan Feng
Pattern Recognit.4
2012 2D representation of facial surfaces for multi-pose 3D face recognition
Yanning Zhang 0001, Yong Xia 0001, Zenggang Lin, David Dagan Feng
Pattern Recognit. Lett.5
2012 SparkMed: A Framework for Dynamic Integration of Multimedia Medical Data Into Distributed m-Health Systems
abstract
With the advent of 4G and other long-term evolution (LTE) wireless networks, the traditional boundaries of patient record propagation are diminishing as networking technologies extend the reach of hospital infrastructure and provide on-demand mobile access to medical multimedia data. However, due to legacy and proprietary software, storage and decommissioning costs, and the price of centralization and redevelopment, it remains complex, expensive, and often unfeasible for hospitals to deploy their infrastructure for online and mobile use. This paper proposes the SparkMed data integration framework for mobile healthcare (m-Health), which significantly benefits from the enhanced network capabilities of LTE wireless technologies, by enabling a wide range of heterogeneous medical software and database systems (such as the picture archiving and communication systems, hospital information system, and reporting systems) to be dynamically integrated into a cloud-like peer-to-peer multimedia data store. Our framework allows medical data applications to share data with mobile hosts over a wireless network (such as WiFi and 3G), by binding to existing software systems and deploying them as m-Health applications. SparkMed integrates techniques from multimedia streaming, rich Internet applications (RIA), and remote procedure call (RPC) frameworks to construct a Self-managing, Pervasive Automated netwoRK for Medical Enterprise Data (SparkMed). Further, it is resilient to failure, and able to use mobile and handheld devices to maintain its network, even in the absence of dedicated server devices. We have developed a prototype of the SparkMed framework for evaluation on a radiological workflow simulation, which uses SparkMed to deploy a radiological image viewer as an m-Health application for telemedical use by radiologists and stakeholders. We have evaluated our prototype using ten devices over WiFi and 3G, verifying that our framework meets its two main objectives: 1) interactive delivery of medical multimedia data to mobile devices; and 2) attaching to non-networked medical software processes without significantly impacting their performance. Consistent response times of under 500 ms and graphical frame rates of over 5 frames per second were observed under intended usage conditions. Further, overhead measurements displayed linear scalability and low resource requirements.
Liviu Constantinescu, Jinman Kim, David Dagan Feng
IEEE Trans. Inf. Technol. Biomed.3
2012 Fuzzy Local Gaussian Mixture Model for Brain MR Image Segmentation
abstract
Accurate brain tissue segmentation from magnetic resonance (MR) images is an essential step in quantitative brain image analysis. However, due to the existence of noise and intensity inhomogeneity in brain MR images, many segmentation algorithms suffer from limited accuracy. In this paper, we assume that the local image data within each voxel's neighborhood satisfy the Gaussian mixture model (GMM), and thus propose the fuzzy local GMM (FLGMM) algorithm for automated brain MR image segmentation. This algorithm estimates the segmentation result that maximizes the posterior probability by minimizing an objective energy function, in which a truncated Gaussian kernel function is used to impose the spatial constraint and fuzzy memberships are employed to balance the contribution of each GMM. We compared our algorithm to state-of-the-art segmentation approaches in both synthetic and clinical data. Our results show that the proposed algorithm can largely overcome the difficulties raised by noise, low contrast, and bias field, and substantially improve the accuracy of brain MR image segmentation.
Zexuan Ji, Yong Xia 0001, Quan-Sen Sun, Qiang Chen 0004, De-Shen Xia, David Dagan Feng
IEEE Trans. Inf. Technol. Biomed.6
2012 A Multistage Discriminative Model for Tumor and Lymph Node Detection in Thoracic Images
abstract
Analysis of primary lung tumors and disease in regional lymph nodes is important for lung cancer staging, and an automated system that can detect both types of abnormalities will be helpful for clinical routine. In this paper, we present a new method to automatically detect both tumors and abnormal lymph nodes simultaneously from positron emission tomography-computed tomography thoracic images. We perform the detection in a multistage approach, by first detecting all potential abnormalities, then differentiate between tumors and lymph nodes, and finally refine the detected tumors for false positive reduction. Each stage is designed with a discriminative model based on support vector machines and conditional random fields, exploiting intensity, spatial and contextual features. The method is designed to handle a wide and complex variety of abnormal patterns found in clinical datasets, consisting of different spatial contexts of tumors and abnormal lymph nodes. We evaluated the proposed method thoroughly on clinical datasets, and encouraging results were obtained.
Yang Song 0001, Tom Weidong Cai, Jinman Kim, David Dagan Feng
IEEE Trans. Medical Imaging4
2012 An Adaptive Recognition Model for Image Annotation
abstract
In this paper, an adaptive recognition model (ARM) is proposed for image annotation. The ARM consists of an adaptive classification network (CFN) and a nonlinear correlation network (CLN). The adaptive CFN aims to annotate an image with keywords, and the CLN is used to unveil the correlative information of keywords for annotation refinement. Image annotation is carried out by an ARM in two stages. In the first stage, the features extracted from regions of the input image are fed to a CFN to produce classification labels. In the second stage, the CLN uses keyword correlations learned from the training images to refine the classification result. The ARM works in a forward-propagating manner, resulting in high efficiency in image annotation. Furthermore, the computational time of an ARM is insensitive to the number of regions of the input image and the vocabulary size. In this paper, the effect of keyword correlation in image annotation is, comprehensively, investigated on a real image dataset and a synthetic image dataset. The exploitation of a controllable synthetic dataset helps to systematically study the function of keyword correlation and effectively analyze the performance of the ARM. Experimental results demonstrate the efficiency and effectiveness of the ARM.
Zenghai Chen, Hong Fu, Zheru Chi, David Dagan Feng
IEEE Trans. Syst. Man Cybern. Part C4
2011 Lung tumor delineation in PET-CT images using a downhill region growing and a Gaussian mixture model
abstract
Combined PET-CT is now increasingly used for the clinical evaluation of cancer and is arguably the best tool to stage non-small cell lung cancer (NSCLC). We propose a framework to better delineate lung tumors which utilizes information from PET and CT images. The framework is based on a downhill region growing technique for PET and a Gaussian mixture model for CT images. We applied our framework in 20 PET-CT studies from patients with NSCLC. Experiments show that our method is able to delineate lung tumors in complex cases where the tumors are located near other organs with similar intensities in PET images or when the tumors extends into the chest wall or the mediastinum. We also compared 10 of the datasets with experts performing manual delineation, which produced a volumetric overlapped fraction of 0.78 ± 0.10.
Cherry G. Ballangan, Xiuying Wang 0001, Michael J. Fulham, Stefan Eberl, David Dagan Feng
ICIP5
2011 Real-time moving object segmentation and tracking for H.264/AVC surveillance videos
abstract
With increased use of H.264/AVC in various applications including video surveillance systems, feature extraction and knowledge representation in compressed domain are becoming attractive. A real-time H.264/AVC compressed domain moving object segmentation and tracking algorithm for surveillance videos is proposed in this paper. This algorithm consists of moving object detection, bounding box matching, spatiotemporal merge and split reasoning and trajectory smoothing, with major innovation in incorporating the information provided by the prediction modes into the framework of motion detection and trajectory construction. The experimental results on both indoor and outdoor surveillance videos demonstrate that the adaptive use of the information from motion vectors, DCT coefficients and prediction modes can substantially improve the performance of moving object segmentation and tracking.
Pei Dong, Yong Xia 0001, Li Zhuo 0001, David Dagan Feng
ICIP4
2011 Nonlinear curvelet diffusion for noisy image enhancement
abstract
Digital image degradation normally arises during image acquisition and processing, which has a direct influence on the visual quality of the image. This paper proposes a combined method for enhancement of noisy image by using the mirror-extended curvelet transform and nonlinear anisotropic diffusion. First, an improved enhancement function is proposed to nonlinearly shrink and stretch the curvelet coefficients. Then, the enhanced results are further processed by the nonlinear diffusion where only the nonsignificant, i.e., nonthresholded, curvelet coefficients are changed by means of a diffusion process in order to reduce the pseudo-Gibbs artifacts. Experimental results indicate the proposed method has better performances to enhance the shape of edges and important detailed features as well as suppress noise, in comparison to the curvelet-based enhancement method without diffusion and the wavelet-based enhancement methods with/without diffusion.
Ying Li 0017, Huijun Ning, Yanning Zhang 0001, David Dagan Feng
ICIP4
2011 Discriminative Pathological Context Detection in Thoracic Images Based on Multi-level Inference
Yang Song 0001, Tom Weidong Cai, Stefan Eberl, Michael J. Fulham, David Dagan Feng
MICCAI (3)5
2011 StoryImaging: a media-rich presentation system for textual stories
abstract
In this demo, we develop the StoryImaging system to illustrate a textual story with both images harvested from the Web and synthesized speech. At the backend, a story is firstly processed to identify key terms such as named entities and to obtain the story summary. With the aid of commercial search engines, images are then collected from the Web for those key terms and re-ranked by taking the summary as context. At last, images are clustered to provide an overview of the story. At the web-based frontend, the user interface has been tailored to both improve information comprehension and provide engaging and explorative experiences for users by closely bridging textual and visual modalities.
Genliang Guan, Zhiyong Wang 0001, Xian-Sheng Hua 0001, David Dagan Feng
ACM Multimedia4
2011 A Time Delay Neural Network model for simulating eye gaze data
abstract
Human eye movement modelling is a new, challenging and promising research topic in computer vision. Human eye movement modelling aims at simulating the scan path in which a human being views an image, a scene or a video. The successful modelling of human eye movements potentially benefits a wide range of applications such as image retrieval, image annotation, medical image diagnosis and human visual perception. This article presents a model based on a Time Delay Neural Network (TDNN) to simulate eye gaze data. First, 120 Hz eye gaze data are acquired by a non-intrusive table-mounted eye tracker. Our proposed model is to simulate the image reading process of a single subject. Seven features are then extracted based on the knowledge of the human oculomotor system and the image contents to train a TDNN. Finally, the trained TDNN combined with a saccade control mechanism is used to simulate the scan path of a human being viewing an image. The proposed model can generate 600 points of raw eye gaze data in a 5-second eye viewing window. Both subjective and objective methods are used to evaluate the model by comparing its behaviour and characteristics with the real eye gaze data collected from an eye tracker. Qualitative assessment shows that the subject can hardly tell the differences between the scan path from the model and that from a human being. By evaluating the coincident probability Cp and coincident significance Cs , quantitative assessment shows that the results from the TDNN model are reasonable and similar to human scan paths.
Hong Fu, Zheru Chi, David Dagan Feng
J. Exp. Theor. Artif. Intell.7
2011 Fast and accuracy extraction of infrared target based on Markov random field
Ying Li 0017, Xingjin Mao, David Dagan Feng, Yanning Zhang 0001
Signal Process.3
2011 An Adaptive Method of Speckle Reduction and Feature Enhancement for SAR Images Based on Curvelet Transform and Particle Swarm Optimization
abstract
This paper proposes an adaptive method based on the mirror-extended curvelet transform and the improved particle swarm optimization (PSO) algorithm, which reduce speckle noise and enhance edge features and contrast of synthetic aperture radar (SAR) images. First, an improved gain function, which integrates the speckle reduction with the feature enhancement, is introduced to nonlinearly shrink and stretch the curvelet coefficients. Then, a novel objective criterion for the quality of the despeckled and enhanced images is proposed in order to adaptively obtain the optimal parameters in the gain function. Finally, the PSO algorithm is employed as a global search strategy for the best despeckled and enhanced image. In order to increase the convergence speed and avoid the premature convergence, two further improvements for the classic PSO algorithm are presented. That is, a new learning scheme and a mutation operator are introduced. Experimental results demonstrate that the proposed method can efficiently reduce the speckle and enhance the edge features and the contrast of SAR images and outperforms the wavelet- and curvelet-based nonadaptive despeckling and enhancement methods.
Ying Li 0017, Hongli Gong, David Dagan Feng, Yanning Zhang 0001
IEEE Trans. Geosci. Remote. Sens.3
2011 Improving Spatial-Spectral Endmember Extraction in the Presence of Anomalous Ground Objects
abstract
Endmember extraction (EE) has been widely utilized to extract spectrally unique and singular spectral signatures for spectral mixture analysis of hyperspectral images. Recently, spatial–spectral EE (SSEE) algorithms have been proposed to achieve superior performance over spectral EE (SEE) algorithms by taking both spectral similarity and spatial context into account. However, these algorithms tend to neglect anomalous endmembers that are also of interest. Therefore, in this paper, an improved SSEE (iSSEE) algorithm is proposed to address such limitation of conventional SSEE algorithms by accounting for both anomalous and normal endmembers. By developing simplex projection and simplex complementary projection, all the hyperspectral pixels are projected into a simplex determined by the normal endmembers extracted in conventional SSEE algorithms. As a result, anomalous endmembers are identified iteratively by utilizing the$l_{2}^{\infty}$norm to find the maximum simplex complementary projection. In order to determine how many anomalous endmembers are to be extracted, a novel Residual-be-Noise Probability-based algorithm is also proposed by elegantly utilizing the spatial-purity map generated in the previous SSEE step. Experimental results on both synthetic and real datasets demonstrate that simplex projection errors can be significantly reduced by identifying both anomalous and normal endmembers in the proposed iSSEE algorithm. It is also confirmed that the performance of the proposed iSSEE algorithm clearly outperforms that of SEE algorithms since both spatial context and spectral similarity are utilized.
Shaohui Mei, Mingyi He, Yifan Zhang 0006, Zhiyong Wang 0001, David Dagan Feng
IEEE Trans. Geosci. Remote. Sens.5
2011 Automated Delineation of Lung Tumors in PET Images Based on Monotonicity and a Tumor-Customized Criterion
abstract
Reliable automated or semiautomated lung tumor delineation methods in positron emission tomography should provide accurate tumor boundary definition and separation of the lung tumor from surrounding tissue or "hot spots" that have similar intensities to the lung tumor. We propose a tumor-customized downhill (TCD) method to achieve these objectives. Our approach includes: 1) automatic formulation of a tumor-customized criterion to improve tumor boundary definition, 2) a monotonic property of the standardized uptake value (SUV) of tumors to separate the tumor from adjacent regions of increased metabolism ("hot spot"), and 3) accounts for tumor heterogeneity. Three simulated lesions and 30 PET-CT studies, grouped into "simple" and "complex" groups, were used for evaluation. Our main findings are that TCD, when compared to the threshold based on 40% and 50% maximum SUV, adaptive threshold, Fuzzy c-means, and watershed techniques achieved the highest Dice's similarity coefficient average for simulation data (0.73) and "complex" group (0.71); the least volumetric error in the "simple" (1.76 mL) and the "complex" group (14.59 mL); and TCD solves the problem of leakage into adjacent tissues when many other techniques fail.
Cherry G. Ballangan, Xiuying Wang 0001, Michael J. Fulham, Stefan Eberl, David Dagan Feng
IEEE Trans. Inf. Technol. Biomed.6
2011 Hybrid Genetic and Variational Expectation-Maximization Algorithm for Gaussian-Mixture-Model-Based Brain MR Image Segmentation
abstract
The expectation-maximization (EM) algorithm has been widely applied to the estimation of gaussian mixture model (GMM) in brain MR image segmentation. However, the EM algorithm is deterministic and intrinsically prone to overfitting the training data and being trapped in local optima. In this paper, we propose a hybrid genetic and variational EM (GA-VEM) algorithm for brain MR image segmentation. In this approach, the VEM algorithm is performed to estimate the GMM, and the GA is employed to initialize the hyperparameters of the conjugate prior distributions of GMM parameters involved in the VEM algorithm. Since GA has the potential to achieve global optimization and VEM can steadily avoid overfitting, the hybrid GA-VEM algorithm is capable of overcoming the drawbacks of traditional EM-based methods. We compared our approach to the EM-based, VEM-based, and GA-EM based segmentation algorithms, and the segmentation routines used in the statistical parametric mapping package and FMRIB Software Library in 20 low-resolution and 17 high-resolution brain MR studies. Our results show that the proposed approach can improve substantially the performance of brain MR image segmentation.
Guangjian Tian, Yong Xia 0001, Yanning Zhang 0001, David Dagan Feng
IEEE Trans. Inf. Technol. Biomed.4
2011 A Hybrid Clustering Method for ROI Delineation in Small-Animal Dynamic PET Images: Application to the Automatic Estimation of FDG Input Functions
abstract
Tracer kinetic modeling with dynamic positron emission tomography (PET) requires a plasma time-activity curve (PTAC) as an input function. Several image-derived input function (IDIF) methods that rely on drawing the region of interest (ROI) in large vascular structures have been proposed to overcome the problems caused by the invasive approach for obtaining the PTAC, especially for small-animal studies. However, the manual placement of ROIs for estimating IDIF is subjective and labor-intensive, making it an undesirable and unreliable process. In this paper, we propose a novel hybrid clustering method (HCM) that objectively delineates ROIs in dynamic PET images for the estimation of IDIFs, and demonstrate its application to the mouse PET studies acquired with [ (18)F]Fluoro-2-deoxy-2-D-glucose (FDG). We begin our HCM using k-means clustering for background removal. We then model the time-activity curves using polynomial regression mixture models in curve clustering for heart structure detection. The hierarchical clustering is finally applied for ROI refinements. The HCM achieved accurate ROI delineation in both computer simulations and experimental mouse studies. In the mouse studies, the predicted IDIF had a high correlation with the gold standard, the PTAC derived from the invasive blood samples. The results indicate that the proposed HCM has a great potential in ROI delineation for automatic estimation of IDIF in dynamic FDG-PET studies.
Xiujuan Zheng, Guangjian Tian, Sung-Cheng Huang, David Dagan Feng
IEEE Trans. Inf. Technol. Biomed.4
2010 Combined Retrieval Strategies for Images with and without Distinct Objects
Hong Fu, Zheru Chi, David Dagan Feng
ACIVS (1)3
2010 Salient-SIFT for Image Retrieval
Hong Fu, Zheru Chi, David Dagan Feng
ACIVS (1)4
2010 Localized multiscale texture based retrieval of neurological image
abstract
The volume and complexity of neurological images have significantly increased, which leads to challenges in efficient data management and retrieval. In this paper, we developed a new content-based image retrieval framework with the localized multiscale Discrete Curvelet Transform (DCvT) features extracted from parametric neurological images. We also compared the performance of three different irregular-to-regular shape padding methods. 142 patient data with neurodegenerative disorders were used in the evaluation. The preliminary results show that our proposed framework supports fast neuroimaging retrieval, and the orthographic projection method can reduce the computational complexity and has a great potential to improve the retrieval for indefinite cases.
Sidong Liu, Tom Weidong Cai, Lingfeng Wen, Stefan Eberl, Michael J. Fulham, David Dagan Feng
CBMS7
2010 A content-based image retrieval framework for multi-modality lung images
abstract
This paper presents a framework for effective and fast content-based image retrieval for multi-modality PET-CT lung scans. PET-CT scans present significant advantages in tumor staging, but also place new challenges in computerized image analysis and retrieval. Our framework comprises 5 major components: lung field estimation, texture feature extraction, feature categorization, refinement using SVM, and similarity measure. Clinical data from lung cancer patients are used as case studies, and effective retrieval performance is demonstrated.
Yang Song 0001, Tom Weidong Cai, Stefan Eberl, Michael J. Fulham, David Dagan Feng
CBMS5
2010 Content-based image retrieval using a combination of visual features and eye tracking data
abstract
Image retrieval technology has been developed for more than twenty years. However, the current image retrieval techniques cannot achieve a satisfactory recall and precision. To improve the effectiveness and efficiency of an image retrieval system, a novel content-based image retrieval method with a combination of image segmentation and eye tracking data is proposed in this paper. In the method, eye tracking data is collected by a non-intrusive table mounted eye tracker at a sampling rate of 120 Hz, and the corresponding fixation data is used to locate the human's Regions of Interest (hROIs) on the segmentation result from the JSEG algorithm. The hROIs are treated as important informative segments/objects and used in the image matching. In addition, the relative gaze duration of each hROI is used to weigh the similarity measure for image retrieval. The similarity measure proposed in this paper is based on a retrieval strategy emphasizing the most important regions. Experiments on 7346 Hemera color images annotated manually show that the retrieval results from our proposed approach compare favorably with conventional content-based image retrieval methods, especially when the important regions are difficult to be located based on visual features.
Hong Fu, Zheru Chi, David Dagan Feng
ETRA5
2010 Eye movement as an interaction mechanism for relevance feedback in a content-based image retrieval system
abstract
Relevance feedback (RF) mechanisms are widely adopted in Content-Based Image Retrieval (CBIR) systems to improve image retrieval performance. However, there exist some intrinsic problems: (1) the semantic gap between high-level concepts and low-level features and (2) the subjectivity of human perception of visual contents. The primary focus of this paper is to evaluate the possibility of inferring the relevance of images based on eye movement data. In total, 882 images from 101 categories are viewed by 10 subjects to test the usefulness of implicit RF, where the relevance of each image is known beforehand. A set of measures based on fixations are thoroughly evaluated which include fixation duration, fixation count, and the number of revisits. Finally, the paper proposes a decision tree to predict the user's input during the image searching tasks. The prediction precision of the decision tree is over 87%, which spreads light on a promising integration of natural eye movement into CBIR systems in the future.
Hong Fu, Zheru Chi, David Dagan Feng
ETRA5
2010 A neural network model with adaptive structure for image annotation
abstract
A neural network model with adaptive structure for image annotation is proposed in this paper. The adaptive structure enables the proposed model to utilize both global and regional visual features, as well as correlative information of annotated keywords for annotation. In order to achieve an approximate global optimum rather than a local optimum, both genetic algorithm and traditional back-propagation algorithm, are combined for model training. The neural network model is experimented on a synthetic image dataset with controllable parameters, which has not been used in previous image annotation experiments. Experimental results demonstrate the effectiveness of the proposed model.
Zenghai Chen, Hong Fu, Zheru Chi, David Dagan Feng
ICARCV4
2010 An improved algorithm for segmenting and recognizing connected handwritten characters
abstract
In this paper, an improved algorithm is proposed for the segmentation and recognition of handwritten character strings. In the method, a gradient descent mechanism is used to weigh the distance measure in applying KNN for segmenting/recognizing connected characters (numerals and Chinese characters) in the left-to-right scanning direction. In recognizing connected characters, a high quality segmentation technique is essential. Conventional approaches attempt to separate the string into individual characters without recognition and apply a recognition algorithm onto each isolated character, resulting improper segmentation and poor recognition results in many situations. Our proposed algorithm simulates the human beings's process in recognizing connected character strings where segmentation and recognition is mingled with each other. Experimental results on 1959 character strings from the USPS database of postal envelopes show that the algorithm works robustly and efficiently.
Zheru Chi, David Dagan Feng
ICARCV3
2010 3D face representation and recognition by Intrinsic Shape Description Maps
abstract
We present a novel method for 3D face recognition, in which the 3D facial surface is first mapped into a 2D domain with specified resolution through a global optimization by constrained conformal geometric maps. The Intrinsic Shape Description Map (ISDM) is then constructed through a modeling technique capable to express geometric and appearance information of the 3D face. Hence the 3D surface matching problem can be simplified to a 2D image matching problem, which greatly reduces the computational complexity. Finally, the Intrinsic Shape Description Feature (ISDF) of ISDM and the discrimination analysis can be calculated. Experimental results implemented on GavabDB demonstrate that our proposed method significantly outperforms the existing methods with respect to pose variation.
Yanning Zhang 0001, Yong Xia 0001, Zenggang Lin, David Dagan Feng
ICASSP5
2010 3D neurological image retrieval with localized pathology-centric CMRGlc patterns
abstract
Functional neuroimaging has an important role in non-invasive diagnosis of neurodegenerative disorders. There are now large volumes of imaging data generated by functional imaging technologies and so there is a need to efficiently manage and retrieve these data. In this paper, we propose a new scheme for efficient 3D content-based neurological image retrieval. 3D pathology-centric masks were adaptively designed and applied for extracting CMRGlc (cerebral metabolic rate of glucose consumption) texture features with volumetric co-occurrence matrices from neurological FDG PET images. Our results, using 93 clinical dementia studies, show that our approach offers a robust and efficient retrieval mechanism for relevant clinical cases and provides advantages in image data analysis and management.
Tom Weidong Cai, Sidong Liu, Lingfeng Wen, Stefan Eberl, Michael J. Fulham, David Dagan Feng
ICIP6
2010 Fully automated liver segmentation for low- and high- contrast CT volumes based on probabilistic atlases
abstract
Automated liver segmentation is problematic due to variations in liver shape / size and because the liver has a similar density distribution to surrounding structures. We propose a method that: 1) utilizes iteratively constructed probabilistic liver and rib cage atlases, 2) conducts the Gaussian distribution analysis to avoid incorrectly classifying the irrelevant surrounding tissues as `liver region' in the conventional probabilistic atlas based method, and maps the intensity range of the input candidate liver region onto the liver atlas, 3) retrieves the `missing parts' of the liver by deformable registration. Our approach is automated and able to segment the liver from high-contrast and low-contrast CT volumes. Forty clinical CT studies were used for atlas construction and validation. Our method outperformed two other probabilistic atlas-based liver segmentation methods.
ChangYang Li, Xiuying Wang 0001, Stefan Eberl, Michael J. Fulham, David Dagan Feng
ICIP6
2010 Multiscale deformable registration using edge preserving scale space for adaptive radiation therapy
abstract
Registration of planning images with daily images is an important component for adaptive radiation therapy (ART). In this paper, a multiscale deformable registration framework is proposed by combining edge preserving scale space with the free form deformation (FFD) for registration of planning computed tomography (CT) images with daily cone beam CT (CBCT) images. The edge preserving scale space which is able to select edges and contours of an image according to their geometric size is derived from the total variation model with the L1 norm (TV-L1). At each scale, the selected edges and contours are sufficiently strong to drive the deformation using the FFD grid, then the deformation fields are gained by a coarse to fine manner. Furthermore, for automated registration we design an optimal estimation of the TV-L1 parameter by minimizing the defined offset. The experiments on CT and CBCT images show accuracy and robustness when compared to traditional methods.
Dengwang Li, Xiuying Wang 0001, David Dagan Feng
ICIP5
2010 Refining a region based attention model using eye tracking data
abstract
Computational visual attention modeling is a topic of increasing importance in machine understanding of images. In this paper, we present an approach to refine a region based attention model with eye tracking data. This paper has three main contributions. (1) A concept of fixation mask is proposed to describe the region saliency of an image by weighting the segmented regions using importance measures obtained in the Human Visual System (HVS) or computational models. (2) A Genetic Algorithm (GA) scheme for refining a region based attention model is proposed. (3) An evaluation method is developed to measure the correlation between the result from the computational model and that from the HVS in terms of fixation mask.
Hong Fu, Zheru Chi, David Dagan Feng
ICIP4
2010 Adaptive reference frame selection for near-duplicate video shot detection
abstract
Near-duplicate video shots provide critical visual link between videos and detecting such video shots efficiently and effectively is of paramount importance in many applications such as detecting copyright infringement. In this paper, we propose an improved near-duplicate video shot detection approach by adaptively selecting reference frames for more effective shot representation. The correlation between adjacent frames is measured with Pearson's Correlation Coefficient (PCC) so that a set of compact yet representative reference frames can be selected adaptively in terms of content variation within video shots. Interest points are further extracted from the selected frames to effectively represent shot contents for similarity matching. Comprehensive experimental results on TRECVID-2008 corpus demonstrate that our proposed approach outperforms the state-of-the-art method effectively.
Shiyang Lu, Zhiyong Wang 0001, Maximilian Ott, David Dagan Feng
ICIP5
2010 Dual-modality 3D brain PET-CT image segmentation based on probabilistic brain atlas and classification fusion
abstract
The increasing prevalence of dual medical imaging modalities, such as PET-CT scanners, poses both challenges and opportunities to image segmentation, as they provide distinct but complementary information. In this paper, we propose a novel segmentation algorithm for 3D brain PET-CT images, which classifies each voxel by fusing the voxel's memberships estimated from four points of view using the PET information, CT information, smoothness prior, and probabilistic brain atlas. All memberships having the same dynamic range greatly facilitates weighting the contribution of the four different information sources. The probabilistic brain atlas estimated for each PET-CT image from a set of training samples provides the anatomical information to the segmentation process. We compared the proposed algorithm to three single-classifier based methods, PET-based SPM algorithm, CT-based Otsu thresholding, and PET-CT based MAP-MRF algorithm. The experimental results in 11 clinical brain PET-CT studies demonstrate that the novel algorithm is capable of providing more accurate and reliable segmentation.
Yong Xia 0001, Stefan Eberl, David Dagan Feng
ICIP3
2010 Two-step similarity matching for Content-Based Video Retrieval in P2P, networks
abstract
Multimedia data, particularly, video data, has dominated peer-to-peer (P2P) networks. Therefore, it is demanding to provide content based retrieval in P2P networks. Similarity matching is one of the challenging issues. In this paper, we present a novel two-step method to reduce computational complexity of similarity matching in P2P networks. In the first step, an efficient maximum matching (MM) technique is employed to obtain an initial set of similar video candidates. In the second step, these candidates are further selected with a more accurate, but more computationally expensive optimal matching (OM) technique. In order to further improve the computational efficiency of the proposed method, four other matching techniques are proposed to replace MM technique. Various experimental results indicate that the proposed approach is more effective for CBVR while achieving significantly computational saving.
Jin Niu, Zhiyong Wang 0001, David Dagan Feng
ICME3
2010 A Novel Shape-Based Image Classification Method by Featuring Radius Histogram of Dilating Discs Filled into Regular and Irregular Shapes
Zheru Chi, David Dagan Feng
ICONIP (1)4
2010 A wearable, wireless electronic interface for textile sensors lin shu
abstract
Electronic interfaces for wearable sensors require wireless connection, appropriate measurement range, small size, simple and robust structure, insensitivity to noise and comfort to wearers in daily life. A novel wearable, wireless electronic interface for resistive textile sensors is presented, which meets the requirements and has the ability to provide real-time measurement. System configuration, accuracy and resolution have been analyzed in-depth and design rules have been defined. Experimental results show that this electronic interface exhibits less than 1% error in a large measurement range for wearable textile resistive sensors. It also shows a good stability to power supply interference. The interface has been successfully applied in a foot pressure measurement syste.
Xiaoming Tao 0002, David Dagan Feng
ISCAS3
2010 Mixture Analysis by Multichannel Hopfield Neural Network
abstract
Due to the spatial-resolution limitation, mixed pixels containing energy reflected from more than one type of ground objects are widely present in remote sensing images, which often results in inefficient quantitative analysis. To effectively decompose such mixtures, a fully constrained linear unmixing algorithm based on a multichannel Hopfield neural network (MHNN) is proposed in this letter. The proposed MHNN algorithm is actually a Hopfield-based architecture which handles all the pixels in an image synchronously, instead of considering a per-pixel procedure. Due to the synchronous unmixing property of MHNN, a noise energy percentage (NEP) stopping criterion which utilizes the signal-to-noise ratio is proposed to obtain optimal results for different applications automatically. Experimental results demonstrate that the proposed multichannel structure makes the Hopfield-based mixture analysis feasible for real-world applications with acceptable time cost. It has also been observed that the proposed MHNN-based mixture-analysis algorithm outperforms the other two popular linear mixture-analysis algorithms and that the NEP stopping criterion can approach optimal unmixing results adaptively and accurately.
Shaohui Mei, Mingyi He, Zhiyong Wang 0001, David Dagan Feng
IEEE Geosci. Remote. Sens. Lett.4
2010 Robust, accurate and efficient face recognition from a single training image: A uniform pursuit approach
Weihong Deng, Jiani Hu, Jun Guo 0002, Tom Weidong Cai, David Dagan Feng
Pattern Recognit.5
2010 Emulating biological strategies for uncontrolled face recognition
Weihong Deng, Jiani Hu, Jun Guo 0002, Tom Weidong Cai, David Dagan Feng
Pattern Recognit.5
2010 Recognition of attentive objects with a concept association network for image annotation
Hong Fu, Zheru Chi, David Dagan Feng
Pattern Recognit.3
2010 Multifractal signature estimation for textured image segmentation
Yong Xia 0001, David Dagan Feng, Rongchun Zhao, Yanning Zhang 0001
Pattern Recognit. Lett.2
2010 An efficient retinex-like brightness normalization method for coding camera flashes and strong brightness variation in videos
Hoi-Kok Cheung, Wan-Chi Siu, David Dagan Feng, Zhiyong Wang 0001
Signal Process. Image Commun.3
2010 Spatial Purity Based Endmember Extraction for Spectral Mixture Analysis
abstract
Spectral mixture analysis (SMA) has been widely utilized to address the mixed-pixel problem in the quantitative analysis of hyperspectral remote sensing images, in which endmember extraction (EE) plays an extremely important role. In this paper, a novel algorithm is proposed to integrate both spectral similarity and spatial context for EE. The spatial context is exploited from two aspects. At first, initial endmember candidates are identified by determining the spatial purity (SP) of pixels in their spatial neighborhoods (SNs). Several SP measurements are investigated at both intensity level and feature level. In order to alleviate local spectra variability, the average of the pixels in pure SNs are voted as endmember candidates. Then, the spatial connectivity is utilized to merge spatially related endmember candidates by finding connection paths in a graph so that the number of endmember candidates is further reduced, which results in computational efficiency and better performance in SMA by alleviating global spectral variability. Experimental results on both synthetic and real hyperspectral images demonstrate that the proposed SP based EE (SPEE) algorithm outperforms the other popular EE algorithms. It is also observed that feature-level SP measurements are more distinguishable than intensity-level SP measurements to discriminate pure SNs from mixed SNs.
Shaohui Mei, Mingyi He, Zhiyong Wang 0001, David Dagan Feng
IEEE Trans. Geosci. Remote. Sens.4
2010 In-shoe plantar pressure measurement and analysis system based on fabric pressure sensing array
abstract
Spatial and temporal plantar pressure distributions are important and useful measures in footwear evaluation, athletic training, clinical gait analysis, and pathology foot diagnosis. However, present plantar pressure measurement and analysis systems are more or less uncomfortable to wear and expensive. This paper presents an in-shoe plantar pressure measurement and analysis system based on a textile fabric sensor array, which is soft, light, and has a high-pressure sensitivity and a long service life. The sensors are connected with a soft polymeric board through conductive yarns and integrated into an insole. A stable data acquisition system interfaces with the insole, wirelessly transmits the acquired data to remote receiver through Bluetooth path. Three configuration modes are incorporated to gain connection with desktop, laptop, or smart phone, which can be configured to comfortably work in research laboratories, clinics, sport ground, and other outdoor environments. A real-time display and analysis software is presented to calculate parameters such as mean pressure, peak pressure, center of pressure (COP), and shift speed of COP. Experimental results show that this system has stable performance in both static and dynamic measurements.
Tao Hua, Yangyong Wang, David Dagan Feng, Xiaoming Tao 0002
IEEE Trans. Inf. Technol. Biomed.5
2009 Eye movement data modeling using a genetic algorithm
abstract
We present a computational model of human eye movements based on a genetic algorithm (GA). The model can generate elemental raw eye movement data in a four-second eye viewing window with a 25 Hz sampling rate. Based on the physiology and psychology characters of human vision system, the fitness function of the GA model is constructed by taking into consideration of five factors including the saliency map, short time memory, saccades distribution, Region of Interest (ROI) map, and a retina model. Our model can produce the scan path of a subject viewing an image, not just several fixations points or artificial ROI's as in the other models. We have also developed both subjective and objective methods to evaluate the model by comparing its behavior with the real eye movement data collected from an eye tracker. Tested on 18 (9 times 2) images from both an obvious-object image group and a non-obvious-object image group, the subjective evaluations shows very close scores between the scan paths generated by the GA model and those real scan paths; for the objective evaluation, experimental results show that the distance between GA's scan paths and human scan paths of the same image has no significant difference by a probability of 78.9% on average.
Hong Fu, Zheru Chi, David Dagan Feng
IEEE Congress on Evolutionary Computation6
2009 A Method Based on Geometric Invariant Feature for 3D Face Recognition
abstract
3D information provides a significant improvement in recognition performance over 2D facial image data. However, the existing 3D approaches show limitations dealing with pose variation, e.g., 3D facial surfaces need to be aligned before the match operation. In this paper, an original framework which has the scale, rotation and expression invariance based on geometric invariant feature is proposed for automatic face recognition without pre-registration. In this study, 3D face scans are first pre-processed, including mesh cropping, holes filling, and mesh regularization; subsequently, the geometric invariant feature combined the local shape variation feature with spatial geometric feature which is invariant to scale and pose is extracted. Experimental results implemented on GavabDB and our purpose-selected database demonstrate that our proposed method significantly outperforms the state-of-the-art methods with respect to pose and facial expression variation.
Yanning Zhang 0001, Zenggang Lin, David Dagan Feng
ICIG4
2009 Hierarchical Gaussian Mixture Model for Image Annotation via PLSA
abstract
In order to mimic the representation of textual documents, some approaches have recently been proposed to represent visual contents in terms of visual words in many applications such as object recognition and image annotation. In this paper, we propose to build an effective visual vocabulary by using Hierarchical Gaussian Mixture model instead of traditional clustering methods. In addition, Probabilistic Latent Semantic Analysis is employed to explore semantic aspects of visual concepts and to discover topic clusters among documents and visual words so that every image is projected on to a lower dimensional topic space for more efficient and effective annotation. Experimental results obtained on TRECVID 2005 dataset demonstrate that the Hierarchical Gaussian Mixture model can achieve better annotation performance than hierarchical k-means clustering even by using simple k-NN annotation scheme.
Zhiyong Wang 0001, Huaibin Yi, David Dagan Feng
ICIG4
2009 A General Image Segmentation Model and its Application
abstract
This paper proposed a general image segmentation model, namely the energy-minimization based image segmentation (EMBIS) model. This model converts image segmentation into a controlled optimization process minimizing the weighted sum of the feature energy and spatial energy, which interpret the homogeneity restriction and spatial constraints, respectively. The EMBIS model provides a unified understanding of various existing segmentation algorithms, and can also serve as a framework for systematic generation of new segmentation algorithms. We provided four examples to illustrate that many existing segmentation algorithms are indeed specialized cases of this model with different instances of both energy functions. We also presented a case study to demonstrate how to use this model to create new algorithms and resulted in the spatial-constrained OTSU (SC-OTSU) algorithm, where segmentation can be achieved by minimizing the feature energy of the OTSU algorithm and spatial energy of the algorithm based on a simple MRF (SMRF) model. Evaluation on both synthetic and real images proved that novel segmentation algorithms derived form the proposed EMBIS model can provide accurate and efficient image segmentation.
Yong Xia 0001, David Dagan Feng
ICIG2
2009 Tree structures with attentive objects for image classification using a neural network
abstract
This paper presents an image classification method based on a neural network model dealing with tree structures of attentive objects. Apart from regions provided by image segmentation, attentive objects, which are extracted from a segmented image by an attention-driven image interpretation algorithm, are used to construct the tree structure to represent an image. Three combinations of tree structures are investigated, including ldquoimage + attentive-object + segmentsrdquo, ldquoimage + attentive-objectsrdquo, as well as ldquoimage + segmentsrdquo. Structure based neural networks are trained to classify the images by using the back propagation through structure (BPTS) algorithm. Experimental results show that the ldquoimage + attentive objectsrdquo structure is more favorable, comparing with both the other two structures proposed by us and a start-of-art tree structure reported in the literature, in terms of classification rate and computational time.
Hong Fu, Shuya Zhang, Zheru Chi, David Dagan Feng
IJCNN4
2009 Improved concept similarity measuring in the visual domain
abstract
Exploring semantic similarity between concepts in visual domain has a wide range of applications such as natural language processing and multimedia retrieval, which in general requires both a large pool of sample images for each concept and a model to capture its visual characteristics. Instead of relying on high quality and large quantity sample data which is very difficult to obtain, in this paper, a novel method is proposed to achieve improvement in measuring concept similarity by incorporating concept modeling technique into data pruning process. At first, a number of sampling concept models are obtained by sampling a subset from the sample dataset of each concept. Then noisy samples are discarded in terms of their probabilities to the sampling concept models. Experimental results on 31,275 Web images of 38 concepts defined in LSCOM indicate that the concept similarity obtained through our proposed approach is more coherent to human cognition. A concept hierarchy tree built from the 38 concepts with their similarity further demonstrates the effectiveness of our proposed method.
Genliang Guan, Zhiyong Wang 0001, Qi Tian 0001, David Dagan Feng
MMSP4
2009 Reliable object recognition using SIFT features
abstract
SIFT (scale invariant feature transform) features have been one of the most efficient descriptors for object recognition. However, the excessive number of key points and high dimensionality has limited its capacity in object recognition. In this paper we present a novel method based on SIFT features for reliable object recognition. At first, a matching tree is constructed to eliminate non-essential key points. In order to achieve viewpoint independence, a 3D model is constructed for each object in the filtered SIFT feature space. Experimental results on both Caltech 101 and COIL 100 datasets indicate the effectiveness of our proposed algorithm.
Florin Alexandru Pavel, Zhiyong Wang 0001, David Dagan Feng
MMSP3
2009 Two-level indexing for high-dimensional range queries in peer-to-peer networks
abstract
Supporting complex and efficient lookup queries in peer-to-peer networks is challenging, though simple keyword based lookup queries are well supported by most deployed systems. This paper presents a two-level indexing structure built on distributed hash table (DHT) aiming to support range queries on high-dimensional feature space in peer-to-peer network. Unlike most existing systems, where every node is responsible for a data partition, our design only utilizes a small part of the nodes to manage partitions. These partition nodes form the first level index. The second level index consists of one or more server nodes, which maintains links to each partition node. Additionally, a merge and split mechanism is designed to dynamically adjust the workload among nodes. Experimental results indicate that our system offers promising performance in terms of workload balance in churn networks. The flexibility to work with any DHT and the capability to support multiple feature spaces further make our proposed approach a feasible extension for file sharing networks.
Lelin Zhang, Zhiyong Wang 0001, David Dagan Feng
MMSP3
2009 Detecting Ghost and Left Objects in Surveillance Video
abstract
This paper proposes an efficient method for detecting ghost and left objects in surveillance video, which, if not identified, may lead to errors or wasted computational power in background modeling and object tracking in video surveillance systems. This method contains two main steps: the first one is to detect stationary objects, which narrows down the evaluation targets to a very small number of regions in the input image; the second step is to discriminate the candidates between ghost and left objects. For the first step, we introduce a novel stationary object detection method based on continuous object tracking and shape matching. For the second step, we propose a fast and robust inpainting method to differentiate between ghost and left objects by reconstructing the real background using the candidate's corresponding regions in the current input and background image. The effectiveness of our method has been validated by experiments over a variety of video sequences and comparisons with existing state-of-art methods.
Sijun Lu, Jian Zhang 0002, David Dagan Feng
Int. J. Pattern Recognit. Artif. Intell.3
2009 An efficient algorithm for attention-driven image interpretation from segments
Hong Fu, Zheru Chi, David Dagan Feng
Pattern Recognit.3
2008 Retinex based motion estimation for sequences with brightness variations and its application to H.264
abstract
Conventional motion estimation does not take inter-frame brightness variations into consideration, which causes inefficient video coding for sequences involving brightness variations. H.264 provides a specific mode called weighted prediction targeting to improve the coding efficiency for this case. In this paper, we propose a Retinex based motion estimation scheme which effectively removes the inter-frame de-correlation factor resulting from brightness variations. We also propose to use some DCT techniques to generate the Retinex images for both current and reference images and apply conventional motion estimation and compensation procedures for coding. We applied the scheme to the H.264 testing the efficiency in the multiple reference frame motion compensation environment. Experimental results show that our proposed scheme outperforms the H.264 system with weighted prediction enabled. It allows the system to use a smaller number of reference frames for coding, e.g. 2, to achieve a similar (or slightly better) compression efficiency of the H.264 system using 5 reference frames.
Hoi-Kok Cheung, Wan-Chi Siu, David Dagan Feng, Zhiyong Wang 0001
ICASSP3
2008 Windowing technique for the DCT based retinex algorithm to handle videos with brightness variations coded using the H.264
abstract
Conventional block based motion estimators assume constant inter-frame object brightness. Pixel discrepancy is resulted primarily from motion factor without considering the influence of brightness changes. In this paper, we propose a simple and efficient windowing technique using the Hamming window and integrate it to our previously proposed algorithm. The algorithm is based on retinex approach using the DCT technique and designed to handle brightness variations. The new technique manages to greatly reduce the influence of ripple effect and further increase the compression efficiency without adding any extra overhead bits to the bit-stream. We applied the scheme to H.264 for testing. Experimental results show that the retinex based approach is an effective technique to handle inter-frame brightness variations and outperforms the H.264 system with weighted prediction enabled for sequence involving brightness variations. With our proposed windowing technique using the Hamming window, the coding efficiency can be further improved by a maximum of 0.17dB.
Hoi-Kok Cheung, Wan-Chi Siu, David Dagan Feng, Zhiyong Wang 0001
ICIP3
2008 Measuring semantic similarity between concepts in visual domain
abstract
Concept similarity has been intensively researched in the natural language processing domain due to its important role in many applications such as language modeling and information retrieval. There are few studies on measuring concept similarity in visual domain, though concept based multimedia information retrieval has attracted a lot of attentions. In this paper, we present a scalable framework for such a purpose, which is different from traditional approaches to exploring correlation among concepts in image/video annotation domain. For each concept, a model based on feature distribution is built using sample images collected from the Internet. And similarity between concepts is measured with the similarity between their models. Hereby, a Gaussian Mixture Model (GMM) is employed to model each concept and two similarity measurements are investigated. Experimental results on 13,974 images of 16 concepts collected through image search engines have demonstrated that the similarity between concepts is very close to human perception. In addition, the entropy of GMM cluster distributions can be a good indication of selecting concepts for image/video annotation.
Zhiyong Wang 0001, Genliang Guan, David Dagan Feng
MMSP4
2008 Image annotation with parametric mixture model based multi-class multi-labeling
abstract
Image annotation, which labels an image with a set of semantic terms so as to bridge the semantic gap between low level features and high level semantics in visual information retrieval, is generally posed as a classification problem. Recently, multi-label classification has been investigated for image annotation since an image presents rich contents and can be associated with multiple concepts (i.e. labels). In this paper, a parametric mixture model based multi-class multi-labeling approach is proposed to tackle image annotation. Instead of building classifiers to learn individual labels exclusively, we model images with parametric mixture models so that the mixture characteristics of labels can be simultaneously exploited in both training and annotation processes. Our proposed method has been benchmarked with several state-of-the-art methods and achieved promising results.
Zhiyong Wang 0001, Wan-Chi Siu, David Dagan Feng
MMSP3
2008 Adaptive fuzzy clustering in constructing parametric images for low SNR functional imaging
abstract
Functional imaging can provide quantitative functional parameters to aid early diagnosis. Low signal to noise ratio (SNR) in functional imaging, especially for single photon emission computed tomography, poses a challenge in generating voxel-wise parametric images due to unreliable or physiologically meaningless parameter estimates. Our aim was to systematically investigate the performance of our recently proposed adaptive fuzzy clustering (AFC) technique, which applies standard fuzzy clustering to sub-divided data. Monte Carlo simulations were performed to generate noisy dynamic SPECT data with quantitative analysis for the fitting using the general linear least square method (GLLS) and enhanced model-aided GLLS methods. The results show that AFC substantially improves computational efficiency and obtains improved reliability as standard fuzzy clustering in estimating parametric images but is prone to slight underestimation. Normalization of tissue time activity curves may lead to severe overestimation for small structures when AFC is applied.
Lingfeng Wen, Stefan Eberl, Michael J. Fulham, David Dagan Feng
MMSP4
2008 Segmentation of dual modality brain PET/CT images using the MAP-MRF model
abstract
Dual modality PET/CT has now essentially replaced PET in clinical practice and provided an opportunity to improve image segmentation through the high resolution, lower noise CT data. Thus far most research efforts have concentrated on segmentation of PET-only data. In this work we propose a systematic solution for the automated segmentation of brain PET/CT images into gray, white matter and CSF regions with the MAP-MRF model. Our approach takes advantage of the full information available from the combined scan. A PET/CT image pair and its segmentation result are modelled as a random field triplet, and segmentation is eventually achieved by solving a maximum a posteriori (MAP) problem using the expectation-maximization (EM) algorithm with simulated annealing. We compared the novel algorithm to two widely used PET-only based segmentation methods in the SPM5 toolbox and the VBM toolbox for simulation and patient data. Our results suggest that using the proposed approach substantially improves the accuracy of the delineation of brain structures.
Yong Xia 0001, Lingfeng Wen, Stefan Eberl, Michael J. Fulham, David Dagan Feng
MMSP5
2008 New Block-Based Motion Estimation for Sequences with Brightness Variation and Its Application to Static Sprite Generation for Video Compression
abstract
In this brief, a new local motion estimator is proposed which can accurately estimate motion activities under varying strong brightness conditions. The proposed estimator makes use of a new block division technique which manages practically to get rid of the adverse influence caused by brightness changes between frames. We also propose a new static sprite coding system using the proposed local motion estimator. The system is characterized not only with the features of accurate motion estimation under varying brightness conditions, but also possesses the capability of coding the brightness variability of the background scene using a single layered sprite image. Experimental results show that the resulting static sprite coding system improves the PSNR by 6.32 dB as compared with the conventional static sprite coding system when the background scenes of the video sequences involve strong brightness variations in the spatial and time domains.
Hoi-Kok Cheung, Wan-Chi Siu, David Dagan Feng, Tom Weidong Cai
IEEE Trans. Circuits Syst. Video Technol.3
2007 An efficient method for detecting ghost and left objects in surveillance video
abstract
This paper proposes an efficient method for detecting ghost and left objects in surveillance video, which, if not identified, may lead to errors or wasted computation in background modeling and object tracking in surveillance systems. This method contains two main steps: the first one is to detect stationary objects, which narrows down the evaluation targets to a very small number of foreground blobs; the second step is to discriminate the candidates between ghost and left objects. For the first step, we introduce a novel stationary object detection method based on continuous object tracking and shape matching. For the second step, we propose a fast and robust inpainting method to differentiate between ghost and left objects by constructing the real background using the candidate 's corresponding regions in the input and the background images. The effectiveness of our method has been validated by experiments over a variety of video sequences.
Sijun Lu, Jian Zhang 0002, David Dagan Feng
AVSS3
2007 Image Recovery from Broken Image Streams
abstract
This paper presents an approach for image recovery from broken image streams based on an image continuity model. The information to be recovered includes the header of an image (width, height, color space etc.) and the data stream which might have been partially damaged. We begin our discussion on common image formats and necessary information for the image recovery. We then define several measures on image contents to facilitate the image recovery. Using these measures, an effective recovery algorithm is proposed. Experimental results show that the algorithm can successfully recover images of different formats, even multi-frame images from broken image streams in most cases.
Zheru Chi, Hong Yan 0001, David Dagan Feng, Gang Chen 0006
ICIP (3)5
2007 Optimization and Implementation of H.264 Encoder on DSP Platform
abstract
Compared with MPEG-4 and other previous standards, H.264 standard has achieved great breakthrough in coding performance. In this paper, the optimization and implementation of H.264 baseline profile encoder on the TMS320DM642 has been presented. Based on the architectural features of TMS320DM642 and the computational complexity analysis of H.264 encoder, the H.264 encoder has been optimized from three aspects: algorithms, data transfer and memory/Cache use. The experimental results demonstrate that, for the video sequences with CIF format, the optimized H.264 encoder can achieve the encoding speed of more than 24 frames per second, which can meet the real-time requirements ofthe applications.
Li Zhuo 0001, David Dagan Feng, Lansun Shen
ICME3
2007 Pre-classification Module for an All-Season Image Retrieval System
abstract
From the study of attention-driven image interpretation and retrieval, we have found that an attention-driven strategy is able to extract important objects from an image and then focus the attentive objects while retrieving images. However, besides the images with distinct objects, there are images which do not show distinct objects. In this paper, the classification of "attentive" and "non-attentive" image is proposed to be a pre-process module in an all-season image retrieval system which can tackle both kinds of images. In this pre-classification module, an image is represented by an adaptive tree structure with each node carrying normalized features that characterize the object/region with visual contrasts and spatial information. Then a neural network is trained to classify an image as an "attentive" or "non-attentive" category by using the Back Propagation Through Structure (BPTS) algorithm. Experimental results indicate the reliability and feasibility of the pre-classification module, which encourages us to conduct further investigations on the all-season image retrieval system.
Hong Fu, Zheru Chi, David Dagan Feng, Weibao Zou, King Chuen Lo
IJCNN3
2007 Concept Constrained Image Region Annotation
abstract
Annotating image regions has been a challenging open issue in many areas such as image content understanding and image retrieval. In this paper, rather than solely rely on visual features of image regions, a novel approach is proposed to improve region annotation by taking concept constraints into account, since high level conceptual information such as image categories can increase the confidence of possible region labels as well as decrease the confidence of impossible region labels. We employ statistical models to learn the relationships among visual features, image concepts, and region labels. As a result, a set of possible region labels can be derived from a set of visual feature vectors of a given image so as to refine the annotation output obtained by using visual feature only. Promising experimental results have been demonstrated on 8462 regions of the University of Washington image dataset with diverse concepts for the proposed approach.
Zhiyong Wang 0001, Kelly Lam, Li Zhuo 0001, David Dagan Feng
MMSP4
2007 Fuzzy vector partition filtering technique for color image restoration
Zhonghua Ma, Hong Ren Wu, David Dagan Feng
Comput. Vis. Image Underst.3
2007 Detecting unattended packages through human activity recognition and object association
Sijun Lu, Jian Zhang 0002, David Dagan Feng
Pattern Recognit.3
2007 Image segmentation by clustering of spatial patterns
Yong Xia 0001, David Dagan Feng, Rongchun Zhao, Yanning Zhang 0001
Pattern Recognit. Lett.2
2007 Real-Time Volume Rendering Visualization of Dual-Modality PET/CT Images With Interactive Fuzzy Thresholding Segmentation
abstract
Three-dimensional (3-D) visualization has become an essential part for imaging applications, including image-guided surgery, radiotherapy planning, and computer-aided diagnosis. In the visualization of dual-modality positron emission tomography and computed tomography (PET/CT), 3-D volume rendering is often limited to rendering of a single image volume and by high computational demand. Furthermore, incorporation of segmentation in volume rendering is usually restricted to visualizing the presegmented volumes of interest. In this paper, we investigated the integration of interactive segmentation into real-time volume rendering of dual-modality PET/CT images. We present and validate a fuzzy thresholding segmentation technique based on fuzzy cluster analysis, which allows interactive and real-time optimization of the segmentation results. This technique is then incorporated into a real-time multi-volume rendering of PET/CT images. Our method allows a real-time fusion and interchangeability of segmentation volume with PET or CT volumes, as well as the usual fusion of PET/CT volumes. Volume manipulations such as window level adjustments and lookup table can be applied to individual volumes, which are then fused together in real time as adjustments are made. We demonstrate the benefit of our method in integrating segmentation with volume rendering in its application to PET/CT images. Responsive frame rates are achieved by utilizing a texture-based volume rendering algorithm and the rapid transfer capability of the high-memory bandwidth available in low-cost graphic hardware.
Jinman Kim, Tom Weidong Cai, Stefan Eberl, David Dagan Feng
IEEE Trans. Inf. Technol. Biomed.4
2007 Fast and Reliable Estimation of Multiple Parametric Images Using an Integrated Method for Dynamic SPECT
abstract
Dynamic single photon emission computed tomography (SPECT) has demonstrated the potential to quantitatively estimate physiological parameters in the brain and the heart. The generalized linear least square (GLLS) method is a well-established method for solving linear compartment models with fast computational speed. However, the high level of noise intrinsic in the SPECT data leads to reliability and instability problems of GLLS for generating parametric images. An integrated method is proposed to restrict the noise in both the temporal and spatial domains to estimate multiple parametric images for dynamic SPECT. This method comprises three steps which are optimum image sampling schedule in the projection space, cluster analysis applied postreconstruction and parametric image generation with GLLS. The simulation and experimental studies for the neuronal nicotine acetylcholine receptor tracer of 5-[123I]-iodo-A-85380 were employed to evaluate the performance of the proposed method. The results of influx rate of K1 and volume of distribution of Vd demonstrated that the integrated method was successful in generating low noise parametric images for high noise SPECT data without enhancing the partial volume effect. Furthermore, the integrated method is computationally efficient for potential clinical applications.
Lingfeng Wen, Stefan Eberl, David Dagan Feng, Tom Weidong Cai
IEEE Trans. Medical Imaging3
2006 A Knowledge-Based Approach for Detecting Unattended Packages in Surveillance Video
abstract
This paper describes a novel approach for detecting unattended packages in surveillance video. Unlike the traditional approach to just detecting stationary objects in monitored scenes, our approach detects unattended packages based on accumulated knowledge about human and non-human objects from continuous object tracking and classification. We design different reasoning rules for detecting different scenarios of the unattended package events. In the case where a package is left unattended by a single person explicitly, a rule using human activity recognition is introduced to decide the package ownership. In the case where a suspicious package is dropped down by a group of humans or under heavy occlusions, a rule based on historic tracking and classification information is proposed. Furthermore, an additional rule is given to reduce false alarms that may happen with traditional stationary object detection methods.
Sijun Lu, Jian Zhang 0002, David Dagan Feng
AVSS3
2006 Annotating Image Regions Using Spatial Context
abstract
Image annotation plays an important role in bridging the semantic gap between low level features and high level semantic contents in image access. In this paper, such a task is tackled by annotating regions which are primitives of a visual scene. We propose a probabilistic model to characterize spatial context for region annotation. Such a model provides a unifying framework integrating both feature distribution models and spatial context models. A wide range of advanced modeling techniques can be utilized to further extend this framework. The approach is also potentially scalable to a large number of semantic concepts and a large number of images. Experimental results based on simple parametric models demonstrate promising results of our approach by investigating the impacts of neighbors, segmentation, and visual features
Zhiyong Wang 0001, David Dagan Feng, Zheru Chi
ISM2
2006 Leaf Vein Extraction Using Independent Component Analysis
abstract
The purpose of this work is to develop an interactive tool which helps botanists to extract the vein system with its hierarchical properties with as little user interaction as possible. In this paper, we present a new venation extraction method using independent component analysis (ICA). The popular and efficient FastICA algorithm is applied to patches of leaf images to learn a set of linear basis functions or features for the images and then the basis functions are used as the pattern map for vein extraction. In our experiments, the training sets are randomly generated from different leaf images. Experimental results demonstrate that ICA is a promising technique for extracting leaf veins and edges of objects. ICA, therefore, can play an important role in automatically identifying living plants.
Yan Li 0002, Zheru Chi, David Dagan Feng
SMC3
2006 Attention-driven image interpretation with application to image retrieval
Hong Fu, Zheru Chi, David Dagan Feng
Pattern Recognit.3
2006 Neural-network approach for optical tomography
Xianwu Huang, David Dagan Feng
Signal Process.4
2006 Partition-based vector filtering technique for suppression of noise in digital color images
abstract
A partition-based adaptive vector filter is proposed for the restoration of corrupted digital color images. The novelty of the filter lies in its unique three-stage adaptive estimation. The local image structure is first estimated by a series of center-weighted reference filters. Then the distances between the observed central pixel and estimated references are utilized to classify the local inputs into one of preset structure partition cells. Finally, a weighted filtering operation, indexed by the partition cell, is applied to the estimated references in order to restore the central pixel value. The weighted filtering operation is optimized off-line for each partition cell to achieve the best tradeoff between noise suppression and structure preservation. Recursive filtering operation and recursive weight training are also investigated to further boost the restoration performance. The proposed filter has demonstrated satisfactory results in suppressing many distinct types of noise in natural color images. Noticeable performance gains are demonstrated over other prior-art methods in terms of standard objective measurements, the visual image quality and the computational complexity.
Zhonghua Ma, Hong Ren Wu, David Dagan Feng
IEEE Trans. Image Process.3
2006 Morphology-based multifractal estimation for texture segmentation
abstract
Multifractal analysis is becoming more and more popular in image segmentation community, in which the box-counting based multifractal dimension estimations are most commonly used. However, in spite of its computational efficiency, the regular partition scheme used by various box-counting methods intrinsically produces less accurate results. In this paper, a novel multifractal estimation algorithm based on mathematical morphology is proposed and a set of new multifractal descriptors, namely the local morphological multifractal exponents is defined to characterize the local scaling properties of textures. A series of cubic structure elements and an iterative dilation scheme are utilized so that the computational complexity of the morphological operations can be tremendously reduced. Both the proposed algorithm and the box-counting based methods have been applied to the segmentation of texture mosaics and real images. The comparison results demonstrate that the morphological multifractal estimation can differentiate texture images more effectively and provide more robust segmentations.
Yong Xia 0001, David Dagan Feng, Rongchun Zhao
IEEE Trans. Image Process.2
2006 Adaptive Segmentation of Textured Images by Using the Coupled Markov Random Field Model
abstract
Although simple and efficient, traditional feature-based texture segmentation methods usually suffer from the intrinsical less inaccuracy, which is mainly caused by the oversimplified assumption that each textured subimage used to estimate a feature is homogeneous. To solve this problem, an adaptive segmentation algorithm based on the coupled Markov random field (CMRF) model is proposed in this paper. The CMRF model has two mutually dependent components: one models the observed image to estimate features, and the other models the labeling to achieve segmentation. When calculating the feature of each pixel, the homogeneity of the subimage is ensured by using only the pixels currently labeled as the same pattern. With the acquired features, the labeling is obtained through solving a maximum a posteriori problem. In our adaptive approach, the feature set and the labeling are mutually dependent on each other, and therefore are alternately optimized by using a simulated annealing scheme. With the gradual improvement of features' accuracy, the labeling is able to locate the exact boundary of each texture pattern adaptively. The proposed algorithm is compared with a simple MRF model based method in segmentation of Brodatz texture mosaics and real scene images. The satisfying experimental results demonstrate that the proposed approach can differentiate textured images more accurately.
Yong Xia 0001, David Dagan Feng, Rongchun Zhao
IEEE Trans. Image Process.2
2006 Segmentation of VOI From Multidimensional Dynamic PET Images by Integrating Spatial and Temporal Features
abstract
Segmentation of multidimensional dynamic positron emission tomography (PET) images into volumes of interest (VOIs) exhibiting similar temporal behavior and spatial features is a challenging task due to inherently poor signal-to-noise ratio and spatial resolution. In this study, we propose VOI segmentation of dynamic PET images by utilizing both the three-dimensional (3-D) spatial and temporal domain information in a hybrid technique that integrates two independent segmentation techniques of cluster analysis and region growing. The proposed technique starts with a cluster analysis that partitions the image based on temporal similarities. The resulting temporal partitions, together with the 3-D spatial information are utilized in the region growing segmentation. The technique was evaluated with dynamic 2-[18F] fluoro-2-deoxy-D-glucose PET simulations and clinical studies of the human brain and compared with the k-means and fuzzy c-means cluster analysis segmentation methods. The quantitative evaluation with simulated images demonstrated that the proposed technique can segment the dynamic PET images into VOIs of different kinetic structures and outperforms the cluster analysis approaches with notable improvements in the smoothness of the segmented VOIs with fewer disconnected or spurious segmentation clusters. In clinical studies, the hybrid technique was only superior to the other techniques in segmenting the white matter. In the gray matter segmentation, the other technique tended to perform slightly better than the hybrid technique, but the differences did not reach significance. The hybrid technique generally formed smoother VOIs with better separation of the background. Overall, the proposed technique demonstrated potential usefulness in the diagnosis and evaluation of dynamic PET neurological imaging studies.
Jinman Kim, Tom Weidong Cai, David Dagan Feng, Stefan Eberl
IEEE Trans. Inf. Technol. Biomed.3
2006 A New Way for Multidimensional Medical Data Management: Volume of Interest (VOI)-Based Retrieval of Medical Images With Visual and Functional Features
abstract
The advances in digital medical imaging and storage in integrated databases are resulting in growing demands for efficient image retrieval and management. Content-based image retrieval (CBIR) refers to the retrieval of images from a database, using the visual features derived from the information in the image, and has become an attractive approach to managing large medical image archives. In conventional CBIR systems for medical images, images are often segmented into regions which are used to derive two-dimensional visual features for region-based queries. Although such approach has the advantage of including only relevant regions in the formulation of a query, medical images that are inherently multidimensional can potentially benefit from the multidimensional feature extraction which could open up new opportunities in visual feature extraction and retrieval. In this study, we present a volume of interest (VOI) based content-based retrieval of four-dimensional (three spatial and one temporal) dynamic PET images. By segmenting the images into VOIs consisting of functionally similar voxels (e.g., a tumor structure), multidimensional visual and functional features were extracted and used as region-based query features. A prototype VOI-based functional image retrieval system (VOI-FIRS) has been designed to demonstrate the proposed multidimensional feature extraction and retrieval. Experimental results show that the proposed system allows for the retrieval of related images that constitute similar visual and functional VOI features, and can find potential applications in medical data management, such as to aid in education, diagnosis, and statistical analysis.
Jinman Kim, Tom Weidong Cai, David Dagan Feng
IEEE Trans. Inf. Technol. Biomed.3
2005 Hybrid registration for two-dimensional gel protein images
Xiuying Wang 0001, David Dagan Feng
APBC2
2005 Classification of Moving Humans Using Eigen-Features and Support Vector Machines
Sijun Lu, Jian Zhang 0002, David Dagan Feng
CAIP3
2005 Texture analysis and retrieval using fractal signature and B-spline wavelet transform with second order derivative
abstract
In the paper, we proposed a novel over-complete B-spline wavelet transform and fractal signature for texture image analysis and retrieval. Traditionally, discrete wavelet frame took the first order derivative of smoothing function into account, which is equivalent to Canny edge detection. The second order derivative spline wavelet has the ability to detect the variation of the edge width, and the finite impulse response is conducted in the paper. Additionally, statistical features in wavelet domain and fractal signature are utilized in the retrieval. Experimental results have shown that the proposed method is reasonable to describe the essences of the textures and can reach the highest retrieval rate (76.54%) comparing with Gabor Filter and first order derivative over-complete wavelet transform.
Qing Wang 0006, David Dagan Feng
ICIP (1)2
2005 An Intelligent Middleware for Dynamic Integration of Heterogeneous Health Care Applications
abstract
In order to accelerate health care information process, to improve patient safety and quality and benefits of healthcare, and to provide comprehensive electronic patient record, we propose an intelligent middleware with two-layer knowledge-based architecture, which can integrate heterogeneous health care applications dynamically. In the first data sources integration layer, standard-based health care ontology and global API metadata are achieved, which make authorized health care applications "understand" and operate heterogeneous data sources automatically and transparently. Based on this unified information, in the second application integration layer, configurable standard-based application workflow integration facilitates the communication and interoperation between the disparate health care applications. The proposed knowledge-based architecture can optimize standard-based medical data exchange, and improves the flexibility, interoperability, maintainability, scalability and extensibility of health care information systems.
David Dagan Feng, Yu S. Lim
MMM2
2005 Trajectory Matching and Classification of Video Moving Objects
abstract
Trajectory matching is an important way to describe and classify behaviors of moving objects in a computer visual system. In this paper, we present two trajectory description methods, time-sampling sequence and space-sampling sequence, which can be used in different matching applications. We then propose two general trajectory matching schemes based on Levenshtein distance and relaxation matching respectively. Trajectory Levenshtein distance scheme is a good way to compare the topological shapes and directions of trajectories, and can be performed quickly. Trajectory relaxation matching scheme can gain the statistical optimal matching. Finally, we propose a top-to-bottom hierarchical clustering algorithm to classify trajectories, and several experiments demonstrate that our schemes are efficient in matching and classifying different shape and direction trajectories
Jiangbin Zheng 0001, David Dagan Feng, Rongchun Zhao
MMSP2
2005 Learning-based algorithm selection for image segmentation
Yong Xia 0001, David Dagan Feng, Rongchun Zhao, Maria Petrou
Pattern Recognit. Lett.2
2005 Guest Editorial Introduction to the Special Issue on Advances in Clinical and Health-Care Knowledge Management
abstract
Clinical and health-care knowledge management (KM) as a discipline has attracted increasing worldwide attention in recent years. The approach encompasses a plethora of interrelated themes including aspects of clinical informatics, clinical governance, artificial intelligence, privacy and security, data mining, genomic mining, information management, and organizational behavior. This paper introduces key manuscripts which detail health-care and clinical KM cases and applications.
Rajeev K. Bali, David Dagan Feng, Frada Burstein, Ashish N. Dwivedi
IEEE Trans. Inf. Technol. Biomed.2
2004 Molecular Imaging and Biomedical Process Modeling
David Dagan Feng
APBC1
2004 Machine learning techniques for ontology-based leaf classification
abstract
Leaf classification, indexing as well as retrieval is an important part of a computerized plant identification system. In this paper, an integrated approach for an ontology-based leaf classification system is proposed, wherein machine learning techniques play a crucial role for the automatization of the system. For the leaf contour classification, a scaled CCD code system is proposed to categorize the basic shape and margin type of a leaf by using the similar taxonomy principle adopted by the botanists. Then a trained neural network is employed to recognize the detailed tooth patterns. The measurement on an unlobed leaf is also conducted automatically according to the method used in botany. For the leaf vein recognition, the vein texture is extracted by employing an efficient combined thresholding and neural network approach so as to obtain more vein details of a leaf. Compared with the past studies, the proposed method integrates low-level features of an image and the specific knowledge in the domain (ontology) of botany, and therefore provides a more practical system for users to comprehend and handle. Primary experiments have shown promising results and proven the feasibility of the proposed system.
Hong Fu, Zheru Chi, David Dagan Feng, Jiatao Song
ICARCV3
2004 Comparison of image partition methods for adaptive image categorization based on structural image representation
abstract
Image categorization is very helpful for organizing large image databases efficiently, however, it is yet very challenging due to lack of effective image representations. Our previous work showed that structural representations were good at characterizing image contents, since image contents could be exploited from coarse to fine scales through the structures representation and fewer visual features are required. In this paper, several popular image partition methods are investigated for adaptive image categorization based on structural representation. Experimental results on seven categories of scenery images show that both the structure and node attributes are important to categorize image contents. In addition, the more similar the structures of each category, the better the categorization performance.
Zhiyong Wang 0001, David Dagan Feng, Zheru Chi
ICARCV2
2004 Automatic hybrid registration for 2-dimensional CT abdominal images
abstract
Registration of abdominal images plays an important role in clinical practice, however, because of the complex deformations of organ structure and volume, abdominal image registration remains a challenge. In this paper, a hybrid registration approach is proposed, which consists of two procedures: intensity-based registration procedure and landmark-based registration procedure. In intensity-based registration, in order to speed up registration convergence and to improve registration computation efficiency, a wavelet-based hierarchical method is proposed, in which the global displacements are corrected using mutual information algorithm. In the landmark-based registration, firstly, the landmark points are selected automatically; and then, the local non-linear deformations are corrected using the elastically thin-plate splines. By combining the advantages of intensity-based registration method and that of landmark-based method, the proposed approach can register the images accurately and efficiently. The registration performance of the proposed algorithm is validated by experiments on clinical computed tomography (CT) abdominal images.
Xiuying Wang 0001, David Dagan Feng
ICIG2
2004 Tradeoff between picture resolution and quantization precision in video coding for embedded systems
abstract
In embedded multimedia applications, improving video quality under constraints of bandwidth and storage is an important problem. In this paper, we discuss the relationship among picture resolution, quantization precision and subjective quality in video coding for embedded systems. Then we propose a principle of tradeoff between picture resolution and quantization precision. Video coding based on the tradeoff principle can achieve higher subjective quality at low bitrates, and significantly reduce the burden of decoders. Experimental results on both MPEG-2 codec and H.264 codec prove that the tradeoff principle is valuable and feasible for embedded systems.
David Dagan Feng, Yuzhuo Zhong
VCIP2
2004 A mixed scheme to improve subjective quality in low bitrate video
abstract
In wireless multimedia applications, low bitrate video takes an important part. However, its poor subjective quality often dissatisfies consumers. This paper presents a mixed scheme to improve subjective quality in low bitrate video. Our scheme is a combination of several original techniques: fast dynamic key-frame/variable frame-rate coding, ringing effect reduction, band effect reduction, and resolution/quantization tradeoff. The scheme can be applied on many popular video-coding standards, such as MPEG-1, MPEG-2, MPEG-4, and H.264. Additional computing costs involved by the scheme are acceptable. Experiments show significant improvement of subjective quality, and prove that the scheme is valuable.
David Dagan Feng, Yuzhuo Zhong
WCNC2
2004 Efficient blind image restoration using discrete periodic Radon transform
abstract
Restoring an image from its convolution with an unknown blur function is a well-known ill-posed problem in image processing. Many approaches have been proposed to solve the problem and they have shown to have good performance in identifying the blur function and restoring the original image. However, in actual implementation, various problems incurred due to the large data size and long computational time of these approaches are undesirable even with the current computing machines. In this paper, an efficient algorithm is proposed for blind image restoration based on the discrete periodic Radon transform (DPRT). With DPRT, the original two-dimensional blind image restoration problem is converted into one-dimensional ones, which greatly reduces the memory size and computational time required. Experimental results show that the resulting approach is faster in almost an order of magnitude as compared with the traditional approach, while the quality of the restored image is similar.
Daniel Pak-Kong Lun, Tommy C. L. Chan, Richard T. C. Hsung, David Dagan Feng, Yuk-Hee Chan
IEEE Trans. Image Process.4
2004 Tracer kinetic modeling of 11C-acetate applied in the liver with positron emission tomography
abstract
It is well known that 40%-50% of hepatocellular carcinoma (HCC) do not show increased 18F-fluorodeoxyglucose (FDG) uptake. Recent research studies have demonstrated that 11C-acetate may be a complementary tracer to FDG in positron emission tomography (PET) imaging of HCC in the liver. Quantitative dynamic modeling is, therefore, conducted to evaluate the kinetic characteristics of this tracer in HCC and nontumor liver tissue. A three-compartment model consisting of four parameters with dual inputs is proposed and compared with that of five parameters. Twelve regions of dynamic datasets of the liver extracted from six patients are used to test the models. Estimation of the adequacy of these models is based on Akaike Information Criteria (AIC) and Schwarz Criteria (SC) by statistical study. The forward clearance K = K1 * k3/(k2 + k3) is estimated and defined as a new parameter called the local hepatic metabolic rate-constant of acetate (LHMRAct) using both the weighted nonlinear least squares (NLS) and the linear Patlak methods. Preliminary results show that the LHMRAct of the HCC is significantly higher than that of the nontumor liver tissue. These model parameters provide quantitative evidence and understanding on the kinetic basis of C-acetate for its potential role in the imaging of HCC using PET.
Sirong Chen, Chilai Ho, David Dagan Feng, Zheru Chi
IEEE Trans. Medical Imaging3
2003 Novel Elastic Registration for 2-D Medical and Gel Protein Images
Xiuying Wang 0001, David Dagan Feng, Hai Hong
APBC2
2003 A new automatic detection approach for hepatocellular carcinoma using C-acetate positron emission tomography
abstract
Functional imaging techniques such as positron emission tomography (PET) has the potential for early diagnosis of malignant tumors. However, 40-50% of hepatocellular carcinoma (HCC), a common malignancy worldwide, can hardly be detected by the widely used F-2-fluoro-2-deoxy-D-glucose (FDG) PET. C-acetate PET has recently been found effective for detecting HCC. To perform quantitative analysis to obtain the diagnosis information, regions of interest (ROls) are needed to be extracted. Manual placement of ROIs is subject to operator's skill and time-consuming. Furthermore, the small sizes of some ROIs make the task even more difficult. In this paper, we propose an approach to segment the dynamic C-acetate PET liver images automatically. The curves extracted from some segmented ROIs are then fitted to the presented C-acetate liver model. Finally, the parameter K, which has been validated as an indicator for detecting HCC, can be calculated.
Sirong Chen, Longkin Wong, David Dagan Feng
ICIP (1)3
2002 Fuzzy integral for leaf image retrieval
abstract
Generally, the more features utilized, the better the retrieval performance. However, it is a very challenging task to combine different feature sets in a way reflecting human perception. This paper presents the combination of different shape based feature sets using fuzzy integral for leaf image retrieval. The feature sets used in our system include centroid-contour distance curve, eccentricity, and angle code histogram. The fuzzy integral approach can release the user's burden from tuning the combination parameters. In order to reduce the matching time in the retrieval process, a thinning based method is proposed to locate the start point of a leaf contour. Experimental results on 440 leaf images from 44 plant species (10 samples from each plant species) show that the fuzzy integral approach can achieve a comparable retrieval performance with the best case of the weighted summation combination. The results also indicate that our approach, which are more efficient, can achieve a better retrieval performance than both the curvature scale space (CSS) method and the modified Fourier descriptor (MFD) method.
Zhiyong Wang 0001, Zheru Chi, David Dagan Feng
FUZZ-IEEE3
2002 A comparative study on the coherent approaches to cooperation between TCP and ATM congestion control algorithms
abstract
Numerous studies have indicated that ATM available bit rate (ABR) service can provide low-delay, fairness, and high throughput, and can handle congestion effectively inside the ATM network. However, network congestion is not really eliminated but rather it is pushed out to the edge of the ATM network, packets from TCP sources competing for the available ATM bandwidth are buffered in the routers or switches at the network edges, causing severe congestion, degraded throughput, and unfairness. This poor performance is mainly due to the uncoordinated interaction between the congestion control mechanism of TCP and ATM. It is well accepted that some form of cooperation at edge device would help to control TCP traffic flow over ATM more effectively. We have previously proposed the fair intelligent explicit window adaptation (FIEWA) scheme and fair intelligent ACK bucket control (FIABC) scheme. The key idea is to combine the feedback information from the receiver, from the underlying ATM network, and from the local information at the edge device intelligently to explicitly/implicitly control the TCP rate. We present a comparative simulation study on our schemes with other established schemes; to identify the characteristics of each different scheme; and to indicate the requirement for a fairer, simpler and more robust coherent approach at the edge device.
Doan B. Hoang, David Dagan Feng
ICCCN3
2002 A fast 2D entropic thresholding method by wavelet decomposition
abstract
Compared with ID grayscale histogram analysis, 2D entropic thresholding makes use of local average as well as pixel gray level. However, it is time consuming to search the threshold vector in the 2D histogram. In the paper, a fast algorithm using wavelet decomposition is proposed, with which a set of candidates of the vector was first obtained in the decomposed histogram. The optimal threshold vector is then obtained without exhaustive searching. Experimental results have shown that our algorithm not only finds the threshold vector as well as Brink's method (1992) but also saves computation costs, using up only 0.53% of the processing time taken by exhaustive searching.
Qing Wang 0006, Qiurang Wang, David Dagan Feng, Rongchun Zhao, Zheru Chi
ICIP (3)3
2002 Content access and distribution of multimedia medical data in E-health
abstract
E-health is greatly impacting on information distribution and availability within the health services, hospitals and to the public. Previous research has addressed the development of system architectures with the aim of integrating the distributed and heterogeneous medical information systems. Easing the difficulties in the sharing and management of multimedia medical data and the timely accessibility to these data are critical needs for health care providers. We have proposed a client-server agent that integrates and allows a portal to every permitted information system of the hospital that consists of picture archiving and communication systems (PACS), radiology information system (RIS) and hospital information system (HIS) via the intranet and the Internet. Our proposed agent enables remote access into the usually closed information system of the hospital and a server that manages all the multimedia medical data and allows for in-depth and complex search queries for content access and automatic creation of patient reports for distribution.
Jinman Kim, David Dagan Feng, Tom Weidong Cai, Stefan Eberl
ICME (2)2
2002 Robust and efficient content-based digital audio watermarking
Changsheng Xu, David Dagan Feng
Multim. Syst.2
2002 Dynamic image data compression in the spatial and temporal domains: clinical issues and assessment
abstract
In our previous work, we developed a novel approach to dynamic image data compression, and demonstrated that very high compression ratios can be achieved while preserving relevant kinetic information. However, the technique has not yet been assessed with clinical data. Many issues need to be addressed to tailor the method for clinical use. In this paper, we apply the compression technique to dynamic [18F] 2-fluoro-deoxy-glucose (FDG) brain positron emission tomography (PET) data, using a five-parameter model to include cerebral blood volume (CBV) and partial volume (PV) effects. Functional images generated from the compressed data are compared with those from the original uncompressed data. We show that the storage requirements for a typical clinical dynamic PET image data set can be reduced by more than 95%, without degradation of image quality. Furthermore, the technique greatly reduces the computational complexity of further clinical image postprocessing such as smoothing and generation of functional images. It is expected that the compression technique will be of benefit in image data management and telemedicine.
David Dagan Feng, Tom Weidong Cai, Roger R. Fulton
IEEE Trans. Inf. Technol. Biomed.1
2001 A novel QoS feedback control for supporting compressed video
abstract
This paper proposes a novel application layer quality of service (QoS) feedback control scheme for supporting compressed video. In this proposal, a QoS packet is sent by the source after each video frame to transfer QoS information. The source employs a simple additive-increase and explicit-decrease bandwidth algorithm to adjust its rate based on the QoS feedback parameters sent from the destination. The simulation results show that the proposed scheme is able to control the jitter, delay and loss very efficiently.
Xiaomei Yu, Doan B. Hoang, David Dagan Feng
GLOBECOM3
2001 Match Between Normalization Schemes and Feature Sets for Handwritten Chinese Character Recognition
abstract
Because of the large number of Chinese characters and many different writing styles involved, the recognition of handwritten Chinese characters remains a very challenging task. It is well recognized that a good feature set plays a key role in a successful recognition system. Shape normalization is as well an essential step toward achieving translation, scale, and rotation invariance in recognition. Many shape normalization methods and different feature sets have been proposed in the literature. We first review five commonly used shape normalization schemes and then discuss various feature extraction techniques usually used in handwritten Chinese character recognition. Based on numerous experiments conducted on 3,755 handwritten Chinese characters (GB2312-80), we discuss the matches made between the normalization schemes and the feature sets and suggest the best match between them in terms of classification performance. The nearest neighbor classifier was adopted in our experiments with templates obtained by using the K-means clustering algorithm.
Qing Wang 0006, Zheru Chi, David Dagan Feng, Rongchun Zhao
ICDAR3
2001 A Robust And Fast Watermarking Scheme For Compressed Audio
abstract
This paper proposes a method to embed and extract the watermark into and from digital compressed audio. The watermark is embedded in partially uncompressed domain and the embedding scheme is high related to audio content. The watermark embedding can be done very fast. The experimental result illustrates that the embedded watermark can survive the decoding and re-encoding process.
Changsheng Xu, Yongwei Zhu, David Dagan Feng
ICME3
2001 Digital audio watermarking based-on multiple-bit hopping and human auditory system
abstract
A novel content-adaptive audio watermarking technique is proposed. To optimally balance in-audibility and robustness when embedding and extracting watermarks, the embedding scheme is high related to audio content by making use of the properties of human auditory system and multiple-bit hopping technique. The experimental results in robustness are provided to support all the novel features in our watermarking scheme.
Changsheng Xu, Yongwei Zhu, David Dagan Feng
ACM Multimedia3
2001 Simultaneous estimation of physiological parameters and the input function - in vivo PET data
abstract
Dynamic imaging with positron emission tomography (PET) is widely used for the in vivo measurement of regional cerebral metabolic rate for glucose (rCMRGlc) with [18F]fluorodeoxy-D-glucose (FDG) and is used for the clinical evaluation of neurological disease. However, in addition to the acquisition of dynamic images, continuous arterial blood sampling is the conventional method to obtain the tracer time-activity curve in blood (or plasma) for the numeric estimation of rCMRGlc in mg glucose/100-g tissue/min. The insertion of arterial lines and the subsequent collection and processing of multiple blood samples are impractical for clinical PET studies because it is invasive, has the remote, but real potential for producing limb ischemia, and it exposes personnel to additional radiation and risks associated with handling blood. In this paper, based on our previously proposed method for extracting kinetic parameters from dynamic PET images, we developed a modified version (post-estimation method) to improve the numerical identifiability of the parameter estimates when we deal with data obtained from clinical studies. We applied both methods to dynamic neurologic FDG PET studies in three adults. We found that the input function and parameter estimates obtained with our noninvasive methods agreed well with those estimated from the gold standard method of arterial blood sampling and that rCMRGlc estimates were highly correlated (r = 0.973). More importantly, no significant difference was found between rCMRGlc estimated by our methods and the gold standard method (P > 0.16). We suggest that our proposed noninvasive methods may offer an advance over existing methods.
Koon-Pong Wong, David Dagan Feng, Steven R. Meikle, Michael J. Fulham
IEEE Trans. Inf. Technol. Biomed.2
2000 Visualization of Biomedical Processes: Local Quantitative Physiological Functions in Living Human Body
abstract
Functional imaging with dynamic positron emission tomography (PET) has been playing a crucial and expanding role in biomedical research and clinical diagnosis, providing image-wide quantitative and qualitative physiological functions in the human body, and supporting visualization of the distribution of these functions corresponding to anatomical structures. A number of parametric imaging algorithms have been developed. We give a brief study on some existing and our recently, developed techniques for generating parametric images. An integrated system for functional image data processing and visualization, and a Web-based application are presented.
David Dagan Feng, Tom Weidong Cai
Computer Graphics International1
2000 Document Image Matching Based on Component Blocks
abstract
Document image matching is the key technique for document registration and retrieval. In this paper, a new matching algorithm based on document component block list and component block tree is proposed. Our method can effectively make use of the local information of each page block and the global information of page layout, while it is also robust to image distortion, filled-in text, and noises. This algorithm is then refined and applied to automatic data extraction of column forms. A demonstrating software package has been developed.
Hanchuan Peng, Fuhui Long, Wan-Chi Siu, Zheru Chi, David Dagan Feng
ICIP5
2000 Multimodal Interface Techniques in Content-Based Multimedia Retrieval
Jinchang Ren, Rongchun Zhao, David Dagan Feng, Wan-Chi Siu
ICMI3
2000 Hidden Markov Random Field Based Approach for Off-Line Handwritten Chinese Character Recognition
abstract
This paper presents a hidden Markov mesh random field (HMMRF) based approach for off-line handwritten Chinese characters recognition using statistical observation sequences embedded in the strokes of a character. Due to a large set of Chinese characters and many different writing styles, the recognition of handwritten Chinese characters is very challenging. In our approach, the binary image is first normalized by a nonlinear shape normalization scheme to adjust the width, length, and the correlation of strokes. Two types of stroke-based features are then extracted to represent the observation sequence. The estimation of model parameters and state sequence decoding algorithms are also discussed in the paper. Experimental results on 470 isolated handwritten Chinese characters demonstrate the effectiveness of our approach.
Qing Wang 0006, Rongchun Zhao, Zheru Chi, David Dagan Feng
ICPR4
2000 Fair intelligent bandwidth allocation for rate-adaptive video traffic
Doan B. Hoang, Xiaomei Yu, David Dagan Feng
Comput. Commun.3
2000 Content-based retrieval of dynamic PET functional images
abstract
The recent information explosion has led to massively increased demand for multimedia data storage in integrated database systems. Content-based retrieval is an important alternative and complement to traditional keyword-based searching for multimedia data and can greatly enhance information management. However, current content-based image retrieval techniques have some deficiencies when applied in the biomedical functional imaging domain. In this paper, we presented a prototype design for a content-based functional image retrieval database system for dynamic positron emission tomography. The system supports efficient content-based retrieval based on physiological kinetic features and reduces image storage requirements. This design makes it possible to maintain a large number of patient data sets online and to rapidly retrieve dynamic functional image sequences for interpretation and generation of physiological parametric images, and offers potential advantages in medical image data management and telemedicine, as well as providing possible opportunities in the statistical and comparative analysis of functional image data.
Tom Weidong Cai, David Dagan Feng, Roger R. Fulton
IEEE Trans. Inf. Technol. Biomed.2
2000 Guest editorial: multimedia information technology in biomedicine
David Dagan Feng
IEEE Trans. Inf. Technol. Biomed.1
2000 Optimal Image Sampling Schedule for Both Image-derived Input and Output Functions in PET Cardiac Studies
abstract
Optimal sampling schedule (OSS) design for both image-derived input and output functions in tracer kinetic modeling with positron emission tomography (PET) is investigated. This problem is very important in noninvasive PET dynamic cardiac studies where both the input function, i.e., the plasma time-activity curve (PTAC), and the output function, i.e., the tissue time-activity curve (TTAC), are obtained simultaneously from the same sequence of PET images. The integral PET measurement is used in this study. The spillover correction for the cross contaminations in cardiac studies is incorporated into the OSS design procedure. A new target function based on the D-optimal criterion involving both the input and output sensitivity functions is proposed. The fluorodeoxyglucose (FDG) model and a six-parameter PTAC model are used to illustrate the simultaneous OSS design for both the PTAC and TTAC. An OSS design consisting of six different scanning intervals is derived. Computer simulations are performed based on the estimated parameters from real studies to evaluate the effectiveness of the OSS. The double modeling approach is used in parameter estimation to simultaneously estimate the parameters involved. The results have shown that, for a wide range of parameter variations, the OSS is as effective as a conventional sampling schedule (CSS) and comparable parameter estimates can be obtained. Compared with the use of the CSS, the use of the OSS leads to an approximately 70% reduction in the storage space and data processing time.
Xianjin Li, David Dagan Feng, Kewei Chen 0001
IEEE Trans. Medical Imaging2
1999 Embedded Singularity Detection Zerotree Wavelet Coding
abstract
We explore the wavelet coefficient selection and denoising by singularity detection (SD) for the embedded zero-tree wavelet (EZW) coding algorithm in this paper. The EZW coding algorithm exploits the relation between the multi-scale wavelet coefficients that finer scale wavelet coefficients are probably to vanish if the coarse scale wavelet coefficient vanishes. It is true for the parts of image that are regular but not the cases for noise-like features. In other words, the performance of the coding algorithm may be greatly degraded for the latter. In this paper, we investigate to arrange the wavelet coefficients according to the local regularity, by using the computed wavelet coefficients from the encoding filters. The advantage is two folds. For normal coding, we can make the encoded bit-stream first appear with the wavelet coefficients that correspond to the most regular part of the image, and the irregular one's follows. For noisy image encoding, we can remove noises before encoding hence increase the image quality as well as the coding efficiency.
Daniel Pak-Kong Lun, Richard T. C. Hsung, David Dagan Feng, Tommy C. L. Chan
ICIP (2)3
1999 Information technology applications in biomedical functional imaging
abstract
In parallel with rapid advances in computer technology, biomedical functional imaging is having an ever-increasing impact on healthcare. Functional imaging allows us to see dynamic processes quantitatively in the living human body. However, as we need to deal with four-dimensional time-varying images, space requirements and computational complexity are extremely high. This makes information management, processing, and communication difficult. Using the minimum amount of data to represent the required information, developing fast algorithms to process the data, organizing the data in such a way as to facilitate information management, and extracting the maximum amount of useful information from the recorded data have become important research tasks in biomedical information technology. For the last ten years, the Biomedical and Multimedia Information Technology (BMIT) Group and, recently, the Center for Multimedia Signal Processing have conducted systematic studies on these topics. Some of the results relating to functional imaging data acquisition, compression, storage, management, processing, modeling, and simulation are briefly reported in this paper.
David Dagan Feng
IEEE Trans. Inf. Technol. Biomed.1
1998 Non-invasive quantification of physiological processes with dynamic PET using blind deconvolution
abstract
Dynamic positron emission tomography (PET) has opened the possibility of quantifying physiological processes within the human body. On performing dynamic PET studies, the tracer concentration in blood plasma has to be measured, and acts as the input function for tracer kinetic modelling. In this paper, we propose an approach to estimate physiological parameters for dynamic PET studies without the need of taking blood samples. The proposed approach comprises two major steps. First, a wavelet denoising technique is used to filter the noise appeared in the projections. The denoised projections are then used to reconstruct the dynamic images using filtered backprojection. Second, an eigen-vector based blind deconvolution technique is applied to the reconstructed dynamic images to estimate the physiological parameters. To demonstrate the performance of the proposed approach, we carried out a Monte Carlo simulation using the fluoro-deoxy-2-glucose model, as applied to tomographic studies of human brain. The results demonstrate that the proposed approach can estimate the physiological parameters with an accuracy comparable to that of invasive approach which requires the tracer concentration in plasma to be measured.
Chi-Hoi Lau, Daniel Pak-Kong Lun, David Dagan Feng
ICASSP3
1998 A Signature for Content-Based Image Retrieval Using a Geometrical Transform
abstract
Article A signature for content-based image retrieval using a geometrical transform Share on Authors: H. Wang Department of Computer Science, The University of Sydney, NSW 2006, Australia Department of Computer Science, The University of Sydney, NSW 2006, AustraliaView Profile , F. Guo Department of Computer Science, The University of Sydney, NSW 2006, Australia Department of Computer Science, The University of Sydney, NSW 2006, AustraliaView Profile , D. D. Feng Department of Computer Science, The University of Sydney, NSW 2006, Australia and DoEE, Hong Kong Polytechnic University, Hong Kong Department of Computer Science, The University of Sydney, NSW 2006, Australia and DoEE, Hong Kong Polytechnic University, Hong KongView Profile , J. S. Jin School of Computer Science and Engineering, University of New South Wales, Sydney 2052, Australia School of Computer Science and Engineering, University of New South Wales, Sydney 2052, AustraliaView Profile Authors Info & Claims MULTIMEDIA '98: Proceedings of the sixth ACM international conference on MultimediaSeptember 1998 Pages 229–234https://doi.org/10.1145/290747.290775Online:01 September 1998Publication History 13citation675DownloadsMetricsTotal Citations13Total Downloads675Last 12 Months3Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
F. Guo, David Dagan Feng, Jesse S. Jin
ACM Multimedia3
1998 Generalized Linear Least Square Method for Fast Generation of Myocardial Blood Flow Parametrick Images with N-13 Ammonia PET
abstract
In this paper, we developed and tested strategies for estimating myocardial blood flow (MBF) and generating MBF parametric images using positron emission tomography (PET), N-13 ammonia, and the generalized linear least square (GLLS) method. GLLS was generalized to the general linear compartment model, modified for the correction of spillover, validated using simulated N-13 ammonia data, and examined using PET data from several patient studies. In comparison to the standard model-fitting procedure, the GLLS method provided similar accuracy and superior computational speed.
Kewei Chen 0001, Michael Lawson, Eric Reiman, Alan Cooper, David Dagan Feng, Sung-Cheng Huang, Daniel Bandy, Dino Ho, Lang-sheng Yun, Anita Palant
IEEE Trans. Medical Imaging5
1998 Minimum Dynamic SPECT Image Acquisition Time Required for T1-201 Tracer Kinetic Modelling
Chi-Hoi Lau, Stefan Eberl, David Dagan Feng, Hidehiro Iida, Daniel Pak-Kong Lun, Wan-Chi Siu, Yoshikazu Tamura, George J. Bautovich, Yukihiko Ono
IEEE Trans. Medical Imaging3
1998 Dynamic Imaging and Tracer Kinetic Modeling for Emission Tomography Using Rotating Detectors
abstract
When performing dynamic studies using emission tomography the tracer distribution changes during acquisition of a single set of projections. This is particularly true for some positron emission tomography (PET) systems which, like single photon emission computed tomography (SPECT), acquire data over a limited angle at any time, with full projections obtained by rotation of the detectors. In this paper, an approach is proposed for processing data from these systems, applicable to either PET or SPECT. A method of interpolation, based on overlapped parabolas, is used to obtain an estimate of the total counts in each pixel of the projections for each required frame-interval, which is the total time to acquire a single complete set of projections necessary for reconstruction. The resultant projections are reconstructed using traditional filtered backprojection (FBP) and tracer kinetic parameters are estimated using a method which relies on counts integrated over the frame-interval rather than instantaneous values. Simulated data were used to illustrate the technique's capabilities with noise levels typical of those encountered in either PET or SPECT. Dynamic datasets were constructed, based on kinetic parameters for fluoro-deoxy-glucose (FDG) and use of either a full ring detector or rotating detector acquisition. For the rotating detector, use of the interpolation scheme provided reconstructed dynamic images with reduced artefacts compared to unprocessed data or use of linear interpolation. Estimates for the metabolic rate of glucose had similar bias to those obtained from a full ring detector.
Chi-Hoi Lau, David Dagan Feng, Brian F. Hutton, Daniel Pak-Kong Lun, Wan-Chi Siu
IEEE Trans. Medical Imaging2
1997 A technique for extracting physiological parameters and the required input function simultaneously from PET image measurements: theory and simulation study
abstract
Positron emission tomography (PET) is an important tool for enabling quantification of human brain function. However, quantitative studies using tracer kinetic modeling require the measurement of the tracer time-activity curve in plasma (PTAC) as the model input function. It is widely believed that the insertion of arterial lines and the subsequent collection and processing of the biomedical signal sampled from the arterial blood are not compatible with the practice of clinical PET, as it is invasive and exposes personnel to the risks associated with the handling of patient blood and radiation dose. Therefore, it is of interest to develop practical noninvasive measurement techniques for tracer kinetic modeling with PET. In this paper, a technique is proposed to extract the input function together with the physiological parameters from the brain dynamic images alone. The identifiability of this method is tested rigorously by using Monte Carlo simulation. The results show that the proposed method is able to quantify all the required parameters by using the information obtained from two or more regions of interest (ROI's) with very different dynamics in the PET dynamic images. There is no significant improvement in parameter estimation for the local cerebral metabolic rate of glucose (LCMRGlc) if the number of ROI's are more than three. The proposed method can provide very reliable estimation of LCMRGlc, which is our primary interest in this study.
David Dagan Feng, Koon-Pong Wong, Chi-Ming Wu, Wan-Chi Siu
IEEE Trans. Inf. Technol. Biomed.1
1997 Dynamic image data compression in spatial and temporal domains: theory and algorithm
abstract
Advanced medical imaging requires storage of large quantities of digitized clinical data. These data must be stored in such a way that their retrieval does not impair the clinician's ability to make a diagnosis. In this paper, we propose the theory and algorithm for near (or diagnostically) lossless dynamic image data compression. Taking advantage of domain-specific knowledge related to medical imaging, the medical practice and the dynamic imaging modality, a compression ratio greater than 80:1 is achieved. The high compression ratios are achieved by the proposed compression algorithm through three stages: 1) addressing temporal redundancies in the data through application of image optimal sampling, 2) addressing spatial redundancies in the data through cluster analysis, and 3) efficient coding of image data using standard still-image compression techniques. To illustrate the practicality of the proposed compression algorithm, a simulated positron emission tomography (PET) study using the fluoro-deoxy-glucose (FDG) tracer is presented. Realistic dynamic image data are generated by "virtual scanning" of a simulated brain phantom as a real PET scanner. These data are processed using the conventional [8] and proposed algorithms as well as the techniques for storage and analysis. The resulting parametric images obtained from the conventional and proposed approaches are subsequently compared to evaluate the proposed compression algorithm. As a result of this study, storage space for dynamic image data is able to be reduced by more than 95%, without loss in diagnostic quality. Therefore, the proposed theory and algorithm are expected to be very useful in medical image database management and telecommunication.
Dino Ho, David Dagan Feng, Kewei Chen 0001
IEEE Trans. Inf. Technol. Biomed.2
1996 An unbiased parametric imaging algorithm for nonuniformly sampled biomedical system parameter estimation
abstract
An unbiased algorithm of generalized linear least squares (GLLS) for parameter estimation of nonuniformly sampled biomedical systems is proposed. The basic theory and detailed derivation of the algorithm are given. This algorithm removes the initial values required and computational burden of nonlinear least regression and achieves a comparable estimation quality in terms of the estimates' bias and standard deviation. Therefore, this algorithm is particular useful in image-wide (pixel-by-pixel based) parameter estimation, e.g., to generate parametric images from tracer dynamic studies with positron emission tomography. An example is presented to demonstrate the performance of this new technique. This algorithm is also generally applicable to other continuous system parameter estimation.
David Dagan Feng, Sung-Cheng Huang, Zhizhong Wang, Dino Ho
IEEE Trans. Medical Imaging1
1996 Optimal image sampling schedule: a new effective way to reduce dynamic image storage space and functional image processing time
abstract
An optimal image sampling schedule for tracer dynamic studies with positron emission tomography (PET) is proposed. This schedule incorporates the characteristics of PET measurement and uses a new cost function and the D-optimal criterion. A detailed case study of the estimation of the local cerebral metabolic rate of glucose (LCMRGLc) using the tracer fluorodeoxyglucose (FDG) and the four-parameter FDG model is presented. As the sampling schedule designed requires only four dynamic images, the storage space and data processing time are greatly reduced, while the precision of the parameter estimates is almost the same as that achieved with a commonly used schedule. The effects of intersubject and intrasubject parameter variations on parameter estimation with the use of this optimal sampling schedule are investigated by computer simulation. The simulation results show that the estimation of parameters is sufficiently robust with respect to these intersubject and intrasubject variations. The optimal sampling schedule is quite suitable therefore for PET regional parameter estimation, as well as for image-wide parameter estimation, for different subjects.
Xianjin Li, David Dagan Feng, Kewei Chen 0001
IEEE Trans. Medical Imaging2
1995 An evaluation of the algorithms for determining local cerebral metabolic rates of glucose using positron emission tomography dynamic data
abstract
Measurement of the local cerebral metabolic rate of glucose (LCMRGlc) and the individual rate constant parameters of the [(18 )F]2-fluoro-2-deoxy-D-glucose (FDG) model can provide a clearer understanding and insight to the physiological processes in the human brain, and a quicker and more accurate means of diagnosis in clinical applications. A systematic study using simulated and clinical tissue time activity data is presented to evaluate several existing and newly developed major algorithms used for determining LCMRGlc and the individual rate constants from positron emission tomography dynamic data. The computational and statistical properties of the autoradiographic approach, weighted and unweighted nonlinear least squares methods, Patlak graphic approach, weighted integration method, linear least squares and generalized linear least squares methods are investigated and discussed in this paper.
David Dagan Feng, Dino Ho, Kewei Chen 0001, Liang-Chih Wu, Jiunn-Kuen Wang, Ren-Shyan Liu, Shin-Hwa Yeh
IEEE Trans. Medical Imaging1
1993 A study on statistically reliable and computationally efficient algorithms for generating local cerebral blood flow parametric images with positron emission tomography
abstract
With the advent of positron emission tomography (PET), a variety of techniques have been developed to measure local cerebral blood flow (LCBF) noninvasively in humans. A potential class of techniques, which includes linear least squares (LS), linear weighted least squares (WLS), linear generalized least squares (GLS), and linear generalized weighted least squares (GWLS), is proposed. The statistical characteristics of these methods are examined by computer simulation. The authors present a comparison of these four methods with two other rapid estimation techniques developed by Huang et al. (1982) and Alpert (1984), and two classical methods, the unweighted and weighted nonlinear least squares regression. The results show that these methods can take full advantage of the contribution from the fine temporal sampling data of modern tomographs, and thus provide statistically reliable estimates that are comparable to those obtained from nonlinear LS regression. These methods also have high computational efficiency, and the parameters can be estimated directly from operational equations in one single step. Therefore, they can potentially be used in image-wide estimation of local cerebral blood flow and distribution volume with PET.
David Dagan Feng, Zhizhong Wang, Sung-Cheng Huang
IEEE Trans. Medical Imaging1