VLDB 2026 Research / reviewers in the wild / expert
Ling Guan
dblp:66/4324
· DBLP profile ↗
262ranked-venue papers
8as first author
28since 2021 · last 2025
0000-0002-2681-2504ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 189 · 5 first-author · 17 since 2021Artificial intelligence and machine learning · 38 · 2 first-author · 5 since 2021Systems, architecture and hardware · 19Computer networks · 9 · 3 since 2021Databases, data management, data science and information retrieval · 7 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Software engineering, systems software and programming languages · 2Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CCAFF: Object Tracking Under Heavy OcclusionabstractDeep features have become a standard in many object-tracking frameworks, replacing traditional handcrafted methods for object representation. However, recent studies have shown that deep features do not outperform handcrafted features in matching under occlusion or re-identification. Many trackers are trained using standard benchmarks under ideal conditions, but they degrade significantly in real-world settings because features deteriorate over time. These features affect the similarity distance and can lead to identity switches. Thus, we propose two novel tracking approaches using only handcrafted features and an extended variant, Contextual Cross Attention Feature Fusion (CCAFF). Both methods use a class-level codebook to capture keypoint cues for feature representation. We evaluate identity preservation for both feature sets using objects under severe occlusion. The CCAFF feature embedding demonstrates an improvement across all metrics on the HOOT dataset with a 4.63% increase in IDF1 score while decreasing identity switches by 7.78% compared to the baseline model. Abdul Bhutta, Naimul Mefraz Khan, Ling Guan |
ISM | 3 |
| 2025 | An Efficient Optimization Criterion for Multi-View Feature Representation LearningabstractThe training of contemporary machine learning (ML) models, particularly deep neural networks (DNNs), often relies on enormous data sources to properly tune model parameters. As a result, achieving competitive results with limited training data and computational resources has been recognized as a significant bottleneck to advance ML. To address these issues, multi-view representation learning has emerged. However, how to efficiently build multi-view learning models remains a big challenge. In this paper, a novel optimization criterion is proposed to tackle this challenge. Specifically, the proposed criterion ensures speedy and effective parameter selection, reducing the effort to reach optimal design of the model while maintaining performance. To validate the efficiency and generalizability of the presented solution, experiments were conducted on face recognition and few-shot learning for image classification using four databases of different scales. Experimental results demonstrate the superiority of the proposed approach, offering an efficient yet robust solution to the data-scarcity challenge. Lei Gao 0001, Kai Liu 0032, Kevin Tang, Ling Guan |
ISM | 4 |
| 2025 | A discriminative multi-modal adaptation neural network model for video action recognition
Lei Gao 0001, Kai Liu 0032, Ling Guan |
Neural Networks | 3 |
| 2025 | Semantic-Guided Flow Matching for Fast and Accurate Remote Sensing Image Super-ResolutionabstractCurrent diffusion-based super-resolution methods for remote sensing images face two critical limitations: (1) excessive computational demands due to iterative sampling; and (2) semantic inconsistency in reconstructed images caused by stochastic denoising. To address these challenges, we propose SfmSR, a Semantics-guided Flow Matching model that achieves fast (1-5 steps) and accurate reconstruction through a novel two-stage framework. Stage 1 performs large-scale pre-training on diverse remote sensing data to learn robust multi-scale representations, while Stage 2 introduces semantic-guided fine-tuning, where extracted high-level features dynamically regulate the flow matching process. This dual-phase approach reduces sampling steps by two orders of magnitude while maintaining stability through deterministic ODE-based generation. Experimental results demonstrate that SfmSR outperforms existing methods on remote sensing image datasets such as Potsdam and Toronto. Moreover, SfmSR exhibits significant advantages in model complexity and inference efficiency, achieving fast inference with fewer parameters and lower memory usage, thus meeting real-time requirements in practical applications. Compared to state-of-the-art (SOTA) methods, SfmSR not only achieves superior reconstruction quality but also demonstrates clear advantages in inference speed and computational efficiency. Zhicheng Gong, Fangzhou Yi, Ling Guan, Chunzhu Dong, Hui Zeng 0003 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | ODMTCNet: An Interpretable Multiview Deep Neural Network Architecture for Feature RepresentationabstractRecently, deep cascade architecture-based algorithms have attracted wide attention and have been applied to numerous application domains successfully. Nevertheless, the black-box structure of such algorithms has always been considered the Achilles' heel by the machine learning community. Moreover, due to its data-driven nature, the deep cascade architecture likely causes over-fitting problems when there is no sufficient data available. In order to solve these pressing issues, this work proposes a novel multiview deep neural network (DNN) model, namely, optimal discriminant multiview tensor convolutional network (ODMTCNet), which integrates statistics-guided optimization (SGO) principles with the DNN architecture. Specifically, a discriminant multiview tensor convolution strategy is proposed and integrated with a deep cascade architecture. Different from the traditional DNN models, the parameters of the convolutional layers in ODMTCNet are determined by solving SGO problems. Based on the SGO principles, the relation between the optimal performance and parameters (e.g., the number of convolutional filters) can be analytically predicted, with each layer generating justified knowledge representations. In addition, information quality (IQ) is adopted to further improve multiview feature representation. Because of its unique design, ODMTCNet is able to handle different types of features (e.g., raw, hand-crafted, prior knowledge-based, and DNN-generated features), forming a general platform for multiview feature representation. To validate the genericness and effectiveness of the ODMTCNet model, we conducted experiments on five datasets of different scales: The Olivetti Research Lab (ORL) database, the Facial Recognition Technology (FERET) database, the ETH-80 database, the Caltech 256 database, and the nanyang technological university (NTU) red green blue-depth (RGB+D) 120 database. Experimental results show the superiority of the presented solution over the state-of-the-art. Implementation codes will be made available in the final version. Lei Gao 0001, Ling Guan |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Mathematics-Inspired Models: A Green and Interpretable Learning Paradigm for Multimedia ComputingabstractThe advances of machine learning (ML), and AI in general, have attracted unprecedented attention in intelligent multimedia computing and many other fields. However, due to the concern for sustainability and black-box nature of ML models, especially deep neural networks (DNNs), green and interpretable learnings have been extensively studied in recent years, despite suspicions on effectiveness, subjectivity of interpretability, and complexity. To address these concerns and suspicions, this article starts with a survey on recent discoveries in green learning and interpretable learning and then presents mathematics-inspired (M-I) learning models. We will demonstrate that the M-I models are green in nature with numerous interpretable properties. Finally, we present several examples in multi-view information computing on both static image-based and dynamic video-based tasks to demonstrate that the M-I methodology promises a plausible and sustainable path for natural evolution of ML, which is worth further investment in. Lei Gao 0001, Kai Liu 0032, Ling Guan |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | Edge-labeling based modified gated graph network for few-shot learning
Peixiao Zheng, Xin Guo 0005, Enqing Chen, Lin Qi 0001, Ling Guan |
Pattern Recognit. | 5 |
| 2024 | An Optimal Edge-weighted Graph Semantic Correlation Framework for Multi-view Feature Representation LearningabstractIn this article, we present an optimal edge-weighted graph semantic correlation (EWGSC) framework for multi-view feature representation learning. Different from most existing multi-view representation methods, local structural information and global correlation in multi-view feature spaces are exploited jointly in the EWGSC framework, leading to a new and high-quality multi-view feature representation. Specifically, a novel edge-weighted graph model is first conceptualized and developed to preserve local structural information in each of the multi-view feature spaces. Then, the explored structural information is integrated with a semantic correlation algorithm, labeled multiple canonical correlation analysis (LMCCA), to form a powerful platform for effectively exploiting local and global relations across multi-view feature spaces jointly. We then theoretically verified the relation between the upper limit on the number of projected dimensions and the optimal solution to the multi-view feature representation problem. To validate the effectiveness and generality of the proposed framework, we conducted experiments on five datasets of different scales, including visual-based (University of California Irvine (UCI) iris database, Olivetti Research Lab (ORL) face database, and Caltech 256 database), text-image-based (Wiki database), and video-based (Ryerson Multimedia Lab (RML) audio-visual emotion database) examples. The experimental results show the superiority of the proposed framework on multi-view feature representation over state-of-the-art algorithms. Lei Gao 0001, Ling Guan |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Sparsity-guided Discriminative Feature Encoding for Robust Keypoint DetectionabstractExisting handcrafted keypoint detectors typically focus on designing specific local structures manually while ignoring whether they have enough flexibility to explore diverse visual patterns in an image. Despite the advancement of learning-based approaches in the past few years, most of them still rely on the availability of the outputs of handcrafted detectors as a part of training. In fact, such dependence limits their ability to discover various visual information. Recently, semi-handcrafted methods based on sparse coding have emerged as a promising paradigm to alleviate the above issue. However, the visual relationships between feature points have not been considered in the encoding stage, which may weaken the discriminative capability of feature representations for keypoint recognition. To tackle this problem, we propose a novel sparsity-guided discriminative feature representation (SDFR) method that attempts to explore the intrinsic correlations of keypoint candidates, thus ensuring the validity of characterizing distinctive and diverse structural information. Specifically, we first incorporate an affinity constraint into the feature representation objective, which jointly encodes all the patches in an image while highlighting the similarities and differences between them. Meanwhile, a smoother sparsity regularization with the Frobenius norm is leveraged to further preserve the similarity relationships of patch representations. Due to the differentiable property of this sparsity, SDFR is computationally feasible and effective for representing dense patches. Finally, we treat the SDFR model as multiple optimization sub-problems and introduce an iterative solver. During comprehensive evaluations on five challenging benchmarks, the proposed method achieves favorable performances compared with the state of the art in the literature. Yurui Xie, Ling Guan |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Towards Efficient Multi-view Representation LearningabstractThe proliferation of deep neural networks (DNNs) has drawn unprecedented interest in the study of various contents such as image, audio, video, to name a few. However, due to the data-driven nature, the high computational requirement and slow running time are considered as Achilles’ heels of DNN-based algorithms, limiting the progress of DNNs in time-sensitive applications. Recently, distinct discriminant canonical correlation analysis network (DDCCANet), a multi-view neural network, has shown great generalizability across multiple application domains, both analytically and experimentally. However, Although the computational requirement and running time by DDCCANet are more manageable than DNN-based algorithms, they can be substantially further improved. This paper proposes two new algorithms for multi-view feature representation learning, namely incremental DDCCANet (IDDCCANet) with substantial save in computational memory and GPU-accelerated DDCCANet (GADDCCANet) with drastically accelerated running time, forming a practically significant platform for multi-view feature representation learning. To validate the power of the proposed algorithms, experiments are conducted on several data sets with different types of inputs (e.g., raw image pixels, classical and DNN-based features). Experimental results clearly show that the proposed algorithms provide promising solutions to address the two longstanding challenges. Kai Liu 0032, Lei Gao 0001, Naimul Mefraz Khan, Ling Guan |
ISM | 5 |
| 2023 | A Discriminant Information Theoretic Learning Framework for Multi-modal Feature RepresentationabstractAs sensory and computing technology advances, multi-modal features have been playing a central role in ubiquitously representing patterns and phenomena for effective information analysis and recognition. As a result, multi-modal feature representation is becoming a progressively significant direction of academic research and real applications. Nevertheless, numerous challenges remain ahead, especially in the joint utilization of discriminatory representations and complementary representations from multi-modal features. In this article, a discriminant information theoretic learning (DITL) framework is proposed to address these challenges. By employing this proposed framework, the discrimination and complementation within the given multi-modal features are exploited jointly, resulting in a high-quality feature representation. According to characteristics of the DITL framework, the newly generated feature representation is further optimized, leading to lower computational complexity and improved system performance. To demonstrate the effectiveness and generality of DITL, we conducted experiments on several recognition examples, including both static cases, such as handwritten digit recognition, face recognition, and object recognition, and dynamic cases, such as video-based human emotion recognition and action recognition. The results show that the proposed framework outperforms state-of-the-art algorithms. Lei Gao 0001, Ling Guan |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2022 | A Semi-Handcrafted Keypoint Detector with Discriminative Feature EncodingabstractMost previous handcrafted keypoint methods focus on designing specific structural patterns using human-defined knowledge. These methods, however, ignore the fact that whether they have enough flexibility to harvest diverse local structures. Recently, the semi-handcrafted approaches based on sparse coding have emerged as a new trend of alleviating the above issue. And yet, the intrinsic relationships of key-points have not been explored actively, which may lead to the ambiguity of feature codes for further analysis. To tackle this problem, in this paper, we introduce a novel semi-handcrafted keypoint detector through a scheme of discriminative feature representations (SDFR). Specifically, we cast keypoint detection as an optimization problem on a visual dictionary that explicitly models the visual relationships of feature points to preserve the consistency of similar features and distance dissimilar ones. Further, we propose an iterative solver for the SDFR model. Experimental results on challenge benchmarks demonstrate that the proposed method performs favorably against state-of-the-art in literature. Yurui Xie, Ling Guan |
ICASSP | 2 |
| 2022 | Characteristic Mode-Based Efficient Broadband Electromagnetic Scattering Analysis Method for Array StructuresabstractA fast and accurate numerical method based on the theory of characteristic mode (TCM) is proposed for analyzing electromagnetic scattering from repetitive multiscale array structures. In order to overcome the low frequency breakdown problem when interacting at short distances between observation and source basis functions, the CMs are extracted from the impedance matrix generated by the augmented electric field integral equation (AEFIE), which is discretized by the method of moments (MoM). Then the CMs of a single unit are used as the global basis functions to reduce the unknowns of the array structures. In addition, the broadband multilevel fast multipole algorithm (MLFMA) based on approximate diagonalization of the Green’s function is implemented to accelerate the matrix-vector multiplications between the impedance matrix and CM vectors. Several numerical examples are given to demonstrate the high efficiency and accuracy of the proposed scheme. Chunlai Jia, Zi He, Zhenhong Fan, Dazhi Ding, Ling Guan, Xia Ai |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2022 | Interpretable Artificial Intelligence through Locality Guided Neural Networks
Randy Tan, Lei Gao 0001, Naimul Mefraz Khan, Ling Guan |
Neural Networks | 4 |
| 2022 | A Discriminative Vectorial Framework for Multi-Modal Feature RepresentationabstractDue to the rapid advancements of sensory and computing technology, multi-modal data sources that represent the same pattern or phenomenon have attracted growing attention. As a result, finding means to explore useful information from these multi-modal data sources has quickly become a necessity. In this paper, a discriminative vectorial framework is proposed for multi-modal feature representation in knowledge discovery by employing multi-modal hashing (MH) and discriminative correlation maximization (DCM) analysis. Specifically, the proposed framework is capable of minimizing the semantic similarity among different modalities by MH and exacting intrinsic discriminative representations across multiple data sources by DCM analysis jointly, enabling a novel vectorial framework of multi-modal feature representation. Moreover, the proposed feature representation strategy is analyzed and further optimized based on canonical and non-canonical cases, respectively. Consequently, the generated feature representation leads to effective utilization of the input data sources of high quality, producing improved, sometimes quite impressive, results in various applications. The effectiveness and generality of the proposed framework are demonstrated by utilizing classical features and deep neural network (DNN) based features with applications to image and multimedia analysis and recognition tasks, including data visualization, face recognition, object recognition; cross-modal (text-image) recognition and audio emotion recognition. Experimental results show that the proposed solutions are superior to state-of-the-art statistical machine learning (SML) and DNN algorithms. Lei Gao 0001, Ling Guan |
IEEE Trans. Multim. | 2 |
| 2021 | 2D-FRFT Based Frequency Shift-Invariant Digital Image EncryptionabstractIn this paper, we study the property of frequency shift in two-dimensional Fractional Fourier Transform (2D-FRFT) domain. Based on the mathematical verification and computer simulations, it is demonstrated that the magnitude of reconstruction from phase information satisfies frequency shift-invariant in 2D-FRFT domain, improving the robustness of image encryption. Experiments are implemented to verify the effectiveness of this property against the frequency shift attack in 2D-FRFT domain, improving the robustness of image encryption. Lei Gao 0001, Lin Qi 0001, Ling Guan |
ICASSP | 3 |
| 2021 | ECG Heart-Beat Classification Using Multimodal Image FusionabstractIn this paper, we present a novel Image Fusion Model (IFM) for ECG heart-beat classification to overcome the weaknesses of existing machine learning techniques that rely either on manual feature extraction or direct utilization of 1D raw ECG signal. At the input of IFM, we first convert the heart-beats of ECG into three different images using Gramian Angular Field (GAF), Recurrence Plot (RP) and Markov Transition Field (MTF) and then fuse these images to create a single imaging modality. We use AlexNet for feature ex-traction and classification and thus employ end-to-end deep learning. We perform experiments on PhysioNet’s MIT-BIH dataset for five different arrhythmias in accordance with the AAMI EC57 standard and on PTB diagnostics dataset for myocardial infarction (MI) classification. We achieved an state-of-an-art results in terms of prediction accuracy, precision and recall. Anika Tabassum, Ling Guan, Naimul Mefraz Khan |
ICASSP | 3 |
| 2021 | Local Feature Descriptors with Deep Hypersphere LearningabstractRecent works have demonstrated the power of L2normalization in local feature descriptor learning. While the descriptors are typically learned in the Euclidean space, the similarity between descriptors is often evaluated on a unit hypersphere due to the post-processing of L2normalization for descriptors, which creates a gap between the training stage and the usage stage of feature descriptors. To bridge the gap, we propose a hyperspherical descriptor learning model, where the whole network is projected onto the hyperspherical space. In addition, a squared angular triplet loss is designed to enable the proposed hyperspherical model to learn angularly discriminative descriptors. Experiments on UBC dataset show that the proposed hyperspherical descriptor outperforms its Euclidean counterparts and the state-of-the-art methods on the feature matching task. Song Wang 0008, Xin Guo 0005, Tie Yun, Lin Qi 0001, Ling Guan |
ICIP | 5 |
| 2021 | Image Matching Using Enhancement Offsets With Adaptive Parameter Selection Via Histogram AnalysisabstractIn this paper, an algorithm for improved image matching is presented using enhancement offsets that are created dynamically with adaptive parameter selection. The algorithm operates by analyzing an input image and building image offsets to improve colour contrast, non-uniform illumination and lack of detail which allows for additional keypoints to be detected by numerous different detectors. Keypoint detection testing on the enhanced images is conducted, as well as image matching using SIFT and SURF. It is quantitatively shown that the proposed algorithm results in visual improvements, as well as in additional, stronger keypoints being detected in all images irrespective of the detector used. Matching experiments are conducted using the Webcam and Oxford/EF datasets wherein the proposed algorithm consistently outputs more correct matches and achieves higher or comparable matching accuracy over other related enhancement algorithms using SIFT and SURF. Jonathan Psaila, Thanh Hong-Phuoc, Ling Guan |
ICIP | 3 |
| 2021 | Discriminative Patch Descriptor Learning With Focal Triplet Loss FunctionabstractThis paper proposes a focal triplet loss function for discriminative patch descriptor learning. The standard triplet loss function usually restrains the distance difference between the matching samples and the non-matching ones. However, along with the training procedure, the majority of triplets in each batch tend to satisfy the constraint of the loss function and produce low loss values, leading to a masquerade that the model is well-trained. To address this problem, the focal triplet loss function is proposed in this paper to weaken the impact of the easy triplets and focus training on the hard ones. By emphasizing the importance of hard triplets on the model training, the proposed loss forces the descriptor vectors with fixed dimension to carry more discriminative information from the patches. With the benefits of the focal mechanism, the proposed method achieves better performance compared to the state-of-the-art on UBC dataset for image matching task. Furthermore, to demonstrate the effectiveness of the proposed method, we extend the focal triplet loss on the cross-model retrieval task. The experimental results indicate that the proposed method can also be used to improve visual-semantic embedding learning. Song Wang 0008, Xin Guo 0005, Tie Yun, Lin Qi 0001, Ling Guan |
ICIP | 5 |
| 2021 | A two-stream heterogeneous network for action recognition based on skeleton and RGB modalitiesabstractRecent years, skeleton based action recognition with graph convolutional network (GCN) has achieved great success. However, since skeleton data only includes human body joints coordinates, other key information on actions is missing such as the subtle motion of hands, the objects the human is interacting, leading to an unsatisfactory performance. In this respect, the RGB data can offer help to recognize actions that skeleton-based methods have limitations on. In this work, we propose a novel two-stream heterogeneous network consisting of GCN and CNN networks for action recognition. Specifically, the GCN network takes the skeletal sequence as input to exploit skeleton information. For the RGB video, the CNN model, ResNet (2+1)D, is adapted to exploit RGB information. Afterwards, the discriminant canonical correlation analysis (DCCA) method is utilized to integrate the output feature maps from the skeleton and RGB streams, resulting in improved performance. Experimental results on the large-scale dataset NTU RGB+D show that the proposed model outperforms state-of-the-art models. Kai Liu 0032, Lei Gao 0001, Naimul Mefraz Khan, Lin Qi 0001, Ling Guan |
ISM | 5 |
| 2021 | Safe Driving of Autonomous Vehicles through State Representation LearningabstractIn this paper, we propose an environment perception framework for autonomous driving using state representation learning (SRL). Unlike existing Q-learning based methods for efficient environment perception and object detection, our proposed method takes the learning loss into account under deterministic as well as stochastic policy gradient. Through a combination of variational autoencoder (VAE), deep deterministic policy gradient (DDPG), and soft actor-critic (SAC), we focus on uninterrupted and reasonably safe autonomous driving without steering off the track for a considerable driving distance. To ensure the effectiveness of the scheme over a sustained period of time, we employ a reward-penalty based system where a higher negative penalty is associated with an unfavourable action and a comparatively lower positive reward is awarded for favourable actions. The results obtained through simulations on DonKey simulator show the effectiveness of our proposed method by examining the variations in policy loss, value loss, reward function, and cumulative reward for `VAE+DDPG' and `VAE+SAC' over the learning process. Abhishek Gupta 0007, Ahmed Shaharyar Khwaja, Alagan Anpalagan, Ling Guan |
IWCMC | 4 |
| 2021 | Integrating vertex and edge features with Graph Convolutional Networks for skeleton-based action recognition
Kai Liu 0032, Lei Gao 0001, Naimul Mefraz Khan, Lin Qi 0001, Ling Guan |
Neurocomputing | 5 |
| 2021 | A discriminant kernel entropy-based framework for feature representation learning
Lei Gao 0001, Lin Qi 0001, Ling Guan |
J. Vis. Commun. Image Represent. | 3 |
| 2021 | The Property of Frequency Shift in 2D-FRFT Domain With Application to Image EncryptionabstractThe Fractional Fourier Transform (FRFT) has been playing a unique and increasingly important role in signal and image processing. In this letter, we investigate the property of frequency shift in two-dimensional FRFT (2D-FRFT) domain. It is shown that the magnitude of image reconstruction from phase information is frequency shift-invariant in 2D-FRFT domain, enhancing the robustness of image encryption, an important multimedia security task. Experiments are conducted to demonstrate the effectiveness of this property against the frequency shift attack, improving the robustness of image encryption. Lei Gao 0001, Lin Qi 0001, Ling Guan |
IEEE Signal Process. Lett. | 3 |
| 2021 | Improving Action Recognition via Temporal and Complementary LearningabstractIn this article, we study the problem of video-based action recognition. We improve the action recognition performance by finding an effective temporal and appearance representation. For capturing the temporal representation, we introduce two temporal learning techniques for improving long-term temporal information modeling, specifically Temporal Relational Network and Temporal Second-Order Pooling-based Network. Moreover, we harness the representation using complementary learning techniques, specifically Global-Local Network and Fuse-Inception Network. Performance evaluation on three datasets (UCF101, HMDB-51, and Mini-Kinetics-200) demonstrated the superiority of the proposed framework compared to the 2D Deep ConvNets-based state-of-the-art techniques. Nour El-Din El-Madany, Ling Guan |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2021 | A Scale and Rotational Invariant Key-point Detector based on Sparse CodingabstractMost popular hand-crafted key-point detectors such as Harris corner, SIFT, SURF aim to detect corners, blobs, junctions, or other human-defined structures in images. Though being robust with some geometric transformations, unintended scenarios or non-uniform lighting variations could significantly degrade their performance. Hence, a new detector that is flexible with context change and simultaneously robust with both geometric and non-uniform illumination variations is very desirable. In this article, we propose a solution to this challenging problem by incorporating Scale and Rotation Invariant design (named SRI-SCK) into a recently developed Sparse Coding based Key-point detector (SCK). The SCK detector is flexible in different scenarios and fully invariant to affine intensity change, yet it is not designed to handle images with drastic scale and rotation changes. In SRI-SCK, the scale invariance is implemented with an image pyramid technique, while the rotation invariance is realized by combining multiple rotated versions of the dictionary used in the sparse coding step of SCK. Techniques for calculation of key-points’ characteristic scales and their sub-pixel accuracy positions are also proposed. Experimental results on three public datasets demonstrate that significantly high repeatability and matching score are achieved. Thanh Hong-Phuoc, Ling Guan |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2021 | A Multi-Stream Graph Convolutional Networks-Hidden Conditional Random Field Model for Skeleton-Based Action RecognitionabstractRecently, Graph Convolutional Network(GCN) methods for skeleton-based action recognition have achieved great success due to their ability to preserve structural information of the skeleton. However, these methods abandon the structural information in the classification stage by employing traditional fully-connected layers and softmax classifier, leading to sub-optimal performance. In this work, a novel Graph Convolutional Networks-Hidden conditional Random Field (GCN-HCRF) model is proposed to solve this problem. The proposed method combines GCN with HCRF to retain the human skeleton structure information even during the classification stage. Our model is trained end-to-end by utilizing the message passing from the belief propagation algorithm on the human structure graph. To further capture spatial and temporal information, we propose a multi-stream framework which takes the relative coordinate of the joints and bone direction as two static feature streams, and the temporal displacements between two consecutive frames as the dynamic feature stream. Experimental results on three challenging benchmarks (NTU RGB+D, N-UCLA, SYSU) show the superior performance of the proposed model over state-of-the-art models. Kai Liu 0032, Lei Gao 0001, Naimul Mefraz Khan, Lin Qi 0001, Ling Guan |
IEEE Trans. Multim. | 5 |
| 2020 | A Distinct Discriminant Canonical Correlation Analysis Network based Deep Information Quality Representation for Image ClassificationabstractIn this paper, we present a distinct discriminant canonical correlation analysis network (DDCCANet) based deep information quality representation with application to image classification. Specifically, to explore the sufficient discriminant information between different data sets, the within-class and between-class correlation matrices are employed and optimized jointly. Moreover, different from the existing canonical correlation analysis network (CCANet) and related algorithms, an information theoretic descriptor, information quality (IQ), is adopted to generate the deep-level feature representation for image classification. Benefiting from the explored discriminant information and IQ descriptor, it is potential to gain a more effective deep-level representation from multi-view data sets, leading to improved performance in classification tasks. To demonstrate the effectiveness of the proposed DDCCANet, we conduct experiments on the Olivetti Research Lab (ORL) face database, ETH80 database and CIFAR10 database. Experimental results show the superiority of the proposed solution on image classification. Lei Gao 0001, Ling Guan |
ICPR | 3 |
| 2020 | Locality Guided Neural Networks for Explainable Artificial IntelligenceabstractIn current deep network architectures, deeper layers in networks tend to contain hundreds of independent neurons which makes it hard for humans to understand how they interact with each other. By organizing the neurons by correlation, humans can observe how clusters of neighbouring neurons interact with each other. In this paper, we propose a novel algorithm for back propagation, called Locality Guided Neural Network (LGNN) for training networks that preserves locality between neighbouring neurons within each layer of a deep network. Heavily motivated by Self-Organizing Map (SOM), the goal is to enforce a local topology on each layer of a deep network such that neighbouring neurons are highly correlated with each other. This method contributes to the domain of Explainable Artificial Intelligence (XAI), which aims to alleviate the black-box nature of current AI methods and make them understandable by humans. Our method aims to achieve XAI in deep learning without changing the structure of current models nor requiring any post processing. This paper focuses on Convolutional Neural Networks (CNNs), but can theoretically be applied to any type of deep learning architecture. In our experiments, we train various VGG and Wide ResNet (WRN) networks for image classification on CIFAR100. In depth analyses presenting both qualitative and quantitative results demonstrate that our method is capable of enforcing a topology on each layer while achieving a small increase in classification accuracy. Randy Tan, Naimul Mefraz Khan, Ling Guan |
IJCNN | 3 |
| 2020 | Negative Label Guided Discriminative Canonical Correlation Analysis for Semi-Supervised and Semi-Paired LearningabstractSemi-supervised learning is a popular trend for learning based methods in recent years, as it fully exploits both the labeled and unlabeled samples in a dataset. This paper sets itself apart from most existing semi-supervised learning algorithms, which only use the exact labels of data already known. We take the negative label as side information to guide the process of semi-supervised learning. Two types of supervision information are regarded as negative label; the first type indicates that a sample definitely does not belong to a specific category, and the second indicates that two samples come from different views, and cannot have a one to one correspondence. By reasonably assuming that nearby points should have similar class indicators, the data labels are propagated under the negative label and the geometric structure revealed by both labeled and unlabeled points. Specifically, we predict one to one pair information by utilizing the neighbor information of samples, under the guidance of the negative pair label. Extensive experiments on several datasets demonstrate the effectiveness of our proposed method. Xin Guo 0005, Song Wang 0008, Tie Yun, Lin Qi 0001, Ling Guan |
ISCAS | 5 |
| 2020 | A Vertex-Edge Graph Convolutional Network for Skeleton-Based Action RecognitionabstractThe Graph Convolutional Network(GCN) methods for skeleton-based action recognition have achieved great success due to their ability of exploiting the joint information from the graph structure of the skeleton data. Recently, as a strong and complementary modality for action recognition, the bone information from skeleton data has attracted more attention. However, most existing GCN-based methods extract the bone and joint features with two separate GCN networks, ignoring the dependencies between joints and bones. In this work, a Vertex-Edge Graph Convolutional Network(VE-GCN) is proposed to reveal the information across joints, bones and their relationships simultaneously. In addition, we learn the additional connections among joints and bones for various action samples besides the natural connections of the skeleton. Then we conduct the convolution operation on joints and their neighbors based on these additional connections. Moreover, the conditional random field (CRF) is utilized as the loss function to achieve improved performance. Experimental results on two large-scale datasets NTU RGB+D and NTU RGB+D 120 show that the proposed model outperforms state-of-the-art models. Kai Liu 0032, Lei Gao 0001, Naimul Mefraz Khan, Lin Qi 0001, Ling Guan |
ISCAS | 5 |
| 2020 | An Effective Rotational Invariant Key-point Detector for Image MatchingabstractTraditional detectors e.g. Harris, SIFT, SFOP... are known inflexible in different contexts as they solely target corners, blobs, junctions or other specific human-designed structures. To account for this inflexibility and additionally their unreliability under non-uniform lighting change, recently, a Sparse Coding based Key-point detector (SCK) relying on no human-designed structures and invariant to non-uniform illumination change was proposed. Yet, geometric transformations such as rotation are not considered in SCK. Thus, a novel Rotational Invariant SCK called RI-SCK is proposed in this paper. To make SCK rotational invariant, an effective use of multiple rotated versions of the original dictionary in the sparse coding step of SCK is proposed. A novel strength measure is also introduced for comparison of key-points across image pyramid levels if scale invariance is required. Experimental results on three public datasets have confirmed that significant gains in repeatability and matching score could be achieved by the proposed detector. Thanh Hong-Phuoc, Ling Guan |
ISM | 2 |
| 2020 | Automatic Sparsity-Aware Recognition for Keypoint DetectionabstractWe present a novel Sparsity-Aware Keypoint detector (SAKD) to localize a set of discriminative keypoints via optimization of group-sparse coding. Unlike most of current handcrafted keypoint detectors that are limited by the manually defined local structures, the proposed method has the capacity to allow flexibility for exploiting diverse structures with the combination of visual atoms from a vocabulary. Another key valuable attribute is that its group-sparsity nature concentrates on discovering sharable structural patterns across keypoints within an image jointly. This main merit facilitates to localize repeatable keypoints and resists against distractors when image undergoes various transformations. Extensive experiments on four challenging benchmark datasets demonstrate that the proposed method achieves favorable performances compared with state-of-the-art in literature. Yurui Xie, Ling Guan |
ISM | 2 |
| 2020 | Weighted hybrid fusion with rank consistency
Song Wang 0008, Xin Guo 0005, Tie Yun, Ivan Lee 0001, Lin Qi 0001, Ling Guan |
Pattern Recognit. Lett. | 6 |
| 2020 | A Complete Discriminative Tensor Representation Learning for Two-Dimensional Correlation AnalysisabstractAs an effective tool for two-dimensional data analysis, two-dimensional canonical correlation analysis (2DCCA) is not only capable of preserving the intrinsic structural information of original two-dimensional (2D) data, but also reduces the computational complexity effectively. However, due to the unsupervised nature, 2DCCA is incapable of extracting sufficient discriminatory representations, resulting in an unsatisfying performance. In this letter, we propose a complete discriminative tensor representation learning (CDTRL) method based on linear correlation analysis for analyzing 2D signals (e.g. images). This letter shows that the introduction of the complete discriminatory tensor representation strategy provides an effective vehicle for revealing, and extracting the discriminant representations across the 2D data sets, leading to improved results. Experimental results show that the proposed CDTRL outperforms state-of-the-art methods on the evaluated data sets. Lei Gao 0001, Ling Guan |
IEEE Signal Process. Lett. | 2 |
| 2020 | A Novel Key-Point Detector Based on Sparse CodingabstractMost popular hand-crafted key-point detectors such as Harris corner, MSER, SIFT, SURF rely on some specific pre-designed structures for detection of corners, blobs, or junctions in an image. The very nature of pre-designed structures can be considered a source of inflexibility for these detectors in different contexts. Additionally, the performance of these detectors is also highly affected by non-uniform change in illumination. To the best of our knowledge, while there are some previous works addressing one of the two aforementioned problems, there currently lacks an efficient method to solve both simultaneously. In this paper, we propose a novel Sparse Coding based Key-point detector (SCK) which is fully invariant to affine intensity change and independent of any particular structure. The proposed detector locates a key-point in an image, based on a complexity measure calculated from the block surrounding its position. A strength measure is also proposed for comparing and selecting the detected key-points when the maximum number of key-points is limited. In this paper, the desirable characteristics of the proposed detector are theoretically confirmed. Experimental results on three public datasets also show that the proposed detector achieves significantly high performance in terms of repeatability and matching score. Thanh Hong-Phuoc, Ling Guan |
IEEE Trans. Image Process. | 2 |
| 2020 | Deep Local Feature Descriptor Learning With Dual Hard Batch ConstructionabstractLocal feature descriptor learning aims to represent distinctive images or patches with the same local features, where their representation is invariant under different types of deformation. Recent studies have demonstrated that descriptor learning based on Convolutional Neural Network (CNN) is able to improve the matching performance significantly. However, they tend to ignore the importance of sample selection during the training process, leading to unstable quality of descriptors and learning efficiency. In this paper, a dual hard batch construction method is proposed to sample the hard matching and non-matching examples for training, improving the performance of the descriptor learning on different tasks. To construct the dual hard training batches, the matching examples with the minimum similarity are selected as the hard positive pairs. For each positive pair, the most similar non-matching example is then sampled from the generated hard positive pairs in the same batch as the corresponding negative. By sampling the hard positive pairs and the corresponding hard negatives, the hard batches are produced to force the CNN model to learn the descriptors with more efforts. In addition, based on the above dual hard batch construction, an ℓ22 triplet loss function is built for optimizing the training model. Specifically, we analyze the superiority of the ℓ22 loss function when dealing with hard examples, and also demonstrate it in the experiments. With the benefits of the proposed sampling strategy and the ℓ22 triplet loss function, our method achieves better performance compared to state-of-the-art on the reference benchmarks for different matching tasks. Song Wang 0008, Xin Guo 0005, Tie Yun, Lin Qi 0001, Ling Guan |
IEEE Trans. Image Process. | 5 |
| 2019 | Discriminative Feature Selection Guided Deep Canonical Correlation AnalysisabstractThis paper proposes a novel Discriminative Feature Selection Guided Deep Canonical Correlation Analysis (D2CCA) for multiview learning. The proposed (D2CCA) enhances the discriminative power of the learned featured representation by imposing the selection of the most discriminative features. Moreover, it learns to maximize the correlations between two views. Also, an alternating iterative learning algorithm is presented to find the sub-optimal solution. The experimental results demonstrated that the proposed (D2CCA) can achieve a higher average accuracy compared to several existing methods. Nour El-Din El-Madany, Ling Guan |
ICASSP | 3 |
| 2019 | Information Fusion via Multimodal Hashing With Discriminant Correlation MaximizationabstractDue to low storage cost and fast query speed, hashing has been applied to similarity search in multimedia data widely. In this paper, an effective information fusion algorithm using multimodal hashing with discriminant correlation maximization is presented. The proposed algorithm not only finds the minimum of the semantic similarity across different modalities by multimodal hashing, but also minimizes the between-class correlation and maximizes the within-class correlation simultaneously to extract discriminant representations for information fusion. More importantly, two solutions with canonical case and non-canonical case are presented, and a novel solution to non-canonical case is proposed. Benefiting from the combination of semantic similarity across different modalities from multimodal hashing information and the discriminant representation strategy, the proposed strategy can achieve improved performance. Experimental results show that the proposed approach outperforms the related methods. Lei Gao 0001, Ling Guan |
ICIP | 2 |
| 2019 | Information Fusion via Deep Cross-Modal Factor AnalysisabstractIn this paper, we introduce Deep Cross-Modal Factor Analysis (DCFA) to identify complex nonlinear transformations of two variables for information fusion. DCFA is able to represent the coupled patterns between two different sets of variables by minimizing the Frobenius norm distance in the transformed domain. Unlike previous kernel methods, the feature mapping of DCFA is achieved with deep networks (DN) instead of the traditional kernel method. Therefore, the representation of DCFA method is not limited by the fixed kernel. Moreover, DCFA can be considered as a nonlinear extension of the linear Cross-Modal Factor Analysis (CFA), and an alternative to the nonparametric method Kernel Cross-Modal Factor Analysis (KCFA) and the recently proposed Deep Canonical Correlation Analysis (Deep CCA) method. The performance of DCFA is evaluated on MNIST handwritten digit dataset and two audio emotion datasets. Experimental results show that the proposed solution outperforms the methods of KCCA, KCFA, Deep CCA and the deep learning based method-Alexnet, in terms of accuracy. Lei Gao 0001, Ling Guan |
ISCAS | 2 |
| 2019 | A Novel Weighted Hybrid Multi-View Fusion Algorithm for Semi-Supervised ClassificationabstractSemi-supervised learning aims to improve the learning performance with very limited label information. To dig more available information from the collected data, we propose a weighted hybrid multi-view feature fusion approach for semi-supervised classification problem. Specifically, under the rank consistency constraint for labels predicted by view-specific learners, the proposed method estimates the optimal fusion weight for each learner to balance the incomparable square losses on different views. In this case, the learners with more powerful prediction capability are pushed to have higher weights during the fusion process. Experimental results on 6 real-world datasets demonstrate the effectiveness of the proposed technique. Song Wang 0008, Xin Guo 0005, Tie Yun, Lin Qi 0001, Ling Guan |
ISCAS | 5 |
| 2019 | Local Feature Descriptor Learning with a Dual Hard Sampling StrategyabstractLocal feature descriptor learning based on Convolutional Neural Network (CNN) has demonstrated its capability to generate descriptors with high quality. While extensive studies focused on mining hard non-matching examples to improve descriptor learning performance, a random sampling strategy is adopted for matching examples. In this paper, a dual hard sampling strategy based on the triplet loss function is proposed to generate the hard matching and non-matching examples for training. To start with, a pair of matching examples with the maximum distance for each class are selected as the positive pair. For each positive pair, their closest non-matching example is then sampled from the generated positive pairs with other classes as the corresponding negative. Based on the above dual hard sampling strategy, a novel triplet loss function is presented for optimization. With the benefits of the proposed sampling strategy and the novel triplet loss function, our method achieves better performance compared to state-of-the-art on the reference benchmark for local feature matching. Song Wang 0008, Xin Guo 0005, Tie Yun, Lin Qi 0001, Ling Guan |
ISM | 5 |
| 2019 | Graph Convolutional Networks-Hidden Conditional Random Field Model for Skeleton-Based Action RecognitionabstractRecently, Graph Convolutional Network(GCN) methods for skeleton-based action recognition have achieved great success due to their ability to preserve structural information of the skeleton. However, these methods abandon the structural information in the classification stage by employing traditional fully-connected layers and softmax classifier, leading to sub-optimal performance. In this work, a novel Graph Convolutional Networks-Hidden conditional Random Field (GCN-HCRF) model is proposed to solve this problem. The proposed method combines GCN and HCRF to retain the human skeleton structure information during the classification stage. The proposed model is trained end-to-end by utilizing the message passing from the belief propagation algorithm on the human structure graph. To further capture spatial and temporal information, we propose a multi-stream framework that takes the relative coordinates of the joints and bone direction as two static feature streams and the temporal displacements as the dynamic feature stream. Experimental results on two challenging benchmarks (NTU RGB+D, N-UCLA) show the superior performance of the proposed model over state-of-the-art models. Kai Liu 0032, Lei Gao 0001, Naimul Mefraz Khan, Lin Qi 0001, Ling Guan |
ISM | 5 |
| 2019 | Multimodal Learning for Human Action Recognition Via Bimodal/Multimodal Hybrid Centroid Canonical Correlation AnalysisabstractIn this paper, we study the problem of human action recognition from multiple feature modalities. We propose bimodal hybrid centroid canonical correlation analysis (BHCCCA) and multimodal hybrid centroid canonical correlation analysis (MHCCCA) to learn the discriminative and informative shared space, by considering the correlation among different classes across two modalities (BHCCCA) and three or more modalities (MHCCCA). We then introduce a new human action recognition framework by using BHCCCA/MHCCCA for fusing different modalities (RGB, depth, skeleton, and accelerometer data). Performance evaluation on four publicly accessible data sets (MSR Action3D, UTD-MHAD, UTD-MHAD-Kinect V2, and Berkeley MHAD) demonstrated the effectiveness of the proposed framework. Nour El-Din El-Madany, Ling Guan |
IEEE Trans. Multim. | 3 |
| 2019 | The Labeled Multiple Canonical Correlation Analysis for Information FusionabstractThe objective of multimodal information fusion is to mathematically analyze information carried in different sources and create a new representation that will be more effectively utilized in pattern recognition and other multimedia information processing tasks. In this paper, we introduce a new method for multimodal information fusion and representation based on the Labeled Multiple Canonical Correlation Analysis (LMCCA). By incorporating class label information of the training samples, the proposed LMCCA ensures that the fused features carry discriminative characteristics of the multimodal information representations and are capable of providing superior recognition performance. We implement a prototype of LMCCA to demonstrate its effectiveness on handwritten digit recognition, face recognition, and object recognition utilizing multiple features, bimodal human emotion recognition involving information from both audio and visual domains. The generic nature of LMCCA allows it to take as input features extracted by any means, including those by deep learning (DL) methods. Experimental results show that the proposed method enhanced the performance of both statistical machine learning methods, and methods based on DL. Lei Gao 0001, Rui Zhang 0010, Lin Qi 0001, Enqing Chen, Ling Guan |
IEEE Trans. Multim. | 5 |
| 2018 | Deep Camera Pose Regression Using Motion VectorsabstractA deep learning based camera pose regression framework is presented in this paper. The major objectives of the proposed method are twofold: enhancing the intra-scene pose regression accuracy and improving the inter-scene inference capability. Unlike other pose regression networks, the proposed framework adopts motion vectors as its input tensor, rather than directly taking the pixel intensities. Such concept is developed from two fundamental facts: the motion vectors are strongly associated with pose transition, and they are less relevant to scene-specific visual cue. Experimental results show that the proposed framework can achieve better performance in terms of intra-scene regression accuracy and inter-scene network inference. Fei Guo 0002, Ling Guan |
ICIP | 3 |
| 2018 | SCK: A Sparse Coding Based Key-Point DetectorabstractAll current popular hand-crafted key-point detectors such as Harris corner, MSER, SIFT, SURF ... rely on some specific pre-designed structures for the detection of corners, blobs, junctions ... in an image. In this paper, a novel sparse coding based key-point detector which requires no particular predesigned structures is presented. The key-point detector is based on measuring the complexity level of each block in an image to decide where a key-point should be. The complexity level of a block is defined as the total number of non-zero components of a sparse representation of that block. Generally, a block constructed with more components is more complex and has greater potential to be a good key-point. Experimental results on Webcam and EF datasets [1], [2] show that the proposed detector achieves significantly high repeatability compared to hand-crafted features, and even outperforms the matching scores of the state-of-the-art learning based detector. Thanh Hong-Phuoc, Ling Guan |
ICIP | 3 |
| 2018 | Integrating Entropy Skeleton Motion Maps and Convolutional Neural Networks for Human Action RecognitionabstractThis paper presents an effective method to represent the information of skeleton sequences as images, referred to skeleton motion maps (SMM) and employ convolutional neural networks to recognize the human actions. The proposed approach employs Entropy SMM which captures the temporal evolution of action leading to more effective and discriminative representation. In order to verify the effectiveness of the proposed method, several experiments were conducted on UTD Multimodal Human Action Dataset (UTD-MHAD), Kinect Action Recognition Dataset (KARD), and Multimodal Action Database (MAD) datasets. The Experimental results show the superiority of the proposed method over the existing work. Nour El-Din El-Madany, Ling Guan |
ICME | 3 |
| 2018 | Discriminative Robust Gaze Estimation Using Kernel-DMCCA FusionabstractThe proposed framework employs discriminative analysis for gaze estimation using kernel discriminative multiple canonical correlation analysis (K-DMCCA), which represents different feature vectors that account for variations of head pose, illumination and occlusion. The feature extraction component of the framework includes spatial indexing, statistical and geometrical elements. Gaze estimation is constructed by feature aggregation and transforming features into a higher dimensional space using the RBF kernel γ and spread factor. The output of fused features through K-DMCCA is robust to illumination, occlusion and is calibration free. Our algorithm is validated on MPII, CAVE, ACS and EYEDIAP datasets. The two main contributions of the framework are the following: Enhancing the performance of DMCCA with the kernel and introducing quadtree as an iris region descriptor. Spatial indexing using quadtree is a robust method for detecting which quadrant the iris is situated, detecting the iris boundary and it is inclusive of statistical and geometrical indexing that are calibration free. Our method achieved an accurate gaze estimation of 4.8° using Cave, 4.6° using MPII, 5.1° using ACS and 5.9° using EYEDIAP datasets respectively. The proposed framework provides insight into the methodology of multi-feature fusion for gaze estimation. Salah Rabba, Matthew J. Kyan, Lei Gao 0001, Azhar Quddus, Ali Shahidi Zandi, Ling Guan |
ISM | 6 |
| 2018 | Deep Reinforcement Learning with Parameterized Action Space for Object DetectionabstractObject detection is a fundamental task in computer vision. With the remarkable progress made in big visual data analytics and deep learning, Reinforcement Learning (RL) is becoming a promising framework to model the object detection problem since the detection procedure can be cast as a Markov decision process (MDP). We propose a Reinforcement Learning system with parameterized action space for image object detection. The proposed system uses an active agent exploring in a scene to identify the location of a target object, and learns a policy to refine the geometry of the agent by taking simple actions in parameterized space, which integrates the discrete actions and its corresponding continuous parameters. We then optimize the representation of the generated region proposals with the discriminative multiple canonical correlation analysis (DMCCA) [11] in preparation for classification with Fast R-CNN. Experiments on PASCAL VOC 2007 and 2012 datasets show the effectiveness of the proposed method. Naimul Mefraz Khan, Lei Gao 0001, Ling Guan |
ISM | 4 |
| 2018 | Incremental generalized multiple maximum scatter difference with applications to feature extraction
Ning Zheng 0003, Xin Guo 0005, Tie Yun, Nan Dong, Lin Qi 0001, Ling Guan |
J. Vis. Commun. Image Represent. | 6 |
| 2018 | Joint intermodal and intramodal correlation preservation for semi-paired learning
Xin Guo 0005, Song Wang 0008, Tie Yun, Lin Qi 0001, Ling Guan |
Pattern Recognit. | 5 |
| 2018 | 3D Human Action Recognition Using a Single Depth Feature and Locality-Constrained Affine Subspace CodingabstractThis paper addresses the problem of recognizing human actions from depth videos. We propose a depth-based local descriptor and affine subspace coding representation with locality-constrained affine subspace coding (LASC) for 3D action recognition. First, each depth video sequence is divided into a set of subsequences (i.e., multi-scale sub-actions) based on the normalized motion energy vector. Next, depth motion map-based gradient local auto-correlation features are employed to capture the shape information and motion cues of each sub-action. In order to obtain discriminative and compact representation, we extract the local high-order information of the depth video using LASC. Through experiments, we show that the use of LASC exhibits better performance compared with existing methods such as locality-constrained linear coding. We compared LASC with the state-of-the-art methods based on similar principle, using features extracted from a single modality, on four datasets, and with those using multiple features or nonlinear recognition machines. The results on four datasets clearly show the effectiveness of the proposed method. Chengwu Liang, Lin Qi 0001, Ling Guan |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | Information Fusion for Human Action Recognition via Biset/Multiset Globality Locality Preserving Canonical Correlation AnalysisabstractIn this paper, we study the problem of human action recognition, in which each action is captured by multiple sensors and represented by multisets. We propose two novel information fusion techniques for fusing the information from multisets. The first technique is biset globality locality preserving canonical correlation analysis (BGLPCCA), which aims to learn the common feature subspace between two sets. The second technique is multiset globality locality preserving canonical correlation analysis (MGLPCCA), which aims to deal with three or more sets. The proposed BGLPCCA and MGLPCCA are able to learn a low-dimensional common subspace that preserves the local and global structures of data samples. Moreover, two novel descriptors are presented for both depth and skeleton. We then propose a new human action recognition framework employing the proposed BGLPCCA or MGLPCCA to learn the shared subspace from multiple sets of features including skeleton, depth, and optical flow. Extensive experiments on five publicly available datasets (MSR Action3D, UTD multimodal human action dataset, multimodal action database, Kinect activity recognition dataset, and SBU Kinect interaction dataset) demonstrate the effectiveness of the proposed framework. Nour El-Din El-Madany, Ling Guan |
IEEE Trans. Image Process. | 3 |
| 2018 | Discriminative Multiple Canonical Correlation Analysis for Information FusionabstractIn this paper, we propose the discriminative multiple canonical correlation analysis (DMCCA) for multimodal information analysis and fusion. DMCCA is capable of extracting more discriminative characteristics from multimodal information representations. Specifically, it finds the projected directions, which simultaneously maximize the within-class correlation and minimize the between-class correlation, leading to better utilization of the multimodal information. In the process, we analytically demonstrate that the optimally projected dimension by DMCCA can be quite accurately predicted, leading to both superior performance and substantial reduction in computational cost. We further verify that canonical correlation analysis (CCA), multiple canonical correlation analysis (MCCA) and discriminative canonical correlation analysis (DCCA) are special cases of DMCCA, thus establishing a unified framework for canonical correlation analysis. We implement a prototype of DMCCA to demonstrate its performance in handwritten digit recognition and human emotion recognition. Extensive experiments show that DMCCA outperforms the traditional methods of serial fusion, CCA, MCCA, and DCCA. Lei Gao 0001, Lin Qi 0001, Enqing Chen, Ling Guan |
IEEE Trans. Image Process. | 4 |
| 2018 | A Novel Image-Centric Approach Toward Direct Volume RenderingabstractTransfer function (TF) generation is a fundamental problem in direct volume rendering (DVR). A TF maps voxels to color and opacity values to reveal inner structures. Existing TF tools are complex and unintuitive for the users who are more likely to be medical professionals than computer scientists. In this article, we propose a novel image-centric method for TF generation where instead of complex tools, the user directly manipulates volume data to generate DVR. The user’s work is further simplified by presenting only the most informative volume slices for selection. Based on the selected parts, the voxels are classified using our novel sparse nonparametric support vector machine classifier, which combines both local and near-global distributional information of the training data. The voxel classes are mapped to aesthetically pleasing and distinguishable color and opacity values using harmonic colors. Experimental results on several benchmark datasets and a detailed user survey show the effectiveness of the proposed method. Naimul Mefraz Khan, Riadh Ksantini, Ling Guan |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2017 | Human action recognition by fusing deep features with Globality Locality Preserving Canonical Correlation AnalysisabstractThis paper proposes a novel Globality Locality Preserving Canonical Correlation Analysis (GLPCCA) for multiview learning. The proposed GLPCCA can preserve the global and local structures. Furthermore, we present a human action recognition framework by using GLPCCA to fuse depth and RGB modalities, which include the proposed Hierarchical Pyramid of Depth Motion Map Deep Convolutional Neural Network (HP-DMM-CNN) for the depth images, and Optical flow CNN for the RGB videos. The proposed framework was evaluated using two datasets, UTD Multimodal Human Action Dataset (UTD-MHAD) and SBU Kinect Interaction data set. The experimental results demonstrated that the proposed GLPCCA can achieve a higher average accuracy compared to several existing methods. Nour El-Din El-Madany, Ling Guan |
ICIP | 3 |
| 2017 | Pupil localization for gaze estimation using unsupervised graph-based modelabstractIn this paper, we propose a graph-based model for pupil localization, which is a step towards gaze detection. The proposed model can differentiate the key points located at the eyelashes, eyebrows and eye white regions. We first crop the eye region with an ellipse and then estimate the pupil center within the ellipse, thus reducing the computational complexity. We also consider the light reflections in the pupil region, which could lead to inaccuracy in pupil localization. We construct an undirected graph in the eye region based on the key points consisting of the corner points in the eye region, the centers of light reflection regions, and the multiple pixels with a high intensity in the pupil region. The pupil center is initially estimated as the weighted center of the revised graph after vertex/edge removal. In addition, we shift the initial pupil center to a revised position based on the line segments in the pupil region. We evaluate the proposed method on 850 eye images from a public database. The experimental results demonstrate that the proposed method can achieve a more accurate result compared to the existing work. Salah Rabba, Matthew J. Kyan, Ling Guan |
ISCAS | 4 |
| 2017 | Heterogeneous Features Fusion with Collaborative Representation Learning for 3D Action RecognitionabstractHuman action recognition of depth sensors has drawn wide attentions in computer vision and multimedia processing areas. In contrast to simple periodic actions, irrelevant actions or sharing sub-actions between different classes of two-person non-periodic interactions make this task challenging. This paper presents heterogeneous features fusion with Collaborative Representation (CR) to address this challenge. Two effective high dimensional low-level features are developed from depth image sequence and skeleton pose sequence respectively. In the Canonical Correlations Analysis (CCA) feature space of these two features, Collaborative Representation (CR) is learned and adopted as the final high-level discriminative representation. Experiments on two depth action datasets (SBU Kinect-Interaction and MSR Action 3D) show that the proposed method is superior to the state-of-the-art methods compared, including some recent deep learning based methods. Chengwu Liang, Enqing Chen, Lin Qi 0001, Ling Guan |
ISM | 4 |
| 2017 | Real-Time System for Human Activity AnalysisabstractWe propose a real-time human activity analysis system, where a user's activity can be quantitatively evaluated with respect to a ground truth recording. We use two Kinects to solve the problem of self-occlusion through extracting optimal joint positions using Singular Value Decomposition (SVD) and Sequential Quadratic Programming (SQP). Incremental Dynamic Time Warping (IDTW) is used to compare the user and expert (ground truth) to quantitatively score the user's performance. Furthermore, the user's performance is displayed through a visual feedback system, where colors on the skeleton represent the user's score. Our experiments use a motion capture suit as ground truth to compare our dual Kinect setup to a single Kinect. We also show that with our visual feedback method, users gain a statistically significant boost to learning as opposed to watching a simple video. Randy Tan, Naimul Mefraz Khan, Ling Guan |
ISM | 3 |
| 2017 | Motion energy guided multi-scale heterogeneous features for 3D action recognitionabstractThis paper is to address the problem of human action recognition in depth sequences. The actions with various speeds and shared sub-actions make the recognition challenging. A new feature set, consisting of two heterogeneous features are proposed to address this challenge. Specifically, we propose an adaptive normalized action motion energy based on the depth video. Guided by this multi-scale energy vector, depth sequence and skeleton pose sequence are divided respectively into two sets of subsequences with multiple scales (i.e., multi-scale sub-actions). Then in depth modality, based on the depth sub-sequence, Depth Motion Maps (DMMs) based Histogram Oriented Gradient (HOG) features are employed to capture the shape information and motion cues. In skeleton modality, based on the pose sub-sequence, pose dynamics using skeleton information are extracted. In order to obtain discriminative and compact representation, the Collaborative Representation (CR) learning scheme based classifier is adopted. Experiments on two datasets show the effectiveness of the proposed method. Chengwu Liang, Lin Qi 0001, Ling Guan |
VCIP | 3 |
| 2017 | Resource Allocation for Energy Harvesting Assisted D2D Communications Underlaying OFDMA Cellular NetworksabstractDevice-to-Device (D2D) communications underlaying cellular communications has been explored in the literature for a while since the benefits of enhanced sum throughput and more efficient spectrum usage have been proven very promising through the activation of direct transmissions between a pair of devices. To achieve better performance in terms of energy preservation, we consider introducing energy harvesting (EH) mechanism into the traditional D2D model. Our aim is to maximize sum throughput for D2D users without compromising the QoS performance of cellular users (CUs) in an EH-aided communications model. D2D transmissions will only be activated at the beginning of a time slot if there remains enough energy, which is set as a lower threshold, for one-slot data transmission in the batteries of D2D users. Otherwise, it will switch into energy harvesting mode until the energy level in the batteries rises back to an upper threshold. The formulated optimization problem is a nonlinear mixed integer problem. Since it is mathematically challenging to get an optimal solution, we aim for a suboptimal solution with an iterative joint resource block and power resource allocation algorithm. Then we compare this heuristic algorithm with a simplified version where the constraints to make sure that every D2D user has at least one RB for communications are slighted. Numerical simulation results show that energy harvesting mechanism can efficiently power D2D communications underlaying cellular networks. They also corroborate higher sum throughput under different parameter settings of our first proposed approach. Shuo Yu 0004, Waleed Ejaz, Ling Guan, Alagan Anpalagan |
VTC Fall | 3 |
| 2017 | vConnect: perceive and interact with real world from CAVE
Xiaoming Nan, Ning Zhang 0023, Fei Guo 0002, Edward Rosales, Ling Guan |
Multim. Tools Appl. | 7 |
| 2017 | Delay-Rate-Distortion Optimization for Cloud Gaming With Hybrid StreamingabstractCloud gaming as the emerging game service has attracted significant attention. However, traditional video streaming approach suffers from high bandwidth consumption, and traditional graphics streaming approach requires a long initial period to download game models. In this paper, we propose a novel hybrid streaming framework, jointly applying video streaming and graphics streaming to provide a high-quality gaming experience. In the proposed framework, cloud servers not only transmit the encoded video frames but also progressively transmit the graphics data, which are used to render a game frame to provide an additional reference to the video encoder. Based on the proposed framework, we investigate the delay-rate-distortion optimization problem, where the source rate between the video stream and the graphics stream is optimized to minimize the overall distortion under the bandwidth and response delay constraints. The experimental results demonstrate that the proposed hybrid streaming can achieve the lowest distortion under the constraints of bandwidth and response delay, compared with the traditional video streaming and graphics streaming. Xiaoming Nan, Xun Guo 0002, Yan Lu 0001, Ling Guan, Shipeng Li 0001, Baining Guo |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2016 | Multiview learning via deep discriminative canonical correlation analysisabstractIn this paper, we propose Deep Discriminative Canonical Correlation Analysis (DDCCA), a method to learn the nonlinear transformation of two data sets such that the within-class correlation is maximized and the inter-class correlation is minimized. Parameters of the two deep transformations are jointly learned. Unlike CCA and Discriminative CCA, the proposed DDCCA does not need inner product. The proposed DDCCA was evaluated in two applications, handwritten digit recognition and speech-based emotion recognition. The experimental results demonstrated that the proposed DDCCA can get a higher recognition accuracy compared to the existing Deep CCA method. Nour El-Din El-Madany, Ling Guan |
ICASSP | 3 |
| 2016 | Information fusion based on kernel entropy component analysis in discriminative canonical correlation space with application to audio emotion recognitionabstractAs an information fusion tool, Kernel Entropy Component Analysis (KECA) is realized by using descriptor of information entropy and optimized by entropy estimation. However, as an unsuper-vised method, it merely puts the information or features from different channels together without considering their intrinsic structures and relations. In this paper, we introduce an enhanced version of KECA for information fusion, KECA in Discriminative Canonical Correlation Space (DCCS). Not only the intrinsic structures and discriminative representations are considered, but also the natural representations of input data are revealed by entropy estimation, leading to improved recognition accuracy. The effectiveness of the proposed solution is evaluated through experiments on two audio emotion databases. Experimental results show that the proposed solution outperforms the existing methods based on similar principles. Lei Gao 0001, Lin Qi 0001, Ling Guan |
ICASSP | 3 |
| 2016 | Human action recognition via multiview discriminative analysis of canonical correlationsabstractThis paper proposes a novel Multiview Discriminative Analysis of Canonical Correlations (MDACC) for multiview learning. The proposed MDACC can capture discriminative features. Furthermore, we present a human action recognition framework by using MDACC to fuse multimodal features, which include the hierarchical Pyramid of Depth Motion Map (HP-DMM) for the depth images, the Histogram of Oriented Displacement (HOD) for the skeleton, and the statistical measurements for the accelerometer. The proposed framework was evaluated using two datasets MSR-Action3D dataset and UTD multimodal human action dataset. The experimental results demonstrated that the proposed framework can achieve a higher average accuracy compared to several existing methods. Nour El-Din El-Madany, Ling Guan |
ICIP | 3 |
| 2016 | Multiview emotion recognition via multi-set locality preserving canonical correlation analysisabstractIn this paper, we propose a novel Multi-set Locality-Preserving Canonical Correlation Analysis (MLPCCA) for multi-view learning and fusion. The proposed MLPCCA captures the intrinsic structure of data while it learns the optimum basis for maximizing the correlation among different sets of data. To verify the effectiveness of the proposed technique, the proposed MLPC A has been applied in audio-based emotion recognition and visual-based emotion recognition, respectively. The experimental results demonstrated that the proposed MLPCCA can achieve a higher recognition accuracy compared to the existing methods including CCA, LPCCA, and MCCA. Nour El-Din El-Madany, Ling Guan |
ISCAS | 3 |
| 2016 | A Novel Discriminative Framework Integrating Kernel Entropy Component Analysis and Discriminative Multiple Canonical Correlation for Information FusionabstractThe effective interpretation and integration of multiple information content are important for the efficacious utilisation of multimedia in a wide variety of application context. The major challenge in information fusion lies in the difficulty of identifying the complementary and discriminatory representations from individual channels or data sources. In this paper, we propose a novel framework integrating kernel entropy-estimation and discriminative multiple canonical correlation (DMCC) to address this challenge. Not only the distribution and complementary representations of input data are revealed by entropy estimation, but also the discriminative representations are considered by DMCC, achieving improved recognition accuracy. The effectiveness of the proposed method is demonstrated on two audio emotion databases. Experimental results show that it outperforms the existing methods based on similar principles. Lei Gao 0001, Ling Guan, Lin Qi 0001, Enqing Chen |
ISM | 2 |
| 2016 | Semi-Supervised and Semi-Paired Graph Regularized Multiset Canonical Correlation AnalysisabstractMultiset canonical correlation analysis (MCCA) plays a key role in analyzing linear correlations among multimodal data. However, when facing semi-supervised and semi-paired multimodal data which widely exist in real world. MCCA normally performs poorly because it requires paired information among different models. At the same time, it fails to exploit the discriminative information due to the fact that it is an unsupervised dimension reduction method. In this paper, we propose a novel algorithm, named semi-supervised semi-paired graph regularized multiset canonical correlation analysis(SSGMCCA). SSGMCCA employs a small amount of paired data to perform MCCA and simultaneously utilize both the global structural information captured from the unlabeled data and the local structural information captured from the labeled data to compensate the limited paired. As a result, SSGMCCA can find the directions which not only maximal correlation among the multimodal data but also maximal separability of the labeled data. Experimental results illustrate the effectiveness of the proposed algorithm. Xin Guo 0005, Lin Qi 0001, Ling Guan |
ISM | 3 |
| 2016 | 3D Action Recognition Using Depth-Based Feature and Locality-Constrained Affine Subspace CodingabstractWe propose a 3D action recognition algorithm which uses depth-based Gradient Local Auto-Correlations (GLAC) feature and Locality-constrained Affine Subspace Coding (LASC) to improve the discriminative ability of human actions in spatio-temporal subsequences of 3D depth videos. First, each entire depth video sequence is divided automatically into a set of subsequences (i.e., multi-scale sub-actions) by the normalized motion energy vector. Next Depth Motion Maps (DMMs) based GLAC features are employed to capture the shape information and motion cues of each sub-action. In order to obtain a more compact and discriminative representation, LASC is then proposed to encode the features extracted from the depth video. We show that the use of LASC exhibits better performance compared to existing methods such as Locality-constrained Linear Coding (LLC). On all three datasets we obtain competitive results compared to fifteen methods, while using fewer features and less complex models. Chengwu Liang, Enqing Chen, Lin Qi 0001, Ling Guan |
ISM | 4 |
| 2016 | Human gesture recognition via bag of angles for 3D virtual city planning in CAVE environmentabstractCave Automatic Virtual Environment (CAVE) provides an immersive virtual environment for 3D city planning. However, the user in the cave has to wear wearable markers for the interaction with the 3D models. To develop a natural interaction, depth camera like Kinect has to be considered. In this paper, we propose a new skeleton joint representation called Bag of Angles (BoA) for human gesture recognition. We evaluated our proposed BoA representation on two dataset UTD-MHAD and UTD-MHAD-KinectV2. The evaluation results demonstrated that the proposed BoA representation can achieve a higher recognition accuracy compared to the other existing representation methods. Nour El-Din El-Madany, Ling Guan |
MMSP | 3 |
| 2016 | Delay-rate-distortion optimization for cloud-based collaborative renderingabstractCloud rendering is emerged as a new cloud service to satisfy user's desire for running sophisticated graphics applications on thin devices. However, traditional cloud rendering approaches, both remote rendering and local rendering, have limitations. Remote rendering shifts intensive rendering tasks to cloud server and streams rendered frames to client, which suffers from high delay and bandwidth usage. Local rendering sends graphics data to client and performs rendering on local devices, which requires initial buffering delay and demands high computation capacity at client. In this paper, we propose a novel cloud based collaborative rendering framework, which adaptively integrates remote rendering and local rendering. Based on the proposed framework, we study the delay-Rate-Distortion (d-R-D) optimization problem, in which the source rates are optimally allocated for streaming encoded video frames and graphics data to minimize the overall distortion under the bandwidth and response delay constraints. Experiment results demonstrate that the proposed collaborative rendering framework can effectively allocate source rates to achieve the minimal distortion compared to the traditional remote rendering and local rendering. Xiaoming Nan, Ling Guan |
MMSP | 3 |
| 2016 | Joint optimization of resource allocation and workload scheduling for cloud based multimedia servicesabstractWith the development of cloud technology, cloud computing has been increasingly used as distributed platforms for multimedia services. However there are two fundamental challenges for service providers: one is resource allocation, and the other is workload scheduling. Due to the rapidly varying workload and strict response time requirement, it is difficult to optimally allocate virtual machines (VMs) and assign workload. In this paper, we study the resource allocation and workload scheduling problem for cloud based multimedia services. Specifically, we introduce a queuing model to quantify the resource demands and service performance, and a directed acyclic graph (DAG) model to characterize the precedence constraints among jobs. Based on the proposed models, we jointly optimize the allocated VMs and the assigned workload to minimize the total resource cost under the response time constraints. Since the formulated problem is mixed integer non-linear programming, a heuristic is proposed to efficiently allocate resources for practical services. Experimental results show that the proposed scheme can effectively allocate VMs and schedule workload to achieve the minimal resource cost. Xiaoming Nan, Ling Guan |
MMSP | 3 |
| 2016 | Improving Action Recognition Using Collaborative Representation of Local Depth Map FeatureabstractBased on depth information, this letter introduces a new local depth map feature describing local spatiotemporal details of human motion and a collaborative representation for classification with regularized least squares. By extracting a multilayered depth motion feature and then applying a multiscale Histograms of Oriented Gradient (HOG) descriptor to it, the proposed feature characterizes the local temporal change of human motion and the local spatial structure (appearance) of an action. Instead of class-specific dictionary, the test action sample is represented collaboratively by the common shared dictionary. Moreover, we present an analytical solution of collaborative representation, which is independent of the query and can be precalculated as a projection matrix, leading to low computational cost in recognition. The evaluations on MSRAction3D and MSRGesture3D datasets demonstrate its effectiveness. Chengwu Liang, Enqing Chen, Lin Qi 0001, Ling Guan |
IEEE Signal Process. Lett. | 4 |
| 2015 | Sparsity preserving multiple canonical correlation analysis with visual emotion recognition to multi-feature fusionabstractSparsity preserving projections (SPP) aim to preserve the sparse reconstructive relationship among the data and have been successfully applied to face recognition. The projections are invariant to rotations, rescalings, and translations of the data, and more importantly, they contain natural discriminating information even without class labels. Based on the concept of SSP, it presents a new method for multi-feature information fusion based on the Sparsity Preserving Multiple Canonical Correlation Analysis (SPMCCA), which can preserve the sparse reconstructive relationship of the data for recognition from multi-feature information representation. We implement a prototype of SPM-CCA with the application to visual-based human emotion recognition. Experimental results show that the proposed method outperforms the traditional methods of serial fusion, Canonical Correlation Analysis (CCA), Multiple Canonical Correlation Analysis (MCCA) and recently proposed Sparsity Preserving Canonical Correlation Analysis (SPCCA). Lei Gao 0001, Lin Qi 0001, Ling Guan |
ICIP | 3 |
| 2015 | A new audiovisual emotion recognition system using entropy-estimation-based multimodal information fusionabstractWe present a novel audiovisual emotion recognition solution using multimodal information fusion based on entropy estimation. Considering the limitations of existing methods, we propose a new dual-level fusion framework which consists of feature level fusion module based on kernel entropy component analysis and score level fusion module based on maximum correntropy criterion. In our system, audio and visual channels are utilized to detect and classify emotional states for intelligent human computer interfaces. Our extensive experimental study on eNTERFACE database and RML database demonstrates the feasibility of the proposed multimodal emotion recognition framework based on integrated analysis of speech and facial expression. The experimental results show that the proposed methods are capable of providing improved performance. The comparison with other methods shows that the proposed two-stage fusion platform outperforms the traditional algorithms in terms of both accuracy and reliability. Zhibing Xie, Tie Yun, Ling Guan |
ISCAS | 3 |
| 2015 | Two-dimensional discriminant multi-manifolds locality preserving projection for facial expression recognitionabstractIn this paper, we assume that samples of different expressions reside on different manifolds and propose a novel human emotion recognition framework named two-dimensional discriminant multi-manifolds locality preserving projection (2D-DMLPP). 2D-DMLPP focuses on salient regions which reflect the significant variation from facial expression images so that it can learn an expression-specific model from salient patches rather than that of subject-specific. Furthermore, conventional manifold learning methods ignore the variation among nearby samples from the same class, leading to serious overfitting. We construct three adjacency graphs to model the margin and information, including diversity and similarity of salient patches from the same expression, and then incorporate the information and margin into dimensionality reduction function. Several experiments show that the proposed method significantly improves the recognition performance of facial expression recognition. Ning Zheng 0003, Xin Guo 0005, Lin Qi 0001, Ling Guan |
ISCAS | 4 |
| 2015 | Human Action Recognition Using Hybrid Centroid Canonical Correlation AnalysisabstractHuman action recognition is a hot research topic in image analysis and computer vision. In this paper, we propose Hybrid Centroid Canonical Correlation Analysis (HCCCA) and multi-set HCCCA for multimodal information analysis and fusion. Furthermore, we present a novel human action recognition framework by using multi-set HCCCA to fuse multimodal features, which include the hierarchal pyramid Depth Motion Map (DMM) for the depth images, the Histogram of Oriented Displacement (HOD) for the skeleton, and the statistical measurements for the accelerometer. The proposed framework was evaluated using two datasets MSR Action 3D dataset and UTD multimodal human action dataset. The experimental results demonstrated that the proposed framework can achieve a higher average accuracy compared to several existing methods. Nour El-Din El-Madany, Ling Guan |
ISM | 3 |
| 2015 | Information Fusion of Audio Emotion Recognition Based on Kernel Entropy Component Analysis in Canonical Correlation SpaceabstractKernel Entropy Component Analysis(KECA), an effective information fusion tool, is realized using descriptor of information entropy and optimized by entropy estimation. However, it merely put the information or data from different channels together to achieve the information fusion without considering their intrinsic structures and relations. In this paper, we enhance the performance of KECA by introducing KECA in Canonical Correlation Space (CCS) or KECA+CCS. Not only the intrinsic structures and relations are considered in CCS, but also the nature of input data are revealed by entropy estimation. It improves the recognition accuracy effectively. The effectiveness of the proposed method is evaluated through experimentation on two audio-based emotion databases. The results show that the proposed method outperforms the existing methods based on similar principles. Lei Gao 0001, Lin Qi 0001, Ling Guan |
ISM | 3 |
| 2015 | A Novel Semi-Supervised Dimensionality Reduction Framework for Multi-manifold LearningabstractIn pattern recognition, traditional single manifold assumption can hardly guarantee the best classification performance, since the data from multiple classes does not lie on a single manifold. When the dataset contains multiple classes and the structure of the classes are different, it is more reasonable to assume each class lies on a particular manifold. In this paper, we propose a novel framework of semi-supervised dimensionality reduction for multi-manifold learning. Within this framework, methods are derived to learn multiple manifold corresponding to multiple classes in a data set, including both the labeled and unlabeled examples. In order to connect each unlabeled point to the other points from the same manifold, a similarity graph construction, based on sparse manifold clustering, is introduced when constructing the neighbourhood graph. Experimental results verify the advantages and effectiveness of this new framework. Xin Guo 0005, Tie Yun, Lin Qi 0001, Ling Guan |
ISM | 4 |
| 2015 | Human action recognition using temporal hierarchical pyramid of depth motion map and KECAabstractHuman action recognition is one of the challenging research problems in computer vision. In this paper, we propose a novel approach for human action recognition. The proposed approach employs a temporal hierarchical pyramid of depth motion map to capture the temporal variations over the time. In addition, Kernel Entropy Component Analysis (KECA) is used to reduce the dimension and to enhance the discriminatory power for action recognition. The proposed method was evaluated using two datasets, MSR-Action 3D dataset and MSR-Gesture 3D dataset. The experimental results demonstrated that the proposed method can achieve a higher average accuracy compared to several existing methods. Nour El-Din El-Madany, Ling Guan |
MMSP | 3 |
| 2015 | An improved ICP registration algorithm with a weight-bootstrap schemeabstractIn this paper, we propose a variant of the iterative closet point (ICP) registration method by introducing a novel weight-bootstrap scheme. The rigid transform between two point sets can be estimated, as long as proper correspondences are given. The accuracy of the estimated transform is mainly determined by the goodness of these matches. However, the confidence of the correspondences is weakened by the observation error. Aiming to address this challenge, the proposed solution parameterizes the pair matching confidence and improves the optimization process. Specifically, we adopt the tangent distance as error metric and introduce an iterative bootstrapped quadratic approximation method to increase the registration accuracy. Compared to the existing methods, experiments show that our method can produce a more accurate transform estimation in 2D case, and yield an acceptable result in 3D case. Fei Guo 0002, Ling Guan |
MMSP | 3 |
| 2015 | Action recognition using multi-layer Depth Motion maps and Sparse Dictionary LearningabstractIn this paper, we propose a new spatio-temporal feature based method for human action recognition using depth image sequence. Fist, Layered Depth Motion maps (LDM) are utilized to capture the temporal motion feature. Next, multi-scale HOG descriptors are computed on LDM to characterize the structural information of actions. Then sparse coding is applied for feature representation. Extending Sparse fisher Discriminative Dictionary Learning (SDDL) model and its corresponding classification scheme are also introduced. In SDDL model, the sub-dictionary is updated class by class, leading to class-specific compact discriminative dictionaries. The proposed method is evaluated on public MSR Action3D datasets and demonstrates great performance, especially in cross subject test. Chengwu Liang, Enqing Chen, Lin Qi 0001, Ling Guan |
MMSP | 4 |
| 2015 | Spatio-Temporal Pyramid Model based on depth maps for action recognitionabstractThis paper presents a novel human action recognition method by using depth maps. Each depth frame in a depth video sequence is projected onto three orthogonal Cartesian planes. Under each projection view, we divide the entire depth maps into several sub-actions. The absolute difference between two consecutive projected maps is accumulated through a depth video (several sub-actions) sequence to form a Depth Motion Map (DMM) to describe the dynamic feature of an action. Also the difference within the threshold between two consecutive projected maps is calculated through the entire depth video to form another kind of Depth Static Map (DSM) to describe the static feature. Collectively, we call them Temporal Pyramid of Depth Model (TPDM). Then Spatial Pyramid Histograms of Oriented Gradient (SPHOG) is computed from the TPDM for the representation of an action. For classification, we apply support vector machine (SVM) to classify the proposed descriptorsbased on MSR Action3D dataset. Experimental results demonstrates the effectiveness of our proposed method. Haining Xu, Enqing Chen, Chengwu Liang, Lin Qi 0001, Ling Guan |
MMSP | 5 |
| 2015 | Intuitive volume exploration through spherical self-organizing map and color harmonization
Naimul Mefraz Khan, Matthew J. Kyan, Ling Guan |
Neurocomputing | 3 |
| 2015 | TapTell: Interactive visual search for mobile task recommendation
Ning Zhang 0023, Tao Mei 0001, Xian-Sheng Hua 0001, Ling Guan, Shipeng Li 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2015 | An Approach to Ballet Dance Training through MS Kinect and Visualization in a CAVE Virtual Reality EnvironmentabstractThis article proposes a novel framework for the real-time capture, assessment, and visualization of ballet dance movements as performed by a student in an instructional, virtual reality (VR) setting. The acquisition of human movement data is facilitated by skeletal joint tracking captured using the popular Microsoft (MS) Kinect camera system, while instruction and performance evaluation are provided in the form of 3D visualizations and feedback through a CAVE virtual environment, in which the student is fully immersed. The proposed framework is based on the unsupervised parsing of ballet dance movement into a structured posture space using the spherical self-organizing map (SSOM). A unique feature descriptor is proposed to more appropriately reflect the subtleties of ballet dance movements, which are represented as gesture trajectories through posture space on the SSOM. This recognition subsystem is used to identify the category of movement the student is attempting when prompted (by a virtual instructor) to perform a particular dance sequence. The dance sequence is then segmented and cross-referenced against a library of gestural components performed by the teacher. This facilitates alignment and score-based assessment of individual movements within the context of the dance sequence. An immersive interface enables the student to review his or her performance from a number of vantage points, each providing a unique perspective and spatial context suggestive of how the student might make improvements in training. An evaluation of the recognition and virtual feedback systems is presented. Matthew J. Kyan, Guoyu Sun, Paisarn Muneesawang, Nan Dong, Bruce Elder, Ling Guan |
ACM Trans. Intell. Syst. Technol. | 8 |
| 2014 | Volume visualization using sparse nonparametric support vector machines and harmoniccolorsabstractIn Direct Volume Rendering (DVR), the Transfer Function (TF) to map voxel values to color and opacity values is difficult to obtain. Existing TF design tools are complex and non-intuitive for the end user, who is more likely to be a medical professional than an expert in image processing. In this paper, we propose a volume visualization method where the user directly works on the volume data to simply select the parts he/she would like to visualize. The user's work is further simplified by presenting only the most informative volume slices for selection. Based on the selected parts, all the voxels are classified using our Sparse Nonparametric Support Vector Machine (SN-SVM) classifier, which combines both local and near-global distributional information of the training data to obtain accurate results. The voxel classes are then mapped to color and opacity values using the concept of harmonic colors, which provides easily distinguishable and aesthetically pleasing results. Experimental results on several benchmark datasets show the effectiveness of the proposed method. Naimul Mefraz Khan, Riadh Ksantini, Ling Guan |
ICASSP | 3 |
| 2014 | Towards optimal resource allocation for differentiated multimedia services in cloud computing environmentabstractCloud-based multimedia services have been widely used in recent years. As the growing scale, users often have quite diverse quality of service (QoS) expectations. A key challenge for differentiated services is how to optimally allocate cloud resources to satisfy different users. In this paper, we study resource allocation problems for differentiated multimedia services. We first propose a queueing model to characterize differentiated services in cloud. Based on the model, we optimize cloud resources in the first-come first-served (FCFS) scenario and priority scenario. In each scenario, we formulate and solve the optimal resource allocation problem to minimize resource cost under response time constraints. We conduct extensive simulations with practical parameters of Amazon EC2. Simulation results demonstrate that the proposed resource allocation schemes can optimally configure resources to provide satisfactory services at the minimal resource cost. Xiaoming Nan, Ling Guan |
ICASSP | 3 |
| 2014 | Towards dynamic resource optimization for cloud-based free viewpoint video serviceabstractRecent years have witnessed the fast development of cloud-based media services. For service providers (SP), a key challenge is how to provide satisfactory services at a low cost. In this paper, we study the resource optimization for cloud-based free viewpoint video (FVV) service. We propose a Two-time-scale Resource Configuration (TRC) model. Specifically, cloud resources are allocated in a mid-long time scale and dynamically reconfigured in a fine-grained time scale. Based on the model, we formulate and solve the resource cost minimization problem and the response time minimization problem, respectively. We have implemented a prototype of cloud-based FVV on a cluster of machines. Experimental results demonstrate that the proposed resource optimization schemes not only allocate resources effectively to achieve the low cost, but also reconfigure resources dynamically to obtain the low response time. Xiaoming Nan, Ling Guan |
ICIP | 3 |
| 2014 | A novel cloud gaming framework using joint video and graphics streamingabstractAs the popularity of smart phones and tablets, users have an increasing desire to enjoy ubiquitous game playing. The emerging cloud gaming turns this desire into reality, enabling users to play games at anywhere on any devices. However, due to the huge amount of data transmission, it is challenging to provide a high quality game experience under the limited bandwidth capacity. In this paper, we propose a novel cloud gaming framework, in which we introduce two synchronized graphics buffers at both the server and the client sides. The server not only streams the compressed frames captured from game scenes, but also progressively transmits graphics data. The received graphics data is used to generate reference frames. When compressing the next frame, the cloud server will choose the reference frame with a lower residual error, from the previous frame and the current frame rendered from the graphics buffer. With the accumulation of graphics data, the frame rendered from the graphics buffer is close to the captured frame, which greatly reduces the transmission bit rates. Based on the proposed framework, we study the rate allocation problem, in which we optimize the allocated bit rates between the compressed frame and the graphics data to minimize the total distortion under the bandwidth constraint. Experimental results demonstrate that the proposed framework can optimally allocate bit rates to achieve a minimal distortion for cloud gaming compared to the traditional video streaming and graphics streaming approaches. Xiaoming Nan, Xun Guo 0002, Yan Lu 0001, Ling Guan, Shipeng Li 0001, Baining Guo |
ICME | 5 |
| 2014 | Incremental GMMSD2 with applications to feature extractionabstractThe generalized MMSD (GMMSD) is considered an efficient implementation of MMSD to extract discriminative information. However, a significant issue with the implementation of GMMSD is the complete recomputation of the training process when new training samples are presented. In this paper, we propose an alternative solution for feature extraction using the principles of GMMSD, which we call GMMSD2. GMMSD2 only requires the computation of centroid matrix, and it can overcome computational cost by applying efficient QR-updating techniques when new training samples are presented. Our experiments on FERET database demonstrate that incremental version of GMMSD2 eliminates the complete recomputation of the training process when new training samples are available, leading to significantly reduced computational cost. Ning Zheng 0003, Lin Qi 0001, Ling Guan |
ISCAS | 3 |
| 2014 | A Visual Evaluation Framework for In-Home Physical RehabilitationabstractWe propose a novel method for in-home physical rehabilitation, where a user can visually evaluate his/her performance compared to that of an expert. Normalized joint coordinates extracted from the Kinect skeleton are used as features. A novel Incremental Dynamic Time Warping (IDTW) algorithm is used to align the user and expert sequences. IDTW extends the classic DTW by providing accurate comparison between incomplete (the user's) and complete (the expert's) sequences while significantly reducing the computational time. Instead of providing a single measurement, the proposed method maps the IDTW measurements to a color-coded skeleton frame. Different colors on the limbs provide the user with an easy-to-interpret evaluation of how he or she is performing. Preliminary analysis involving different users and exercises and comparisons against the classic DTW algorithm show the effectiveness of the proposed method. Naimul Mefraz Khan, Stephen Lin 0001, Ling Guan, Baining Guo |
ISM | 3 |
| 2014 | An Advanced Computational Intelligence System for Training of Ballet Dance in a Cave Virtual Reality EnvironmentabstractThis paper presents a computer-based system for assessment and training of ballet dance in a CAVE virtual reality environment. The system utilizes Kinect sensor to capture student's dance and extracts features from skeleton joints. This system depends on a structured posture space, which comprises a set of dance elements that represent key moments -- "postures", that typically will be so briefly held as to experience as a fleeting moment in a flux -- in the dance movements whose performance we are attempting to assess. The recording captured from the Kinect allows the parsing of dance movement into a structured posture space using the spherical self-organizing map (SSOM). From this, a unique descriptor can be obtained by following gesture trajectories through posture space on the SSOM, which appropriately reflects the subtleties of ballet dance movements. Consequently, the system can recognize the category of movement the student is attempting, and this allows us make a quantitative assessment of individual movements. Based on the experimental results, the proposed system appears to be very effective for recognition and offering generalization across instances of movement. Thus, it is possible for the construction of assessment and visualization of ballet dance movements performed by the student in an instructional, virtual reality setting. Guoyu Sun, Paisarn Muneesawang, Matthew J. Kyan, Nan Dong, Bruce Elder, Ling Guan |
ISM | 8 |
| 2014 | Stereo correspondence using an assisted discrete cosine transform methodabstractIn this paper, a stereo matching algorithm using a window based frequency comparison method is formulated. The algorithm works with a local matching stereo model where a normalized cost function between frequency components and intensity values is used. The algorithm determines matching points in a stereo pair and uses a weighted cost function to determine the true disparity of the stereo pair. Unlike classical stereo correspondence algorithms that determine initial disparity maps through window based color intensity comparisons, the proposed algorithm uses window based frequency comparisons to exemplify the ability of frequency components to accurately find high detailed segments of the image. The algorithm is evaluated on the Middlebury data sets, and shows that it is noise and distortion resistant similar to the work in [1], thus allowing for higher reliability during comparisons. Additionally, this provides an advantage over typical color intensity comparisons as noise present in an image may cause mismatching when color intensity comparisons are executed. Edward Rosales, Ling Guan |
VCIP | 2 |
| 2014 | Queueing model based resource optimization for multimedia cloud
Xiaoming Nan, Ling Guan |
J. Vis. Commun. Image Represent. | 3 |
| 2014 | Generalized multiple maximum scatter difference feature extraction using QR decomposition
Ning Zheng 0003, Lin Qi 0001, Ling Guan |
J. Vis. Commun. Image Represent. | 3 |
| 2014 | Covariance-guided One-Class Support Vector Machine
Naimul Mefraz Khan, Riadh Ksantini, Imran Ahmad 0001, Ling Guan |
Pattern Recognit. | 4 |
| 2013 | Incorporating covariance information in one class support vector classificationabstractUnlike multi-class problems, the low variance directions in the training data are important for one-class classification. However, projecting in these directions before classification will result in loss of important data properties. This paper introduces a Covariance-guided One-Class Support Vector Machine (COSVM) classification method which emphasizes the low variance projectional directions of the training data without compromising any important characteristics. COSVM combines the global information from the covariance matrix of the training data with the local information of Support Vectors. Our proposed method is a convex optimization problem resulting in one global solution, which can be found efficiently with the help of existing numerical methods. The method also keeps the principal structure of the OSVM method intact, and can be implemented easily with the existing OSVM applications. Comparative experimental results with contemporary one-class classifiers on numerous benchmark datasets verify that our method results in significantly better performance. Naimul Mefraz Khan, Riadh Ksantini, Imran Ahmad 0001, Ling Guan |
ICASSP | 4 |
| 2013 | Optimal task-level scheduling for cloud based multimedia applicationsabstractAs an emerging computing paradigm, cloud computing has been increasingly used in multimedia applications. One fundamental challenge for application providers is how to effectively schedule multimedia tasks to multiple virtual machines for distributed processing. In this paper, we study task-level scheduling problem for cloud based multimedia applications. Specifically, we introduce a directed acyclic graph to model precedence constraints among tasks. Based on the model, we study the optimal task scheduling problem for the sequential, the parallel, and the mixed structures, respectively. Moreover, we propose a heuristic to perform the near optimal task scheduling in a practical way. Experimental results demonstrate that the proposed scheduling scheme can optimally assign tasks to virtual machines to minimize the execution time. Xiaoming Nan, Ling Guan |
ICASSP | 3 |
| 2013 | Multimodal information fusion of audiovisual emotion recognition using novel information theoretic toolsabstractThis paper aims at providing general theoretical analysis for the issue of multimodal information fusion and implementing novel information theoretic tools in multimedia application. The most essential issues for information fusion include feature transformation and reduction of feature dimensionality. Most previous solutions are based on the second order statistics, which is only optimal for Gaussian-like distribution, while in this paper we describe kernel entropy component analysis (KECA) which utilizes descriptor of information entropy and achieves improved performance by entropy estimation. We present a new solution based on the integration of information fusion theory and information theoretic tools in this paper. The proposed method has been applied to audiovisual emotion recognition. Information fusion has been implemented for audio and video channels at feature level and decision level. Experimental results demonstrate that the proposed algorithm achieves improved performance in comparison with the existing methods, especially when the dimension of feature space is substantially reduced. Zhibing Xie, Ling Guan |
ICME | 2 |
| 2013 | Human emotion recognition using the adaptive sub-layer-compensation based facial edge detectionabstractIn this paper, we propose an adaptive sub-layer compensation (ASLC) based facial edge detection for human emotion recognition. We modify the Marr-Hildreth edge detector with Wiener filtering, sub-layer compensation and hysteresis analysis to compensate the negative effects of the Laplacian of Gaussian (LoG) operator such as image degradation, high response to unwanted details, and disconnected edges. We investigate a Discriminative Isomap (D-Isomap) based approach that combines the ASLC feature and deformable elastic body spline (EBS) feature for the final decision. RML emotion database and Cohn-Kanade database are used for the experiment and the results demonstrate the effectiveness of the proposed method. Tie Yun, Anastasios N. Venetsanopoulos, Ling Guan |
ISCAS | 4 |
| 2013 | Optimal resource allocation for multimedia application providers in multi-site cloudabstractCloud-based multimedia applications have been widely used in recent years. With the advance of globalization, application providers hope to offer services at multiple sites serving users all over the world. The key challenge is how to effectively manage resource allocation and workload balancing. In this paper, we study the optimal resource allocation problem in multi-site cloud. Specifically, we jointly optimize the global workload assignment and the local VM allocation to minimize the resource cost under the service response time requirements. Moreover, we propose a greedy algorithm to efficiently apply our study in a practical way. Experimental results demonstrate that the proposed optimal resource allocation scheme can optimally allocate resources and balance workload to achieve the minimal resource cost for multimedia application providers. Xiaoming Nan, Ling Guan |
ISCAS | 3 |
| 2013 | Optimization of workload scheduling for multimedia cloud computingabstractThe cloud based multimedia applications have been widely adopted in recent years. Due to the large-scale and time-varying workload, an effective workload scheduling scheme is becoming a challenge faced by multimedia application providers. In this paper, we study the workload scheduling schemes for multimedia cloud. Specifically, we examine and solve the response time minimization problem and the resource cost minimization problem, respectively. Moreover, we propose a greedy algorithm to efficiently schedule workload for practical multimedia cloud. Simulation results demonstrate that the proposed workload scheduling schemes can optimally balance workload to achieve the minimal response time or the minimal resource cost for multimedia application providers. Xiaoming Nan, Ling Guan |
ISCAS | 3 |
| 2013 | Human emotional state recognition using real 3D visual features from Gabor library
Tie Yun, Ling Guan |
Pattern Recognit. | 2 |
| 2013 | A Deformable 3-D Facial Expression Model for Dynamic Human Emotional State RecognitionabstractAutomatic emotion recognition from facial expression is one of the most intensively researched topics in affective computing and human-computer interaction. However, it is well known that due to the lack of 3-D feature and dynamic analysis the functional aspect of affective computing is insufficient for natural interaction. In this paper, we present an automatic emotion recognition approach from video sequences based on a fiducial point controlled 3-D facial model. The facial region is first detected with local normalization in the input frames. The 26 fiducial points are then located on the facial region and tracked through the video sequences by multiple particle filters. Depending on the displacement of the fiducial points, they may be used as landmarked control points to synthesize the input emotional expressions on a generic mesh model. As a physics-based transformation, elastic body spline technology is introduced to the facial mesh to generate a smooth warp that reflects the control point correspondences. This also extracts the deformation feature from the realistic emotional expressions. Discriminative Isomap-based classification is used to embed the deformation feature into a low dimensional manifold that spans in an expression space with one neutral and six emotion class centers. The final decision is made by computing the nearest class center of the feature space. Tie Yun, Ling Guan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2013 | Monocular Human Motion Tracking by Using DE-MC Particle FilterabstractTracking human motion from monocular video sequences has attracted significantly increased interests in recent years. A key to accomplishing this task is to efficiently explore a high-dimensional state space. However, the traditional particle filter method and many of its variants have not been able to meet expectations as they lack a strategy to do efficiently sampling or stochastic search. We present a novel approach, namely differential evolution-Markov chain (DE-MC) particle filtering. By taking the advantage of the DE-MC algorithm's ability to approximate complicated distributions, substantial improvement can be made to the traditional structure of the particle filter. As a result, an efficient stochastic search can be performed to locate the modes of likelihoods. Furthermore, we apply the proposed algorithm to solve the 3D articulated model-based human motion tracking problem. A reliable image likelihood function is built for visual tracker design. Based on the proposed DE-MC particle filter and the image likelihood function, we perform a variety of monocular human motion tracking experiments. Experimental results, including the comparison with the performance of other particle filtering methods demonstrate the reliable tracking performance of the proposed approach. Xiaoming Nan, Ling Guan |
IEEE Trans. Image Process. | 3 |
| 2012 | A Sparse Support Vector Machine Classifier with Nonparametric Discriminants
Naimul Mefraz Khan, Riadh Ksantini, Imran Ahmad 0001, Ling Guan |
ICANN (2) | 4 |
| 2012 | Optimal resource allocation for multimedia cloud in priority service schemeabstractMultimedia cloud is an emerging computing paradigm that can effectively process multimedia applications and provide multi-QoS provisions for customers. Two major challenges exist in multimedia cloud: the resource cost and the service response time. In this paper, we employ the proposed queuing model to optimize the resource allocation for multimedia cloud in priority service scheme. Specifically, we formulate and solve the resource cost minimization problem and the service response time minimization problem respectively. The simulation results demonstrate that the proposed optimal resource allocation method can greatly enhance the performance of multimedia cloud data center in terms of resource cost and service response time. Xiaoming Nan, Ling Guan |
ISCAS | 3 |
| 2012 | Human emotion recognition using a deformable 3D facial expression modelabstractAutomatic emotion recognition from facial expression is one of the most intensively researched topics in affective computing and human-computer interaction. However, due to the lack of 3D feature and dynamic analysis the functional aspect of affective computing is insufficient for natural interaction. This paper presents an automatic emotion recognition approach from video sequences based on a fiducial point controlled 3D facial model. As a physics-based transformation, elastic body spline technology is applied on a facial mesh to generate a smooth warp that reflects the control point corresponding to the displacement of fiducial points. It also extracts the deformation feature from the realistic emotional expressions. Discriminative Isomap based classification is used to embed the deformation feature into a low dimensional manifold that spans in an expression space with one neutral and six emotion class centers. The final decision is made by computing the Nearest Class Center of the feature space. Tie Yun, Ling Guan |
ISCAS | 2 |
| 2012 | Discriminative Multiple Canonical Correlation Analysis for Multi-feature Information FusionabstractThis paper presents a novel approach for multi-feature information fusion. The proposed method is based on the Discriminative Multiple Canonical Correlation Analysis (DMCCA), which can extract more discriminative characteristics for recognition from multi-feature information representation. It represents the different patterns among multiple subsets of features identified by minimizing the Frobenius norm. We will demonstrate that the Canonical Correlation Analysis (CCA), the Multiple Canonical Correlation Analysis (MCCA), and the Discriminative Canonical Correlation Analysis (DCCA) are special cases of the DMCCA. The effectiveness of the DMCCA is demonstrated through experimentation in speaker recognition and speech-based emotion recognition. Experimental results show that the proposed approach outperforms the traditional methods of serial fusion, CCA, MCCA and DCCA. Lei Gao 0001, Lin Qi 0001, Enqing Chen, Ling Guan |
ISM | 4 |
| 2012 | 2D-FRFT Based Rotation Invariant Digital Image WatermarkingabstractThe extraction of rotation invariant representation is important for many signal processing problems such as image analysis, computer vision, and pattern recognition. In this paper, we present a systematic analysis of the Two-Dimensional Fractional Fourier Transform (2D-FRFT), and show that under certain conditions, the 2D-FRFT technique possesses the attractive property of rotation invariance. Based on our analysis, we proposed a novel digital image watermarking method which combines 2D chirp signal with the addition and rotation invariant properties of 2D-FRFT to achieve improved robustness and security. The effectiveness of the proposed solution is demonstrated through experiments. Lei Gao 0001, Lin Qi 0001, Shouyi Yang, Yongjin Wang, Tie Yun, Ling Guan |
ISM | 6 |
| 2012 | Multimodal Information Fusion of Audio Emotion Recognition Based on Kernel Entropy Component AnalysisabstractThis paper focuses on the application of novel information theoretic tools in the area of information fusion. Feature transformation and fusion is critical for the performance of information fusion, however the majority of the existing works depend on the second order statistics, which is only optimal for Gaussian-like distribution. In this paper, the integration of information fusion techniques and kernel entropy component analysis provides a new information theoretic tool. The fusion of features is realized using descriptor of information entropy and optimized by entropy estimation. A novel multimodal information fusion strategy of audio emotion recognition based on kernel entropy component analysis (KECA) has been presented. The effectiveness of the proposed solution is evaluated though experimentation on two audiovisual emotion databases. Experimental results show that the proposed solution outperforms the existing methods, especially when the dimension of feature space is substantially reduced. The proposed method offers general theoretical analysis which gives us an approach to implement information theory into multimedia research. Zhibing Xie, Ling Guan |
ISM | 2 |
| 2012 | Optimal allocation of virtual machines for cloud-based multimedia applicationsabstractWith the emergence of cloud computing, cloud-based multimedia applications have been increasingly adopted in recent years. There are two major challenges for multimedia application providers: the round trip time (RTT) requirement and the resource cost. In this paper, we study the virtual machine (VM) allocation problem for multimedia application providers to minimize the resource cost under RTT requirements. Specifically, we propose the optimal VM allocation schemes for single-site cloud and multi-site cloud, respectively. Moreover, we propose the greedy algorithms to efficiently allocate VMs in each case. Simulation results demonstrate that the proposed optimal VM allocation schemes can optimally allocate VMs to achieve a minimal resource cost. Xiaoming Nan, Ling Guan |
MMSP | 3 |
| 2012 | Interactive mobile visual search for social activities completion using query image contextual modelabstractMobile devices are ubiquitous. People use their phones as a personal concierge not only discovering information but also searching for particular interest on-the-go and making decisions. This brings a new horizon for multimedia retrieval on mobile. While existing efforts have predominantly focused on understanding textual or a voice query, this paper presents a new perspective which understands visual queries captured by the built-in camera such that mobile-based social activities can be recommended for users to complete. In this work, a query image-based contextual model is proposed for visual search. A mobile user can take a photo and naturally indicate an object-of-interest within the photo via circle based gesture called “O” gesture. Both selected object-of-interest region as well as surrounding visual context in photo are used in achieving a search-based recognition by retrieving similar images based on a large-scale of visual vocabulary tree. Consequently, social activities such as visiting contextually relevant entities (i.e., local businesses) are recommended to the users based on their visual queries and GPS location. Along with the proposed method, an exemplary real application has been developed on Windows Phone 7 devices and evaluated with a wide variety of scenarios on million-scale image database. To test the performance of proposed mobile visual search model, extensive experimentation has been conducted and compared with state-of-the-art algorithms in content-based image retrieval (CBIR) domain. Ning Zhang 0023, Tao Mei 0001, Xian-Sheng Hua 0001, Ling Guan, Shipeng Li 0001 |
MMSP | 4 |
| 2012 | Optimization of resource reconfiguration for cloud-based multimedia applicationsabstractThe cloud-based multimedia applications have been increasingly adopted in recent years. The key challenge for the application providers is how to optimally reconfigure resources to cope with the time-varying workload. In this paper, we study the optimal resource reconfiguration for cloud-based multimedia applications to minimize the average round-trip-time (RTT). Specifically, we propose the optimal resource reconfiguration schemes for single-site cloud and multi-site cloud, respectively. In each case, we formulate and solve the average RTT minimization problem. Simulation results demonstrate that the proposed optimal resource reconfiguration schemes can optimally utilize cloud resources to achieve the minimal average RTT. Xiaoming Nan, Ling Guan |
VCIP | 3 |
| 2012 | Generalized MMSD feature extraction using QR decompositionabstractMultiple Maximum scatter difference (MMSD) discriminant criterion is an effective feature extraction method that computes the discriminant vectors from both the range of the between-class scatter matrix and the null space of the within-class scatter matrix. However, singular value decomposition (SVD) of two times is involved in MMSD, making this method impractical for high dimensional data. In this paper, we propose a novel method for feature extraction and classification based on MMSD criterion, called generalized MMSD (GMMSD), which employs QR decomposition rather than SVD. Unlike MMSD, GMMSD does not require the computation of the whole scatter matrix. Instead, it computes the discriminant vectors from both the range of whitenizated input data matrix and the null space of the within-class scatter matrix. We evaluate the effectiveness of the GMMSD method in terms of classification accuracy in the reduced dimensional space. Our experiments on two facial expression databases demonstrate that the GMMSD method provides favorable performance in terms of both recognition accuracy and computational efficiency. Ning Zheng 0003, Lin Qi 0001, Lei Gao 0001, Ling Guan |
VCIP | 4 |
| 2012 | A hybrid approach for personalized recommendation of news on the Web
Liping Fang, Ling Guan |
Expert Syst. Appl. | 3 |
| 2012 | A Generic Approach for Systematic Analysis of Sports VideosabstractVarious innovative and original works have been applied and proposed in the field of sports video analysis. However, individual works have focused on sophisticated methodologies with particular sport types and there has been a lack of scalable and holistic frameworks in this field. This article proposes a solution and presents a systematic and generic approach which is experimented on a relatively large-scale sports consortia. The system aims at the event detection scenario of an input video with an orderly sequential process. Initially, domain knowledge-independent local descriptors are extracted homogeneously from the input video sequence. Then the video representation is created by adopting a bag-of-visual-words (BoW) model. The video’s genre is first identified by applying the k-nearest neighbor (k-NN) classifiers on the initially obtained video representation, and various dissimilarity measures are assessed and evaluated analytically. Subsequently, an unsupervised probabilistic latent semantic analysis (PLSA)-based approach is employed at the same histogram-based video representation, characterizing each frame of video sequence into one of four view groups, namely closed-up-view, mid-view, long-view, and outer-field-view. Finally, a hidden conditional random field (HCRF) structured prediction model is utilized for interesting event detection. From experimental results, k-NN classifier using KL-divergence measurement demonstrates the best accuracy at 82.16% for genre categorization. Supervised SVM and unsupervised PLSA have average classification accuracies at 82.86% and 68.13%, respectively. The HCRF model achieves 92.31% accuracy using the unsupervised PLSA based label input, which is comparable with the supervised SVM based input at an accuracy of 93.08%. In general, such a systematic approach can be widely applied in processing massive videos generically. Ning Zhang 0023, Ling-Yu Duan, Lingfang Li, Qingming Huang, Wen Gao 0001, Ling Guan |
ACM Trans. Intell. Syst. Technol. | 7 |
| 2012 | Kernel Cross-Modal Factor Analysis for Information Fusion With Application to Bimodal Emotion RecognitionabstractIn this paper, we investigate kernel based methods for multimodal information analysis and fusion. We introduce a novel approach, kernel cross-modal factor analysis, which identifies the optimal transformations that are capable of representing the coupled patterns between two different subsets of features by minimizing the Frobenius norm in the transformed domain. The kernel trick is utilized for modeling the nonlinear relationship between two multidimensional variables. We examine and compare with kernel canonical correlation analysis which finds projection directions that maximize the correlation between two modalities, and kernel matrix fusion which integrates the kernel matrices of respective modalities through algebraic operations. The performance of the introduced method is evaluated on an audiovisual based bimodal emotion recognition problem. We first perform feature extraction from the audio and visual channels respectively. The presented approaches are then utilized to analyze the cross-modal relationship between audio and visual features. A hidden Markov model is subsequently applied for characterizing the statistical dependence across successive time segments, and identifying the inherent temporal structure of the features in the transformed domain. The effectiveness of the proposed solution is demonstrated through extensive experimentation. Yongjin Wang, Ling Guan, Anastasios N. Venetsanopoulos |
IEEE Trans. Multim. | 2 |
| 2011 | Robust Signal Generation and Analysis of Rat Embryonic Heart Rate in Vitro Using Laplacian Eigenmaps and Empirical Mode Decomposition
M. Khalid Khan, Muhammad Talal Ibrahim, Mats F. Nilsson, Anna-Carin Sköld, Ling Guan, Ingela Nyström |
CAIP (2) | 5 |
| 2011 | Kernel cross-modal factor analysis for multimodal information fusionabstractThis paper presents a novel approach for multimodal information fusion. The proposed method is based on kernel cross-modal factor analysis (KCFA), in which the optimal transformations that represent the coupled patterns between two different subsets of features are identified by minimizing the Frobenius norm in the transformed domain. It generalizes the linear cross-modal factor analysis (CFA) method via the kernel trick to model the nonlinear relationship between two multidimensional variables. The effectiveness of the introduced solution is demonstrated through experimentation on an audiovisual based emotion recognition problem. Experimental results show that the proposed approach outperforms the concatenation based feature level fusion, the linear CFA, as well as the canonical correlation analysis (CCA) and kernel CCA methods. Yongjin Wang, Ling Guan, Anastasios N. Venetsanopoulos |
ICASSP | 2 |
| 2011 | Optimal source rate allocation in Body Sensor Networks with energy harvestingabstractBody Sensor Networks (BSNs) are promising in pervasive health monitoring applications. One of the major challenges in BSNs is the sustainable power supply since each body sensor has limited battery capacity. In this paper, we optimize the source rate of the sensor to provide an uninterrupted service for BSNs with energy harvesting. First, we employ a discrete-time Markov chain to model the energy harvesting process at each sensor, and then theoretically analyze the relationship between the source rate and the uninterrupted lifetime of the sensor. Second, we formulate a steady-rate optimization problem, which minimizes the rate fluctuation with respect to the average sustainable rate via optimal source rate allocation at each sensor, under the requirement of the uninterrupted service. We propose an analytical solution to solve the optimization problem. In the simulations, we demonstrate that the proposed optimal solution enables the BSN to maintain an uninterrupted service with a steady output rate at each sensor. Wenwu Zhu 0001, Ling Guan |
ICME | 3 |
| 2011 | Optimal resource allocation to provide QOS guarantee in pervasive health monitoring systemsabstractPervasive health monitoring is an eHealth service, which plays an important role in prevention and early detection of diseases. In health monitoring systems, a loss or an excessive delay of the critical data may cause a fatal accident. Therefore it is important to provide Quality of Service (QoS) guarantee for the delivery of data streams. In this paper, we formulate the QoS optimization problem, which jointly optimizes the transmission power and the transmission rate at each aggregator to provide QoS guarantee to data delivery. The optimization problem is converted into a convex optimization problem, which can be solved efficiently. We demonstrate in the simulations that the proposed optimized scheme improves the service quality of the health monitoring system in terms of delay and Packet Loss Rate (PLR). Wenwu Zhu 0001, Ling Guan |
ICME | 3 |
| 2011 | Audiovisual emotion recognition via cross-modal association in kernel spaceabstractIn this paper, we introduce a new method for audiovisual based multimodal emotion recognition. The proposed method identifies the optimal transformations that are capable of representing the coupled patterns between audio and visual information through cross-modal association. Specifically, kernel machine technique is utilized for capturing the nonlinear relationship between two different subsets of features. A hidden Markov model is subsequently applied for characterizing the statistical dependence across successive time segments, and identifying the inherent temporal structure of the features in the transformed domain. Information fusion at the feature and score levels are examined and compared. The effectiveness of the introduced solution is demonstrated through extensive experimentation. Yongjin Wang, Ling Guan, Anastasios N. Venetsanopoulos |
ICME | 2 |
| 2011 | TapTell: understanding visual intents on-the-goabstractThis demonstration presents a mobile-based visual recognition and recommendation application on Windows Phone 7 called TapTell. This is different from other mobile-based visual search mechanisms which merely focus on the search process. TapTell firstly discovers and understands users' visual intents via a circle based natural user interaction called "O" gestures. Following, a Tap action is operated to choose the "O" gestured regions. The context-aware visual search mechanism is utilized for recognizing the intents and associating them with indexed metadata. Finally, the "Tell" action recommends relevant entities utilizing contextual information. The TapTell system has been evaluated at different scenarios on million scale images. Ning Zhang 0023, Tao Mei 0001, Xian-Sheng Hua 0001, Ling Guan, Shipeng Li 0001 |
ACM Multimedia | 4 |
| 2011 | Multi-feature pLSA for combining visual features in image annotationabstractWe study in this paper the problem of combining low-level visual features for image region annotation. The problem is tackled with a novel method that combines texture and color features via a mixture model of their joint distribution. The structure of the presented model can be considered as an extension of the probabilistic latent semantic analysis (pLSA) in that it handles data from two different visual feature domains by attaching one more leaf node to the graphical structure of the original pLSA. Therefore, the proposed approach is referred to as multi-feature pLSA (MF-pLSA). The supervised paradigm is adopted to classify a new image region into one of a few pre-defined object categories using the MF-pLSA. To evaluate the performance, the VOC2009 and LabelMe databases were employed in our experiments, along with various experimental settings in terms of the number of visual words and mixture components. Evaluated based on the average recall and precision, the MF-pLSA is demonstrated superior to seven other approaches, including other schemes for visual feature combination. Rui Zhang 0010, Lei Zhang 0001, Xin-Jing Wang, Ling Guan |
ACM Multimedia | 4 |
| 2011 | Optimal resource allocation for multimedia cloud based on queuing modelabstractMultimedia cloud, as a specific cloud paradigm, addresses how cloud can effectively process multimedia services and provide QoS provisioning for multimedia applications. There are two major challenges in multimedia cloud. The first challenge is the service response time in multimedia cloud, and the second challenge is the cost of cloud resources. In this paper, we optimize resource allocation for multimedia cloud based on queuing model. Specifically, we optimize the resource allocation in both single-class service case and multiple-class service case. In each case, we formulate and solve the response time minimization problem and resource cost minimization problem, respectively. Simulation results demonstrate that the proposed optimal allocation scheme can optimally utilize the cloud resources to achieve a minimal mean response time or a minimal resource cost. Xiaoming Nan, Ling Guan |
MMSP | 3 |
| 2011 | Tap-to-search: Interactive and contextual visual search on mobile devicesabstractMobile visual search has been an emerging topic for both researching and industrial communities. Among various methods, visual search has its merit in providing an alternative solution, where text and voice searches are not applicable. This paper proposes an interactive “tap-to-search” approach utilizing both individual's intention in selecting interested regions via “tap” actions on the mobile touch screen, as well as a visual recognition by search mechanism in a large-scale image database. Automatic image segmentation technique is applied in order to provide region candidates. Visual vocabulary tree based search is adopted by incorporating rich contextual information which are collected from mobile sensors. The proposed approach has been conducted on an image dataset with the scale of two million. We demonstrated that using GPS contextual information, such an approach can further achieve satisfactory results with the standard information retrieval evaluation. Ning Zhang 0023, Tao Mei 0001, Xian-Sheng Hua 0001, Ling Guan, Shipeng Li 0001 |
MMSP | 4 |
| 2011 | Optimal Resource Allocation for Pervasive Health Monitoring Systems with Body Sensor NetworksabstractPervasive health monitoring is an eHealth service, which plays an important role in prevention and early detection of diseases. There are two major challenges in pervasive health monitoring systems with Body Sensor Networks (BSNs). The first challenge is the sustainable power supply for BSNs. The second challenge is Quality of Service (QoS) guarantee for the delivery of data streams. In this paper, we optimize the resource allocations to provide a sustainable and high-quality service in health monitoring systems. Specifically, we formulate and solve two resource optimization problems, respectively. In the first optimization problem, steady-rate optimization problem, we optimize the source rate at each sensor to minimize the rate fluctuation with respect to the average sustainable rate, subject to the requirement of uninterrupted service. The first optimization problem is solved by a proposed analytical solution. The second optimization problem is formulated based on the optimal source rates of the sensors obtained in the steady-rate optimization problem. In the second optimization problem, we jointly optimize the transmission power and the transmission rate at each aggregator to provide QoS guarantee to data delivery. The second optimization problem is converted into a convex optimization problem, which is then solved efficiently. In the simulations, we demonstrate that the proposed optimized scheme enables the pervasive health monitoring system to provide a sustainable service with guaranteed low delay and low Packet Loss Rate (PLR) to subscribers. Wenwu Zhu 0001, Ling Guan |
IEEE Trans. Mob. Comput. | 3 |
| 2010 | Fiducial point tracking for facial expression using multiple particle filters with kernel correlation analysisabstractDetecting and tracking fiducial points successfully can generate necessary dynamic and deformable information for facial image interpretation tasks with numerous potential applications. In this paper we propose an automatic fiducial points tracking method using multiple Differential Evolution Markov Chain (DE-MC) particle filters with kernel correlation techniques. Fiducial points are initialized through the scale invariant feature based detectors. By taking the advantage of the ability to approximate complicated proposal distributions, multiple DE-MC particle filters are applied for fiducial points tracking by building a path connecting sampling with measurements, based on the fact that the posteriori depends on both the previous state and the current observation. A Kernel correlation analysis approach is proposed to find the detection likelihood with maximization of the similarity criterion between the target points and the candidate points. Sampling efficiency is improved and computational time is substantially reduced by making use of the intermediate results obtained in particle allocation. Tie Yun, Ling Guan |
ICIP | 2 |
| 2010 | On-Line Signature Verification Using 1-D Velocity-Based Directional AnalysisabstractIn this paper, we propose a novel approach for identity verification based on the directional analysis of velocity-based partitions of an on-line signature. First, inter-feature dependencies in a signature are exploited by decomposing the shape (horizontal trajectory, vertical trajectory) into two partitions based on the velocity profile of the base-signature for each signer, which offers the flexibility of analyzing both low and high-curvature portions of the trajectory independently. Further, these velocity-based shape partitions are analyzed directionally on the basis of relative angles. Support Vector Machine (SVM) is then used to find the decision boundary between the genuine and forgery class. Experimental results demonstrate the superiority of our approach in on-line signature verification in comparison with other techniques. Muhammad Talal Ibrahim, Matthew J. Kyan, Muhammad A. Khan 0002, Ling Guan |
ICPR | 4 |
| 2010 | Streaming capacity in multi-channel P2P VoD systemsabstractPeer-to-Peer (P2P) Video-on-Demand (VoD) systems with multiple channels are called multi-channel P2P VoD systems, which can be categorized into independent-channel P2P VoD systems and correlated-channel P2P VoD systems. Most of the existing P2P VoD systems are independent-channel P2P systems, in which the peers share resources with each other within the same channel. In this paper, we examine the cross-channel resource sharing in correlated-channel P2P VoD systems. We optimize the server upload allocations among channels to maximize the average streaming capacity. Furthermore, we introduce bandwidth amplifiers to establish cross-channel links, thus enabling cross-channel sharing of peer upload bandwidths. We demonstrate in the simulations that the correlated-channel P2P VoD systems with cross-channel resource sharing can achieve a higher average streaming capacity compared to the independent-channel P2P VoD systems without cross-channel resource sharing. Ling Guan |
ISCAS | 2 |
| 2010 | Human emotion recognition using real 3D visual features from Gabor libraryabstractEmotional state recognition is an important component for efficient human-computer interaction. Most existing works address this problem using 2D features, but they are sensitive to head pose, clutter, and variations in lighting conditions. The general 3D based methods only consider geometric information for feature extraction. In this paper, we present a real 3D visual features based method for human emotion recognition. 3D geometric information plus colour/density information of the facial expressions are extracted by 3D Gabor library to construct visual feature vectors. The filter's scale, orientation, and shape of the library are specified according to the appearance patterns of the 3D facial expressions. An improved kernel canonical correlation analysis (IKCCA) algorithm is proposed for final decision. From training samples, the semantic ratings that describe the different facial expressions are computed by IKCCA to generate a seven dimensional semantic expression vector. It is applied for learning the correlation with different testing samples. According to this correlation, we estimate the associated expression vector and perform expression classification. From experiment results, our proposed method demonstrates impressive performance. Tie Yun, Ling Guan |
MMSP | 2 |
| 2010 | An efficient framework on large-scale video genre classificationabstractEfficient data mining and indexing is important for multimedia analysis and retrieval. In the field of large-scale video analysis, effective genre categorization plays an important role and serves one of the fundamental steps for data mining. Existing works utilize domain-knowledge dependent feature extraction, which is limited from genre diversification as well as data volume scalability. In this paper, we propose a systematic framework for automatically classifying video genres using domain-knowledge independent descriptors in feature extraction, and a bag-of-visualwords (BoW) based model in compact video representation. Scale invariant feature transform (SIFT) local descriptor accelerated by GPU hardware is adopted for feature extraction. BoW model with an innovative codebook generation using bottom-up two-layer K-means clustering is proposed to abstract the video characteristics. Besides the histogram-based distribution in summarizing video data, a modified latent Dirichlet allocation (mLDA) based distribution is also introduced. At the classification stage, a k-nearest neighbor (k-NN) classifier is employed. Compared with state of art large-scale genre categorization in, the experimental results on a 23-sports dataset demonstrate that our proposed framework achieves a comparable classification accuracy with 27% and 64% expansion in data volume and diversity, respectively. Ning Zhang 0023, Ling Guan |
MMSP | 2 |
| 2010 | A Bayesian image annotation framework integrating search and contextabstractConventional approaches to image annotation tackle the problem based on the low-level visual information. Considering the importance of the information on the constrained interaction among the objects in a real world scene, contextual information has been utilized to recognize scene and object categories. In this paper, we propose a Bayesian approach to region-based image annotation, which integrates the content-based search and context into a unified framework. The content-based search selects representative keywords by matching an unlabeled image with the labeled ones followed by a weighted keyword ranking, which are in turn used by the context model to calculate the a prior probabilities of the object categories. Finally, a Bayesian framework integrates the a priori probabilities and the visual properties of image regions. The framework was evaluated using two databases and several performance measures, which demonstrated its superiority to both visual content-based and context-based approaches. Rui Zhang 0010, Kui Wu 0002, Kim-Hui Yap, Ling Guan |
MMSP | 4 |
| 2010 | Automatic segmentation of pupil using local histogram and standard deviationabstractThis paper presents a novel approach for automatic pupil segmentation. The proposed algorithm uses local histogram and standard deviation based adaptive thresholding method that looks for the region that has the highest probability of having the pupil. We have tested our proposed algorithm on two public databases namely: CASIA v1.0 and MMU v1.0. Experimental results show that the proposed method has satisfying performance and good robustness against the reflection in the pupil. Muhammad Talal Ibrahim, Tariq Mahmood Khan, Muhammad A. Khan 0002, Ling Guan |
VCIP | 4 |
| 2010 | Velocity and pressure-based partitions of horizontal and vertical trajectories for on-line signature verification
Muhammad Talal Ibrahim, Muhammad A. Khan 0002, Khurram Saleem Alimgeer, M. Khalid Khan, Imtiaz A. Taj, Ling Guan |
Pattern Recognit. | 6 |
| 2010 | Solving Streaming Capacity Problems in P2P VoD SystemsabstractPeer-to-peer (P2P) video-on-demand (VoD) is a popular Internet service for a large number of concurrent users. Streaming capacity in a P2P VoD system is defined as the maximum streaming rate that can be received by every user. In this letter, we study the streaming capacity problem in P2P VoD systems. We formulate the streaming capacity problem into an optimization problem which maximizes the streaming rate subject to peer bandwidth constraints, and then solve it with a distributed algorithm. From the study on streaming capacity, we find that the streaming capacity is limited by the over-demanded video segments. Therefore we introduce helpers, the peers who are willing to contribute their remaining upload bandwidths to help other peers, into P2P VoD systems. We optimize helper assignment and rate allocation to improve the streaming capacity. In the simulations, we demonstrate that the streaming capacity can be obtained in a distributed manner by optimizing the resource allocation in the P2P VoD system. Ling Guan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2009 | A collaborative Bayesian image retrieval frameworkabstractIn this paper, an image retrieval framework combining content-based and content-free methods is proposed, which employs both short-term relevance feedback (STRF) and long-term relevance feedback (LTRF) as the means of user interaction. The STRF refers to iterative query-specific model learning during a retrieval session, and the LTRF is the estimation of a user history model from the past retrieval results approved by previous users. The framework is formulated based on the Bayes' theorem, in which the results from STRF and LTRF play the roles of refining the likelihood and the a priori information, respectively, and the images are ranked according to the a posteriori probability. Since the estimation of the user history model is based on the principle of collaborative filtering, the system is referred to as a collaborative Bayesian image retrieval (CLBIR) framework. To evaluate the effectiveness of the proposed framework, nearest neighbor CLBIR (NN-CLBIR) and support vector machine active learning CLBIR (SVMAL-CLBIR) were implemented. Experimental results showed the improvement over content-based methods in terms of both accuracy and ranking due to the integration in the proposed framework. Rui Zhang 0010, Ling Guan |
ICASSP | 2 |
| 2009 | On-Line Signature Verification: Directional Analysis of a Signature Using Weighted Relative Angle Partitions for Exploitation of Inter-Feature DependenciesabstractIn this paper, we propose a new directional analysis tool for On-line signatures that decomposes the given input signature into directional bands on the basis of relative angles. Our directional analysis tool takes the independent trajectories (horizontal and vertical) as an input and then decomposes them into directional bands on the basis of relative angles. We have used both user-dependent and user-independent thresholds for selecting an optimal number of partitions for each signer. By decomposing signature trajectories based upon relative angles of an individualpsilas signature, the resulting process can be thought of as one that exploits inter-feature dependencies . In the verification phase, distances of each partitioned trajectory of a test signature are calculated against a similarly partitioned template trajectory for a known signer. Each partition is then weighted based on its quality and quantity. Experimental results demonstrate the superiority of our approach to On-line signature verification in comparison with other techniques. Muhammad Talal Ibrahim, Matthew J. Kyan, Muhammad A. Khan 0002, Khurram Saleem Alimgeer, Ling Guan |
ICDAR | 5 |
| 2009 | Multimedia multimodal methodologiesabstractThis paper outlines several multimedia systems that utilize a multimodal approach. These systems include audiovisual based emotion recognition, image and video retrieval, and face and head tracking. Data collected from diverse sources/sensors are employed to improve the accuracy of correctly detecting, classifying, identifying, and tracking of a desired object or target. It is shown that the integration of multimodality data will be more efficient and potentially more accurate than if the data was acquired from a single source. A number of cutting-edge applications for multimodal systems will be discussed. An advanced assistance robot using the multimodal systems will be presented. Ling Guan, Paisarn Muneesawang, Yongjin Wang, Rui Zhang 0010, Tie Yun, Adrian Bulzacki, Muhammad Talal Ibrahim |
ICME | 1 |
| 2009 | Toward natural and efficient human computer interactionabstractEffective detection, recognition, interpretation, and analysis of human physiological and behavioral characteristics are of fundamental importance in the design and development of intelligent human computer interaction (HCI) systems. This paper illustrates the issues and challenges in the design of such systems through two real examples, emotion recognition and face detection. In particular, we focus on audiovisual based bimodal emotion recognition, face detection in crowded scene, and facial fiducial points detection. The integration of these systems is expected to produce more robust and stronger performance, and provide more natural and friendly man-machine interaction. Ling Guan, Yongjin Wang, Tie Yun |
ICME | 1 |
| 2009 | Improving the streaming capacity in P2P VoD systems with helpersabstractPeer-to-peer (P2P) video-on-demand (VoD) is a promising solution to provide video service to a large number of users. Streaming capacity in a P2P VoD system is defined as the maximal streaming rate that every user can receive. Due to the upload bottleneck, the streaming capacity in the P2P VoD system is limited. In this paper, we introduce helpers in the P2P VoD system and then optimize the helper resources to improve the streaming capacity. Specifically, we first optimize the helper assignment using a greedy algorithm. Then we develop a proximal distributed algorithm to maximize the streaming capacity by optimizing the link rates. Through simulations, we demonstrate that the P2P VoD system with optimized helpers can obtain a much higher streaming capacity compared to the P2P VoD system without any helper or the one with randomly assigned helpers. Ling Guan |
ICME | 2 |
| 2009 | Optimal resource allocation for video communication over distributed systemsabstractMany multimedia applications involve real-time video communication over distributed systems, in which there is no centralized controller. Examples of such distributed systems are peer-to-peer (P2P) networks, wireless ad hoc networks, and wireless sensor networks. In this paper, we provide a review of recent advances on optimal resource allocation for video communication over some major distributed systems including P2P streaming systems, wireless ad hoc networks, and wireless visual sensor networks. In P2P streaming systems, we review the scheduling optimization problem, streaming capacity problem, routing optimization problem, and the prefetching optimization problem. In wireless ad hoc networks, we present the routing optimization problem, joint optimization of the source rate and the routing scheme, joint optimization of sender selection and the routing scheme, and joint optimization of the source rate, the routing scheme and the power. In wireless visual sensor networks, we discuss the network lifetime maximization problem, optimal power allocation, maximization of accumulative visual information (AVI). Illustrative simulation results are provided to demonstrate the performance improvement brought by the optimal resource allocation in the distributed systems. Finally, we give our vision on the future work in the area of video communication over distributed systems. Ling Guan |
ICME | 2 |
| 2009 | Multimodal image retrieval via bayesian information fusionabstractIn this paper, a multimodal image retrieval framework integrating the information in both audio and visual domain via Bayesian decision level fusion is proposed. In both domains, a statistical model for each semantic class is learned. Based on the Bayes' theorem, the a posteriori probability of each class given a query is calculated in the audio domain, which is propagated to the images classified into the corresponding semantic class in the visual domain. These probabilistic measures are utilized as the a priori probability in the overall framework, which is combined with the likelihood evaluated based on nearest neighbor content-based image retrieval. Through the Bayes' theorem again, the images are ranked based on their a posteriori probabilities given the audio and visual feature of a query. To further improve the system, we also propose a relevance feedback scheme in the audio domain. Experimental results demonstrate the advantage of the proposed method over the retrieval simply based on visual features. Rui Zhang 0010, Ling Guan |
ICME | 2 |
| 2009 | Streaming Capacity in P2P VoD SystemsabstractPeer-to-Peer (P2P) Video-on-Demand (VoD) has become a popular service on the Internet. However current P2P VoD systems can only provide a video at a low streaming rate. The upper bound of the streaming rate that a P2P VoD system can support becomes an attractive topic. In this paper, we investigate the streaming capacity in P2P VoD systems, which is defined as the maximum supported streaming rate that can be received by every receiver. We formulate the streaming capacity problem considering the bandwidth constraints, and find the optimal solution to this problem. Furthermore, we develop a proximal decomposition algorithm to achieve the streaming capacity in a distributed manner. The simulation results demonstrate that the proposed proximal decomposition algorithm can obtain the streaming capacity very close to the optimal result. Ling Guan |
ISCAS | 2 |
| 2009 | Automatic sports genre categorization and view-type classification over large-scale datasetabstractThis paper presents a framework with two automatic tasks targeting large-scale and low quality sports video archives collected from online video streams. The framework is based on the bag of visual-words model using speeded-up robust features (SURF). The first task is sports genre categorization based on hierarchical structure. Following on the second task which is based on automatically obtained genre, views are classified using support vector machines (SVMs). As a consequence, the views classification result can be used in video parsing and highlight extraction. As compared with state-of-the-art methods, our approach is fully automatic as well as domain knowledge free and thus provides a better extensibility. Furthermore, our dataset consists of 14 sport genres with 6850 minutes in total. Both sport genre categorization and view type classification have more than 80% accuracy rates, which validate this framework's robustness and potential in web-based applications. Lingfang Li, Ning Zhang 0023, Ling-Yu Duan, Qingming Huang, Ling Guan |
ACM Multimedia | 6 |
| 2009 | Content-free image orientation detection using Local Binary PatternabstractImage orientation detection is a basic, but important subject in image processing. It is also a difficult task because of the uncertainty of the context information of the image and the illumination. In this paper, we present a theoretically and computationally simple yet efficient content-free image orientation detection approach. The proposed approach is based on Local Binary Patterns which is recognized as gray scale and rotation invariant. By using the binary mode and the histogram computed over the overlapping region of two images, the experiments show that it can get excellent results in both artificial and real images. Furthermore, the method can be used in different images with similar contents. Through optimization, the run time of the algorithm reduces by a large amount and makes it applicable in real applications. Nan Dong, Ling Guan |
MMSP | 2 |
| 2009 | Automatic fiducial points detection for facial expressions using scale invariant featureabstractDetecting fiducial points successfully in facial images or video sequences can play an important role in numerous facial image interpretation tasks such as face detection and identification, facial expression recognition, emotion recognition, and face image database management. In this paper we propose an automatic and robust method of facial fiducial point's detection for facial expressions analysis in video sequences using scale invariant feature based Adaboost classifiers. Face region is first located using the face detector with local normalization and optimal adaptive correlation technique. Candidate points are then selected over the face region using local scale-space extrema detection. The scale invariant feature for each candidate point is extracted for further examination. We choose 26 fiducial points on the face region from training samples to build the fiducial point detectors with Adaboost classifiers. All the candidate points in the test samples are examined through these detectors. Finally, all the 26 facial fiducial points are located on each frame of the test samples. Cohn-Kanade database and Mind Reading DVD are used for experiment. The results show that our method achieves a good performance of 90.69% average recognition rate. Tie Yun, Ling Guan |
MMSP | 2 |
| 2009 | Automatic face detection in video sequences using local normalization and optimal adaptive correlation techniques
Tie Yun, Ling Guan |
Pattern Recognit. | 2 |
| 2009 | Distributed Algorithms for Network Lifetime Maximization in Wireless Visual Sensor NetworksabstractNetwork lifetime maximization is a critical issue in wireless sensor networks since each sensor has a limited energy supply. In contrast with conventional sensor networks, video sensor nodes compress the video before transmission. The encoding process demands a high power consumption, and thus raises a great challenge to the maintenance of a long network lifetime. In this paper, we examine a strategy for maximizing the network lifetime in wireless visual sensor networks by jointly optimizing the source rates, the encoding powers, and the routing scheme. Fully distributed algorithms are developed using the Lagrangian duality to solve the lifetime maximization problem. We also examine the relationship between the collected video quality and the maximal network lifetime. Through extensive numerical simulations, we demonstrate that the proposed algorithm can achieve a much longer network lifetime compared to the scheme optimized for the conventional wireless sensor networks. Ivan Lee 0001, Ling Guan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2009 | Optimized Video Multicasting Over Wireless Ad Hoc Networks Using Distributed AlgorithmabstractRecently there has been a compelling need to support real-time video multicast from a single source to multiple receivers in wireless ad hoc networks. The existing work uses tree-based schemes to perform video multicast. The optimization of those schemes typically requires a centralized computation, which is not suitable for wireless ad hoc networks. In this paper, we propose an optimized video multicast scheme over wireless ad hoc networks. First, we apply a prioritized coding scheme to enable the heterogeneous receivers to reconstruct the video at different quality levels. Then we formulate the video multicasting problem using the network model, the packet loss model, and the video distortion model. To solve the optimization problem, we propose a distributed algorithm to jointly optimize the source rate, the routing scheme, and the power allocation using hierarchical dual decompositions. The distributed nature of the proposed algorithm makes it very appropriate for wireless ad hoc networks. Through extensive simulations, we demonstrate that the proposed video multicast scheme can achieve much higher video quality compared to the uniform-power scheme or the tree-based routing schemes. Ivan Lee 0001, Ling Guan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2009 | Distributed Throughput Maximization in P2P VoD ApplicationsabstractIn peer-to-peer (P2P) video-on-demand (VoD) systems, a scalable source coding is a promising solution to provide heterogeneous peers with different video quality. In this paper, we present a systematic study on the throughput maximization problem in P2P VoD applications. We apply network coding to scalable P2P systems to eliminate the delivery redundancy. Since each peer receives distinct packets, a peer with a higher throughput can reconstruct the video at a higher quality. We maximize the throughput in the existing buffer-forwarding P2P VoD systems using a fully distributed algorithm. We demonstrate in the simulations that the proposed distributed algorithm achieves a higher throughput compared to the proportional allocation scheme or the equal allocation scheme. The existing buffer-forwarding architecture has a limitation in total upload capacity. Therefore we propose a hybrid-forwarding P2P VoD architecture to improve the throughput by combining the buffer-forwarding approach with the storage-forwarding approach. The throughput maximization problem in the hybrid-forwarding architecture is also solved using a fully distributed algorithm. We demonstrate that the proposed hybrid-forwarding architecture greatly improves the throughput compared to the existing buffer-forwarding architecture. In addition, by adjusting the priority weight at each peer, we can implement the differentiated throughput among different users within a video session in the buffer-forwarding architecture, and the differentiated throughput among different video sessions in the hybrid-forwarding architecture. Ivan Lee 0001, Ling Guan |
IEEE Trans. Multim. | 3 |
| 2009 | Optimal Prefetching Scheme in P2P VoD Applications With Guided SeeksabstractMost existing peer-to-peer (P2P) video-on-demand (VoD) systems have been designed and optimized for the sequential playback. In practice, users often want to seek to the positions they are interested in. Such frequent seeks raise greater challenges to the design of the prefetching scheme. In this work, we first propose the concept of guided seeks. With the guidance, users can perform more efficient seeks to the desired positions. The guidance can be obtained from collective seeking statistics of other peers who have watched the same title in the previous and/or concurrent sessions. However, it is very challenging to aggregate the statistics efficiently, timely and in a completely distributed way. We design the hybrid sketches that not only capture the seeking statistics at significantly reduced space and time complexity, but also adapt to the popularity of the video. From the collected seeking statistics, we estimate the segment access probability, based on which we further develop an optimal prefetching scheme and an optimal cache replacement policy to minimize the expected seeking delay at every viewing position. Through extensive simulations, we demonstrate that the proposed prefetching framework significantly reduces the seeking delay compared to the sequential prefetching scheme. Guobin Shen, Yongqiang Xiong, Ling Guan |
IEEE Trans. Multim. | 4 |
| 2008 | Distributed throughput maximization in hybrid-forwarding P2P VoD applicationsabstractIn peer-to-peer (P2P) video-on-demand (VoD) systems, a scalable source coding is a promising solution to provide heterogeneous peers with different video quality. In a scalable P2P VoD system, the received video quality at each peer is dependent on the throughput if the packet redundancy has been addressed with coding techniques. In this paper, we propose a hybrid-forwarding P2P VoD architecture, which integrates both the buffer-forwarding approach and storage-forwarding approach. Furthermore, we develop a distributed algorithm to maximize the throughput. Through simulations, we demonstrate that the hybrid-forwarding architecture can obtain a much higher throughput compared to the architecture with only buffer-forwarding links or with only storage-forwarding links. Ivan Lee 0001, Ling Guan |
ICASSP | 3 |
| 2008 | Use hierarchical genetic particle filter to figure articulated human trackingabstractUsing particle filter to track human movement, a key problem is how to draw samples in high-dimensional state space. In this paper, we present a novel framework of particle filtering, namely Hierarchical Genetic Particle Filter (HGPF), to improve the efficiency of samples by a hierarchical evolutionary detection. As a result, we can obtain reasonably distributed samples thus translating into reliable tracking performance. Finally, we apply the technique to 2D articulated human movement tracking. Result demonstrates the effectiveness of HGPF in solving the tracking problem like self-occlusion and cluttered background. Long Ye, Qin Zhang 0009, Ling Guan |
ICME | 3 |
| 2008 | On-line signature verification by using most discriminating pointsabstractIn this paper, we propose a novel approach that looks for the most discriminating points in horizontal, vertical, velocity and pressure trajectories. Then based on these discriminating points, we have made two composite feature sets i.e. (horizontal; vertical) and (velocity; pressure). Out of these two composite feature sets, one will be declared as the most discriminating composite feature set during the training phase. Finally the verification based on the selected discriminating composite feature will be performed. We believe that every signer has some discriminating points that a forger cannot mimic with certain pressure and velocity whereas these points are maintained by the genuine signer. So the verification based on these discriminating points can lead us to a better performance of the verification system. For a reliable verification system, comparison between forgery and genuine signer should be made on the basis of discriminating points rather than using all the sample points of each trajectory in the verification which means that all the points are given equal weights and there is always a possibility that the non-discriminating points can overshadow the effect of discriminating points. Experimental results demonstrate superiority of our approach in On-line signature verification in comparison with other techniques. Muhammad Talal Ibrahim, Ling Guan |
ICPR | 2 |
| 2008 | Probabilistic prefetching scheme for P2P VoD applications with frequent seeksabstractIn Peer-to-Peer Video-on-Demand (P2P VoD) applications, users tend to seek to the positions that they are interested in. The frequent seeks raise a great challenge to the design of the prefetching scheme. In this paper, we propose a probabilistic prefetching framework to reduce the seeking distance. Each peer performs prefetching based on the segment access probability, which is estimated from the seeking statistics in the previous sessions. It is a challenging task to collect the seeking statistics in a distributed P2P network. In the proposed framework, we employ FM sketches to represent the seeking statistics, thus greatly reducing the space and time complexity. The simulation results show that the proposed prefetching scheme can approach closer to the desired seeking positions compared to the prefetching scheme neglecting the user viewing pattern. Guobin Shen, Yongqiang Xiong, Ling Guan |
ISCAS | 4 |
| 2008 | Graph cut video object segmentation using histogram of oriented gradientsabstractThis paper introduces a novel way to implement Graph Cut for video object segmentation with shape information. Graph Cut is a very efficient algorithm for image segmentation and Histogram of Oriented Gradients (HOG) is useful in detecting humans. We combine the HOG feature to incorporate a shape prior into Graph Cut algorithm as a new way to enhance video object segmentation accuracy. In previous work, we used a fully connected 3-D that is slow and is subject to weak edges, inconsistent luminance. The new method is compared with old methods to show that it helps by introducing a shape prior for segmentation of pre-trained objects such as humans. Chun-Hao Wang, Ling Guan |
ISCAS | 2 |
| 2008 | Robust Fingerprint Image Enhancement: An Improvement to Directional Analysis of Fingerprint Image Using Directional Gaussian Filtering and Non-subsampled Contourlet TransformabstractIn this paper, a new fingerprint image enhancement method based on the integration of directional Gaussian filtering and non-subsampled contourlet transform (NSCT) has been proposed. As the fingerprint images suffers from non-uniform illumination, so there is a need to completely remove or reduce the non-uniform illumination present in the fingerprint image before applying any enhancement technique. In this paper, we have used bandpass filter for the elimination of non-uniform illumination and for the creation of frequency ridge image. Further the directional noise present in the fingerprint image have been removed by using the directional Gaussian filter and the noise-free image was decomposed into non-subsampled directional images by using NSCT. Lately, we construct an enhanced image through a block-by-block process which compares the energy of all the directional images and picks the one that provides maximum energy. Further, binarization and thinning are also applied on the enhanced image. Muhammad Talal Ibrahim, Tariq Bashir, Ling Guan |
ISM | 3 |
| 2008 | On-line Signature Verification Using Most Discriminating Features and Fisher Linear Discriminant Analysis (FLD)abstractIn this work, we employ a combination of strategies for partitioning and detecting abnormal fluctuations in the horizontal and vertical trajectories of an on-line generated signature profile. Alternative partitions of these spatial trajectories are generated by splitting each of the related angle, velocity and pressure profiles into two regions representing both high and low activity. The overall process can be thought of as one that exploits inter-feature dependencies by decomposing signature trajectories based upon angle, velocity and pressure - information quite characteristic to an individualpsilas signature. In the verification phase, distances of each partitioned trajectory of a test signature are calculated against a similarly partitioned template trajectory for a known signer. Finally, these distances become inputs to Fisherpsilas Linear Discriminant Analysis (FLD). Experimental results demonstrate the superiority of our approach in On-line signature verification in comparison with other techniques. Muhammad Talal Ibrahim, Matthew J. Kyan, Ling Guan |
ISM | 3 |
| 2008 | Local Normalization with Optimal Adaptive Correlation for Automatic and Robust Face Detection on Video SequencesabstractThis paper proposes an automatic and robust method to detect human faces from video sequences that combines feature extraction and face detection based on local normalization, Gabor wavelets transform and Adaboost algorithm. The key step and the main contribution of this work is the incorporation of a normalization technique based on local histograms with optimal adaptive correlation (OAC) technique to alleviate a common problem in conventional face detection methods: inconsistent performance due to the sensitivity to illumination variations such as local shadowing, noise and occlusion. This approach uses a cascade of classifiers to adopt a coarse-to-fine strategy to achieve higher detection rate with lower false positives. The experimental results demonstrate a significant performance improvement by local normalization over method without normalizations in real video sequences with a wide range of facial variations in color, position, scale, and varying lighting conditions. Tie Yun, Ling Guan |
ISM | 2 |
| 2008 | Modelling an individual's Web search interests by utilizing navigational dataabstractAn approach to model and quantify a user’s Web search interests using the user’s navigational data is presented. The approach is based on the premise that frequently visiting certain types of content indicates that the user is interested in that content. The proposed approach can be divided into three steps: monitoring the user’s navigational data; using the cumulative weight to determine a Web page’s content; and employing the Naïve Bayes Model for updating the user’s interest model. In order to demonstrate the effectiveness of the proposed model, experimental software is developed to analyze a user’s interests in sports. The experimental results demonstrate that the approach can effectively model the user’s interest. The proposed model could be integrated with personalized Web services. Liping Fang, Ling Guan |
MMSP | 3 |
| 2008 | A new framework of relevance feedback for content-free image retrievalabstractHuman beings recognize similarity in scene perception based on their available high-level knowledge about the low-level visual features, which is gradually accumulated throughout their entire lives. Once there is not enough knowledge they tend to rely on low-level visual content. Inspired by this observation, we proposed a new framework of relevance feedback for content-free image retrieval to tackle the problem of sample sparseness. The framework is composed of two components, i.e. short-term feedback and long-term feedback. The former refers to an operation of query conversion and/or refinement during a retrieval session by incorporating a content-aware module, while the latter consists of incrementally updating the system model using the accumulated retrieval results since the last system update. 10000 images from 200 categories of the COREL image collection were employed for evaluating the performance of the framework using the criterion of averaged precision as a function of the number of relevance feedback needed. Experimental results demonstrated a human-like behavior of the proposed framework in that while long-term update helps the system accumulate more knowledge, the content-aware short-term relevance feedback further boosts its performance when the amount of knowledge is limited. Rui Zhang 0010, Ling Guan |
MMSP | 2 |
| 2008 | The HCM for perceptual image segmentation
Jonathan Randall, Ling Guan, Wanqing Li 0001 |
Neurocomputing | 2 |
| 2008 | Recognizing Human Emotional State From Audiovisual SignalsabstractMachine recognition of human emotional state is an important component for efficient human-computer interaction. The majority of existing works address this problem by utilizing audio signals alone, or visual information only. In this paper, we explore a systematic approach for recognition of human emotional state from audiovisual signals. The audio characteristics of emotional speech are represented by the extracted prosodic, Mel-frequency Cepstral Coefficient (MFCC), and formant frequency features. A face detection scheme based on HSV color model is used to detect the face from the background. The visual information is represented by Gabor wavelet features. We perform feature selection by using a stepwise method based on Mahalanobis distance. The selected audiovisual features are used to classify the data into their corresponding emotions. Based on a comparative study of different classification algorithms and specific characteristics of individual emotion, a novel multiclassifier scheme is proposed to boost the recognition performance. The feasibility of the proposed system is tested over a database that incorporates human subjects from different languages and cultural backgrounds. Experimental results demonstrate the effectiveness of the proposed system. The multiclassifier scheme achieves the best overall recognition rate of 82.14%. Yongjin Wang, Ling Guan |
IEEE Trans. Multim. | 2 |
| 2008 | Recognizing Human Emotional State From Audiovisual SignalsabstractMachine recognition of human emotional state is an important component for efficient human-computer interaction. The majority of existing works address this problem by utilizing audio signals alone, or visual information only. In this paper, we explore a systematic approach for recognition of human emotional state from audiovisual signals. The audio characteristics of emotional speech are represented by the extracted prosodic, Mel-frequency Cepstral Coefficient (MFCC), and formant frequency features. A face detection scheme based on HSV color model is used to detect the face from the background. The visual information is represented by Gabor wavelet features. We perform feature selection by using a stepwise method based on Mahalanobis distance. The selected audiovisual features are used to classify the data into their corresponding emotions. Based on a comparative study of different classification algorithms and specific characteristics of individual emotion, a novel multiclassifier scheme is proposed to boost the recognition performance. The feasibility of the proposed system is tested over a database that incorporates human subjects from different languages and cultural backgrounds. Experimental results demonstrate the effectiveness of the proposed system. The multiclassifier scheme achieves the best overall recognition rate of 82.14%. Yongjin Wang, Ling Guan |
IEEE Trans. Multim. | 2 |
| 2007 | Combination of Recognizers and Fusion of Features Approach to Missing Data ASR Under Non-Stationary Noise ConditionsabstractThe difficulty of ASR under non-stationary noise conditions is a major contributing factor hindering the widespread deployment of ASR systems. Bottom up techniques such as speech noise separation and top down methods to adapt the acoustic model to the environment have been applied to address the issue. The missing data approach to ASR improves upon existing techniques basing recognition solely on the reliable components of the signal and has been demonstrated as an effective method to handle non-stationarity. Proposed in this paper is a novel technique whereby ASR using missing data theory under non-stationary noise conditions is improved by use of a fusion of models at the decision level. This fused model introduces more resilient features to the missing data decode process. The fused decoder is found to significantly increase recognition performance over conventional missing data techniques. A major finding in this paper is when the fused decoder exhibits the fusion of bottom up and top down processes. Under this condition, the proposed combination of recognizers technique is found to outperform all other tested ASR systems. Neil Joshi, Ling Guan |
ICASSP (4) | 2 |
| 2007 | Wavelet-Based Texture Retrieval using Independent Component AnalysisabstractIn this paper, a novel approach to texture retrieval using independent component analysis (ICA) in wavelet domain is proposed. It is well recognized that the wavelet coefficients in different subbands are statistically correlated, resulting in the fact that the product of the marginal distributions of wavelet coefficients is not accurate enough to characterize the stochastic properties of texture images. To tackle this problem, we employ (ICA) in feature extraction to decorrelate the analysis coefficients in different subbands, followed by modeling the marginal distributions of the separated sources using generalized Gaussian density (GGD), and perform similarity measure based on the maximum likelihood criterion. It is demonstrated by simulation results on a database consisting of 1776 texture images that the proposed method improve the accuracy of texture image retrieval in terms of average retrieval rate, compared with the traditional method using GGD for feature extraction and Kullback-Leibler divergence for similarity measure. Rui Zhang 0010, Xiao-Ping Zhang 0002, Ling Guan |
ICIP (6) | 3 |
| 2007 | Distributed Rate Allocation in P2P StreamingabstractIn peer-to-peer (P2P) streaming, each peer contributes its upload bandwidth to redistributing the data stream to its downstream peers. How to optimally utilize the upload bandwidth of each peer is an important issue. In this paper, we propose a fully distributed algorithm to optimize the link rate allocation in P2P streaming. We use dual decomposition to separate the optimization problem into multiple sub-problems, which are solved at each individual peer respectively. Through simulations, we demonstrate that the proposed rate allocation scheme can quickly converge, and achieve a higher throughput compared to proportional or equal allocation scheme. Ivan Lee 0001, Ling Guan |
ICME | 3 |
| 2007 | Network Lifetime Maximization in Wireless Visual Sensor Networks using a Distributed AlgorithmabstractNetwork lifetime maximization is a critical issue in wireless sensor networks since each sensor has a limited energy supply. Different from conventional sensors, video sensors compress the captured video before transmission. The encoding processing demands high power consumption, thus raises challenges to maintain a long network lifetime. In this paper, we formulate the network lifetime maximization problem in wireless visual sensor networks, and propose a fully distributed algorithm to solve this problem. The proposed algorithm maximizes the network lifetime by jointly optimizing the encoding powers, the source rates, and the link rates. Ivan Lee 0001, Ling Guan |
ICME | 3 |
| 2007 | Special Effects in Film Making with Object Based TransformationsabstractThe new media initiative project of Ryerson University had been successful in applying image and video processing techniques in creating special effects in film making. Our system utilizes a graph cut image segmentation and snakes active contour approach to obtain object cut outs in 3-D with tracking. It uses exemplar-based inpainting for background filling for missing regions uncovered by object transformation. The two methods combined allows for easy video object editing with reduction in user input. Previously implemented shot detection by twin window amplification method and steerable pyramids texture generation is also part of the system. Lastly, a set of transformation was implemented that serves as a basis for testing the system. Chun-Hao Wang, Xiaoming Fan, Bruce Elder, Xiaoou Tang, Ling Guan |
ICME | 6 |
| 2007 | Optimized multi-path routing using dual decomposition for wireless video streamingabstractFor video streaming over wireless ad hoc networks, source rate allocation and routing scheme are two important issues. In this paper, we propose a fully distributed algorithm to jointly optimize the source rate and routing scheme. We use dual decomposition to separate the optimization problem into multiple subproblems, which are then solved in parallel. The distributed nature of the proposed algorithm is extremely adequate for wireless ad hoc networks. Simulation results show that the proposed routing scheme outperforms some existing multi-path routing schemes. Ivan Lee 0001, Ling Guan |
ISCAS | 3 |
| 2007 | Application of Laplacian Mixture Model to Image and Video RetrievalabstractIn this paper, we study the peaky nature of wavelet coefficient distributions. The study shows that the wavelet coefficients cannot be effectively modeled by a single distribution. We then propose a new modeling scheme based on a Laplacian mixture model and apply it to the indexing and retrieval of image and video databases. In this work, the parameters of the model are first used to represent texture information in image retrieval. Then we explore its application to video retrieval. Traditionally, visual information is used for video indexing and retrieval. However, in some cases audio information is more helpful for finding clues to the video events. The proposed feature extraction scheme is based on the fundamental property of the wavelet transform. Therefore, it can also be adopted to analyze the audio contents of the video data. The experimental evaluation indicates the high discriminatory power of the proposed feature set. The dimension of the extracted feature vector is low, which is important for the retrieval efficiency of the system in terms of response time. User feedback is used to enhance the retrieval performance by modifying the system parameters according to the users' behavior. A nonlinear approach for defining the similarity between the two images is also explored in this work. Tahir Amin, Mehmet Zeytinoglu, Ling Guan |
IEEE Trans. Multim. | 3 |
| 2006 | Monocular Human Motion Tracking with the DE-MC Particle FilterabstractA key to accomplish articulated human motion tracking and other high-dimensional visual tracking tasks is to have an efficient way to draw samples from the state space. The typical particle filter method and most of its variants do not perform well in achieving this goal. To solve the problem we present a novel algorithm, namely the differential evolution-Markov chain (DE-MC) particle filtering. It substantially improves the core of traditional particle filter, i.e. the sampling strategy. As a result, we can obtain reasonably distributed samples in an efficient way thus translating into reliable tracking performance. Experimental results demonstrate the power of the proposed approach Ling Guan |
ICASSP (2) | 2 |
| 2006 | Substream Allocation in Layered P2P StreamingabstractIn layered P2P streaming system, how to allocate number of the copies for each layer is a challenging problem. In this paper, we present a substream allocation scheme in layered P2P streaming. The proposed allocation scheme is adaptive to the request rate and number of the qualified peers. The simulation results show that the proposed allocation scheme enables the system to achieve an overall better quality compared to the general allocation schemes with fixed allocation percentages. In addition, the proposed allocation scheme can accelerate the growth of peer population in the initial stage of hybrid P2P streaming systems Ivan Lee 0001, Ling Guan |
ICME | 3 |
| 2006 | Computational Intellegence Techniques and their Applications in Content-Based Image RetrievalabstractThe main focus of this paper is to present a methodology for optimizing relevance identification in content-based image retrieval (CBIR) systems through the principle of feature weight detection. The purpose of relevance identification is to find a collection of images that are statistically similar to, or match with, an original query image within a large visual database. The novelty of this scheme is two-fold: using a base-10 genetic algorithm method to accurately determine the contribution of individual feature vectors for a successful retrieval in the so-called feature weight detection process, and defining a new unsupervised learning algorithm, the directed self-organizing tree map (DSOTM), for the purpose of classification in the automatic relevance identification module of the search engine. Comprehensive experiments demonstrate feasibility of the proposed methodology Kambiz Jarrah, Matthew J. Kyan, Sridhar Krishnan 0001, Ling Guan |
ICME | 4 |
| 2006 | Special Effects in Film/Video Making: A New Media Initiative ProjectabstractWe present a system and a set of tools for producing special effects in film/video making by applying image processing and human centered computing techniques. A combination of shot detection, object segmentation, background generation, and image warping techniques are used. The user selects the image frame or the object of interest, and the image warp transformation to be used from the GUI. Wrapping can be performed either on a whole image sequence or on an object of interest in the sequence. In the latter, the object is first segmented and its motion tracked. Object segmentation is achieved either by snakes or graph cuts. Steerable pyramid background generation is then used to fill in the portion cut from the foreground Chun-Hao Wang, Yongjin Wang, Meifeng Lian, Bruce Elder, Xiaoou Tang, Ling Guan |
ICME | 6 |
| 2006 | Automatic Content-Based Image Retrieval Using Hierarchical Clustering AlgorithmsabstractThe overall objective of this paper is to present a methodology for guiding adaptations of an RBF based relevance feedback network, embedded in automatic content-based image retrieval (CBIR) systems, through the principle of unsupervised hierarchical clustering. The self organizing tree map (SOTM) is essentially attractive for our approach since it not only extracts global intuition from an input pattern space but also injects some degree of localization into the discriminative process such that maximal discrimination becomes a priority at any given resolution. The main focus of this paper is two-fold: introducing a new member of SOTM family, the Directed SOTM (DSOTM) that not only provides a partial supervision on duster generation by forcing divisions away from the query class, but also presents a flexible verdict on resemblance of the input pattern as its tree structure grows; and modifying the current structure of the normalised graph cuts (Ncut) process by enabling the algorithm to determine appropriate number of clusters within an unknown dataset prior to its recursive clustering scheme through the principle of self-organizing normalized graph cuts (SONcut). Comprehensive comparisons with the Self-Organizing feature Map (SOFM), SOTM, and Ncut algorithms demonstrate feasibility of the proposed methods. Kambiz Jarrah, Sridhar Krishnan 0001, Ling Guan |
IJCNN | 3 |
| 2006 | The Self-Organising Hierarchical Variance MapabstractThe Self-Organising Hierarchical Variance Map (SOHVM), a novel unsupervised clustering technique is proposed. Based on both Kohonen and Hebbian principles of self-organisation, the algorithm works to dynamically conauct a topology preserved mapping of dominant prototype clusters from within an unknown data source. Each neuron in the network consists of a dual memory element that tracks information regarding a discovered prototype. In addition to position, Hebbian based Maximum Eigenfilters (HME) simultaneously estimate the maximal variance of local data. Competitive Hebbian Learning (CHL) is used to dynamically associate prototypes such that an accurate topology is maintained throughout the discovery process. Knowledge may then be progressively imparted to the network through appropriate neighbouring memory elements. Vigilance is assessed via interplay between local variances such that more informed decisions control and naturally limit network growth. The approach is closely related to Self-organizing Tree Maps (SOTM), Growing Neural Gas (GNG) and their variants. Matthew J. Kyan, Ling Guan |
IJCNN | 2 |
| 2006 | Missing data ASR with fusion of features and combination of recognizersabstractSpeech recognition under noisy conditions has been actively researched and effective techniques have been developed to handle stationary noise. Under circumstances where the stationary assumption is not valid, the performance of speech recognizers is extremely poor. Missing data theory provides a method for the development of robust speech recognition under any noisy condition. A limitation to ASR with missing data theory techniques is the choice of features used in the model. There exist alternative feature representations that have been demonstrated to be much more effective for signal recognition purposes. This paper presents a novel method to incorporate the use of alternative feature sets within the realm of ASR with missing data theory techniques. Using the proposed combination of recognizers, or fusion of features, an ASR decoding process is developed based upon the coupling of spectral features using missing data techniques and traditional MFCC based features. The proposed technique is demonstrated to increase recognition performance under all experimented noise conditions over traditional missing data techniques. Neil Joshi, Ling Guan |
SLT | 2 |
| 2005 | Indexing of NFL Video using MPEG-7 Descriptors and MFCC featuresabstractIn this paper, we propose an application system to classify American football (NFL) video shots into 4 categories, namely: pass plays, run plays, field goal/extra point plays (FG/XP) and kickoff/punt plays (K/P). The proposed system consists of two stages. The first stage is responsible for play event localization and the latter stage is responsible for feature mapping and classification. For play event localization we have proposed an algorithm that uses MPEG-7 motion activity descriptor and mean of the magnitudes of motion vectors, in a collaborative manner to detect the starting point of a play event within a video shot with 83% accuracy. The indexing and classification stage uses MPEG-7 motion and audio descriptors along with Mel Frequency Cepstrum Coefficients (MFCC) features to classify the events into 4 categories using Fisher's LDA. We obtain indexing accuracy of 92.5% by using a leave-one-out classification technique on a database of 200 video shots taken from 4 different games obtained from 4 different networks. Syed G. Quadri, Sridhar Krishnan 0001, Ling Guan |
ICASSP (2) | 3 |
| 2005 | Recognizing human emotion from audiovisual informationabstractIn this paper, we present an emotion recognition system to classify human emotional state from audiovisual signals. We extract prosodic, mel-frequency cepstral coefficient (MFCC), and formant frequency features to represent the audio characteristics of the emotional speech. A face detection scheme, based on the HSV color model, is used to detect the face from the background. The facial expressions are represented by Gabor wavelet features. We perform feature selection by using a stepwise method based on Mahalanobis distance. A classification scheme involving the analysis of individual class and combinations of different classes is proposed. Our emotion recognition system is tested over a language and race independent database, and an overall recognition accuracy of 82.14% is achieved. Yongjin Wang, Ling Guan |
ICASSP (2) | 2 |
| 2005 | Centralized Peer-to-Peer Video Streaming Over Hybrid Wireless NetworkabstractVideo streaming over wireless network has drawn a great interest. In traditional wireless local area networks (WLANs), when the number of the users and the number of the flows increases, the contention for the wireless channel will lead to packet loss and packet delay. In this paper, we propose a centralized peer-to-peer video streaming over hybrid wireless network to improve the performance of the video transport over wireless Internet. The base layer of the video is transported from the server via the WLAN mode, which benefits the centralized management of the video distribution, while the enhancement layers are delivered over the multiple paths via the ad hoc mode, which can reduce the congestion in the access point (AP). The simulation results show that our proposed scheme can achieve a better perceptual video quality compared to the WLAN deployment. Ivan Lee 0001, Xijia Gu, Ling Guan |
ICME | 4 |
| 2005 | Reliable video communication with multi-path streaming using MDCabstractVideo streaming demands high data rates and hard delay constraints, and it raises several challenges on today's packet-based and best-effort Internet. In this paper, we propose an efficient multiple-description coding (MDC) technique based on video frame sub-sampling and cubic-spline interpolation to provide spatial diversity, such that no additional buffering delay or storage is required. The frame dropping rate due to packet loss and drifting error under the multi-path streaming environment is analyzed in this paper. Ivan Lee 0001, Ling Guan |
ICME | 2 |
| 2005 | Dynamic feature fusion in the self organising tree map-applied to the segmentation of biofilm imagesabstractThe self organising tree map (SOTM) neural network is investigated as a means of segmenting microorganisms from confocal microscope image data. Features describing pixel & regional intensities, phase congruency and spatial proximity are explored in terms of their impact on the segmentation of bacteria and other micro-organisms. The significance of individual features is investigated, and it is proposed that, within the context of micro-biological image segmentation, better object delineation can be achieved if certain features dominate the initial stages of learning. In this way, other features are allowed to become more/less significant as learning progresses: as the network gains more knowledge about the data being segmented. The efficiency and flexibility of the SOTM in adapting to, and preserving the topology of input space, makes it an appropriate candidate for implementing this idea. Preliminary experiments are presented and it is found that favouring intensity characteristics in the early phases of learning, whilst relaxing proximity constraints in later phases of learning, offers a general mechanism through which we can improve the segmentation of microbial constituents. Matthew J. Kyan, Ling Guan, Steven Liss |
IJCNN | 2 |
| 2005 | Using knowledge of the region of interest (ROI) in automatic image retrieval learningabstractIn this paper, we propose an automatic relevance feedback retrieval system using perceptually important features extracted from regions of interest. The system is implemented via self-learning using a self-organizing tree map (SOTM) neural network. Our proposed method involves the construction of regions of interest from retrieved images using edge flow model, and the grouping of the regions into a single perceptually significant entity. This knowledge is fed into a set of unsupervised relevance feedback learning modules based on the SOTM to guide the adaptation of relevance feedback parameters through a machine learning approach without user interaction. Optimal tradeoff between the user workload in the interactive process and user subjectivity is then be explored by incorporating a semi-automatic retrieval strategy. Experimental results indicate that this system, with automatic and semiautomatic adaptations, can minimize user interaction, optimize precision, as well as reduce performance errors caused user subjectivity. Paisarn Muneesawang, Ling Guan |
IJCNN | 2 |
| 2005 | Centralized P2P Streaming with MDCabstractPeer-to-peer networking technique represents a vast potential to overcome many constraints in the conventional content distribution networks. In this paper, we propose centralized peer-to-peer (P2P) video streaming with multiple-description coding. Centralized P2P streaming with single and multiple forwarding peers are studied in this paper. We compare between the network loads at the bottleneck link of the client/server framework and that of the proposed framework. The reconstructed video qualities are evaluated in several experiments. We also analyze the frame dropout rates for the proposed system Ivan Lee 0001, Ling Guan |
MMSP | 3 |
| 2005 | Transformation of Compressed Domain Features for Content-Based Image Indexing and Retrieval
Hau-San Wong, Horace Ho-Shing Ip, Lawrence P. L. Iu, Kent K. T. Cheung, Ling Guan |
Multim. Tools Appl. | 5 |
| 2005 | Refining competition in the self-organising tree map for unsupervised biofilm image segmentation
Matthew J. Kyan, Ling Guan, Steven Liss |
Neural Networks | 2 |
| 2005 | Adaptive video indexing and automatic/semi-automatic relevance feedbackabstractThis paper presents adaptive methods for content-based retrieval in video database applications. A new adaptive video indexing (AVI) technique based on a template-frequency model, together with a self-training retrieval architecture, is proposed, to allow full use of temporal information. AVI takes into account spatio-temporal information for relevance feedback analysis of the dynamic content of video data. The AVI indexing method can be effectively adapted for video shot, scene, and story queries, in order to facilitate multiple-level access to a video database. Our system incorporates this indexing structure to a self-training neural network which implements automatic adaptive retrieval, through its signal propagation process. This greatly reduces the search time for video transmissions over the Web because relevance feedback is implemented in automatic and semi-automatic fashions. The AVI structure not only works well in fully automatic mode, but is also effective in the user-interaction interface system, to achieve a user-friendly environment. Experimentally, we demonstrated the proposed indexing technique and automatic relevance feedback for retrieval of CNN news videos. We also investigated the resilience of the system with a user-controlled interaction process and applied this to an automatically indexed database of 20 h of Hollywood movies. P. Munesawang, Ling Guan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2004 | Continuous human activity recognitionabstractEffectively recognizing human activities requires at least 32 joint related degrees of freedom to be estimated so as to reliably track the human body in 3D. The particle filter is robust to distracting clutter by maintaining multiple hypotheses for each of these joint angles. Real-time tracking is difficult however with the computational overhead of such a large search space. This paper optimizes this search space utilizing feedback from a continuous human activity recognition (CHAR) system and improves the robustness and efficiency of each particle calculation using a novel body model. The joint angles are estimated for the next frame using a particle filter with forward smoothing. A new paradigm enables the temporal segmentation of continuous motion into dynemes. Using HMM, the CHAR system attempts to infer the human movement skill that could have produced the observed sequence of dynemes. Hundreds of movement skills, from gait to saltos, are successfully tracked and recognized. Richard D. Green, Ling Guan |
ICARCV | 2 |
| 2004 | Interactive video retrieval using embedded audio contentabstractAudio is a rich source of information in the digital videos that can provide useful descriptors for indexing the video databases. In this paper, we model the shape of the distribution of wavelet coefficients of embedded audio with a Laplacian mixture. The distributions of wavelet coefficients are very peaky in nature. The shape of these distributions can be modeled with only two components in the Laplacian mixture with low computational complexity. The parameters of this mixture model form a low dimensional feature vector representing global similarity of the audio content of the video clips. An interactive approach involving the feature vector updating scheme is used to adapt the retrieval system to the users' needs. This relevance feedback (RF) increases the retrieval ratio substantially. A comprehensive experimental evaluation using the CNN news database has been performed. Tahir Amin, Mehmet Zeytinoglu, Ling Guan |
ICASSP (3) | 3 |
| 2004 | Semi-automated relevance feedback for distributed content based image retrievalabstractRetrieving images according to the semantic meanings is a challenging problem, mainly due to the complexity of mapping semantic meanings to low-level descriptors. Such complexity raises the scalability issue, especially when the database is distributed over multiple servers such as the peer-to-peer network. To address the scalability issue, we present an approach for content-based image retrieval (CBIR) over a distributed peer-to-peer network. The proposed system features: (1) improved retrieval precision; (2) decentralized database for high availability; (3) decentralized processing to utilize the computation resources. On the proposed peer-to-peer retrieval system, we present (1) query node based and (2) agent based approaches for on-demand advanced-feature calculation. Finally, we present the analysis for semi-automated relevance feedback over the peer-to-peer CBIR framework. Ling Guan |
ICME | 2 |
| 2004 | iARM -an interactive video retrieval systemabstractThis work presents the iARM system for content-based video retrieval in an interactive framework. This system explores a new model-based video indexing technique to improve the effectiveness of relevance feedback and make interactive video retrieval a user-friendly environment. The system emphasizes the accuracy in modeling spatio-temporal information in a video clip, so that relevance feedback analysis needs only a few cycles and a few training samples, which greatly reduce the search time for video transmissions over the network. We investigate the resilience of the system, and apply the interactive content-based retrieval method to an automatically indexed database of 20 hours of video. Paisarn Muneesawang, Ling Guan |
ICME | 2 |
| 2004 | An investigation of speech-based human emotion recognitionabstractThis paper presents our recent work on recognizing human emotion from the speech signal. The proposed recognition system was tested over a language, speaker, and context independent emotional speech database. Prosodic, Mel-frequency cepstral coefficient (MFCC), and formant frequency features are extracted from the speech utterances. We perform feature selection by using the stepwise method based on Mahalanobis distance. The selected features are used to classify the speeches into their corresponding emotional classes. Different classification algorithms including maximum likelihood classifier (MLC), Gaussian mixture model (GMM), neural network (NN), K-nearest neighbors (K-NN), and Fisher's linear discriminant analysis (FLDA) are compared in this study. The recognition results show that FLDA gives the best recognition accuracy by using the selected features. Yongjin Wang, Ling Guan |
MMSP | 2 |
| 2004 | Quantifying and recognizing human movement patterns from monocular video Images-part I: a new framework for modeling human motionabstractResearch into tracking and recognizing human movement has so far been mostly limited to gait or frontal posing. Part I of this paper presents a continuous human movement recognition (CHMR) framework which forms a basis for the general biometric analysis of continuous human motion as demonstrated through tracking and recognition of hundreds of skills from gait to twisting saltos. Part II of this paper presents CHMR applications to the biometric authentication of gait, anthropometric data, human activities, and movement disorders. In Part I of this paper, a novel three-dimensional color clone-body-model is dynamically sized and texture mapped to each person for more robust tracking of both edges and textured regions. Tracking is further stabilized by estimating the joint angles for the next frame using a forward smoothing particle filter with the search space optimized by utilizing feedback from the CHMR system. A new paradigm defines an alphabet of dynemes, units of full-body movement skills, to enable recognition of diverse skills. Using multiple hidden Markov models, the CHMR system attempts to infer the human movement skill that could have produced the observed sequence of dynemes. The novel clone-body-model and dyneme paradigm presented in this paper enable the CHMR system to track and recognize hundreds of full-body movement skills, thus laying the basis for effective biometric authentication associated with full-body motion and body proportions. Richard D. Green, Ling Guan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2004 | Quantifying and recognizing human movement patterns from monocular video images-part II: applications to biometricsabstractBiometric authentication of gait, anthropometric data, human activities, and movement disorders are presented in this paper using the continuous human movement recognition (CHMR) framework introduced in Part I. A novel biometric authentication of anthropometric data is presented based on the realization that no one is average-sized in as many as ten dimensions. These body part dimensions are quantified using the CHMR body model. Gait signatures are then evaluated using motion vectors, temporally segmented by gait dynemes, and projected into a gait space for an eigengait-based biometric authentication. Left-right asymmetry of gait is also evaluated using robust CHMR left-right labeling of gait strides. Accuracy of the gait signature is further enhanced by incorporating the knee-hip angle-angle relationship popular in biomechanics gait research, together with other gait parameters. These gait and anthropometric biometrics are fused to further improve accuracy. The next biometric identifies human activities which require a robust segmentation of the many skills encompassed. For this reason, the CHMR activity model is used to identify various activities from making coffee to using a computer. Finally, human movement disorders were evaluated by studying patients with dopa-responsive Parkinsonism and age-matched normals who were videotaped during several gait cycles to determine a robust metric for classifying movement disorders. The results suggest that the CHMR system enabled successful biometric authentication of anthropometric data, gait signatures, human activities, and movement disorders. Richard D. Green, Ling Guan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2004 | Retrieval for color artistry conceptsabstractThis paper presents a work on the retrieval of artworks for color artistry concepts. First we affirm the view that the Query-by-Example paradigm fundamental to the current content-based retrieval systems is able to extend only limited usefulness. We then propose a concept-based retrieval engine based on the generative grammar of elemental concepts methodology. In the latter, the language by which color artistry concepts are communicated in artworks is used to operate semantic searches. The color artistry language is explicated into elemental concepts and the associated generative grammar. The elemental concepts are used to index the artworks, while the generative grammar is used to facilitate post-coordinate expression of color artistry concept queries by using the elemental concepts. Jose A. Lay, Ling Guan |
IEEE Trans. Image Process. | 2 |
| 2004 | An interactive approach for CBIR using a network of radial basis functionsabstractAn important requirement for constructing effective content-based image retrieval (CBIR) systems is accurate characterization of visual information. Conventional nonadaptive models, which are usually adopted for this task in simple CBIR systems, do not adequately capture all aspects of the characteristics of the human visual system. An effective way of addressing this problem is to adopt a "human-computer" interactive approach, where the users directly teach the system about what they regard as being significant image features and their own notions of image similarity. We propose a machine learning approach for this task, which allows users to directly modify query characteristics by specifying their attributes in the form of training examples. Specifically, we apply a radial-basis function (RBF) network for implementing an adaptive metric which progressively models the notion of image similarity through continual relevance feedback from users. Experimental results show that the proposed methods not only outperform conventional CBIR systems in terms of both accuracy and robustness, but also previously proposed interactive systems. Paisarn Muneesawang, Ling Guan |
IEEE Trans. Multim. | 2 |
| 2003 | Tracking human movement patterns using particle filteringabstractAt least 32 joint related degrees of freedom need to be estimated to reliably track the human body in 3D. The particle filter is robust to distracting clutter by maintaining multiple hypotheses for each of these joint angles. Real-time tracking is difficult however with the computational overhead of such a large search space. This paper optimizes this search space utilizing feedback from a Continuous Human Movement Recognition (CHMR) system and improves the robustness and efficiency of each particle calculation using a novel body model. ne joint angles are estimated for the next frame using a Particle filter with forward smoothing. A new paradigm enables the temporal segmentation of continuous motion into dynemes. Using HMM, the CHMR system attempts to infer the human movement skill that could have produced the observed sequence of dynemes. Hundreds of movement skills, from gait to saltos, are successfully tracked and recognized. Richard D. Green, Ling Guan |
ICASSP (3) | 2 |
| 2003 | Automatic relevance feedback for video retrievalabstractThe paper presents an automatic relevance feedback method for improving retrieval accuracy in video databases. We first demonstrate a representation based on a template-frequency model (TFM) that allows the full use of the temporal dimension. We then integrate the TFM with a self-training neural network structure to capture adaptively different degrees of visual importance in a video sequence. Forward and backward signal propagation is the key in this automatic relevance feedback method in order to enhance retrieval accuracy. Paisarn Muneesawang, Ling Guan |
ICASSP (3) | 2 |
| 2003 | Tracking human movement patterns using particle filteringabstractAt least 32 joint related degrees of freedom need to be estimated to reliably track the human body in 3D. The particle filter is robust to distracting clutter by maintaining multiple hypotheses for each of these joint angles. Real-time tracking is difficult however with the computational overhead of such a large search space. This paper optimizes this search space utilizing feedback from a continuous human movement recognition (CHMR) system and improves the robustness and efficiency of each particle calculation using a novel body model. The joint angles are estimated for the next frame using a particle filter with forward smoothing. A new paradigm enables the temporal segmentation of continuous motion into dynemes. Using HMM, the CHMR system attempts to infer the human movement skill that could have produced the observed sequence of dynemes. Hundreds of movement skills, from gait to saltos, are successfully tracked and recognized. Richard D. Green, Ling Guan |
ICME | 2 |
| 2003 | Centralized peer-to-peer streaming with layered videoabstractFaster computational power and higher network bandwidth facilitates Internet applications to improve the way we live, work and play. For video streaming applications, it is a challenge to provide the scalability while maintaining the central manageability and the robustness. In this paper, we propose the centralized peer-to-peer streaming protocol (P2PSP), which features the centralized management, guaranteed perceived quality of service (PQoS), while decentralized the traffic loads. To evaluate the protocol, a layered video coding technique based on 3D discrete cosine transform (DCT) is designed. This codec applies self-organizing tree map (SOTM) to generate cascaded vector quantization codebooks. We demonstrate the effectiveness of the centralized P2PSP by comparing the video streaming simulations under the proposed architecture and the client/server architecture. Ivan Lee 0001, Ling Guan |
ICME | 2 |
| 2003 | Image retrieval with embedded sub-class information using Gaussian mixture modelsabstractThis paper describes content-based image retrieval techniques within the relevance feedback framework. The Gaussian mixture model (GMM) is used to characterize sub-class information to increase retrieval accuracy and reduce number of interactions during a query session. The implementation of GMM is based on the radial basis function using a new learning algorithm that can cope with small training samples in the relevance feedback cycle. The proposed retrieval system is successfully applied to image databases of very large sizes, and experimental results show that the proposed system competes favorably with the other recently proposed interactive systems. Paisarn Muneesawang, Ling Guan |
ICME | 2 |
| 2003 | Automatic relevance feedback for video retrievalabstractThis paper presents an automatic relevance feedback method for improving retrieval accuracy in video database. We first demonstrate a representation based on a template-frequency model (TFM) that allows the full use of the temporal dimension. We then integrate the TFM with a self-training neural network structure to adaptively capture different degrees of visual importance in a video sequence. Forward and backward signal propagation is the key in this automatic relevance feedback method in order to enhance retrieval accuracy. Paisarn Muneesawang, Ling Guan |
ICME | 2 |
| 2003 | Content-based image retrieval using a Gaussian mixture model in the wavelet domain
Xiao-Ping Zhang 0002, Ling Guan |
VCIP | 3 |
| 2002 | A scalable video codec design for streaming over distributed peer-to-peer networkabstractWith the growing popularity of the Internet, information is easier to distribute and obtain than ever. Among the efforts of developing practical applications over the Internet, video streaming has drawn a lot of attention both from academia and industry. However, the best-effort nature of the Internet poses many challenges to meet the bandwidth, delay, loss, and QoS requirements for transmitting video. In this paper, we address these problems by proposing new approaches to maximize the perceptual quality for streaming video. Our contributions in this paper are: (1) our proposal of video streaming over the Peer-to-Peer network to complement the existing client-server model (2) our investigations on an innovative scalable video coding technique based on the inter-subband redundancy removal using the neural network techniques. Ling Guan |
GLOBECOM | 2 |
| 2002 | A fuzzy blur algorithm to adaptive blind image deconvolutionabstractThis paper proposes a new approach to blind image deconvolution based on fuzzy blur interference algorithm. Conventional blind algorithms require a crisp decision to be made on the structure of the blurring function prior to formulation. This creates a dilemma as the complete blur information is usually unknown a priori. Most blind algorithms either employ absolute but inflexible parametric modeling or ignore the parametric blur knowledge completely. This paper presents a fuzzy approach to resolve this difficulty by constructing a soft model set consisting of parametric estimates of the current blur. The relevance of these estimates are evaluated, and integrated to form a fuzzy blur using compositional inference rule. The main feature of the technique lies in its ability to incorporate domain knowledge while preserving the flexibility of the scheme. Experimental results show that the technique is effective in restoring blurred, noisy images without prior knowledge of the blur. Kim-Hui Yap, Ling Guan |
ICARCV | 2 |
| 2002 | Classification of images as a pre-processing step for image retrievalabstractThe indexing and retrieval of images in a multimedia database can be a very time consuming process, considering the amount of visual data stored in digital video libraries (DVLs). This paper proposes a semiautomated method to categorize images into groups as a pre-processing step in order to enable faster retrieval from a DVL. The wavelet transform is used to extract low-level features such as colour, texture and shape from images. The extracted features are employed to categorize the images using a novel neural network, the self-organizing tree map. The classification results are then used to train a feed forward neural network to identify the image classes an incoming query image belongs to. Initial experiments show promising results. Tat Loong Chan, Ling Guan |
ICASSP | 2 |
| 2002 | Hierarchical cluster model for perceptual image processingabstractWe propose a method for extracting object symmetries from a digital image. To achieve this we examine the way in which the human visual system processes and organises visual information. Psychological evidence is combined with physiological processing. The evidence is based on image structure and the processing is based on the Hierarchical Cluster Model (HCM) which is used to model the human brain. Jonathan Randall, Ling Guan, Wanqing Li 0001 |
ICASSP | 2 |
| 2002 | Scene change detection using DC coefficientsabstractWe propose an algorithm to detect shot transitions in MPEG video data in the compressed domain. The algorithm is based on the conventional solution, where energy histograms of DC (discrete cosine) coefficients are used to calculate the distance between consecutive I/P frames, and the DC coefficients of the P-frames are obtained by frame conversion. We enhance the result by using the ratio between two sliding windows to attenuate the low-pass filtered frame distance, and this results in the amplification of the transitional regions. The advantage of our method is that it achieves high detection rates with low computational complexity. The algorithm is equally applicable in detecting sharp transitions and dissolved transitions. Oliver K. Bao, Ling Guan |
ICIP (2) | 2 |
| 2002 | Video coding algorithm using 3-D DCT and vector quantizationabstractThis paper describes a compression algorithm based on 3-D discrete cosine transform (DCT) and vector quantization (VQ) in DCT domain. 3-D DCT is the extension of the conventional block-based 2-D DCT to explore the temporal correlation of video sequences in time. Instead of using scalar quantization which is normally used in popular standards, we employ VQ on the zig-zag sequence of DCT coefficients. Preliminary experimental results show that the proposed algorithm has significantly out-performed a standard MPEG-1 encoder. T. H. Lai, Ling Guan |
ICIP (1) | 2 |
| 2002 | Minimizing user interaction by automatic and semi-automatic relevance feedback for image retrievalabstractThis paper describes the unsupervised-interactive learning method, using the self-organizing tree map (SOTM) architecture, for the automation of relevance feedback (RF) in content-based image retrieval. The SOTM is shown to exhibit good behavior in relevance classification; providing a possible solution to minimizing user interactions in both fully automatic and semiautomatic domains, while achieving high retrieval accuracy in the context of adaptive retrieval. Computer simulation shows this system is very effective when applied to compressed domain retrieval systems for texture retrieval and the JPEG photograph database applications. Paisarn Muneesawang, Ling Guan |
ICIP (2) | 2 |
| 2002 | Approximating color consistency in retrieval using energy histograms of DCT coefficientsabstractThe paper presents the use of the energy histogram of DCT coefficients to approximate color consistency in retrieving gamma-corrected and filter-edited JPEG images. The work has been used for coarse searching of global color similarity in the SoloArt project. Jose A. Lay, Ling Guan |
ICME (1) | 2 |
| 2002 | Approximating color consistency in retrieval using energy histograms of DCT coefficientsabstractThis paper presents the use of the energy histogram of DCT coefficients to approximate color consistency in retrieving gamma-corrected and filter-edited JPEG images. The work has been used for coarse searching of global color similarity in the SoloArt project. Jose A. Lay, Ling Guan |
ICME (1) | 2 |
| 2002 | The hierarchical cluster model for image region segmentationabstractThe hierarchical cluster model (HCM), a neural network inspired by the human brain (see Sutton, J., Harvard Medical School, MIT, Neural Systems Group, Technical Report, 1995), is demonstrated for the purpose of region segmentation in digital images. Starting with an over segmented image, regions are merged based on evidence of a valid edge between the two regions. Unlike Sutton's work, in which the HCM is used to recall a set of pre-trained memory patterns, the HCM in our work demonstrates unsupervised decision making capabilities. Jonathan Randall, Ling Guan, Wanqing Li 0001 |
ICME (1) | 2 |
| 2002 | A computational reinforced learning scheme to blind image deconvolutionabstractThis paper presents a new approach to adaptive blind image deconvolution based on computational reinforced learning in an attractor-embedded solution space. The new technique develops an evolutionary strategy that generates the improved blur and image populations progressively. A dynamic attractor space is constructed by integrating the knowledge domain of the blur structures into the algorithm. The attractors are predicted using a maximum a posteriori estimator and their relevance is evaluated with respect to the computed blurs. We develop a novel reinforced mutation scheme that combines stochastic search and pattern acquisition throughout the blur identification. It enhances the algorithmic convergence and reduces the computational cost significantly. The new technique is robust in alleviating the constraints and difficulties encountered by most conventional methods. Experimental results show that the new algorithm is effective in restoring the degraded images and identifying the blurs. Kim-Hui Yap, Ling Guan |
IEEE Trans. Evol. Comput. | 2 |
| 2002 | Guest editorial special issue on intelligent multimedia processing
Ling Guan, Tülay Adali, Shigeru Katagiri, Jan Larsen, José C. Príncipe |
IEEE Trans. Neural Networks | 1 |
| 2002 | Automatic machine interactions for content-based image retrieval using a self-organizing tree map architectureabstractIn this paper, an unsupervised learning network is explored to incorporate a self-learning capability into image retrieval systems. Our proposal is a new attempt to automate recursive content-based image retrieval. The adoption of a self-organizing tree map (SOTM) is introduced, to minimize the user participation in an effort to automate interactive retrieval. The automatic learning mode has been applied to optimize the relevance feedback (RF) method and the single radial basis function-based RF method. In addition, a semiautomatic version is proposed to support retrieval with different user subjectivities. Image similarity is evaluated by a nonlinear model, which performs discrimination based on local analysis. Experimental results show robust and accurate performance by the proposed method, as compared with conventional noninteractive content-based image retrieval (CBIR) systems and user controlled interactive systems, when applied to image retrieval in compressed and uncompressed image databases. Paisarn Muneesawang, Ling Guan |
IEEE Trans. Neural Networks | 2 |
| 2001 | Blind adaptive detection for CDMA systems based on regularized independent component analysisabstractWe present a new approach to blind adaptive detection for CDMA systems based on regularized independent component analysis (ICA). Classical ICA algorithms are effective in separating linearly weighted signal mixtures consisting of subGaussian and superGaussian signals. However, they do not incorporate any information of the weighting matrix, in this case, the user's signature sequence into the formulation. This results in underutilization of the information available. To address this difficulty, we propose a new ICA algorithm that combines a contrast function and a regularization functional to integrate the information of the user's signature. A blind adaptive detector based on stochastic gradient optimization of the new cost function is derived. Simulation results show that the new technique provides good interference suppression, fast convergence and low BER performance when compared with other blind detectors. Kim-Hui Yap, Ling Guan, Jamie S. Evans |
GLOBECOM | 2 |
| 2001 | Interactive CBIR using RBF-based relevance feedback for WT/VQ coded imagesabstractPowerful interfaces provide great potential for retrieval systems to adapt to dynamic user needs and allow a more accurate modeling of image similarity from the users' point of view. We propose a novel method within the interactive framework. It allows the users to directly modify the system characteristics by specifying their desired image attributes in the form of training samples. More specifically, we have adopted a radial basis function (RBF) method for implementing an adaptive metric which progressively models the notion of image similarity through continual feedback from the users. The proposed approach has been integrated into an image retrieval system using images compressed by wavelet transform and vector quantization coders. Comparisons with some of the recent systems using the standard texture database indicate that the proposed method provides the more favorable retrieval result. Paisarn Muneesawang, Ling Guan |
ICASSP | 2 |
| 2001 | An attractor space approach to blind image deconvolutionabstractWe present a new approach to adaptive blind image deconvolution based on computational reinforced learning in attractor-embedded solution space. A new subspace optimization technique is developed to restore the image and identify the blur. Conjugate gradient optimization is employed to provide an adaptive image restoration while a new evolutionary scheme is devised to generate the high-performance blur estimates. The new technique is flexible as it does not suffer from various image or blur constraints imposed by most traditional blind methods. Experimental results show that the new algorithm is effective in blind deconvolution of images degraded under different blur structures and noise levels. Kim-Hui Yap, Ling Guan |
ICASSP | 2 |
| 2001 | Automatic similarity learning using SOTM for CBIR of the WT/VQ coded imagesabstractThe unsupervised learning network is explored to incorporate self-learning capability into image retrieval systems. More specifically, we propose the adoption of a self organizing tree map (SOTM) to implement a self-learning methodology that allows minimization of the role of users in an effort to automate interactive retrieval. This automatic-learning mode is applied to interactive retrieval strategies such as the radial basis function method and the relevance feedback method. The proposed method has been applied to retrieve the images compressed by wavelet transform and vector quantization coders. Retrieval performances are compared with conventional retrieval systems employing both non-interactive and user controlled interactive retrieval using the MIT texture database. The results obtained are compared favorably with preceding methods. Paisarn Muneesawang, Ling Guan |
ICIP (2) | 2 |
| 2001 | A neural network approach for learning image similarity in adaptive CBIRabstractThe adoption of neural network techniques is studied for the purpose of image retrieval. More specifically, we propose an adaptive retrieval system which incorporates learning capability into the image retrieval module where the network weights represent the adaptivity. This system can learn users' notions of similarity between images through the continual relevance feedback from the users. Accordingly it makes the proper adjustment to improve performance. This retrieval system has demonstrated its effectiveness in performance. It is confirmed by simulations conducted for applications such as texture retrieval and retrieval of DCT compressed images. Paisarn Muneesawang, Ling Guan |
MMSP | 2 |
| 2001 | Dynamic resource allocation via video content and short-term traffic statisticsabstractThe reliable and efficient transmission of high-quality variable bit rate (VBR) video through the Internet generally requires network resources be allocated in a dynamic fashion. This includes the determination of when to renegotiate for network resources, as well as how much to request at a given time. The accuracy of any resource request method depends critically on its prediction of future traffic patterns. Such a prediction can be performed using the content and traffic information of short video segments. This paper presents a systematic approach to select the best features for prediction, indicating that while content is important in predicting the bandwidth of a video hit stream, the use of both content and available short-term bandwidth statistics can yield significant improvements. A new framework for traffic prediction is proposed in this paper; experimental results show a smaller mean-square resource prediction error and higher overall link utilization. Min Wu 0001, Robert A. Joyce, Hau-San Wong, Ling Guan, Sun-Yuan Kung |
IEEE Trans. Multim. | 4 |
| 2001 | A neural learning approach for adaptive image restoration using a fuzzy model-based network architectureabstractWe address the problem of adaptive regularization in image restoration by adopting a neural-network learning approach. Instead of explicitly specifying the local regularization parameter values, they are regarded as network weights which are then modified through the supply of appropriate training examples. The desired response of the network is in the form of a gray level value estimate of the current pixel using weighted order statistic (WOS) filter. However, instead of replacing the previous value with this estimate, this is used to modify the network weights, or equivalently, the regularization parameters such that the restored gray level value produced by the network is closer to this desired response. In this way, the single WOS estimation scheme can allow appropriate parameter values to emerge under different noise conditions, rather than requiring their explicit selection in each occasion. In addition, we also consider the separate regularization of edges and textures due to their different noise masking capabilities. This in turn requires discriminating between these two feature types. Due to the inability of conventional local variance measures to distinguish these two high variance features, we propose the new edge-texture characterization (ETC) measure which performs this discrimination based on a scalar value only. This is then incorporated into a fuzzified form of the previous neural network which determines the degree of membership of each high variance pixel in two fuzzy sets, the EDGE and TEXTURE fuzzy sets, from the local ETC value, and then evaluates the appropriate regularization parameter by appropriately combining these two membership function values. Hau-San Wong, Ling Guan |
IEEE Trans. Neural Networks | 2 |
| 2000 | Congestion control of compressed video traffic over ATM networksabstractVideo traffic is expected to account for a significant share of the traffic volume in future asynchronous transfer mode (ATM) networks. MPEG-2, proposed by the Moving Picture Expert Group, is one of the most promising compression techniques for such applications. One of the critical issues in MPEG-2 is to realize effective variable bit rate (VBR) video transfer thorough ATM networks. The leaky bucket (LB) scheme has been widely accepted as the usage parameter control (UPC) mechanism to police the VBR sources. We propose a new adaptive dynamic leaky bucket (ADLB) congestion control mechanism, which is based on the LB scheme. Unlike the conventional LB, the leak rate of the ADLB is controlled using delayed feedback information of available bandwidth sent by the network. The simulation results show that overall cell loss and delay are reduced significantly at the ATM switch node. Gajendra Sisodia, Ling Guan, Subrata De, Mehran Dowlatshahi |
ICCCN | 2 |
| 2000 | Human Gait and Posture Analysis for Diagnosing Neurological DisordersabstractThis paper describes a number of new techniques to enhance the performance of a video analysis system, free from motion markers and complicated setup procedures, for the purpose of quantitatively identifying gait abnormalities in static human posture analysis. Visual features are determined from still frame images out of the entire walking sequence. The features are used as a guide to train a neural network, in an attempt to providing assistance to clinicians in diagnosing patients with neurological disorders. Howard Lee, Ling Guan, John A. Burne |
ICIP | 2 |
| 2000 | Multiresolution-Histogram Indexing and Relevance Feedback Learning for Image RetrievalabstractTwo fundamental aspects for content-based image retrieval system are studied: visual feature extraction and retrieval system design. (i) Feature extraction shares some common properties with image compression where the multiresolution nature of wavelet decomposition can be exploited. When wavelet coefficients are vector quantized, the information content in each spatial-frequency subband is mapped onto coding labels. Thus, the statistics of these labels will reflect the subband characteristics of an image. This constitutes a feature vector that is then used for image matching. (ii) The proposed retrieval system supports queries based on system-user interaction that utilise a non-Euclidean similarity measure. A non-linear function based on a radial basis function (RBF) is adopted for characterising the behaviour of human users in an interactive section where relevance feedback is applied. Experimental results show that the retrieval efficiency is considerably improved by implementing the proposed approach. Paisarn Muneesawang, Ling Guan |
ICIP | 2 |
| 2000 | Characterization of Perceptual Importance for Object-Based Image SegmentationabstractWe propose a machine learning approach for characterizing the perceptual importance of particular regions in an image. A modular neural network architecture is adopted for encoding our usual notion of a perceptually important region in such a way that generalization of this knowledge to previously unseen images is possible. Specifically, users are allowed to specify examples of perceptually significant regions in images, which are then incorporated as training data for the network. An important characteristic of this approach is its provision for grouping distinct regions into a single perceptually significant area through the previous user guidance, unlike conventional segmentation approaches which partition the image into homogeneous regions without further specifying the relationship between these regions. Hau-San Wong, Ling Guan |
ICIP | 2 |
| 2000 | A Recursive Soft-Decision PSF and Neural Network Approach to Adaptive Blind Image RegularizationabstractWe present a new approach to adaptive blind image regularization based on a neural network and soft-decision blur identification. We formulate blind image deconvolution into a recursive scheme by projecting and optimizing a novel cost function with respect to its image and blur subspaces. The new algorithm provides a continual blur adaptation towards the best-fit parametric structure throughout the restoration. It integrates the knowledge of real-life blur structures without compromising its flexibility in restoring images degraded by other nonstandard blurs. A nested neural network, called the hierarchical cluster model is employed to provide an adaptive, perception-based restoration. On the other hand, conjugate gradient optimization is adopted to identify the blur. Experimental results show that the new approach is effective in restoring the degraded image without the prior knowledge of the blur. Kim-Hui Yap, Ling Guan |
ICIP | 2 |
| 2000 | A model-based neural network for edge characterization
Hau-San Wong, Terry Caelli, Ling Guan |
Pattern Recognit. | 3 |
| 2000 | Application of evolutionary programming to adaptive regularization in image restorationabstractImage restoration is a difficult problem due to the ill-conditioned nature of the associated inverse filtering operation, which requires regularization techniques. The choice of the corresponding regularization parameter is thus an important issue since an incorrect choice would either lead to noisy appearances in the smooth regions or excessive blurring of the textured regions. In addition, this choice has to be made adaptively across, different image regions to ensure the best subjective quality for the restored image. We employ evolutionary programming (EP) to solve this adaptive regularization problem by generating a population of potential regularization strategies, and allowing them to compete under a new error measure which characterizes a large class of images in terms of their local correlational properties. The nonavailability of explicit gradient information for this measure motivates the adoption of EP techniques for its optimization, which allows efficient search at multiple error surface points. The adoption of EP also allows the broadening of the range of possible cost functions for image processing so that we can choose the most relevant function rather than the most tractable one for a particular image processing application. Hau-San Wong, Ling Guan |
IEEE Trans. Evol. Comput. | 2 |
| 2000 | A CAD System for the Automatic Detection of Clustered Microcalcification in Digitized Mammogram FilmsabstractClusters of microcalcifications in mammograms are an important early sign of breast cancer. This paper presents a computer-aided diagnosis (CAD) system for the automatic detection of clustered microcalcifications in digitized mammograms. The proposed system consists of two main steps. First, potential microcalcification pixels in the mammograms are segmented out by using mixed features consisting of wavelet features and gray level statistical features, and labeled into potential individual microcalcification objects by their spatial connectivity. Second, individual microcalcifications are detected by using a set of 31 features extracted from the potential individual microcalcification objects. The discriminatory power of these features is analyzed using general regression neural networks via sequential forward and sequential backward selection methods. The classifiers used in these two steps are both multilayer feedforward neural networks. The method is applied to a database of 40 mammograms (Nijmegen database) containing 105 clusters of microcalcifications. A free-response operating characteristics (FROC) curve is used to evaluate the performance. Results show that the proposed system gives quite satisfactory detection performance. In particular, a 90% mean true positive detection rate is achieved at the cost of 0.5 false positive per image when mixed features are used in the first step and 15 features selected by the sequential backward selection method are used in the second step. However, we must be cautious when interpreting the results, since the 20 training samples are also used in the testing step. Songyang Yu, Ling Guan |
IEEE Trans. Medical Imaging | 2 |
| 2000 | Weight assignment for adaptive image restoration by neural networksabstractThis paper presents a scheme for adaptively training the weights, in terms of varying the regularization parameter, in a neural network for the restoration of digital images. The flexibility of neural-network-based image restoration algorithms easily allow the variation of restoration parameters such as blur statistics and regularization value spatially and temporally within the image. This paper focuses on spatial variation of the regularization parameter.We first show that the previously proposed neural-network method based on gradient descent can only find suboptimal solutions, and then introduce a regional processing approach based on local statistics. A method is presented to vary the regularization parameter spatially. This method is applied to a number of images degraded by various levels of noise, and the results are examined. The method is also applied to an image degraded by spatially variant blur. In all cases, the proposed method provides visually satisfactory results in an efficient way. Stuart W. Perry, Ling Guan |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 1999 | Image retrieval based on energy histograms of the low frequency DCT coefficientsabstractWith the increasing popularity of the use of compressed images, an intuitive approach for lowering computational complexity towards a practically efficient image retrieval system is to propose a scheme that is able to perform retrieval computation directly in the compressed domain. In this paper, we investigate the use of energy histograms of the low frequency DCT coefficients as features for the retrieval of DCT compressed images. We propose a feature set that is able to identify similarities on changes of image-representation due to several lossless DCT transformations. We then use the features to construct an image retrieval system based on the real-time image retrieval model. We observe that the proposed features are sufficient for performing high level retrieval on medium size image databases. And by introducing transpositional symmetry, the features can be brought to accommodate several lossless DCT transformations such as horizontal and vertical mirroring, rotating, transposing, and transversing. Jose A. Lay, Ling Guan |
ICASSP | 2 |
| 1999 | Edge characterization using a model-based neural networkabstractIn this paper, we investigate the feasibility of characterizing significant image edges using a model-based neural network with modular architecture. Instead of employing traditional mathematical models for characterization, we ask human users to select what they regard as significant features on an image, and then incorporate these selected edges directly as training examples for the network. Unlike conventional edge detection schemes where decision thresholds have to be specified, the current NN-based edge characterization scheme implicitly represents these decision parameters in the form of network weights which are updated during the training process. Experiments have confirmed that the resulting network is capable of generalizing this previously acquired knowledge to identify important edges in images not included in the training set. Most importantly, the current approach is very robust against noise contaminations, such that no re-training of the network is required when it is applied to noisy images. Hau-San Wong, Terry Caelli, Ling Guan |
ICASSP | 3 |
| 1999 | Feature selection using general regression neural networks for the automatic detection of clustered microcalcificationsabstractGeneral regression neural networks (GRNNs) are proposed for selecting the most discriminating features for the automatic detection of clustered microcalcifications in digital mammograms. Previously, We have designed an image processing system for detecting clustered microcalcifications. The system uses wavelet coefficients and feed forward neural networks to identify possible microcalcification pixels and a set of structure features to locate individual microcalcifications. In this work, more features are extracted, and the most discriminating features are selected through the analysis of the GRNNs. The selected features are incorporated into our image processing system and applied to a database of 40 mammograms (Nijmegen database) containing 105 clusters of microcalcifications. Free response operating characteristics (FROG) curves are used to evaluate the performance. Results show that, by incorporating the proposed feature selection scheme, the performance of our system is improved significantly. Songyang Yu, Ling Guan |
ICASSP | 2 |
| 1999 | Feature Extraction of Chromosomes from 3D Confocal Microscope ImagesabstractUse of a confocal light microscope enables biologists to observe dividing cells (living or presented) within a 3D volume that can be visualised from multiple aspects. The Nomarski differential interference contrast (DIG) mode used for imaging translucent specimens, such as chromosomes, produces images not suitable for volume rendering. Segmentation of the chromosomes from this data is thus necessary. Kohonen's self-organising feature map (SOFM) was used to perform segmentation, based on a collection of various statistics or features defining the image. In the past, classical features such as the mean and variance of pixel intensities have been used, providing reasonable extraction of chromosome bodies, while only mildly resolving surface detail. In this investigation, a local energy feature detector was implemented, producing an alternative image statistic based on phase congruency in the image. This, along with combinations of other image statistics, was applied to the SOFM, producing a series of resultant 3D images exhibiting vast improvements in the level of detail defining the internal structure of the specimen chromosomes. Matthew J. Kyan, Ling Guan, Matthew R. Arnison, Carol J. Cogswell |
ICIP (2) | 2 |
| 1999 | A Fuzzy Model-Based Neural Network for Adaptive Regularization in Image RestorationabstractWe address the problem of adaptive regularization in image restoration by adopting a neural network learning approach. The local regularization parameter values are regarded as network weights which are then modified through the supply of appropriate training examples. We also consider the separate regularization of edges and textures due to their different noise masking capabilities, which in turn requires discrimination between these two feature types. A new edge-texture characterization (ETC) measure is derived and incorporated into a fuzzified form of the previous NN for the above purpose. Hau-San Wong, Ling Guan |
ICIP (1) | 2 |
| 1999 | Image Restoration Based on Hierarchical Cluster Model with Evolutionary OptimizationabstractIn this paper, a new approach to adaptive image regularization based on Hierarchical Cluster Model (HCM) with evolutionary optimization is proposed. HCM is a hierarchical neural network with distributed clusters. Its sparse synaptic connections and parallel structure reduce the computational cost of restoration. Adaptive restoration is achieved by assigning entries of an optimized regularization vector to each homogeneous cluster of the image. The clusters are restored in the order of smooth to texture and edge regions to minimize the regularization error. An evolutionary scheme is employed to improve the performance profile of the restored image and optimize the regularization vector. Experimental results show that the new approach is superior in suppressing noise and ringing at the smooth background while preserving fine details at the texture and edge regions effectively. Kim-Hui Yap, Ling Guan |
ICIP (1) | 2 |
| 1999 | Modularity in neural computingabstractThis paper considers neural computing models for information processing in terms of collections of subnetwork modules. Two approaches to generating such networks are studied. The first approach includes networks with functionally independent subnetworks, where each subnetwork is designed to have specific functions, communication, and adaptation characteristics. The second approach is based on algorithms that can actually generate network and subnetwork topologies, connections, and weights to satisfy specific constraints. Associated algorithms to attain these goals include evolutionary computation and self-organizing maps. We argue that this modular approach to neural computing is more in line with the neurophysiology of the vertebrate cerebral cortex, particularly with respect to sensation and perception. We also argue that this approach has the potential to aid in solutions to large-scale network computational problems - an identified weakness of simply defined artificial neural networks. Terry Caelli, Ling Guan, Wilson Wen |
Proc. IEEE | 2 |
| 1999 | Scanning the issue/technologyabstractProvides an overview of the technical articles and features presented in this issue. David B. Fogel, Toshio Fukuda, Ling Guan |
Proc. IEEE | 3 |
| 1998 | Neural vision system and applications in image processing and analysisabstractWe present a computer vision system based on an integrated neural network architecture. In the low level vision subsystem, a network of networks-a biologically inspired network is used to recursively perform filtering, segmentation and edge detection; in the intermediate level and the high level, hierarchically structured arrays of self-organizing tree maps-extension of the popular self-organizing map are utilized to carry out image/feature analysis. The system has been applied to solve a number of real world problems. Some interesting and encouraging results are reported. Ling Guan, Stuart W. Perry, Raniero Romagnoli, Hau-San Wong, Haosong Kong |
ICASSP | 1 |
| 1998 | Perception based adaptive image restorationabstractThis paper presents an image restoration technique which uses a cost function based on a novel image error measure. The cost function presented here takes into account local statistical information of the image when performing restoration. It is shown that this technique compares favourably with other techniques, especially when applied to colour images. Stuart W. Perry, Ling Guan |
ICASSP | 2 |
| 1998 | New Statistical Model for VBR Video Traffic in ATM NetworksabstractWe present a new approach to modelling variable bit rate (VBR) coded video sources in asynchronous transfer mode (ATM) networks. Unlike the existing methods which model the number of cells generated by the coder for a sequence of video frames, the new approach improves modelling accuracy by considering the characteristics of different cells generated by the coder, and modelling the number of cells in each type of macroblocks of a frame separately. The model is tested by comparing the cell loss rate in simulation of an ATM switch to cell loss rate produced when traces generated by the model are used as the source. Comparison with the existing models are performed in the area of measuring the quality of their predictions for network performance and the grade of service experienced by a user. Gajendra Sisodia, Mark Hedley, Ling Guan, Subrata De |
ICCCN | 3 |
| 1998 | Self-organizing map for segmenting 3D biological imagesabstractAn image processing method for features extraction and segmentation from three-dimensional (3D) image datasets is presented. Kohonen's self-organizing map (SOM) is used to perform segmentation. Previously, the segmentation method worked on a 2D dataset based on a projection of the three-dimensional dataset (Nguyen et al., 1998). Our 3D approach to segment biological images preserves the 3D object orientations with respect to the surrounding cell volume. A few examples from genetics and brain analysis are provided in order to demonstrate the performance of the proposed method. Luigi Cinque, Raniero Romagnoli, Stefano Levialdi, P. T. A. Nguyen, Ling Guan |
ICPR | 5 |
| 1998 | Source Model for VBR Coded Video Traffic in ATM NetworksabstractWe present a new approach to modelling variable bit rate (VBR) coded video sources in asynchronous transfer mode (ATM) networks. Unlike the existing methods which model the number of cells generated by the coder for a sequence of video frames, the new approach improves the modelling accuracy by considering the characteristics of different cells generated by the coder, and modelling the number of cells in each type of macroblocks of a frame separately. The model is tested by comparing the autocorrelation function and mean queue size in the buffer by simulation of an ATM switch to the mean queue size produced when traces generated by the model are used as the source. Comparisons with the existing models are performed in the area of measuring the quality of their predictions for the network performance and the grade of service experienced by a user. Gajendra Sisodia, Mark Hedley, Subrata De, Ling Guan |
LCN | 4 |
| 1998 | A combined evolution method for associative memory networks
Andrew C. C. Cheng, Ling Guan |
Neural Networks | 2 |
| 1997 | A network of networks processing model for image regularizationabstractWe introduce a network of networks (NoN) model to solve image regularization problems. The method is motivated by the fact that natural image formation involves both local processing and globally coordinated parallel processing. Both forms are readily implemented using an NoN architecture. The modeling is very powerful in that it achieves high-quality adaptive processing, and it reduces the computational difference between inhomogeneous and homogeneous conditions. This method is able to provide fast, quality imaging in early vision, and its replicating structure and sparse connectivity readily lend themselves to hardware implementations. Ling Guan, James A. Anderson, Jeffrey P. Sutton |
IEEE Trans. Neural Networks | 1 |
| 1996 | An adaptive approach for removing impulsive noise in digital imagesabstractA new approach is proposed to eliminate impulsive noise with Gaussian or uniform random distribution in digital images. The method is based on impulse noise detection by means of a self-organizing neural network and a class of noise-exclusive adaptive filters. The filtering scheme presented can suppress impulse noise effectively as well as preserving image edges and fine details. Experimental results demonstrated that the performance of the noise-exclusive adaptive filters are superior to that of traditional median filter, central weighted median filter and multistage median filter. Haosong Kong, Ling Guan |
ICASSP | 2 |
| 1996 | Enhancement and real-time analysis of an adaptive impulsive noise removal methodabstractThis paper presents the enhancement and real-time processing features of an adaptive filter for the removal of impulse noise in TV picture transmission. The basic method, which was tested on synthetic data, is first enhanced to deal with real TV pictures suffering impulse noise. Then the suitability of the method in term of real-time processing is investigated. A numerical example shows that the proposed method provides high quality solutions in much shorter time than the traditional median-type filters. Haosong Kong, Ling Guan |
ICECCS | 2 |
| 1996 | A neural network adaptive filter for the removal of impulse noise in digital images
Haosong Kong, Ling Guan |
Neural Networks | 2 |
| 1996 | An optimal neuron evolution algorithm for constrained quadratic programming in image restorationabstractAn optimal neuron evolution algorithm for the restoration of linearly distorted images is presented in this paper. The proposed algorithm is motivated by the symmetric positive-definite quadratic programming structure inherent in restoration. Theoretical analysis and experimental results show that the algorithm not only significantly increases the convergence rate of processing, but also produces good restoration results. In addition, the algorithm provides a genuine parallel processing structure which ensures computationally feasible spatial domain image restoration. Ling Guan |
IEEE Trans. Syst. Man Cybern. Part A | 1 |
| 1995 | Design and implementation of a distributed real-time image processing systemabstractA new development in the design and implementation of a distributed, real-time image processing system is presented. The system uses an IBM personal computer as the front end to a remote computer via the Internet. The standard TCP/IP networking protocols are utilised to link the IBM-PC and high performance remote devices such as transputer networks and supercomputers. The access to the powerful remote computers enables the system to complete complex image processing tasks in real-time. During processing, the image is transferred to the remote machine and then transferred back to the PC for display. The system serves as a prototype for a full-feature image processing and analysis package, as well as a programming platform for the research and development of new image processing algorithms. D. M. Wu, Ling Guan, G. Lau, D. Rahija |
ICECCS | 2 |
| 1995 | Compression of digital mammogram databases using a near-lossless schemeabstractThe authors introduce a near-lossless scheme for the compression of digital mammogram databases. In the scheme a self-organizing neural network is first used to separate the breast area from the background. Then an optimized JPEG coding algorithm is introduced to code the segmented breast area only. The combined segmentation/compression procedure is motivated by the massive storage requirement of mammograms. The proposed scheme exploits the fact that a large proportion of the mammogram consists of uninteresting background, and the breast region occupies only a small area. The experimental results have confirmed that the scheme is capable of extending beyond the compression limits of conventional transform coding methods and achieving a far lower bit rate. As a result, the current approach provides an efficient means for the storage and transmission of digital mammogram databases. Hau-San Wong, Ling Guan, H. Hong |
ICIP | 2 |
| 1992 | Restoration of randomly blurred images via the maximum a posteriori criterionabstractThe maximum a posteriori (MAP) estimation technique is applied to the problem of restoring images distorted by noisy point spread functions and additive noise. The resulting MAP estimator is nonlinear and is obtained by numerically maximizing a conditional probability density function. The energy nonnegativity constraint is incorporated in the optimization process. Although the deblurring results are slightly inferior to those obtained by applying the Wiener criterion, the advantage of the MAP estimator lies in its significant suppression of noise. Ling Guan, Rabab K. Ward |
IEEE Trans. Image Process. | 1 |
| 1988 | A maximum a posteriori approach to the restoration of randomly distorted signalsabstractA maximum a posterior (MAP) approach to the restoration of randomly distorted signals is introduced. Since random distortion is due to the uncertainty in the impulse response of the physical system, it cannot be modelled as additive. The theoretical derivation of the MAP filter shows that it is necessary to solve a nonlinear optimization problem. The performance of the method is compared to that of a modified Wiener filter, and a numerical example shows that the MAP filter outperforms the modified Wiener filter.> Ling Guan, Rabab K. Ward |
ICASSP | 1 |