EDBT 2026 Demo / reviewers in the wild / expert
Liang-Tien Chia
dblp:05/2952
· DBLP profile ↗
105ranked-venue papers
0as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 76Artificial intelligence and machine learning · 27 · 1 since 2021Databases, data management, data science and information retrieval · 10Computer networks · 3Software engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
18 papers |
Multimedia analysis and retrieval · 46% Image and video processing · 34% Image and video coding · 8% | |
| Artificial intelligence
13 papers |
Representation and self-supervised learning · 46% Image recognition and object detection · 26% 3D vision · 12% | |
| Databases, data mining, and information retrieval
3 papers |
Information retrieval · 89% Knowledge graphs · 11% | |
| Computer networks
1 paper |
Routing and switching · 67% Internet of things and sensor networks · 33% |
Topics — the 30 heaviest of 60, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Image recognition and object detection
image classification |
0.8 | 7 | 2014 | Concurrent Single-Label Image Classification and Annotation via Efficient Multi-Layer Group Sparse Coding · IEEE Trans. Multim. 2014 Multi-layer group sparse coding - For concurrent image classification and annotation · CVPR 2011 Image-to-Class Distance Metric Learning for Image Classification · ECCV (1) 2010 |
Multimedia analysis and retrieval
image annotation |
0.4 | 3 | 2014 | Concurrent Single-Label Image Classification and Annotation via Efficient Multi-Layer Group Sparse Coding · IEEE Trans. Multim. 2014 Multi-layer group sparse coding - For concurrent image classification and annotation · CVPR 2011 Automatic image tagging via category label and web data · ACM Multimedia 2010 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding
kernel sparse representation |
0.3 | 2 | 2013 | Sparse Representation With Kernels · IEEE Trans. Image Process. 2013 Kernel Sparse Representation for Image Classification and Face Recognition · ECCV (4) 2010 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
sparse coding |
0.3 | 2 | 2013 | Sparse Representation With Kernels · IEEE Trans. Image Process. 2013 Local features are not lonely - Laplacian sparse coding for image classification · CVPR 2010 |
Image and video processing
saliency detection |
0.3 | 2 | 2013 | Regularized Feature Reconstruction for Spatio-Temporal Saliency Detection · IEEE Trans. Image Process. 2013 Improved saliency detection based on superpixel clustering and saliency propagation · ACM Multimedia 2010 |
Machine learning › Representation and self-supervised learning › representation learning
metric learning |
0.3 | 2 | 2012 | Learning Class-to-Image Distance via Large Margin and L1-Norm Regularization · ECCV (2) 2012 Image-to-Class Distance Metric Learning for Image Classification · ECCV (1) 2010 |
Multimedia analysis and retrieval
video analysis |
0.2 | 2 | 2013 | Background subtraction via coherent trajectory decomposition · ACM Multimedia 2013 Automatic extraction of motion trajectories in compressed sports videos · ACM Multimedia 2004 |
Machine learning › Representation and self-supervised learning › visual representation › image representation
bag of visual words |
0.2 | 2 | 2010 | Local features are not lonely - Laplacian sparse coding for image classification · CVPR 2010 Image classification using tensor representation · ACM Multimedia 2007 |
Information retrieval
image retrieval |
0.2 | 3 | 2009 | Coherent Phrase Model for Efficient Image Near-Duplicate Retrieval · IEEE Trans. Multim. 2009 Does ontology help in image retrieval?: a comparison between keyword, text ontology and multi-modality ontology approaches · ACM Multimedia 2006 Attention region selection with information from professional digital camera · ACM Multimedia 2005 |
Image and video processing
background subtraction |
0.2 | 1 | 2013 | Background subtraction via coherent trajectory decomposition · ACM Multimedia 2013 |
Multimedia analysis and retrieval › feature coding
bag-of-visual-words |
0.2 | 1 | 2013 | Laplacian Sparse Coding, Hypergraph Laplacian Sparse Coding, and Applications · IEEE Trans. Pattern Anal. Mach. Intell. 2013 |
Multimedia analysis and retrieval
image classification |
0.2 | 1 | 2013 | Laplacian Sparse Coding, Hypergraph Laplacian Sparse Coding, and Applications · IEEE Trans. Pattern Anal. Mach. Intell. 2013 |
Image and video processing › motion analysis
moving object detection |
0.2 | 1 | 2013 | Background subtraction via coherent trajectory decomposition · ACM Multimedia 2013 |
Image and video processing
sparse representation |
0.2 | 1 | 2013 | Laplacian Sparse Coding, Hypergraph Laplacian Sparse Coding, and Applications · IEEE Trans. Pattern Anal. Mach. Intell. 2013 |
Image and video coding
image quality assessment |
0.1 | 1 | 2012 | Fourier Transform-Based Scalable Image Quality Measure · IEEE Trans. Image Process. 2012 |
Image and video coding › image quality assessment › objective image quality assessment
reduced-reference quality assessment |
0.1 | 1 | 2012 | Fourier Transform-Based Scalable Image Quality Measure · IEEE Trans. Image Process. 2012 |
Audio and music processing
speech quality assessment |
0.1 | 1 | 2012 | Nonintrusive Quality Assessment of Noise Suppressed Speech With Mel-Filtered Energies and Support Vector Regression · IEEE Trans. Speech Audio Process. 2012 |
Image and video processing
subspace analysis |
0.1 | 2 | 2007 | Scale adaptive visual attention detection by subspace analysis · ACM Multimedia 2007 Robust subspace analysis for detecting visual attention regions in images · ACM Multimedia 2005 |
Image and video processing › saliency detection
visual attention detection |
0.1 | 2 | 2007 | Scale adaptive visual attention detection by subspace analysis · ACM Multimedia 2007 Robust subspace analysis for detecting visual attention regions in images · ACM Multimedia 2005 |
Computer vision › 3D vision
feature matching |
0.1 | 1 | 2011 | Spatially-coherent pyramid matching based on max-pooling · ACM Multimedia 2011 |
Machine learning › Learning paradigms
multi-label classification |
0.1 | 1 | 2011 | Multi-layer group sparse coding - For concurrent image classification and annotation · CVPR 2011 |
Machine learning › Representation and self-supervised learning › representation learning › visual representation learning
region representation |
0.1 | 1 | 2011 | Spatially-coherent pyramid matching based on max-pooling · ACM Multimedia 2011 |
Machine learning › Representation and self-supervised learning › representation learning › feature extraction
spatial pyramid matching |
0.1 | 1 | 2011 | Spatially-coherent pyramid matching based on max-pooling · ACM Multimedia 2011 |
Multimedia analysis and retrieval › image annotation
tag propagation |
0.1 | 1 | 2011 | Multi-layer group sparse coding - For concurrent image classification and annotation · CVPR 2011 |
Multimedia analysis and retrieval › multimedia feature representation
video representation |
0.1 | 1 | 2011 | Stratification-Based Keyframe Cliques for Effective and Efficient Video Representation · IEEE Trans. Multim. 2011 |
Multimedia analysis and retrieval
video retrieval |
0.1 | 1 | 2011 | Stratification-Based Keyframe Cliques for Effective and Efficient Video Representation · IEEE Trans. Multim. 2011 |
Machine learning › Representation and self-supervised learning › visual representation
image representation |
0.1 | 2 | 2009 | Image near-duplicate retrieval using local dependencies in spatial-scale space · ACM Multimedia 2008 Coherent Phrase Model for Efficient Image Near-Duplicate Retrieval · IEEE Trans. Multim. 2009 |
Computer vision › 3D vision
3d scene understanding |
0.1 | 1 | 2010 | Estimating camera pose from a single urban ground-view omnidirectional image and a 2D building outline map · CVPR 2010 |
Computer vision › 3D vision
camera pose estimation |
0.1 | 1 | 2010 | Estimating camera pose from a single urban ground-view omnidirectional image and a 2D building outline map · CVPR 2010 |
Computer vision › Face, body and person analysis
face recognition |
0.1 | 1 | 2010 | Kernel Sparse Representation for Image Classification and Face Recognition · ECCV (4) 2010 |
Methods — techniques the papers use, named apart from their topics
group sparse coding · 0.6reproducing kernel hilbert space · 0.4kernel methods · 0.2kNN · 0.2k-NN · 0.2spatial pyramid matching · 0.2sparse reconstruction · 0.2similarity preserving term · 0.2regularized feature reconstruction · 0.2markov random field · 0.2low-rank decomposition · 0.2laplacian smoothing · 0.2kernel trick · 0.2hypergraph · 0.2histogram intersection kernel · 0.2large margin regularization · 0.1l1-norm regularization · 0.1RKHS · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Deep residual pooling network for texture recognition
Shangbo Mao, Deepu Rajan, Liang-Tien Chia |
Pattern Recognit. | 3 |
| 2019 | Texture Recognition on Metal Surface using Order-Less Scale Invariant GLACabstractInspection of metal surface textures using computer vision and machine learning techniques plays an important role in Automated Visual Inspection (AVI) systems. Texture recognition on metal surface is challenging because the characteristics of each texture type are dependent on the properties of the metal surface when captured under different lighting conditions. Since these textures have no obvious repetitive patterns like general textures, this results in high intra-class diversities. Prior knowledge has shown that surface properties such as surface curvature and depth are discriminant to different texture types on metal surface. Since scale, shapes and location of textures within the same type are not fixed, scale property and spatial ordering information are less important for differentiating between texture types. There-fore, surface property, scale invariance and order-less property should be considered when exploring a suitable image feature for metal surface texture recognition. This paper proposes Order-less Scale Invariant Gradient Local Auto-Correlation (OS-GLAC) which meets all three requirements for robust texture recognition. The experiment results show that OS-GLAC is robust to separate different metal surface texture types. In addition, we observed that OS-GLAC is not only useful for texture recognition on metal surface but also for general texture recognition when combined with pre-trained deep learning features as these two features capture complimentary information. The experiment results show that such a combination of OS-GLAC achieves competitive results on three well-established general texture datasets i.e., KTH-TIP-2a, KTH-TIPS-2b and FMD. Shangbo Mao, Vidhya Natarajan, Liang-Tien Chia, Guang-Bin Huang |
ICTAI | 3 |
| 2019 | Salient Textural Anomaly Proposals and Classification for Metal Surface AnomaliesabstractWith a drastic growth in the development of Automated Visual Inspection (AVI) systems in the industries, the capabilities of such applications to aid human inspectors for anomaly localization and identification have also increased. However, some issues with anomaly detection and classification in AVI systems are that such anomalies are rare in occurrence and exhibit behaviours that are unique to the application. Hence, these anomaly datasets are small and imbalanced, and a robust framework is required for such datasets. This paper proposes a Salient Textural Anomaly Proposal (STAP) framework to generate and classify salient textural proposals of regions of anomalies on metal surfaces. These anomalies have both salient and texture characteristics that are dependent on the properties of the metal surface. Furthermore, when observed across different lighting conditions, the anomalies in this AVI anomaly dataset have a small inter-class variance and large intraclass variance. The proposed STAP framework uses a Fourier transformation based technique to generate proposals of salient and textural anomaly regions. Transfer learning of Convolutional Neural Network (CNN) learned from a large dataset is used to train a linear Support Vector Machine (SVM) to classify the generated proposals. The proposed STAP framework performs the best when compared to state-of-the-art object recognition techniques on an AVI anomaly industrial dataset that has salient and texture anomalies on metal surfaces. Vidhya Natarajan, Shangbo Mao, Liang-Tien Chia |
ICTAI | 3 |
| 2017 | Phase Fourier Reconstruction for Anomaly Detection on Metal Surface Using Salient Irregularity
Tzu-Yi Hung, Sriram Vaikundam, Vidhya Natarajan, Liang-Tien Chia |
MMM (1) | 4 |
| 2016 | Anomaly region detection and localization in metal surface inspectionabstractVisual inspection and identification of anomalies are important steps in the manufacturing process. A data set containing images of metallic components at different orientations and lighting conditions have been used for this experiment. Our aim is to find the anomaly in an image by subtracting it from an equivalent reference image with no anomalies. This is challenging due to the availability of multiple and slightly changing viewpoints for the reference image. The anomalies can only be detected in the residue obtained by using an identical reference image with no anomalies in the subtraction process. The proposed system is based on finding the best reference from a pool of similar images. A method for ranking the images based on similarity is also introduced. The proposed method scales well for high contrast images with dark colored anomalies. Sriram Vaikundam, Tzu-Yi Hung, Liang-Tien Chia |
ICIP | 3 |
| 2014 | Region-Based Saliency Detection and Its Application in Object RecognitionabstractThe objective of this paper is twofold. First, we introduce an effective region-based solution for saliency detection. Then, we apply the achieved saliency map to better encode the image features for solving object recognition task. To find the perceptually and semantically meaningful salient regions, we extract superpixels based on an adaptive mean shift algorithm as the basic elements for saliency detection. The saliency of each superpixel is measured by using its spatial compactness, which is calculated according to the results of Gaussian mixture model (GMM) clustering. To propagate saliency between similar clusters, we adopt a modified PageRank algorithm to refine the saliency map. Our method not only improves saliency detection through large salient region detection and noise tolerance in messy background, but also generates saliency maps with a well-defined object shape. Experimental results demonstrate the effectiveness of our method. Since the objects usually correspond to salient regions, and these regions usually play more important roles for object recognition than background, we apply our achieved saliency map for object recognition by incorporating a saliency map into sparse coding-based spatial pyramid matching (ScSPM) image representation. To learn a more discriminative codebook and better encode the features corresponding to the patches of the objects, we propose a weighted sparse coding for feature coding. Moreover, we also propose a saliency weighted max pooling to further emphasize the importance of those salient regions in feature pooling module. Experimental results on several datasets illustrate that our weighted ScSPM framework greatly outperforms ScSPM framework, and achieves excellent performance for object recognition. Zhixiang Ren, Shenghua Gao, Liang-Tien Chia, Ivor W. Tsang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2014 | Concurrent Single-Label Image Classification and Annotation via Efficient Multi-Layer Group Sparse CodingabstractWe present a multi-layer group sparse coding framework for concurrent single-label image classification and annotation. By leveraging the dependency between image class label and tags, we introduce a multi-layer group sparse structure of the reconstruction coefficients. Such structure fully encodes the mutual dependency between the class label, which describes image content as a whole, and tags, which describe the components of the image content. Therefore we propose a multi-layer group based tag propagation method, which combines the class label and subgroups of instances with similar tag distribution to annotate test images. To make our model more suitable for nonlinear separable features, we also extend our multi-layer group sparse coding in the Reproducing Kernel Hilbert Space (RKHS), which further improves performances of image classification and annotation. Moreover, we also integrate our multi-layer group sparse coding with kNN strategy, which greatly improves the computational efficiency. Experimental results on the LabelMe, UIUC-Sports and NUS-WIDE-Object databases show that our method outperforms the baseline methods, and achieves excellent performances in both image classification and annotation tasks. Shenghua Gao, Liang-Tien Chia, Ivor W. Tsang, Zhixiang Ren |
IEEE Trans. Multim. | 2 |
| 2013 | Background subtraction via coherent trajectory decompositionabstractBackground subtraction, the task to detect moving objects in a scene, is an important step in video analysis. In this paper, we propose an efficient background subtraction method based on coherent trajectory decomposition. We assume that the trajectories from background lie in a low-rank subspace, and foreground trajectories are sparse outliers in this background subspace. Meanwhile, the Markov Random Field (MRF) is used to encode the spatial coherency and trajectory consistency. With the low-rank decomposition and the MRF, our method can better handle videos with moving camera and obtain coherent foreground. Experimental results on a video dataset show our method achieves very competitive performance. Zhixiang Ren, Liang-Tien Chia, Deepu Rajan, Shenghua Gao |
ACM Multimedia | 2 |
| 2013 | Laplacian Sparse Coding, Hypergraph Laplacian Sparse Coding, and ApplicationsabstractSparse coding exhibits good performance in many computer vision applications. However, due to the overcomplete codebook and the independent coding process, the locality and the similarity among the instances to be encoded are lost. To preserve such locality and similarity information, we propose a Laplacian sparse coding (LSc) framework. By incorporating the similarity preserving term into the objective of sparse coding, our proposed Laplacian sparse coding can alleviate the instability of sparse codes. Furthermore, we propose a Hypergraph Laplacian sparse coding (HLSc), which extends our Laplacian sparse coding to the case where the similarity among the instances defined by a hypergraph. Specifically, this HLSc captures the similarity among the instances within the same hyperedge simultaneously, and also makes the sparse codes of them be similar to each other. Both Laplacian sparse coding and Hypergraph Laplacian sparse coding enhance the robustness of sparse coding. We apply the Laplacian sparse coding to feature quantization in Bag-of-Words image representation, and it outperforms sparse coding and achieves good performance in solving the image classification problem. The Hypergraph Laplacian sparse coding is also successfully used to solve the semi-auto image tagging problem. The good performance of these applications demonstrates the effectiveness of our proposed formulations in locality and similarity preservation. Shenghua Gao, Ivor W. Tsang, Liang-Tien Chia |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2013 | Sparse Representation With KernelsabstractRecent research has shown the initial success of sparse coding (Sc) in solving many computer vision tasks. Motivated by the fact that kernel trick can capture the nonlinear similarity of features, which helps in finding a sparse representation of nonlinear features, we propose kernel sparse representation (KSR). Essentially, KSR is a sparse coding technique in a high dimensional feature space mapped by an implicit mapping function. We apply KSR to feature coding in image classification, face recognition, and kernel matrix approximation. More specifically, by incorporating KSR into spatial pyramid matching (SPM), we develop KSRSPM, which achieves a good performance for image classification. Moreover, KSR-based feature coding can be shown as a generalization of efficient match kernel and an extension of Sc-based SPM. We further show that our proposed KSR using a histogram intersection kernel (HIK) can be considered a soft assignment extension of HIK-based feature quantization in the feature coding process. Besides feature coding, comparing with sparse coding, KSR can learn more discriminative sparse codes and achieve higher accuracy for face recognition. Moreover, KSR can also be applied to kernel matrix approximation in large scale learning tasks, and it demonstrates its robustness to kernel matrix approximation, especially when a small fraction of the data is used. Extensive experimental results demonstrate promising results of KSR in image classification, face recognition, and kernel matrix approximation. All these applications prove the effectiveness of KSR in computer vision and machine learning tasks. Shenghua Gao, Ivor W. Tsang, Liang-Tien Chia |
IEEE Trans. Image Process. | 3 |
| 2013 | Regularized Feature Reconstruction for Spatio-Temporal Saliency DetectionabstractMultimedia applications such as image or video retrieval, copy detection, and so forth can benefit from saliency detection, which is essentially a method to identify areas in images and videos that capture the attention of the human visual system. In this paper, we propose a new spatio-temporal saliency detection framework on the basis of regularized feature reconstruction. Specifically, for video saliency detection, both the temporal and spatial saliency detection are considered. For temporal saliency, we model the movement of the target patch as a reconstruction process using the patches in neighboring frames. A Laplacian smoothing term is introduced to model the coherent motion trajectories. With psychological findings that abrupt stimulus could cause a rapid and involuntary deployment of attention, our temporal model combines the reconstruction error, regularizer, and local trajectory contrast to measure the temporal saliency. For spatial saliency, a similar sparse reconstruction process is adopted to capture the regions with high center-surround contrast. Finally, the temporal saliency and spatial saliency are combined together to favor salient regions with high confidence for video saliency detection. We also apply the spatial saliency part of the spatio-temporal model to image saliency detection. Experimental results on a human fixation video dataset and an image saliency detection dataset show that our method achieves the best performance over several state-of-the-art approaches. Zhixiang Ren, Shenghua Gao, Liang-Tien Chia, Deepu Rajan |
IEEE Trans. Image Process. | 3 |
| 2013 | Learning image-to-class distance metric for image classificationabstractImage-To-Class (I2C) distance is a novel distance used for image classification and has successfully handled datasets with large intra-class variances. However, it uses Euclidean distance for measuring the distance between local features in different classes, which may not be the optimal distance metric in real image classification problems. In this article, we propose a distance metric learning method to improve the performance of I2C distance by learning per-class Mahalanobis metrics in a large margin framework. Our I2C distance is adaptive to different classes by combining with the learned metric for each class. These multiple per-class metrics are learned simultaneously by forming a convex optimization problem with the constraints that the I2C distance from each training image to its belonging class should be less than the distances to other classes by a large margin. A subgradient descent method is applied to efficiently solve this optimization problem. For efficiency and scalability to large-scale problems, we also show how to simplify the method to learn a diagonal matrix for each class. We show in experiments that our learned Mahalanobis I2C distance can significantly outperform the original Euclidean I2C distance as well as other distance metric learning methods in several prevalent image datasets, and our simplified diagonal matrices can preserve the performance but significantly speed up the metric learning procedure for large-scale datasets. We also show in experiment that our method is able to correct the class imbalance problem, which usually leads the NN-based methods toward classes containing more training images. Zhengxiang Wang, Yiqun Hu, Liang-Tien Chia |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2012 | Learning Class-to-Image Distance via Large Margin and L1-Norm Regularization
Zhengxiang Wang, Shenghua Gao, Liang-Tien Chia |
ECCV (2) | 3 |
| 2012 | Spatiotemporal Saliency Detection via Sparse RepresentationabstractMultimedia applications like retrieval, copy detection etc. can gain from saliency detection, which is essentially a method to identify areas in images and videos that capture the attention of the human visual system. In this paper, we propose a new spatiotemporal saliency framework for videos based on sparse representation. For temporal saliency, we model the movement of the target patch as a reconstruction process, and the overlapping patches in neighboring frames are used to reconstruct the target patch. The learned coefficients encode the positions of the matched patches, which are able to represent the motion trajectory of the target patch. We also introduce a smoothing term into our sparse coding framework to learn coherent motion trajectories. Based on the psychological findings that abrupt stimulus could cause a rapid and involuntary deployment of attention, our temporal model combines the reconstruction error, sparsity regularizer, and local trajectory contrast to measure the motion saliency. For spatial saliency, a similar sparse reconstruction process is adopted to capture the regions with high center-surround contrast. Finally, the temporal saliency and spatial saliency are combined by agreement to favor the salient regions with high confidence. Experimental results on a human fixation video dataset show our method achieved the best performance over five state-of-the-art approaches. Zhixiang Ren, Shenghua Gao, Deepu Rajan, Liang-Tien Chia |
ICME | 4 |
| 2012 | Video saliency detection with robust temporal alignment and local-global spatial contrastabstractVideo saliency detection, the task to detect attractive content in a video, has broad applications in multimedia understanding and retrieval. In this paper, we propose a new framework for spatiotemporal saliency detection. To better estimate the salient motion in temporal domain, we take advantage of robust alignment by sparse and low-rank decomposition to jointly estimate the salient foreground motion and the camera motion. Consecutive frames are transformed and aligned, and then decomposed to a low-rank matrix representing the background and a sparse matrix indicating the objects with salient motion. In the spatial domain, we address several problems of local center-surround contrast based model, and demonstrate how to utilize global information and prior knowledge to improve spatial saliency detection. Individual component evaluation demonstrates the effectiveness of our temporal and spatial methods. Final experimental results show that the combination of our spatial and temporal saliency maps achieve the best overall performance compared to several state-of-the-art methods. Zhixiang Ren, Liang-Tien Chia, Deepu Rajan |
ICMR | 2 |
| 2012 | Topic Based Query Suggestions for Video Search
Kong-Wah Wan, Ah-Hwee Tan, Joo-Hwee Lim, Liang-Tien Chia |
MMM | 4 |
| 2012 | A non-parametric visual-sense model of images - extending the cluster hypothesis beyond text
Kong-Wah Wan, Ah-Hwee Tan, Joo-Hwee Lim, Liang-Tien Chia |
Multim. Tools Appl. | 4 |
| 2012 | Nonintrusive Quality Assessment of Noise Suppressed Speech With Mel-Filtered Energies and Support Vector RegressionabstractObjective speech quality assessment is a challenging task which aims to emulate human judgment in the complex and time consuming task of subjective assessment. It is difficult to perform in line with the human perception due the complex and nonlinear nature of the human auditory system. The challenge lies in representing speech signals using appropriate features and subsequently mapping these features into a quality score. This paper proposes a nonintrusive metric for the quality assessment of noise-suppressed speech. The originality of the proposed approach lies primarily in the use of Mel filter bank energies (FBEs) as features and the use of support vector regression (SVR) for feature mapping. We utilize the sensitivity of FBEs to noise in order to obtain an effective representation of speech towards quality assessment. In addition, the use of SVR exploits the advantages of kernels which allow the regression algorithm to learn complex data patterns via nonlinear transformation for an effective and generalized mapping of features into the quality score. Extensive experiments conducted using two third party databases with different noise-suppressed speech signals show the effectiveness of the proposed approach. Manish Narwaria, Weisi Lin, Ian McLoughlin 0001, Sabu Emmanuel, Liang-Tien Chia |
IEEE Trans. Speech Audio Process. | 5 |
| 2012 | Fourier Transform-Based Scalable Image Quality MeasureabstractWe present a new image quality assessment (IQA) algorithm based on the phase and magnitude of the 2D (twodimensional) Discrete Fourier Transform (DFT). The basic idea is to compare the phase and magnitude of the reference and distorted images to compute the quality score. However, it is well known that the Human Visual Systems (HVSs) sensitivity to different frequency components is not the same. We accommodate this fact via a simple yet effective strategy of nonuniform binning of the frequency components. This process also leads to reduced space representation of the image thereby enabling the reduced-reference (RR) prospects of the proposed scheme. We employ linear regression to integrate the effects of the changes in phase and magnitude. In this way, the required weights are determined via proper training and hence more convincing and effective. Lastly, using the fact that phase usually conveys more information than magnitude, we use only the phase for RR quality assessment. This provides the crucial advantage of further reduction in the required amount of reference image information. The proposed method is therefore further scalable for RR scenarios. We report extensive experimental results using a total of 9 publicly available databases: 7 image (with a total of 3832 distorted images with diverse distortions) and 2 video databases (totally 228 distorted videos). These show that the proposed method is overall better than several of the existing fullreference (FR) algorithms and two RR algorithms. Additionally, there is a graceful degradation in prediction performance as the amount of reference image information is reduced thereby confirming its scalability prospects. To enable comparisons and future study, a Matlab implementation of the proposed algorithm is available at http://www.ntu.edu.sg/home/wslin/reduced_phase.rar. Manish Narwaria, Weisi Lin, Ian McLoughlin 0001, Sabu Emmanuel, Liang-Tien Chia |
IEEE Trans. Image Process. | 5 |
| 2011 | Multi-layer group sparse coding - For concurrent image classification and annotationabstractWe present a multi-layer group sparse coding framework for concurrent image classification and annotation. By leveraging the dependency between image class label and tags, we introduce a multi-layer group sparse structure of the reconstruction coefficients. Such structure fully encodes the mutual dependency between the class label, which describes the image content as a whole, and tags, which describe the components of the image content. Then we propose a multi-layer group based tag propagation method, which combines the class label and subgroups of instances with similar tag distribution to annotate test images. Moreover, we extend our multi-layer group sparse coding in the Reproducing Kernel Hilbert Space (RKHS) which captures the nonlinearity of features, and further improves performances of image classification and annotation. Experimental results on the LabelMe, UIUC-Sport and NUS-WIDE-Object databases show that our method outperforms the baseline methods, and achieves excellent performances in both image classification and annotation tasks. Shenghua Gao, Liang-Tien Chia, Ivor W. Tsang |
CVPR | 2 |
| 2011 | Spatially-coherent pyramid matching based on max-poolingabstractThis paper presents a method of max-pooling spatially-coherent pyramid matching (MpScPM). Higher-layer representations are generated from lower-layer subregions, by a biologically-inspired max pooling strategy. Second, instead of reshaping the pyramid representation into a vector (used in generic SPM), the layer and location information of each subregion are kept and weak geometrical correspondences between matched subregions are explored to enhance our pyramid matching method. To enhance the possibility of finding the best matches at different scales and locations, cross-layer region similarities are computed, while the correspondences (either spatial neighbors or adjacent layers) are also incorporated. We evaluate our proposed MpScPM method on several existing benchmark datasets and it achieves excellent performances. Xiangang Cheng, Liang-Tien Chia |
ACM Multimedia | 2 |
| 2011 | Exploiting local dependencies with spatial-scale space (S-Cube) for near-duplicate retrieval
Xiangang Cheng, Yiqun Hu, Liang-Tien Chia |
Comput. Vis. Image Underst. | 3 |
| 2011 | Improved learning of I2C distance and accelerating the neighborhood search for image classification
Zhengxiang Wang, Yiqun Hu, Liang-Tien Chia |
Pattern Recognit. | 3 |
| 2011 | Stratification-Based Keyframe Cliques for Effective and Efficient Video RepresentationabstractAs there is an exponential increase of web videos, it is time-consuming to get a query result from the tremendous data. An effective and efficient video management system is in urgent need. To increase the efficiency of video retrieval and storage, the most widely used methods are indexing schemes, such as locality sensitive hashing (LSH). However, it is more essential to represent the video itself compactly. In this paper, we propose a strategy to generate stratification-based keyframe cliques (SKCs) for video description, which are more compact and informative than frames or keyframes. The new representations are scalable for different retrieval tasks due to the ranking of SKCs. To further accelerate the retrieval speed, only top SKCs will be used; meanwhile, the searching results are still satisfactory. Experiments are conducted on TRECVID dataset as well as web video dataset. Results show that our proposed SKCs are more succinct and informative for video retrieval and management. Xiangang Cheng, Liang-Tien Chia |
IEEE Trans. Multim. | 2 |
| 2010 | Salient Region Detection by Jointly Modeling Distinctness and Redundancy of Image Content
Yiqun Hu, Zhixiang Ren, Deepu Rajan, Liang-Tien Chia |
ACCV (2) | 4 |
| 2010 | Estimating camera pose from a single urban ground-view omnidirectional image and a 2D building outline mapabstractA framework is presented for estimating the pose of a camera based on images extracted from a single omnidirectional image of an urban scene, given a 2D map with building outlines with no 3D geometric information nor appearance data. The framework attempts to identify vertical corner edges of buildings in the query image, which we term VCLH, as well as the neighboring plane normals, through vanishing point analysis. A bottom-up process further groups VCLH into elemental planes and subsequently into 3D structural fragments modulo a similarity transformation. A geometric hashing lookup allows us to rapidly establish multiple candidate correspondences between the structural fragments and the 2D map building contours. A voting-based camera pose estimation method is then employed to recover the correspondences admitting a camera pose solution with high consensus. In a dataset that is even challenging for humans, the system returned a top-30 ranking for correct matches out of 3600 camera pose hypotheses (0.83% selectivity) for 50.9% of queries. Tat-Jen Cham, Arridhana Ciptadi, Wei-Chian Tan, Minh-Tri Pham, Liang-Tien Chia |
CVPR | 5 |
| 2010 | Local features are not lonely - Laplacian sparse coding for image classificationabstractSparse coding which encodes the original signal in a sparse signal space, has shown its state-of-the-art performance in the visual codebook generation and feature quantization process of BoW based image representation. However, in the feature quantization process of sparse coding, some similar local features may be quantized into different visual words of the codebook due to the sensitiveness of quantization. In this paper, to alleviate the impact of this problem, we propose a Laplacian sparse coding method, which will exploit the dependence among the local features. Specifically, we propose to use histogram intersection based kNN method to construct a Laplacian matrix, which can well characterize the similarity of local features. In addition, we incorporate this Laplacian matrix into the objective function of sparse coding to preserve the consistence in sparse representation of similar local features. Comprehensive experimental results show that our method achieves or outperforms existing state-of-the-art results, and exhibits excellent performance on Scene 15 data set. Shenghua Gao, Ivor W. Tsang, Liang-Tien Chia, Peilin Zhao |
CVPR | 3 |
| 2010 | Kernel Sparse Representation for Image Classification and Face Recognition
Shenghua Gao, Ivor W. Tsang, Liang-Tien Chia |
ECCV (4) | 3 |
| 2010 | Image-to-Class Distance Metric Learning for Image Classification
Zhengxiang Wang, Yiqun Hu, Liang-Tien Chia |
ECCV (1) | 3 |
| 2010 | Learning to combine multi-resolution spatially-weighted co-occurrence matrices for image representationabstractBag-of-Words is widely used to describe images for image classification. However, this approach is limited because the spatial relation over visual words is not well exploited and also it is difficult to generate a single comprehensive vocabulary. In this paper, we propose novel effective schemes to handle these two issues. First, we propose a structure propagation technique to build more reasonable co-occurrence matrices of visual words to exploit the spatial information, which assigns a higher weight to the co-occurrence over two patches that lie in the same object part. Second, we build the multiple-histogram representation over hierarchical vocabularies to avoid the ambiguity of single vocabulary, and particularly present a learning approach to combine the multiple histograms to integrate both within-vocabulary and cross-vocabulary information. We evaluate our proposed method using the Princeton sports event dataset. Compared to the state-of-the-art results, our proposed approach has shown promising results. Xiangang Cheng, Jingdong Wang 0001, Liang-Tien Chia, Xian-Sheng Hua 0001 |
ICME | 3 |
| 2010 | Faceted topic retrieval of news video using joint topic modeling of visual features and speech transcriptsabstractBecause of the inherent ambiguity in user queries, an important task of modern retrieval systems is faceted topic retrieval (FTR), which relates to the goal of returning diverse or novel information elucidating the wide range of topics or facets of the query need. We introduce a generative model for hypothesizing facets in the (news) video domain by combining the complementary information in the visual keyframes and the speech transcripts. We evaluate the efficacy of our multimodal model on the standard TRECVID-2005 video corpus annotated with facets. We find that: (1) the joint modeling of the visual and text (speech transcripts) information can achieve significant F-score improvement over a text-alone system; (2) our model compares favorably with standard diverse ranking algorithms such as the MMR. Our FTR model has been implemented on a news search prototype that is undergoing commercial trial. Kong-Wah Wan, Ah-Hwee Tan, Joo-Hwee Lim, Liang-Tien Chia |
ICME | 4 |
| 2010 | Image Retargeting in Compressed DomainabstractA simple algorithm for image retargeting in the compressed domain is proposed. Most existing retargeting algorithms work directly in the spatial domain of the raw image. Here, we work on the DCT coefficients of a JPEG-compressed image to generate a gradient map that serves as an importance map to help identify those parts in the image that need to be retained during the retargeting process. Each 8×8 block of DCT coefficients is scaled based on the least importance value. Retargeting can be done both in the horizontal and vertical directions with the same framework. We also illustrate image enlargement using the same method. Experimental results show that the proposed algorithm produces less distortion in the retargeted image compared to some other algorithms reported recently. O. V. Ramana Murthy, Karthik Muthuswamy, Deepu Rajan, Liang-Tien Chia |
ICPR | 4 |
| 2010 | Automatic image tagging via category label and web dataabstractImage tagging is an important technique for the image content understanding and text based image processing. Given a selection of images, how to tag these images efficiently and effectively is an interesting problem. In this paper, a novel semi-auto image tagging technique is proposed: By assigning each image a category label first, our method can automatically recommend those promising tags to each image by utilizing existing vast web data. The main contributions of our paper can be highlighted as follows: (i) By assigning each image a category label, our method can automatically recommend other tags to the image, thus reducing the human annotation efforts. Meanwhile, our method guarantee tags' diversity due to abundant web data. (ii) We use sparse coding to automatically select those semantically related images for tag propagation. (iii) Local & global ranking agglomeration will make our method robust to noisy tags. We use Event dataset as the images to be tagged, and crawled Flickr images with their associated tags according to the category label in Event dataset as the auxiliary web data. Experimental results show that our method achieves promising performance for image tagging, which proves the effectiveness of our method. Shenghua Gao, Zhengxiang Wang, Liang-Tien Chia, Ivor W. Tsang |
ACM Multimedia | 3 |
| 2010 | Improved saliency detection based on superpixel clustering and saliency propagationabstractSaliency detection is useful for high level applications such as adaptive compression, image retargeting, object recognition, etc. In this paper, we introduce an effective region-based solution for saliency detection. We first use the adaptive mean shift algorithm to extract superpixels from the input image, then apply Gaussian Mixture Model (GMM) to cluster superpixels based on their color similarity, and finally calculate the saliency value for each cluster using compactness metric together with modified PageRank propagation. This solution is able to represent the image in a perceptually meaningful way and is robust to over-segmentation. It highlights salient regions with full resolution, well-defined boundary. Experimental results show that both the adaptive mean shift and the modified PageRank algorithm contribute substantially to the saliency detection result. In addition, the ROC analysis demonstrates that our approach significantly outperforms five existing popular methods. Zhixiang Ren, Yiqun Hu, Liang-Tien Chia, Deepu Rajan |
ACM Multimedia | 3 |
| 2010 | Discovering Class-Specific Informative Patches and Its Application in Landmark Charaterization
Shenghua Gao, Xiangang Cheng, Liang-Tien Chia |
MMM | 3 |
| 2010 | Non-intrusive Speech Quality Assessment with Support Vector Regression
Manish Narwaria, Weisi Lin, Ian McLoughlin 0001, Sabu Emmanuel, Liang-Tien Chia |
MMM | 5 |
| 2010 | Scene classification using multiple features in a two-stage probabilistic classification framework
Deepu Rajan, Liang-Tien Chia |
Neurocomputing | 3 |
| 2010 | Web image concept annotation with better understanding of tags and visual features
Shenghua Gao, Liang-Tien Chia, Xiangang Cheng |
J. Vis. Commun. Image Represent. | 2 |
| 2010 | Cross-media retrieval using query dependent search methods
Yi Yang 0001, Fei Wu 0001, Dong Xu 0001, Yueting Zhuang, Liang-Tien Chia |
Pattern Recognit. | 5 |
| 2009 | A Latent Model for Visual Disambiguation of Keyword-based Image SearchabstractThe problem of polysemy in keyword-based image search arises mainly from the inherent ambiguity in user queries. We propose a latent model based approach that resolves user search ambiguity by allowing sense specific diversity in search results. Given a query keyword and the images retrieved by issuing the query to an image search engine, we first learn a latent visual sense model of these polysemous images. Next, we use Wikipedia to disambiguate the word sense of the original query, and issue these Wiki-senses as new queries to retrieve sense specific images. A sense-specific image classifier is then learnt by combining information from the latent visual sense model, and used to cluster and re-rank the polysemous images from the original query keyword into its specific senses. Results on a ground truth of 17K image set returned by 10 keyword searches and their 62 word senses provides empirical indications that our method can improve upon existing keyword based search engines. Our method learns the visual word sense models in a totally unsupervised manner, effectively filters out irrelevant images, and is able to mine the long tail of image search. Kong-Wah Wan, Ah-Hwee Tan, Joo-Hwee Lim, Liang-Tien Chia, Sujoy Roy |
BMVC | 4 |
| 2009 | Hierarchicalword image representation for parts-based object recognitionabstractMany multimedia applications can benefit from recognizing image content. It requires a robust and discriminative representation of objects, especially in the situation of only a few training samples available. In this paper, we present a new approach to integrate the advantages of bag-of-words model and part-based model for image recognition. Each image is encoded as a hierarchical word image (HWI), which contains not only visual appearance but also spatial information. The object parts are then located and represented in HWI. Finally, the part-based star model (SM) is used to learn the object model and recognize the test images. It is shown that our proposed approach can detect more accurate part candidates and significantly improve the performance of original part-based model for object recognition. Xiangang Cheng, Yiqun Hu, Liang-Tien Chia |
ICIP | 3 |
| 2009 | Concept model-based unsupervised web image re-rankingabstractCurrent large scale image retrieval engines rely heavily on the surrounding text information, which inevitably includes some irrelevant images in the retrieval results due to the noisy environment. To improve the retrieval performance, we propose an unsupervised web image re-ranking method by incorporating images' visual information. Our method can automatically select a set of representative images from the original image pool as concept model, which is highly related to the query concept and critically important for the re-ranking result. With a similarity graph constructed by top results given by text based retrieval, we utilize Normalized Cut to select the part with the highest similarity density as concept model. We re-rank the rest images according to their similarities to the concept model. The advantages of our method are (i): Our method is unsupervised, and it doesn't need any pre-prepared query/training image or user's feedback, Thus it greatly facilitates users' retrieval. (ii): By finding a set of images rather than single image, we are able to give a more complete and more robust model for the query concept. (iii): Multi-ranking Integration Strategy is adopted to re-rank the rest images. Experiments show that our method can achieve satisfying results. Shenghua Gao, Xiangang Cheng, Huan Wang 0004, Liang-Tien Chia |
ICIP | 4 |
| 2009 | Learning instance-to-class distance for human action recognitionabstractIn this paper, we propose a large margin framework to learn the local instance-to-class distance function using local patch-based feature vectors, which satisfies the property that distance from instance to its own class should be less than the distance to other class. This instance-to-class distance is modeled as the weighted combination of the distance from every patch in test image to its nearest patch in training class, where the weight is learned through the above learning phase. We evaluate the proposed method on human action datasets and compare with related methods. It is shown that the proposed method achieves promising performance and improves the efficiency. Zhengxiang Wang, Yiqun Hu, Liang-Tien Chia |
ICIP | 3 |
| 2009 | A Bayesian approach integrating regional and global features for image semantic learningabstractIn content-based image retrieval, the ldquosemantic gaprdquo between visual image features and user semantics makes it hard to predict abstract image categories from low-level features. We present a hybrid system that integrates global features (G-features) and region features (R-features) for predicting image semantics. As an intermediary between image features and categories, we introduce the notion of mid-level concepts, which enables us to predict an image's category in three steps. First, a G-prediction system uses G-features to predict the probability of each category for an image. Simultaneously, a R-prediction system analyzes R-features to identify the probabilities of mid-level concepts in that image. Finally, our hybrid H-prediction system based on a Bayesian network reconciles the predictions from both R-prediction and G-prediction to produce the final classifications. Results of experimental validations show that this hybrid system outperforms both G-prediction and R-prediction significantly. Luong-Dong Nguyen, Ghim-Eng Yap, Ying Liu 0026, Ah-Hwee Tan, Liang-Tien Chia, Joo-Hwee Lim |
ICME | 5 |
| 2009 | LT codes decoding: Design and analysisabstractLT codes provide an efficient way to transfer information over erasure channels. Past research has illustrated that LT codes can perform well for a large number of input symbols. However, it is shown that LT codes have poor performance when the number of input symbols is small. We notice that the poor performance is due to the design of the LT decoding process. In this respect, we present a decoding algorithm called full rank decoding that extends the decodability of LT codes by usingWiedemann algorithm.We provide a detailed mathematical analysis on the rank of the random coefficient matrix to evaluate the probability of successful decoding for our proposed algorithm. Our studies show that our proposed method reduces the overhead significantly in the cases of small number of input symbols yet preserves the simplicity of the original LT decoding process. Chuan Heng Foh, Jianfei Cai 0001, Liang-Tien Chia |
ISIT | 4 |
| 2009 | Attention-from-motion: A factorization approach for detecting attention objects in motion
Yiqun Hu, Deepu Rajan, Liang-Tien Chia |
Comput. Vis. Image Underst. | 3 |
| 2009 | Coherent Phrase Model for Efficient Image Near-Duplicate RetrievalabstractThis paper presents an efficient and effective solution for retrieving image near-duplicate (IND) from image database. We introduce the coherent phrase model which incorporates the coherency of local regions to reduce the quantization error of the bag-of-words (BoW) model. In this model, local regions are characterized byvisual phraseof multiple descriptors instead of visual word of single descriptor. We propose two types of visual phrase to encode the coherency in feature and spatial domain, respectively. The proposed model reduces the number of false matches by using this coherency and generates sparse representations of images. Compared to other method, the local coherencies among multiple descriptors of every region improve the performance and preserve the efficiency for IND retrieval. The proposed method is evaluated on several benchmark datasets for IND retrieval. Compared to the state-of-the-art methods, our proposed model has been shown to significantly improve the accuracy of IND retrieval while maintaining the efficiency of the standard bag-of-words model. The proposed method can be integrated with other extensions of BoW. Yiqun Hu, Xiangang Cheng, Liang-Tien Chia, Xing Xie 0001, Deepu Rajan, Ah-Hwee Tan |
IEEE Trans. Multim. | 3 |
| 2008 | Motion Context: A New Representation for Human Action Recognition
Yiqun Hu, Syin Chan, Liang-Tien Chia |
ECCV (4) | 4 |
| 2008 | Image near-duplicate retrieval using local dependencies in spatial-scale spaceabstractThis paper presents an efficient and effective solution for retrieving Image Near-Duplicate (IND). Different from traditional methods, we analyze the local dependencies among region descriptors in a spatial-scale space. Such local dependencies in spatial-scale space(LDSS) encodes not only visual appearance but also the spatial and scale co-occurrence of them. The local dependencies are integrated over all spatial locations and multiple scales to form the image representation, which is invariant to spatial transformation and scale change. We evaluate our proposed LDSS method for IND retrieval using an existing benchmark as well as a new dataset extracted from the keyframes of TRECVID corpus. Compared to the state-of-the-art results, local dependencies in spatial-scale space(LDSS) approach has been shown to significantly improve the accuracy of IND retrieval. Xiangang Cheng, Yiqun Hu, Liang-Tien Chia |
ACM Multimedia | 3 |
| 2008 | NBgossip: An Energy-Efficient Gossip Algorithm for Wireless Sensor Networks
Liang-Tien Chia, Kok-Leong Tay, Wai-Hoe Chong |
J. Comput. Sci. Technol. | 2 |
| 2008 | Detection of visual attention regions in images using robust subspace analysis
Yiqun Hu, Deepu Rajan, Liang-Tien Chia |
J. Vis. Commun. Image Represent. | 3 |
| 2008 | Study on the distribution of DCT residues and its application to R-D analysis of video coding
Liang-Tien Chia |
J. Vis. Commun. Image Represent. | 2 |
| 2008 | Image retrieval with a multi-modality ontology
Huan Wang 0004, Song Liu 0001, Liang-Tien Chia |
Multim. Syst. | 3 |
| 2008 | Image retrieval ++ - web image retrieval with an enhanced multi-modality ontology
Huan Wang 0004, Liang-Tien Chia, Song Liu 0001 |
Multim. Tools Appl. | 2 |
| 2007 | Discriminative Signatures for Image ClassificationabstractBag-of-words representation has shown to be a powerful technique for image classification. In this paper, we propose a new approach to discover the discriminability of each visual word (image feature) in the codebook for each image category. A general linear model (GLM) is employed to construct new histograms of the images which are the basis for image classification. We also discuss the relations between our approach and boosting approaches and non-negative matrix factorization (NMF). Syin Chan, Liang-Tien Chia |
ICIP (2) | 3 |
| 2007 | MobileMaps@SG - Mappedia Version 1.1abstractTechnology has always been moving. Throughout the decades, improvements in various technological areas have led to a greater sense of convenience to ordinary people, whether it is cutting down time in accessing normal-to-day activities or getting privileged services. One of the technological areas that had been moving very rapidly is that of mobile computing. The common mobile device now has the mobility, provides entertainment via multimedia, connects to the Internet and is powered by intelligent and powerful chips. This paper will touch on an idea that is currently in the works, an integration of a recent technology that has netizens talking all over the world; Google Maps, that provide street and satellite images via the internet to the PC and Wikipedia 's user content support idea, the biggest free-content encyclopedia on the Internet. We will hit on how it is able to integrate such a technology with the idea of free form editing into one application in a small mobile device. The new features provided by this application will work toward supporting the development of multimedia application and computing. Daniel Dengyang Gan, Liang-Tien Chia |
ICME | 2 |
| 2007 | Semantic Retrieval with Enhanced Matchmaking and Multi-Modality OntologyabstractThis paper describes a semantic retrieval system that allows matchmaking with ranked output and the use of multi-modality ontology to retrieve animal images. Our multi-modality ontology, which integrates both image features and text information, is extended to provide a ranking mechanism. Ranking is calculated from correlation in each modality and is used to refine the semantic matchmaking result. To benchmark our results, we use the top 200 images of Google Image Search for each category to do the experimental comparison. Google Image Search claims to be the most comprehensive on the Web, with billions of images indexed and available for viewing. For different categories of animals in the canine family, we found averages of about 60% of top 200 images are correct images. Google returns even more false results outside this range. Therefore any bigger image set will become meaningless in our experiment. The medium size of the data set is fine since we are testing the retrieval performance on web images and concerned mainly with the precision of top retrievals. We believe the canine domain is challenging as demonstrated by the visual variance of objects and backgrounds. Twenty animal categories, containing animal images and corresponding web pages, are collected to form a systematic animal family. Results show that we can classify perceptually close animal species which share similar appearances as we can infer their hidden relationships from the canine family graph. By assigning a ranking to the semantic relationships we show unequivocal evidence that our improved model achieves good accuracy. Huan Wang 0004, Liang-Tien Chia, Song Liu 0001 |
ICME | 2 |
| 2007 | Click4BuildingID@NTU: Click for Building Identification with GPS-enabled Camera Cell PhoneabstractA working prototype of a building identification service which can be used on any camera cell phones equipped with GPS capability has been developed. Users can simply snap photos of architectures and send them, together with the corresponding GPS coordinates, via MMS to a remote server. The server will match the photos with the stored, GPS-tagged images using a combination of scale saliency algorithm for feature matching and earth movers distance measure for scene matching. The estimated location and other information are then sent back to the users via MMS. This prototype will have better accuracy than systems which rely solely on photo recognition given the exploitation of GPS information. Moreover, it is computationally lighter since the recognition engine only needs to compare stored images which lie within the GPS coordinates error range. It is relatively inexpensive as no special phones or subscriptions to telecommunication providers for provision of GPS equivalent location data (i.e.cell location) are needed. Chai Kiat Yeo, Liang-Tien Chia, Tat-Jen Cham, D. Rajon |
ICME | 2 |
| 2007 | A ZGPCA Algorithm for Subspace EstimationabstractWe propose a new algorithm called the ZGPCA algorithm for subspace estimation based on the GPCA (Generalized Principal Component Analysis) algorithm. It is formulated within an FIR filter framework so that the norm vectors of the subspaces correspond to filter coefficients. It is shown that such an approach leads to a more accurate and computationally efficient method compared to the GPCA algorithm. We extend the ZGPCA algorithm to make it recursive so that subspaces with possibly different dimensions can be obtained. We also propose a new distance measure that can be used for k-means clustering of sample points within a subspace. Experimental results on synthetic data and applications on face clustering and sports video clustering show good performance of the proposed algorithm. Haoran Yi, Deepu Rajan, Liang-Tien Chia |
ICME | 3 |
| 2007 | Codebook+: A New Module for Creating Discriminative CodebooksabstractIn this paper, we introduce a new module,Codebook+, into a classical framework which combines bag-of-words image representation with probabilistic latent semantic analysis (pLSA) for unsupervised object categorization. This new module makes the framework less sensitive to the image sampling methods as well as improves its performance. In this module, we create a new codebook based on the discriminability of each codeword in the original codebook for different categories. In our experiments, we compare the classification results of the framework with and withoutCodebook+ using five different image sampling methods. Syin Chan, Liang-Tien Chia |
ICME | 3 |
| 2007 | NBgossip - Neighborhood Gossip with Network Coding Based Message AggregationabstractGossip based algorithms for information dissemination have recently received significant attention for sensor and ad hoc network applications due to their simplicity and robustness. However, a common drawback of many gossip based protocols is the waste of energy in passing redundant information over the network. Thus gossip algorithms need to be re-engineered in order to be applicable for energy constrained networks. In this paper, we consider a scenario where each node in the network holds a piece of information (message) at the beginning, and the objective is to simultaneously disseminate all information (messages) among all nodes quickly and cheaply. To provide a practical solution to this problem for ad hoc and sensor networks, NBgossip algorithm is proposed, which is based on network coding and neighborhood gossip. In NBgossip, nodes do not simply forward messages they received, instead, the linear combinations of the messages are sent out. In addition, every node exchanges messages with its neighboring nodes only. Mathematical proof and simulation studies show that the proposed NBgossip terminates in the optimal O(n)-order rounds and outperforms the existing gossip based approaches in terms of energy incurred to spread all information. Liang-Tien Chia, Kok-Leong Tay |
MASS | 2 |
| 2007 | Scale adaptive visual attention detection by subspace analysisabstractWe describe a method to extract visual attention regions in images by robust subspace analysis from simple feature like intensity endowed with scale adaptivity in order to represent textured areas in an image. The scale adaptive descriptor is mapped onto clusters in linear spaces. A new subspace estimation algorithm based on the Generalized Principal Component Analysis (GPCA) is proposed to estimate multiple linear subspaces. The visual attention of each region is calculated using a new region attention measure that considers feature contrast and spatial geometric properties. Compared with existing visual attention detection methods, the proposed method directly measures global visual attention at the region level as opposed to pixel level. Yiqun Hu, Deepu Rajan, Liang-Tien Chia |
ACM Multimedia | 3 |
| 2007 | Image classification using tensor representationabstractWe propose a new approach to exploit the different discriminability of image features at different scales simultaneously. By modifying the Bag-of-words model, we represent an image as a matrix whose elements are the occurrences of a set of codewords within different scale ranges. In this way, we can represent an image collection using a 3rd-order tensor. Then a new classification method, tensor-pLSA, which is an extension of Probabilistic Latent Semantic Analysis (pLSA), is introduced to classify these images based on this tensor representation. Finally, we compare the tensor representation with the original matrix representation to show the effectiveness of our approach. Syin Chan, Liang-Tien Chia |
ACM Multimedia | 3 |
| 2007 | Multimedia Web Services for an Object Tracking and Highlighting Application
Liang-Tien Chia |
MMM (2) | 2 |
| 2007 | Secure multi-path in sensor networksabstractWireless sensor network has been identified as being useful in a variety of domains including the battlefield and perimeter defense. These mission critical applications raise the concern for security in sensor network. Typical security problems identified include passive information gathering, subversion of a node, legitimate addition of a node to an existing sensor network, and so forth [1]. Under traditional routing concept, information security is usually ensured through high-level security protocols, but limited computational power and memory space in sensor nodes make traditional cryptographical techniques cumbersome to be implemented in sensor networks. Lijuan Geng, Liang-Tien Chia, Ying-Chang Liang |
SenSys | 3 |
| 2007 | Mapping, indexing and querying of MPEG-7 descriptors in RDBMS with IXMDB
Yang Chu 0002, Liang-Tien Chia, Sourav S. Bhowmick |
Data Knowl. Eng. | 2 |
| 2007 | Efficient sampling of training set in large and noisy multimedia dataabstractAs the amount of multimedia data is increasing day-by-day thanks to less expensive storage devices and increasing numbers of information sources, machine learning algorithms are faced with large-sized and noisy datasets. Fortunately, the use of a good sampling set for training influences the final results significantly. But using a simple random sample (SRS) may not obtain satisfactory results because such a sample may not adequately represent the large and noisy dataset due to its blind approach in selecting samples. The difficulty is particularly apparent for huge datasets where, due to memory constraints, only very small sample sizes are used. This is typically the case for multimedia applications, where data size is usually very large. In this article we propose a new and efficient method to sample of large and noisy multimedia data. The proposed method is based on a simple distance measure that compares the histograms of the sample set and the whole set in order to estimate the representativeness of the sample. The proposed method deals with noise in an elegant manner which SRS and other methods are not able to deal with. We experiment on image and audio datasets. Comparison with SRS and other methods shows that the proposed method is vastly superior in terms of sample representativeness, particularly for small sample sizes although time-wise it is comparable to SRS, the least expensive method in terms of time. Surong Wang, Manoranjan Dash, Liang-Tien Chia, Min Xu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2006 | A Repository Adapter for Resource Management InformationabstractIntegrated network management frameworks for self-managing systems in a grid environment consisting of disparate applications, devices and subsystems require the use of a common definition of these managed resources. The common information model (CIM) provides such a standard for their description. However, the CIM specification lacks formalism which limits its use in knowledge aggregation and reasoning. This paper discusses the design of a repository adapter for resource information modeled in CIM. The adapter translates CIM constructs to an ontology-based language, the Data Centre Markup Language (DCML), thereby formalizing the model. Issues encountered during this process are identified and areas for future work are discussed. T. M. Ong, Liang-Tien Chia, Bu-Sung Lee |
CCGRID | 2 |
| 2006 | An Event-Driven Sports Video Adaptation for the MPEG-21 DIA FrameworkabstractWe present an event-driven video adaptation system in this paper. Events are detected by audio/video analysis and annotated by the description schemes (DSs) provided by MPEG-7 multimedia description schemes (MDSs). And then, adaptation take account of users' preference of events and network characteristic to adapt video by event selection and frame dropping as following three steps: 1) the event information is parsed from MPEG-7 annotation XML file together with bitstream to generate generic bitstream syntax description (gBSD), 2) users' preference, network characteristic and adaptation QoS (AQoS) are considered for making adaptation decision, 3) adaptation engine automatically parses adaptation decisions and gBSD to achieve adaptation. Different from most existing adaptation work, the system adapts video by interesting events according to users' preference. To achieve a generic adaptation solution, the system is developed following MPEG-7 and MPEG-21 standards. gBSD based adaptation avoids complex video computation. 30 students from various departments test the system with satisfaction. Although, the system is tested on basketball video adaptation so far, it is easy to extend to other video domains Min Xu 0001, Jiaming Li 0003, Yiqun Hu, Liang-Tien Chia, Bu-Sung Lee, Deepu Rajan, Jianfei Cai 0001 |
ICME | 4 |
| 2006 | Does ontology help in image retrieval?: a comparison between keyword, text ontology and multi-modality ontology approachesabstractOntologies are effective for representing domain concepts and relations in a form of semantic network. Many efforts have been made to import ontology into information matchmaking and retrieval. This trend is further accelerated by the convergence of various high-level concepts and low-level features supported by ontologies. In this paper we propose a comparison between traditional keyword based image retrieval and the promising ontology based image retrieval. To be complete, we construct the ontologies not only on text annotation, but also on a combination of text annotation and image feature. The experiments are conducted on a medium-sized data set including about 4000 images. The result proved the efficacy of utilizing both text and image features in a multi modality ontology to improve the image retrieval. Huan Wang 0004, Song Liu 0001, Liang-Tien Chia |
ACM Multimedia | 3 |
| 2006 | Event on demand with MPEG-21 video adaptation systemabstractIn this paper, we present an event-on-demand (EoD)video adaptation system. The proposed system supports users in deciding their events of interest and considers network conditions to adapt video source by event selection and frame dropping.Firstly, events are detected by audio/video analysis and annotated by the description schemes (DSs)provided by MPEG-7 Multimedia Description Schemes (MDSs). And then, to achieve a generic adaptation solution, the adaptation is developed following MPEG-21 Digital Item Adaptation (DIA)framework. We look at early release of the MPEG-21 Reference Software on XML generation and develop our own system for EoD video adaptation in three steps:1) the event information is parsed from MPEG-7 annotation XML file together with bitstream to generate generic Bitstream Syntax Description (gBSD). 2) Users' preference, Network Characteristic and Adaptation QoS (AQoS) are considered for making adaptation decision. 3) adaptation engine automatically parses adaptation decisions and gBSD to achieve adaptation.Unlike most existing adaptation work, the system adapts video of events with interest according to users' preference. Implementation following MPEG-7 and MPEG-21 standards provides a generic video adaptation solution. gBSD based adaptation avoids complex video computation. 30 students from various departments were invited to test the system and their responses has been positive. Min Xu 0001, Jiaming Li 0003, Liang-Tien Chia, Yiqun Hu, Bu-Sung Lee, Deepu Rajan, Jesse S. Jin |
ACM Multimedia | 3 |
| 2006 | Image model based on salient regions and its applicationsabstractDetection, representation, and training are the three main issues that need to be resolved in an object recognition or classification system. One possible method is using collection of regions to represent object categories where each region has a distinctive feature. In this paper we present a region-based image model which learn and classify objects by training the image model with variant of the objects within the same category. Each object category is represented by a constellation of representative parts. These regions are detected by salient region detector over suitable scales. The standard disjunction rule is applied to construct the image model. During the learning procedure the distance between any two regions is calculated and accumulated as a measure which is inversely proportional to the probability of a match. The regions with large distances are removed from the image model iteratively. Finally, a small set of regions is kept as the image model. This image model can be used to retrieve similar images or for object classification. Experimental results show the method is easy to calculate and efficient. Surong Wang, Liang-Tien Chia |
MMM | 2 |
| 2006 | An improved distortion model for rate control of DCT-based video codingabstractThis paper presents a rate control algorithm for the dominant discrete cosine transform (DCT)-based video coding. It is developed based on a more accurate rate-distortion (R-D) model, specifically, a new distortion-quantization (D-Q) model. Different from previous work that employs a uniform D-Q model or an empirical distortion model, our work proposes an accurate distortion model, which can quantitatively describe the relationship of distortion with respect to video source information and the selected quantization resolution. Based on understanding the distribution of source frequency coefficients and the quantization theory, our distortion model is proposed. This distortion model is combined with the classical R-D theory to generate a new rate model. Finally, the proposed model is implemented on an MPEG-4 encoder to perform rate control for a low-delay visual communication system. We also compare the proposed rate control with the VM18 rate control and it has been shown to be more efficient Liang-Tien Chia, Bu-Sung Lee |
MMM | 2 |
| 2006 | Affective content detection in sitcom using subtitle and audioabstractFrom a personalized media point of view, many users favor a flexible tool to quickly browse the affective content in a video. Such affective content may cause audiences' strong reactions or special emotional experiences, such as anger, sadness, fear, joy and love. This paper attempts to extract affective content for digital videos by analyzing the subtitle files of DVD/DivX videos and utilize audio event to assist affective content detection. Firstly, videos are segmented by dialogue script partition. Compared to traditional video shot, video segmented by scripts is not affected by camera changes and shooting angles and easy to include video segments with compact content. Secondly, emotion-related vocabularies in video script are detected to locate affective video content. Using script to directly access video content avoids complex video analysis. Thirdly, audio event detection is utilized to assist affective content detection. Compared with traditional video semantic analysis, affective content analysis puts much more emphasis on the audience's reactions and emotions. Initial experiments are carried on sitcom videos because its simple video structure provides useful domain knowledge. The experimental results demonstrate that subtitle file analysis and audio event detection provides effective and efficient clues to determine the emotional content of the videos. Min Xu 0001, Liang-Tien Chia, Haoran Yi, Deepu Rajan |
MMM | 2 |
| 2006 | Efficient data reduction in multimedia data
Surong Wang, Manoranjan Dash, Liang-Tien Chia, Min Xu 0001 |
Appl. Intell. | 3 |
| 2006 | FRACTURE mining: Mining frequently and concurrently mutating structures from historical XML documents
Ling Chen 0006, Sourav S. Bhowmick, Liang-Tien Chia |
Data Knowl. Eng. | 3 |
| 2006 | A motion-based scene tree for browsing and retrieval of compressed videos
Haoran Yi, Deepu Rajan, Liang-Tien Chia |
Inf. Syst. | 3 |
| 2006 | A motion-based scene tree for compressed video content management
Haoran Yi, Deepu Rajan, Liang-Tien Chia |
Image Vis. Comput. | 3 |
| 2006 | Dynamic Programming-Based Reverse Frame Selection for VBR Video Delivery Under Constrained ResourcesabstractIn this paper, we investigate optimal frame-selection algorithms based on dynamic programming for delivering stored variable bit rate (VBR) video under both bandwidth and buffer size constraints. Our objective is to find a feasible set of frames that can maximize the video's accumulated motion values without violating any constraint. It is well known that dynamic programming has high complexity. In this research, we propose to eliminate nonoptimal intermediate frame states, which can effectively reduce the complexity of dynamic programming. Moreover, we propose a reverse frame selection (RFS) algorithm, where the selection starts from the last frame and ends at the first frame. Compared with the conventional dynamic programming-based forward frame selection, the RFS is able to find all of the optimal results for different preloads in one round. We further extend the RFS scheme to solve the problem of frame selection for VBR channels. In particular, we first perform the RFS algorithm offline, and the complexity is modest and scalable with the aids of frame stuffing and nonoptimal state elimination. During online streaming, we only need to retrieve the optimal frame-selection path from the pregenerated offline results, and it can be applied to any VBR channels as long as the VBR channels can be modeled as piecewise CBR channels. Experimental results show good performance of our proposed algorithms Dayong Tao, Jianfei Cai 0001, Haoran Yi, Deepu Rajan, Liang-Tien Chia, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2005 | Mining Positive and Negative Association Rules from XML Query Patterns for Caching
Ling Chen 0006, Sourav S. Bhowmick, Liang-Tien Chia |
DASFAA | 3 |
| 2005 | SM3+: An XML Database Solution for the Management of MPEG-7 Descriptions
Yang Chu 0002, Liang-Tien Chia, Sourav S. Bhowmick |
DEXA | 2 |
| 2005 | Global Motion Compensated Key Frame Extraction from Compressed VideosabstractA key frame extraction approach, based on change detection of DC images extracted from compressed video, is proposed in this paper. We define a simple pixel change map that captures additional information in a frame with respect to its adjacent frames. Since global motion contributes to pixel changes, falsely indicating the presence of key frames, it is compensated by adaptively filtering the pixel change map using a modified version of the least mean square (LMS) algorithm. The prediction errors thus obtained are used to subsequently select the key frames. The key frames are selected so that the cumulative prediction error is partitioned into equal amounts in each segment. The entire procedure is computationally simple and flexible. Experimental results illustrate the good performance of the proposed algorithm. Haoran Yi, Deepu Rajan, Liang-Tien Chia |
ICASSP (2) | 3 |
| 2005 | Adaptive local context suppression of multiple cues for salient visual attention detectionabstractVisual attention is obtained through determination of contrasts of low level features or attention cues like intensity, color etc. We propose a new texture attention cue that is shown to be more effective for images where the salient object regions and background have similar visual characteristics. Current visual attention models do not consider local contextual information to highlight attention regions. We also propose a feature combination strategy by suppressing saliency based on context information that is effective in determining the true attention region. We compare our approach with other visual attention models using a novel average discrimination ratio measure. Yiqun Hu, Deepu Rajan, Liang-Tien Chia |
ICME | 3 |
| 2005 | Semantic Knowledge Building for Image Database by Analyzing Web Page ContentsabstractIn this paper, we present a method of semantic knowledge building for image database by extracting semantic meanings from Web page contents. The novelty of our method is that it is able to effectively extract media with a high degree of relevancy to a specific topic by incorporating word similarity and ontologies. The method is implemented in our Web image crawler and analysis system (WICAS). The system downloads Web pages and media automatically and further analyzes the semantic meanings of page contents to build up semantic knowledge for media entities. Subsequently, our system accepts high-level query terms and returns relevant media efficiently. Our experiment results show that with this new method of high-level content abstraction, media retrieval accuracy can be improved tremendously over traditional methods Yung-Kwang Lai, Song Liu 0001, Liang-Tien Chia, Syin Chan |
ICME | 3 |
| 2005 | Adaptive hierarchical multi-class SVM classifier for texture-based image classificationabstractIn this paper, we present a new classification scheme based on support vector machines (SVM) and a new texture feature, called texture correlogram, for high-level image classification. Originally, SVM classifier is designed for solving only binary classification problem. In order to deal with multiple classes, we present a new method to dynamically build up a hierarchical structure from the training dataset. The texture correlogram is designed to capture spatial distribution information. Experimental results demonstrate that the proposed classification scheme and texture feature are effective for high-level image classification task and the proposed classification scheme is more efficient than the other schemes while achieving almost the same classification accuracy. Another advantage of the proposed scheme is that the underlying hierarchical structure of the SVM classification tree manifests the interclass relationships among different classes. Song Liu 0001, Haoran Yi, Liang-Tien Chia, Deepu Rajan |
ICME | 3 |
| 2005 | EASIER Sampling for Audio Event IdentificationabstractAn audio event refers to some specific audio sound which plays important role for video content analysis. In our previous work [3], we have established audio event identification as an audio classification task. Due to the large size of audio database, representative samples are necessary for training the classifier. However, the commonly used random selection of training samples is often not adequate in selecting representative samples. In this paper we present EASIER sampling algorithm to select those data which more efficiently represent audio data characters for audio event identifier training. EASIER deterministically produces a subsample whose “distance” from the complete database is minimal. Experiments in the context of audio event identification show that EASIER outperforms simple random sampling significantly. Surong Wang, Min Xu 0001, Liang-Tien Chia, Manoranjan Dash |
ICME | 3 |
| 2005 | Affective content analysis in comedy and horror videos by audio emotional event detectionabstractWe study the problem of affective content analysis. In this paper, we think of affective contents as those video/audio segments, which may cause an audience's strong reactions or special emotional experiences, such as laughing or fear. Those emotional factors are related to the users' attention, evaluation, and memories of the content. The modeling of affective effects depends on the video genres. In this work, we focus on comedy and horror films to extract the affective content by detecting a set of so-called audio emotional events (AEE) such as laughing, horror sounds, etc. Those AEE can be modeled by various audio processing techniques, and they can directly reflect an audience's emotion. We use the AEE as a clue to locate corresponding video segments. Domain knowledge is more or less employed at this stage. Our experimental dataset consists of 40-minutes comedy video and 40-minutes horror film. An average recall and precision of above 90% is achieved. It is shown that, in addition to rich visual information, an appropriate usage of special audios is an effective way to assist affective content analysis. Min Xu 0001, Liang-Tien Chia, Jesse S. Jin |
ICME | 2 |
| 2005 | Robust subspace analysis for detecting visual attention regions in imagesabstractDetecting visually attentive regions of an image is a challenging but useful issue in many multimedia applications. In this paper, we describe a method to extract visual attentive regions in images using subspace estimation and analysis techniques. The image is represented in a 2D space using polar transformation of its features so that each region in the image lies in a 1D linear subspace. A new subspace estimation algorithm based on Generalized Principal Component Analysis (GPCA) is proposed. The robustness of subspace estimation is improved by using weighted least square approximation where weights are calculated from the distribution of K nearest neighbors to reduce the sensitivity of outliers. Then a new region attention measure is defined to calculate the visual attention of each region by considering both feature contrast and geometric properties of the regions. The method has been shown to be effective through experiments to be able to overcome the scale dependency of other methods. Compared with existing visual attention detection methods, it directly measures the global visual contrast at the region level as opposed to pixel level contrast and can correctly extract the attentive region. Yiqun Hu, Deepu Rajan, Liang-Tien Chia |
ACM Multimedia | 3 |
| 2005 | Attention region selection with information from professional digital cameraabstractThe attentive region extraction is a challenging issue for semantic interpretation of image and video content. The successful attentive region extraction greatly facilitates image classification, adaptation, compression and retrieval. Different from the traditional visual attention detection models, we propose a new attentive region extraction method based on out-of-focus blurring (OFB) technique used by professional photographers. Firstly, we combine metadata in Exchangeable Image File Format (EXIF) with visual features to quickly select professional photographs from image database. After that, an algorithm is implemented to automatically extract the attentive region from these photographs. This algorithm measures the saliency for individual pixels based on edge distribution of the images. The experimental results on OFB images have proved that our approach is able to overcome the contrast map selection problem of traditional visual attention methods and extract the attentive region using OFB information. The attentive region generated by our algorithm has similar shape and size with the subject of photographs which is a useful information for searching and retrieving the high-level semantic meaningful objects. Song Liu 0001, Liang-Tien Chia, Deepu Rajan |
ACM Multimedia | 2 |
| 2005 | Efficient Sampling: Application to Image Data
Surong Wang, Manoranjan Dash, Liang-Tien Chia |
PAKDD | 3 |
| 2005 | Enhancement layer rate control for high bitrate SNR scalable video coding
Liang-Tien Chia |
J. Vis. Commun. Image Represent. | 2 |
| 2005 | Automatic Generation of MPEG-7 Compliant XML Document for Motion Trajectory Descriptor in Sports Video
Haoran Yi, Deepu Rajan, Liang-Tien Chia |
Multim. Tools Appl. | 3 |
| 2005 | A new motion histogram to index motion content in video segments
Haoran Yi, Deepu Rajan, Liang-Tien Chia |
Pattern Recognit. Lett. | 3 |
| 2004 | Mining Maximal Frequently Changing Subtree Patterns from XML Documents
Ling Chen 0006, Sourav S. Bhowmick, Liang-Tien Chia |
DaWaK | 3 |
| 2004 | Coefficient thresholding and optimized selection of the lagrangian multiplier for non-reference frames in H.264 video codingabstractA new strategy to select the Lagrangian multiplier for nonreference frames is proposed, with the objective to maximize the average PSNR. The selection is based on the observation that nonreference frames should be optimized using a Lagrangian multiplier equivalent to the negative slope of the global R-D curve. This curve is not known a priori and a simple approximation is presented. Further, a new criterion for transform coefficient thresholding is proposed based on the current frame type and Lagrangian optimization. The new scheme is shown to improve the PSNR between 0.35 and 1.12 dB compared to the H.264 Test Model. The average improvement for all sequences is 0.63 dB, or equivalently a bit rate reduction of 11%. Pontus Carlsson, F. Pan, Liang-Tien Chia |
ICIP | 3 |
| 2004 | MPEG-21 digital item adaptation by applying perceived motion energy to H.264 video
Zhao Gang, Liang-Tien Chia, Yang Zongkai |
ICIP | 2 |
| 2004 | Optimum bit allocation for fgs video codingabstractThis paper proposes a new bit allocation scheme for fine-granular-scalability (FGS) video coding, through which we can achieve better video quality. Different from traditional rate-distortion (R-D) optimization schemes, we consider the characteristics of the bit-plane (BF) coding method. To be specific, we first find the approximate linear relationship between the bit rate of the FGS-layer and the percentage of nonzero binary-scaled coefficients (NZBC) in each BF; second, with mathematical justification, we derive an optimal strategy by analyzing the overall distortion with respect to NZBC. Finally, we perform our optimum bit allocation (OBA) on a FGS coder. Experimental results prove that our scheme can achieve smooth video quality with a higher average PSNR gain compared with uniform bit allocation (UBA). And for certain frames with lower PSNR, it has a gain of up to 3 dB. It is highly source-independent and more robust compared with previous bit allocation schemes. Liang-Tien Chia, Bu-Sung Lee |
ICIP | 2 |
| 2004 | DAML-QoS Ontology for Web ServicesabstractAs more and more Web services are deployed, Web service's discovery mechanisms become essential. Similar services can have quite different QoS levels. For service selection and management purpose, it is necessary to explicitly, precisely, and unambiguously specify various constraints and QoS metrics for Web services descriptions. This paper provides a novel DAML-QoS ontology as a complement for DAML-S ontology to provide a better QoS metrics model. Three layers are defined together with clear role descriptions for developments. Cardinality constraints are utilized to describe the QoS property constraints. Basic profile is presented for general Web service's description and the speed startup of ontology definition. Matchmaking algorithm for QoS property constraints is presented and different matching degrees are described. When incorporated with DAML-S, multiple service levels can be described through attaching multiple QoS profiles to one service profile. Well-defined Metrics can be further utilized by measurement organizations to guarantee the promised service level. Liang-Tien Chia, Bu-Sung Lee |
ICWS | 2 |
| 2004 | Region-of-interest based image resolution adaptation for MPEG-21 digital itemabstractThe upcoming MPEG-21 standard proposes a general framework for augmented use of multimedia services in different network environments, for various users with various terminal devices. In the context of image adaptation, terminals with different screen size limitation require the multimedia adaptation engine to adapt image resources intelligently. Saliency map based visual attention analysis provides some intelligence for finding the attention area within the image. In this paper, we improved the standard MPEG-21 metadata driven adaptation engine by using enhanced saliency map based visual attention model which provides a mean to intelligently adapt JPEG2000 image resolution for different terminal devices with varying screen size according to human visual attention. Yiqun Hu, Liang-Tien Chia, Deepu Rajan |
ACM Multimedia | 2 |
| 2004 | Audio keyword generation for sports video analysisabstractSemantic sports video analysis has attracted many research interests and audio cues have been shown to play an important role in semantics inference. To facilitate event detection using audio information, we have introduced the concept of audio keyword (e.g. excited/plain commentator speech, excited/plain audience sound, etc.) to describe the game-specific sound associated with an event. In our previous work, we have designed a hierarchical Support Vector Machine (SVM) classifier for audio keyword identification. However, there are two inherent weaknesses: 1) a frame-based SVM classifier does not incorporate any contextual information; 2) a robust recognizer relies on large amounts of training data in the case of different sports games videos. In this demo, we present a flexible Hidden Markov Model (HMM)-based audio keyword generation system. This is motivated by the successful story of applying HMM in speech recognition. Unlike the frame-based SVM classification followed by a major voting, our HMM-based system treats an audio keyword as a continuous time series data and employs hidden states transition to capture contexts. Moreover, our system introduces an adaptation mechanism to tune the initial HMM models (obtained from available training data) to improve performance by a small number of data from a new sports game video. Promising results has been demonstrated on the tennis, soccer and basketball videos with the total length of 2 hours. Min Xu 0001, Ling-Yu Duan, Liang-Tien Chia, Changsheng Xu |
ACM Multimedia | 3 |
| 2004 | Automatic extraction of motion trajectories in compressed sports videosabstractThis paper presents an algorithm for automatically extracting significant motion trajectories in sports videos. Our approach consists of four stages: global motion estimation, motion blob detection, trajectory evolution and trajectory refinement. Global motion is estimated from the motion vectors in the compressed video using an iterative algorithm with robust outlier rejection. A statistical hypothesis test is carried out within the Block Rejection Map(BRM), which is the by-product of the global motion estimation, for the detection of motion blobs. Trajectory evolution is the process in which the motion blobs are either appended to an existing trajectory or are considered to be the beginning of a new trajectory based on its distance to an adaptive trajectory description. Finally, the extracted motion trajectories are refined using a Kalman filter. Experimental results on both indoor and outdoor sports videos demonstrate the effectiveness and efficiency of the proposed method. Haoran Yi, Deepu Rajan, Liang-Tien Chia |
ACM Multimedia | 3 |
| 2004 | Mining Association Rules from Structural Deltas of Historical XML Documents
Ling Chen 0006, Sourav S. Bhowmick, Liang-Tien Chia |
PAKDD | 3 |
| 2003 | Efficient image retrieval using MPEG-7 descriptorsabstractIn this paper, a new method to calculate the similarity among images using dominant color descriptor is discussed. Using earth mover's distance (EMD), better retrieval results can be obtained compared with those obtained from the original MPEG-7 reference software (XM) [Text of ISO/IEC 15938-6/FDIS Information Technology-Multimedia content description interface-Part 6: Reference Software]. To further improve the retrieval accuracy, texture information from edge histogram descriptor is added. In order to reduce the retrieval time, two different methods which can prune the images far from the query image are discussed. One is the lower bound of EMD, while the other is the M-tree index based on EMD distance. Experiments show that the lower bound is easier to implement and more efficient than the M-tree. Surong Wang, Liang-Tien Chia, Deepu Rajan |
ICIP (3) | 2 |
| 2003 | A unified approach to detection of shot boundaries and subshots in compressed videoabstractThis paper describes a method to partition a video sequence into shots and subshots. By subshots, we mean one or a combination of the three camera motions of pan, tilt and zoom. The proposed technique detects both hard cuts and gradual transitions in MPEG compressed video using a single technique. We also present a motion estimation algorithm to compute the dominant motion represented by an affine model. The motion information is used to refine the location of dissolves as well as to subdivide the shot into subshots, thus providing a characterization of camera motion. We consider the dissimilarity between the I-, P- and B-frames with respect to the type of macroblocks used for encoding. Unlike previous algorithms reported, our method requires minimal decompression of the video sequence and uses very loose thresholds. The algorithm is evaluated on several types of video sequences to demonstrate its effectiveness. Haoran Yi, Deepu Rajan, Liang-Tien Chia |
ICIP (2) | 3 |
| 2003 | UX- An Architecture Providing QoS-Aware and Federated Support for UDDI
Liang-Tien Chia, Bilhanan Silverajan, Bu-Sung Lee |
ICWS | 2 |