VLDB 2026 Research / reviewers in the wild / expert
Yap-Peng Tan
dblp:93/4472
· DBLP profile ↗
233ranked-venue papers
9as first author
35since 2021 · last 2026
0000-0002-0645-9109ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 170 · 9 first-author · 19 since 2021Artificial intelligence and machine learning · 61 · 21 since 2021Systems, architecture and hardware · 23Security and privacy · 6 · 2 since 2021Databases, data management, data science and information retrieval · 4Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive memory refinement and perception enhancement for exo-to-ego video generation
Weipeng Hu, Jiun Tian Hoe, Ping Hu 0001, Xudong Jiang 0001, Yap-Peng Tan |
Neurocomputing | 7 |
| 2026 | PointVDP: Learning view-dependent projection by fireworks rays for 3D point cloud segmentation
Yueqi Duan, Haowen Sun 0004, Ziwei Wang 0010, Jiwen Lu, Yap-Peng Tan |
Pattern Recognit. | 6 |
| 2026 | Nonuniform low-light image enhancement via noise-aware decomposition and adaptive correction
Jiancai Huang, Zhaohui Jiang 0001, Xingjian Liu, Yap-Peng Tan, Weihua Gui 0001 |
Pattern Recognit. | 4 |
| 2026 | HP-Gaussian: Head Prior-Guided Gaussian Splatting for Personalized Talking Head Synthesis From Few-Second VideoabstractGaussian Splatting-based talking head synthesis has made significant progress in recent years, yet existing methods often struggle with generalization beyond specific training identity. In this paper, we propose Head Prior guided Gaussian Splatting for personalized talking head synthesis (HP-Gaussian) that can generalize to new identities with only few training data. Unlike traditional optimization-based Gaussian Splatting methods, our approach directly predicts Gaussian parameters from multi-modal inputs, including audio and visual cues. This feed-forward design enables multiple identities pre-training, allowing the model to learn shared head priors from large-scale datasets, while supporting flexible speaker-specific adaptation. To further enhance Gaussian feature learning, we introduce a Spatial Gaussian Transformer that captures correlations among neighboring Gaussians, improving parameter estimation accuracy. Additionally, recognizing the critical importance of personalized speaking styles in the synthesis of high-quality talking videos, a two-stage training strategy is implemented. A base model is initially trained across diverse identities to establish the foundational head prior knowledge. Subsequently, we introduce the short-video personalized adaptation phase for more realistic customized talking video generation. Extensive experiments demonstrate that our HP-Gaussian can synthesize high-fidelity and personalized talking videos with remarkably few training examples, setting a new benchmark for efficiency and quality in talking head synthesis. We highly recommend viewing our demonstration video at https://youtu.be/RpjWdvikKhU for intuitive visual comparisons and qualitative results. Shuai Shen, Wanhua Li 0001, Weipeng Hu, Jiwen Lu, Yap-Peng Tan |
IEEE Trans. Image Process. | 6 |
| 2026 | Learning Action Distribution Flow for Open-Set Temporal Action SegmentationabstractIn this paper, we tackle the open-set temporal action segmentation task, which aims to identify unknown frames while ensuring accurate segmentation of known actions in the temporal domain. Existing open-set methods struggle with identifying unknown frames due to their indistinguishability against ambiguous known frames during action transitions, resulting in significant performance degradation. To address this, we propose the action distribution flow, which models transitions between action sequences to capture the inherent feature discrepancies between unknown and known frames. Specifically, our method first models the distributions of known actions using the training data, and then interpolates these distributions along the optimal transport path for consecutive actions in the testing videos. By evaluating the likelihood of testing frames against the modeled action distribution flow, our approach effectively identifies unknown frames without requiring additional training or prior knowledge of the unknown data. Extensive experiments on open-set versions of the GTEA, 50Salads, and Breakfast datasets demonstrate the superiority of the proposed method across all evaluation metrics. Runzhong Zhang, Fengrui Tian, Yueqi Duan, Ziwei Wang 0010, Weipeng Hu, Peijun Bao, Suchen Wang, Yap-Peng Tan |
IEEE Trans. Image Process. | 10 |
| 2026 | Ambiguity-Aware Point Cloud Segmentation by Adaptive Margin Contrastive LearningabstractThis paper proposes an adaptive margin contrastive learning method for 3D semantic segmentation on point clouds. Most existing methods use equally penalized objectives, which ignore the per-point ambiguities and less discriminated features stemming from transition regions. However, as highly ambiguous points may be indistinguishable even for humans, their manually annotated labels are less reliable, and hard constraints over these points would lead to sub-optimal models. To address this, we first design AMContrast3D, a method comprising contrastive learning into an ambiguity estimation framework, tailored to adaptive objectives for individual points based on ambiguity levels. As a result, our method promotes model training, which ensures the correctness of low-ambiguity points while allowing mistakes for high-ambiguity points. As ambiguities are formulated based on position discrepancies across labels, optimization during inference is constrained by the assumption that all unlabeled points are uniformly unambiguous, lacking ambiguity awareness. Inspired by the insight of joint training, we further propose AMContrast3D++ integrating with two branches trained in parallel, where a novel ambiguity prediction module concurrently learns point ambiguities from generated embeddings. To this end, we design a masked refinement mechanism that leverages predicted ambiguities to enable the ambiguous embeddings to be more reliable, thereby boosting segmentation performance and enhancing robustness. Experimental results on 3D indoor scene datasets, S3DIS and ScanNet, demonstrate the effectiveness of the proposed method. Code is available athttps://github.com/YangChenApril/AMContrast3D. Yueqi Duan, Haowen Sun 0004, Jiwen Lu, Yap-Peng Tan |
IEEE Trans. Multim. | 5 |
| 2025 | Backdoor Attacks Against No-Reference Image Quality Assessment Models via a Scalable TriggerabstractNo-Reference Image Quality Assessment (NR-IQA), responsible for assessing the quality of a single input image without using any reference, plays a critical role in evaluating and optimizing computer vision systems, e.g., low-light enhancement. Recent research indicates that NR-IQA models are susceptible to adversarial attacks, which can significantly alter predicted scores with visually imperceptible perturbations. Despite revealing vulnerabilities, these attack methods have limitations, including high computational demands, untargeted manipulation, limited practical utility in white-box scenarios, and reduced effectiveness in black-box scenarios. To address these challenges, we shift our focus to another significant threat and present a novel poisoning-based backdoor attack against NR-IQA (BAIQA), allowing the attacker to manipulate the IQA model's output to any desired target value by simply adjusting a scaling coefficient alpha for the trigger. We propose to inject the trigger in the discrete cosine transform (DCT) domain to improve the local invariance of the trigger for countering trigger diminishment in NR-IQA models due to widely adopted data augmentations. Furthermore, the universal adversarial perturbations (UAP) in the DCT space are designed as the trigger, to increase IQA model susceptibility to manipulation and improve attack effectiveness. In addition to the heuristic method for poison-label BAIQA (P-BAIQA), we explore the design of clean-label BAIQA (C-BAIQA), focusing on alpha sampling and image data refinement, driven by theoretical insights we reveal. Extensive experiments on diverse datasets and various NR-IQA models demonstrate the effectiveness of our attacks. Yi Yu 0011, Song Xia, Xun Lin, Wenhan Yang, Shijian Lu, Yap-Peng Tan, Alex Chichung Kot |
AAAI | 6 |
| 2025 | MTL-UE: Learning to Learn Nothing for Multi-Task LearningabstractMost existing unlearnable strategies focus on preventing unauthorized users from training single-task learning (STL) models with personal data. Nevertheless, the paradigm has recently shifted towards multi-task data and multi-task learning (MTL), targeting generalist and foundation models that can handle multiple tasks simultaneously. Despite their growing importance, MTL data and models have been largely neglected while pursuing unlearnable strategies. This paper presents MTL-UE, the first unified framework for generating unlearnable examples for multi-task data and MTL models. Instead of optimizing perturbations for each sample, we design a generator-based structure that introduces label priors and class-wise feature embeddings which leads to much better attacking performance. In addition, MTL-UE incorporates intra-task and inter-task embedding regularization to increase inter-class separation and suppress intra-class variance which enhances the attack robustness greatly. Furthermore, MTL-UE is versatile with good supports for dense prediction tasks in MTL. It is also plug-and-play allowing integrating existing surrogate-dependent unlearnable methods with little adaptation. Extensive experiments show that MTL-UE achieves superior attacking performance consistently across 4 MTL datasets, 3 base UE methods, 5 model backbones, and 5 MTL task-weighting strategies. Code is available at https://github.com/yuyi-sd/MTL-UE. Yi Yu 0011, Song Xia, Siyuan Yang 0001, Chenqi Kong, Wenhan Yang, Shijian Lu, Yap-Peng Tan, Alex Chichung Kot |
ICML | 7 |
| 2025 | E3RG: Building Explicit Emotion-driven Empathetic Response Generation System with Multimodal Large Language Model
Ronghao Lin, Shuai Shen, Weipeng Hu, Qiaolin He, Aolin Xiong, Haifeng Hu 0001, Yap-Peng Tan |
ACM Multimedia | 8 |
| 2025 | CyIN: Cyclic Informative Latent Space for Bridging Complete and Incomplete Multimodal LearningabstractMultimodal machine learning, mimicking the human brain’s ability to integrate various modalities has seen rapid growth. Most previous multimodal models are trained on perfectly paired multimodal input to reach optimal performance. In real‑world deployments, however, the presence of modality is highly variable and unpredictable, causing the pre-trained models in suffering significant performance drops and fail to remain robust with dynamic missing modalities circumstances. In this paper, we present a novel Cyclic INformative Learning framework (CyIN) to bridge the gap between complete and incomplete multimodal learning. Specifically, we firstly build an informative latent space by adopting token- and label-level Information Bottleneck (IB) cyclically among various modalities. Capturing task-related features with variational approximation, the informative bottleneck latents are purified for more efficient cross-modal interaction and multimodal fusion. Moreover, to supplement the missing information caused by incomplete multimodal input, we propose cross-modal cyclic translation by reconstruct the missing modalities with the remained ones through forward and reverse propagation process. With the help of the extracted and reconstructed informative latents, CyIN succeeds in jointly optimizing complete and incomplete multimodal learning in one unified model. Extensive experiments on 4 multimodal datasets demonstrate the superior performance of our method in both complete and diverse incomplete scenarios. Ronghao Lin, Qiaolin He, Sijie Mai, Aolin Xiong, Yap-Peng Tan, Haifeng Hu 0001 |
NeurIPS | 7 |
| 2025 | Cascaded Dynamic Memory Refinement and Semantic Alignment for Exo-to-Ego Cross-View Video GenerationabstractCross-view video generation from exocentric (third-person) to egocentric (first-person) perspectives poses a challenging task, due to the significant viewpoint gap and limited overlap between these two views. Previous methods exhibit limitations in capturing long-range temporal context and overlook egocentric semantic priors, leading to degraded performance in cross-view synthesis. To address these challenges, we propose a cue-free video-based approach termed cascaded Dynamic memory Refinement and Semantic Alignment (DRSA), which integrates temporal knowledge over extended periods and learns egocentric semantic information to generate videos. The Dynamic Memory Refinement (DMR) exploits long horizon temporal dynamics to learn salient information that compensates for the limited overlap between views. Specifically, we devise a dynamic memory that serves as a knowledge repository, and utilize a sliding window to locate the corresponding long-term temporal information, which is subsequently processed with adaptive weighting and cross-attention transformer to refine feature representations. Furthermore, aware of the considerable viewpoint divergence that hinder semantic learning of target view, we propose Viewpoint-aware Semantic Alignment (VSA) with dual encoder-decoder learning and semantic alignment, which transfer egocentric semantic details from the egocentric synthesis pipeline to the exocentric synthesis pipeline. In particular, the VSA module narrows the semantic gap between views, further promoting long-range temporal modeling in DMR under alignment constraints. By extending this into a cascaded fashion, the Cascaded Alignment and Refinement (CAR) progressively aligns semantic features and performs feature refinement to facilitate viewpoint learning at different levels of granularity. To overcome the limitations of existing databases known for their limited static scenes and scarcity of interacting objects, we create a new dataset with dynamic exocentric scenes and rich interacting objects to further promote the task. Thorough experimental analysis reveals that our method surpasses current state-of-the-art techniques in terms of both quantitative metrics and qualitative evaluations. Weipeng Hu, Jiun Tian Hoe, Haifeng Hu 0001, Xudong Jiang 0001, Yap-Peng Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Robust and Transferable Backdoor Attacks Against Deep Image Compression With Selective Frequency PriorabstractRecent advancements in deep learning-based compression techniques have demonstrated remarkable performance surpassing traditional methods. Nevertheless, deep neural networks have been observed to be vulnerable to backdoor attacks, where an added pre-defined trigger pattern can induce the malicious behavior of the models. In this paper, we propose a novel approach to launch a backdoor attack with multiple triggers against learned image compression models. Drawing inspiration from the widely used discrete cosine transform (DCT) in existing compression codecs and standards, we propose a frequency-based trigger injection model that adds triggers in the DCT domain. In particular, we design several attack objectives that are adapted for a series of diverse scenarios, including: 1) attacking compression quality in terms of bit-rate and reconstruction quality; 2) attacking task-driven measures, such as face recognition and semantic segmentation in downstream applications. To facilitate more efficient training, we develop a dynamic loss function that dynamically balances the impact of different loss terms with fewer hyper-parameters, which also results in more effective optimization of the attack objectives with improved performance. Furthermore, we consider several advanced scenarios. We evaluate the resistance of the proposed backdoor attack to the defensive pre-processing methods and then propose a two-stage training schedule along with the design of robust frequency selection to further improve resistance. To strengthen both the cross-model and cross-domain transferability on attacking downstream CV tasks, we propose to shift the classification boundary in the attack loss during training. Extensive experiments also demonstrate that by employing our trained trigger injection models and making slight modifications to the encoder parameters of the compression model, our proposed attack can successfully inject multiple backdoors accompanied by their corresponding triggers into a single image compression model. Yi Yu 0011, Yufei Wang 0006, Wenhan Yang, Lanqing Guo, Shijian Lu, Ling-Yu Duan, Yap-Peng Tan, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | Boundary Voting Network for Ambiguity-Aware Timestamp-Supervised Action SegmentationabstractTimestamp-supervised action segmentation aims to segment and classify actions in untrimmed videos with a random frame annotated per action. Precisely localizing action boundaries from timestamp annotations is crucial for this setting, as it enables generating framewise pseudo-labels and applying the well-explored fully-supervised training. However, prevailing methods struggle with intrinsic uncertainty in boundary localization due to less discriminative features in action-transiting regions. This imprecise boundary estimation significantly reduces the stability and reliability of the generated pseudo-labels in ambiguous action-transiting regions, consequently resulting in performance deterioration of the trained segmentation models. In our paper, we introduce the boundary voting network that mitigates feature ambiguity by hierarchically propagating video-level global prior knowledge into local action-transiting regions. By generating key action representations as votes throughout the video and targeting action-transiting regions, all votes collaboratively contribute to action-transiting feature enhancement and boundary localization refinement. Extensive experiments demonstrate the effectiveness of our method on GTEA, 50Salads, and Breakfast datasets. Runzhong Zhang, Yueqi Duan, Weipeng Hu, Suchen Wang, Yap-Peng Tan |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | Prompt-Guided Transformer and MLLM Interactive Learning for Text-Based Pedestrian SearchabstractAiming to retrieve pedestrian images based on a textual description query, Text-Based Pedestrian Search (TBPS) gains increasingly attention due to its applications in security surveillance. As a fine-grained classification task, TBPS requires identifying images of individuals with different semantic contexts yet the same identity, as well as distinguishing images of individuals who share similar appearances but distinct identities. Consequently, TBPS is challenged by semantic variations in positive pairs and appearance similarity between negative pairs. To tackle these challenges, we propose the Prompt-guided Transformer and MLLM Interactive learning (PTMI) model to learn identity-discriminative representations across different modalities. PTMI consists of three components: the Prompt-guided Transformer (Promformer), MLLM Interactive Learning (MIL) and Dual-branch Cross-modal Learning (DCL). Firstly, the Promformer is designed to handle semantic variations in positive pairs by introducing learnable prompts, composing of three types: instance-shared, instance-specific and layer-specific. Optimized by Cross-modal Intra-class Consistency (CIC) loss, these prompts minimize intra-class variations and retrieve positive images with various semantics. Secondly, the MIL component is introduced to address appearance similarity between negative pairs by focusing on key image patches and description words filtering by the local discriminator. Powered by Multimodal Large Language Model (MLLM), the local discriminator adopts soft attention to highlight important image regions and descriptive words, which preserves semantic information while emphasize discriminative details. Lastly, the DCL integrates global and local branches to bridge modality discrepancies. The global branch employs SDM loss for heterogeneous distribution alignment, while the local branch applies Anchor-Based Contrastive (ABC) loss for instance-level contrastive learning. Unlike conventional contrastive loss, ABC loss leverages MLLM features as anchors to decouple modality and semantic differences, enhancing alignment efficiency. Extensive experiments on three TBPS datasets have validated the effectiveness of PTMI. Zefeng Lu, Ronghao Lin, Yap-Peng Tan, Haifeng Hu 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Toward Model Resistant to Transferable Adversarial Examples via Trigger ActivationabstractAdversarial examples, characterized by imperceptible perturbations, pose significant threats to deep neural networks by misleading their predictions. A critical aspect of these examples is their transferability, allowing them to deceive unseen models in closed-box scenarios. Despite the widespread exploration of defense methods, including those on transferability, they show limitations: inefficient deployment, ineffective defense, and degraded performance on clean images. In this work, we introduce a novel training paradigm aimed at enhancing robustness against transferable adversarial examples (TAEs) in a more efficient and effective way. We propose a model that exhibits random guessing behavior when presented with clean data$\boldsymbol {x}$as input, and generates accurate predictions when with triggered data$\boldsymbol {x}+\boldsymbol {\tau }$. Importantly, the trigger$\boldsymbol {\tau }$remains constant for all data instances. We refer to these models as models with trigger activation. We are surprised to find that these models exhibit certain robustness against TAEs. Through the consideration of first-order gradients, we provide a theoretical analysis of this robustness. Moreover, through the joint optimization of the learnable trigger and the model, we achieve improved robustness to transferable attacks. Extensive experiments conducted across diverse datasets, evaluating a variety of attacking methods, underscore the effectiveness and superiority of our approach. Yi Yu 0011, Song Xia, Xun Lin, Chenqi Kong, Wenhan Yang, Shijian Lu, Yap-Peng Tan, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2025 | Residual Quotient Learning for Zero-Reference Low-Light Image EnhancementabstractRecently, neural networks have become the dominant approach to low-light image enhancement (LLIE), with at least one-third of them adopting a Retinex-related architecture. However, through in-depth analysis, we contend that this most widely accepted LLIE structure is suboptimal, particularly when addressing the non-uniform illumination commonly observed in natural images. In this paper, we present a novel variant learning framework, termed residual quotient learning, to substantially alleviate this issue. Instead of following the existing Retinex-related decomposition-enhancement-reconstruction process, our basic idea is to explicitly reformulate the light enhancement task as adaptively predicting the latent quotient with reference to the original low-light input using a residual learning fashion. By leveraging the proposed residual quotient learning, we develop a lightweight yet effective network called ResQ-Net. This network features enhanced non-uniform illumination modeling capabilities, making it more suitable for real-world LLIE tasks. Moreover, due to its well-designed structure and reference-free loss function, ResQ-Net is flexible in training as it allows for zero-reference optimization, which further enhances the generalization and adaptability of our entire framework. Extensive experiments on various benchmark datasets demonstrate the merits and effectiveness of the proposed residual quotient learning, and our trained ResQ-Net outperforms state-of-the-art methods both qualitatively and quantitatively. Furthermore, a practical application in dark face detection is explored, and the preliminary results confirm the potential and feasibility of our method in real-world scenarios. Linfeng Fei, Huanjie Tao, Yaocong Hu, Wei Zhou 0042, Jiun Tian Hoe, Weipeng Hu, Yap-Peng Tan |
IEEE Trans. Image Process. | 8 |
| 2025 | Snippet-Inter Difference Attention Network for Weakly-Supervised Temporal Action LocalizationabstractThe purpose of weakly-supervised temporal action localization (WTAL) task is to simultaneously classify and localize action instances in untrimmed videos with only video-level labels. Previous works fail to extract multi-scale temporal features to identify action instances with different durations, and they do not fully use the temporal cues of action video to learn discriminative features. In addition, the classifiers trained by current methods usually focus on easy-to-distinguish snippets while ignoring other semantically ambiguous features, which leads to incomplete and over-complete localization. To address these issues, we introduce a new Snippet-inter Difference Attention Network (SDANet) for WTAL, which can be trained end-to-end. Specifically, our model presents three modules, with primary contributions lying in the snippet-inter difference attention (SDA) module and potential feature mining (PFM) module. Firstly, we construct a simple multi-scale temporal feature fusion (MTFF) module to generate multi-scale temporal feature representation, so as to help the model better detect short action instances. Secondly, we consider the temporal cues of video features and design SDA module based on the Transformer to capture global discriminative features for each modality based on multi-scale features. It calculates the differences between temporal neighbor snippets in each modality to explore salient-difference features, and then utilizes them to guide correlation modeling. Thirdly, after learning discriminative features, we devise PFM module to excavate potential action and background snippets from ambiguous features. By contrastive learning, potential actions are forced closer to discriminative actions and away from the background, thereby learning more accurate action boundaries. Finally, two losses (i.e., similarity loss and reconstruction loss) are further developed to constrain the consistency between two modalities and help the model retain original feature information for better localization results. Extensive experiments show that our model achieves better performance against current WTAL methods on three datasets, i.e., THUMOS14, ActivityNet1.2 and ActivityNet1.3. Wei Zhou 0042, Kang Lin, Weipeng Hu, Haifeng Hu 0001, Yap-Peng Tan |
IEEE Trans. Multim. | 7 |
| 2025 | Class-Specific Prompt Learning for Vision-Language ModelsabstractThe use of learning prompts to adapt pretrained vision-language models (VLMs) for downstream tasks has gained significant attention due to its potential to reduce training costs compared to model fine-tuning through few-shot learning. Most existing methods rely on a universal prompt for all classes, as it generally delivers consistent performance across various datasets. However, a universal prompt cannot capture class-specific discriminative information. To overcome this limitation, we propose class-specific prompt learning (CPL). CPL represents the context of a prompt using two components: a base vector shared among all classes and a class-specific vector designed for individual classes. This method combines the generalization ability of the base context with the adaptability of the class-specific context. Furthermore, we introduce contrastive CPL, which enhances the ability of the prompt to capture discriminative features unique to each class. Also, we adopt the self-consistency loss to regularize the base context, enhancing its generalization ability. As a result, CPL effectively learns tailored prompts for each class. Extensive experiments demonstrate that CPL achieves superior performance over existing methods in both base-class classification and new class generalization. Runhao Li, Yongming Chen, Zhenyu Weng, Zhiping Lin 0001, Yap-Peng Tan |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Cross-Domain Few-Shot Segmentation via Iterative Support-Query Correspondence MiningabstractCross-Domain Few-Shot Segmentation (CD-FSS) poses the challenge of segmenting novel categories from a distinct domain using only limited exemplars. In this paper, we undertake a comprehensive study of CD-FSS and uncover two crucial insights: (i) the necessity of a fine-tuning stage to effectively transfer the learned meta-knowledge across domains, and (ii) the overfitting risk during the naive fine-tuning due to the scarcity of novel category examples. With these insights, we propose a novel cross-domain fine-tuning strategy that addresses the challenging CD-FSS tasks. We first design Bi-directional Few-shot Prediction (BFP), which establishes support-query correspondence in bi-directional manner, crafting augmented supervision to reduce the overfitting risk. Then we further extend BFP into Iterative Few-shot Adaptor (IFA), which is a recursive framework to capture the support-query correspondence iteratively, targeting maximal exploitation of supervisory signals from the sparse novel category samples. Extensive empirical evaluations show that our method significantly outperforms the state-of-the-arts (+7.8%), which verifies that IFA tackles the cross-domain challenges and mitigates the overfitting simultaneously. Jiahao Nie 0002, Yun Xing 0001, Gongjie Zhang, Pei Yan, Aoran Xiao, Yap-Peng Tan, Alex Chichung Kot, Shijian Lu |
CVPR | 6 |
| 2024 | InteractDiffusion: Interaction Control in Text-to-Image Diffusion ModelsabstractLarge-scale text-to-image (T2I) diffusion models have showcased incredible capabilities in generating coherent images based on textual descriptions, enabling vast applications in content generation. While recent advancements have introduced control over factors such as object localization, posture, and image contours, a crucial gap remains in our ability to control the interactions between objects in the generated content. Well-controlling interactions in generated images could yield meaningful applications, such as creating realistic scenes with interacting characters. In this work, we study the problems of conditioning T2I diffusion models with Human-Object Interaction (HOI) information, consisting of a triplet label (person, action, object) and corresponding bounding boxes. We propose a pluggable interaction control model, called InteractDiffusion that extends existing pre-trained T2I diffusion models to enable them being better conditioned on interactions. Specifically, we tokenize the HOI information and learn their relationships via interaction embeddings. A conditioning self-attention layer is trained to map HOI tokens to visual tokens, thereby conditioning the visual tokens better in existing T2I diffusion models. Our model attains the ability to control the interaction and location on existing T2I diffusion models, which outperforms existing baselines by a large margin in HOI detection score, as well as fidelity in FID and KID. Project page: https://jiuntian.github.io/interactdiffusion. Jiun Tian Hoe, Xudong Jiang 0001, Chee Seng Chan, Yap-Peng Tan, Weipeng Hu |
CVPR | 4 |
| 2024 | Unlearnable Examples Detection via Iterative Filtering
Yi Yu 0011, Qichen Zheng, Siyuan Yang 0001, Wenhan Yang, Jun Liu 0036, Shijian Lu, Yap-Peng Tan, Kwok-Yan Lam, Alex Chichung Kot |
ICANN (10) | 7 |
| 2024 | Adaptive Margin Contrastive Learning for Ambiguity-aware 3D Semantic SegmentationabstractIn this paper, we propose an adaptive margin contrastive learning method for 3D point cloud semantic segmentation, namely AMContrast3D. Most existing methods use equally penalized objectives, which ignore per-point ambiguities and less discriminated features stemming from transition regions. However, as highly ambiguous points may be indistinguishable even for humans, their manually annotated labels are less reliable, and hard constraints over these points would lead to sub-optimal models. To address this, we design adaptive objectives for individual points based on their ambiguity levels, aiming to ensure the correctness of low-ambiguity points while allowing mistakes for high-ambiguity points. Specifically, we first estimate ambiguities based on position embeddings. Then, we develop a margin generator to shift decision boundaries for contrastive feature embeddings, so margins are narrowed due to increasing ambiguities with even negative margins for extremely high-ambiguity points. Experimental results on large-scale datasets, S3DIS and ScanNet, demonstrate that our method outperforms state-of-the-art methods. Yueqi Duan, Runzhong Zhang, Yap-Peng Tan |
ICME | 4 |
| 2024 | Purify Unlearnable Examples via Rate-Constrained Variational AutoencodersabstractUnlearnable examples (UEs) seek to maximize testing error by making subtle modifications to training examples that are correctly labeled. Defenses against these poisoning attacks can be categorized based on whether specific interventions are adopted during training. The first approach is training-time defense, such as adversarial training, which can mitigate poisoning effects but is computationally intensive. The other approach is pre-training purification, e.g., image short squeezing, which consists of several simple compressions but often encounters challenges in dealing with various UEs. Our work provides a novel disentanglement mechanism to build an efficient pre-training purification method. Firstly, we uncover rate-constrained variational autoencoders (VAEs), demonstrating a clear tendency to suppress the perturbations in UEs. We subsequently conduct a theoretical analysis for this phenomenon. Building upon these insights, we introduce a disentangle variational autoencoder (D-VAE), capable of disentangling the perturbations with learnable class-wise embeddings. Based on this network, a two-stage purification approach is naturally developed. The first stage focuses on roughly eliminating perturbations, while the second stage produces refined, poison-free results, ensuring effectiveness and robustness across various scenarios. Extensive experiments demonstrate the remarkable performance of our method across CIFAR-10, CIFAR-100, and a 100-class ImageNet-subset. Code is available at https://github.com/yuyi-sd/D-VAE. Yi Yu 0011, Yufei Wang 0006, Song Xia, Wenhan Yang, Shijian Lu, Yap-Peng Tan, Alex Chichung Kot |
ICML | 6 |
| 2024 | Joint-Neighborhood Product Quantization for Unsupervised Cross-Modal RetrievalabstractProduct quantization (PQ) is a technique that transforms high-dimensional data into compact binary codes to reduce data storage and improve search efficiency. However, existing PQ methods separate the learning of modality-specific features from the learning of quantization codewords, resulting in suboptimal performance in cross-modal retrieval tasks. In this paper, we propose a joint-neighborhood product quantization (JNPQ) method to simultaneously learn modality-specific features and quantization codewords. To achieve this, we first introduce a cross-modal quantization contrastive learning module that preserves the inter-modal neighborhood of the original data and reduces the quantization error. Then, we design a self-neighbor contrastive learning module that enhances the intra-modal neighborhood within individual modalities. Extensive experiments demonstrate that JNPQ achieves state-of-the-art results in crossmodal retrieval when compared with other unsupervised crossmodal quantization methods. Runhao Li, Zhenyu Weng, Yongming Chen, Huiping Zhuang, Yap-Peng Tan, Zhiping Lin 0001 |
VCIP | 5 |
| 2023 | Backdoor Attacks Against Deep Image Compression via Adaptive Frequency TriggerabstractRecent deep-learning-based compression methods have achieved superior performance compared with traditional approaches. However, deep learning models have proven to be vulnerable to backdoor attacks, where some specific trigger patterns added to the input can lead to malicious behavior of the models. In this paper, we present a novel backdoor attack with multiple triggers against learned image compression models. Motivated by the widely used discrete cosine transform (DCT) in existing compression systems and standards, we propose a frequency-based trigger injection model that adds triggers in the DCT domain. In particular, we design several attack objectives for various attacking scenarios, including: 1) attacking compression quality in terms of bit-rate and reconstruction quality; 2) attacking task-driven measures, such as downstream face recognition and semantic segmentation. Moreover, a novel simple dynamic loss is designed to balance the influence of different loss terms adaptively, which helps achieve more efficient training. Extensive experiments show that with our trained trigger injection models and simple modification of encoder parameters (of the compression model), the proposed attack can successfully inject several backdoors with corresponding triggers in a single image compression model. Yi Yu 0011, Yufei Wang 0006, Wenhan Yang, Shijian Lu, Yap-Peng Tan, Alex Chichung Kot |
CVPR | 5 |
| 2023 | Temporal Coherent Test Time Optimization for Robust Video Classification
Chenyu Yi, Siyuan Yang 0001, Yufei Wang 0006, Haoliang Li, Yap-Peng Tan, Alex Chichung Kot |
ICLR | 5 |
| 2023 | HOI-aware Adaptive Network for Weakly-supervised Action SegmentationabstractIn this paper, we propose an HOI-aware adaptive network named AdaAct for weakly-supervised action segmentation. Most existing methods learn a fixed network to predict the action of each frame with the neighboring frames. However, this would result in ambiguity when estimating similar actions, such as pouring juice and pouring coffee. To address this, we aim to exploit temporally global but spatially local human-object interactions (HOI) as video-level prior knowledge for action segmentation. The long-term HOI sequence provides crucial contextual information to distinguish ambiguous actions, where our network dynamically adapts to the given HOI sequence at test time. More specifically, we first design a video HOI encoder that extracts, selects, and integrates the most representative HOI throughout the video. Then, we propose a two-branch HyperNetwork to learn an adaptive temporal encoder, which automatically adjusts the parameters based on the HOI information of various videos on the fly. Extensive experiments on two widely-used datasets including Breakfast and 50Salads demonstrate the effectiveness of our method under different evaluation metrics. Runzhong Zhang, Suchen Wang, Yueqi Duan, Yansong Tang, Yue Zhang 0065, Yap-Peng Tan |
IJCAI | 6 |
| 2023 | POAR: Towards Open Vocabulary Pedestrian Attribute RecognitionabstractPedestrian attribute recognition (PAR) aims to predict the attributes of a target pedestrian. Recent methods often address the PAR problem by training a multi-label classifier with predefined attribute classes, but they can hardly exhaust all possible pedestrian attributes in the real world. To tackle this problem, we propose a novel Pedestrian Open-Attribute Recognition (POAR) approach by formulating the problem as a task of image-text search. Our approach employs a Transformer-based Encoder with a Masking Strategy (TEMS) to focus on the attributes of specific pedestrian parts (e.g., head, upper body, lower body, feet, etc.), and introduces a set of attribute tokens to encode the corresponding attributes into visual embeddings. Each attribute category is described as a natural language sentence and encoded by the text encoder. Then, we compute the similarity between the visual and text embeddings to find the best attribute descriptions for the input images. To handle multiple attributes of a single pedestrian, we propose a Many-To-Many Contrastive (MTMC) loss with masked tokens. In addition, we propose a Grouped Knowledge Distillation (GKD) method to minimize the disparity between visual embeddings and unseen attribute text embeddings. We evaluate our proposed method on three PAR datasets with an open-attribute setting. The results demonstrate the effectiveness of our method as a strong baseline for the POAR task. Our code is available at https://github.com/IvyYZ/POAR. Yue Zhang 0065, Suchen Wang, Shichao Kan, Zhenyu Weng, Yi-Gang Cen, Yap-Peng Tan |
ACM Multimedia | 6 |
| 2023 | Crowd counting from single images using recursive multi-pathway zooming and foreground enhancement
Zhiyang Jia, Yap-Peng Tan, Jun Liu 0036 |
Pattern Recognit. | 5 |
| 2023 | RATIR-Net: Adaptive SAR Image Reconstruction Based on Transformer ArchitectureabstractDespite its widespread use in Earth remote sensing, synthetic aperture radar (SAR) image reconstruction remains challenging. The difficulties mainly lie in the handling of diverse scenes and motion errors with sparsely sampled data. Existing matched filtering (MF)-based methods cannot handle sparsely sampled data, while regularization-based methods lack adaptability to scene diversity. Although deep learning-based SAR methods can deal with these two issues, their performance will be degraded by motion errors. To address this, we propose a Transformer-based SAR image reconstruction method called RATIR-Net. The proposed method can obtain SAR images of various scenes under sparse sampling and motion errors by learning the correlations between echo data. In RATIR-Net, CNN-based encoding and decoding blocks are constructed to implement azimuth processes of range profiles (RP) in the range-Doppler domain according to the MF-based method. Meanwhile, a Residual Attention Transformer (RAT) block is designed to extract correlations between RPs, compensating for information loss caused by sparse sampling and suppressing non-correlated perturbations caused by motion errors. The CNN-based encoding and decoding blocks help reduce computing costs, and the RAT block mitigates the dependence on scene features and the influence of motion errors. These make RATIR-Net efficient and effective. Simulation experiments have been conducted to verify the proposed method. Min Li 0031, Weibo Huo, Yap-Peng Tan, Junjie Wu 0001, Jianyu Yang 0001, Huiyong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Learning Transferable Human-Object Interaction Detector with Natural Language SupervisionabstractIt is difficult to construct a data collection including all possible combinations of human actions and interacting objects due to the combinatorial nature of human-object interactions (HOI). In this work, we aim to develop a transferable HOI detector for unseen interactions. Existing HOI detectors often treat interactions as discrete labels and learn a classifier according to a predetermined category space. This is inherently inapt for detecting unseen interactions which are out of the predefined categories. Conversely, we treat independent HOI labels as the natural language supervision of interactions and embed them into a joint visual-and-text space to capture their correlations. More specifically, we propose a new HOI visual encoder to detect the interacting humans and objects, and map them to a joint feature space to perform interaction recognition. Our visual encoder is instantiated as a Vision Transformer with new learnable HOI tokens and a sequence parser to generate unique HOI predictions. It distills and leverages the transferable knowledge from the pretrained CLIP model to perform the zero-shot interaction detection. Experiments on two datasets, SWIG-HOI and HICO-DET, validate that our proposed method can achieve a notable mAP improvement on detecting both seen and unseen HOIs. Our code is available at https://github.com/scwangdyd/promting_hoi. Suchen Wang, Yueqi Duan, Henghui Ding, Yap-Peng Tan, Kim-Hui Yap, Junsong Yuan 0001 |
CVPR | 4 |
| 2022 | Towards Robust Rain Removal Against Adversarial Attacks: A Comprehensive Benchmark Analysis and BeyondabstractRain removal aims to remove rain streaks from images/videos and reduce the disruptive effects caused by rain. It not only enhances image/video visibility but also allows many computer vision algorithms to function properly. This paper makes the first attempt to conduct a comprehensive study on the robustness of deep learning-based rain removal methods against adversarial attacks. Our study shows that, when the image/video is highly degraded, rain removal methods are more vulnerable to the adversarial attacks as small distortions/perturbations become less noticeable or detectable. In this paper, we first present a comprehensive empirical evaluation of various methods at different levels of attacks and with various losses/targets to generate the perturbations from the perspective of human perception and machine analysis tasks. A systematic evaluation of key modules in existing methods is performed in terms of their robustness against adversarial attacks. From the insights of our analysis, we construct a more robust deraining method by integrating these effective modules. Finally, we examine various types of adversarial attacks that are specific to deraining problems and their effects on both human and machine vision tasks, including 1) rain region attacks, adding perturbations only in the rain regions to make the perturbations in the attacked rain images less visible; 2) object-sensitive attacks, adding perturbations only in regions near the given objects. Code is available at https://github.com/yuyi-sd/Robust_Rain_Removal. Yi Yu 0011, Wenhan Yang, Yap-Peng Tan, Alex Chichung Kot |
CVPR | 3 |
| 2022 | Generalized and Discriminative Collaborative Representation for Multiclass ClassificationabstractThis article presents a generalized collaborative representation-based classification (GCRC) framework, which includes many existing representation-based classification (RC) methods, such as collaborative RC (CRC) and sparse RC (SRC) as special cases. This article also advances the GCRC theory by exploring theoretical conditions on the general regularization matrix. A key drawback of CRC and SRC is that they fail to use the label information of training data and are essentially unsupervised in computing the representation vector. This largely compromises the discriminative ability of the learned representation vector and impedes the classification performance. Guided by the GCRC theory, we propose a novel RC method referred to as discriminative RC (DRC). The proposed DRC method has the following three desirable properties: 1) discriminability: DRC can leverage the label information of training data and is supervised in both representation and classification, thus improving the discriminative ability of the representation vector; 2) efficiency: it has a closed-form solution and is efficient in computing the representation vector and performing classification; and 3) theory: it also has theoretical guarantees for classification. Experimental results on benchmark databases demonstrate both the efficacy and efficiency of DRC for multiclass classification. Yulong Wang 0002, Yap-Peng Tan, Yuan Yan Tang, Hong Chen 0004, Cuiming Zou, Luoqing Li |
IEEE Trans. Cybern. | 2 |
| 2021 | Discovering Human Interactions with Large-Vocabulary Objects via Query and Multi-Scale DetectionabstractIn this work, we study the problem of human-object interaction (HOI) detection with large vocabulary object categories. Previous HOI studies are mainly conducted in the regime of limit object categories (e.g., 80 categories). Their solutions may face new difficulties in both object detection and interaction classification due to the increasing diversity of objects (e.g., 1000 categories). Different from previous methods, we formulate the HOI detection as a query problem. We propose a unified model to jointly discover the target objects and predict the corresponding interactions based on the human queries, thereby eliminating the need of using generic object detectors, extra steps to associate human-object instances, and multi-stream interaction recognition. This is achieved by a repurposed Transformer unit and a novel cascade detection over multi-scale feature maps. We observe that such a highly-coupled solution brings benefits for both object detection and interaction classification in a large vocabulary setting. To study the new challenges of the large vocabulary HOI detection, we assemble two datasets from the publicly available SWiG and 100 Days of Hands datasets. Experiments on these datasets validate that our proposed method can achieve a notable mAP improvement on HOI detection with a faster inference speed than existing one-stage HOI detectors. Our code is available at https://github.com/scwangdyd/large_vocabulary_hoi_detection. Suchen Wang, Kim-Hui Yap, Henghui Ding, Jiyan Wu, Junsong Yuan 0001, Yap-Peng Tan |
ICCV | 6 |
| 2021 | Adversarial Multi-Label Variational HashingabstractIn this paper, we propose an adversarial multi-label variational hashing (AMVH) method to learn compact binary codes for efficient image retrieval. Unlike most existing deep hashing methods which only learn binary codes from specific real samples, our AMVH learns hash functions from both synthetic and real data which make our model effective for unseen data. Specifically, we design an end-to-end deep hashing framework which consists of a generator network and a discriminator-hashing network by enforcing simultaneous adversarial learning and discriminative binary codes learning to learn compact binary codes. The discriminator-hashing network learns binary codes by optimizing a multi-label discriminative criterion and minimizing the quantization loss between binary codes and real-value codes. The generator network is learned so that latent representations can be sampled in a probabilistic manner and used to generate new synthetic training sample for the discriminator-hashing network. Experimental results on several benchmark datasets show the efficacy of the proposed approach. Jiwen Lu, Venice Erin Liong, Yap-Peng Tan |
IEEE Trans. Image Process. | 3 |
| 2020 | Discovering Human Interactions With Novel Objects via Zero-Shot LearningabstractWe aim to detect human interactions with novel objects through zero-shot learning. Different from previous works, we allow unseen object categories by using its semantic word embedding. To do so, we design a human-object region proposal network specifically for the human-object interaction detection task. The core idea is to leverage human visual clues to localize objects which are interacting with humans. We show that our proposed model can outperform existing methods on detecting interacting objects, and generalize well to novel objects. To recognize objects from unseen categories, we devise a zero-shot classification module upon the classifier of seen categories. It utilizes the classifier logits for seen categories to estimate a vector in the semantic space, and then performs nearest search to find the closest unseen category. We validate our method on V-COCO and HICO-DET datasets, and obtain superior results on detecting human interactions with both seen and unseen objects. Suchen Wang, Kim-Hui Yap, Junsong Yuan 0001, Yap-Peng Tan |
CVPR | 4 |
| 2020 | Deep historical long short-term memory network for action recognition
Junlin Hu 0001, Tzu-Yi Hung, Yap-Peng Tan |
Neurocomputing | 5 |
| 2020 | Detecting spatiotemporal irregularities in videos via a 3D convolutional autoencoder
Mengjia Yan 0003, Jingjing Meng, Chunluan Zhou, Zhigang Tu 0001, Yap-Peng Tan, Junsong Yuan 0001 |
J. Vis. Commun. Image Represent. | 5 |
| 2020 | Motion-Guided Cascaded Refinement Network for Video Object SegmentationabstractIn this work, we propose a motion-guided cascaded refinement network for video object segmentation. By assuming the foreground objects show different motion patterns from the background, for each video frame we apply an active contour model on optical flow to coarsely segment the foreground. The proposed Cascaded Refinement Network (CRN) then takes as guidance the coarse segmentation to generate an accurate segmentation in full resolution. In this way, the motion information and the deep CNNs can complement each other well to accurately segment the foreground objects from video frames. To deal with multi-instance cases, we extend our method with a spatial-temporal instance embedding model that further segments the foreground regions into instances and propagates instance labels. We further introduce a single-channel residual attention module in CRN to incorporate the coarse segmentation map as attention, which makes the network effective and efficient in both training and testing. We perform experiments on popular benchmarks and the results show that our method achieves state-of-the-art performance with high time efficiency. Ping Hu 0001, Gang Wang 0012, Xiangfei Kong, Jason Kuen, Yap-Peng Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2020 | Deep Variational and Structural HashingabstractIn this paper, we propose a deep variational and structural hashing (DVStH) method to learn compact binary codes for multimedia retrieval. Unlike most existing deep hashing methods which use a series of convolution and fully-connected layers to learn binary features, we develop a probabilistic framework to infer latent feature representation inside the network. Then, we design a struct layer rather than a bottleneck hash layer, to obtain binary codes through a simple encoding procedure. By doing these, we are able to obtain binary codes discriminatively and generatively. To make it applicable to cross-modal scalable multimedia retrieval, we extend our method to a cross-modal deep variational and structural hashing (CM-DVStH). We design a deep fusion network with a struct layer to maximize the correlation between image-text input pairs during the training stage so that a unified binary vector can be obtained. We then design modality-specific hashing networks to handle the out-of-sample extension scenario. Specifically, we train a network for each modality which outputs a latent representation that is as close as possible to the binary codes which are inferred from the fusion network. Experimental results on five benchmark datasets are presented to show the efficacy of the proposed approach. Venice Erin Liong, Jiwen Lu, Ling-Yu Duan, Yap-Peng Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2020 | Pruning 3D Filters For Accelerating 3D ConvNetsabstractMany methods have been proposed to accelerate 2D ConvNets by removing redundant parameters. However, few efforts are devoted to the problem of accelerating 3D Convolutional Networks. The 3D ConvNets, which are mainly designed for extracting spatiotemporal features, have been widely used in many video analytics tasks, such as action recognition and scene analysis. In this paper, we focus on accelerating 3D ConvNets for two motivations: (1) Fast video processing techniques are in dire need due to the explosive growth of video data; (2) Compared with individual images, video data consist of consecutively similar frames, thus are inherently more redundant. In this paper, we present a novel algorithm to dramatically accelerate 3D ConvNets by pruning redundant convolutional filters, while preserving the discriminative power of the networks. Specifically, we formulate the filter pruning from 3D ConvNets as a subset selection problem where each filter is regarded as a candidate. Determinantal Point Processes (DPPs) are employed to discriminatively select the filter candidates which are informative and yet diverse. We evaluate our method using two popular 3D networks, C3D and Pseudo-3D, on Sports-1 M dataset for video classification. Extensive experimental results demonstrate both the efficiency and performance advantages of our method. We also show that the proposed method can be easily generalized to 2D ConvNets pruning with promising experimental results on VGGnet and ResNet. Weixiang Hong 0001, Yap-Peng Tan, Junsong Yuan 0001 |
IEEE Trans. Multim. | 3 |
| 2019 | Joint Representative Selection and Feature Learning: A Semi-Supervised ApproachabstractIn this paper, we propose a semi-supervised approach for representative selection, which finds a small set of representatives that can well summarize a large data collection. Given labeled source data and big unlabeled target data, we aim to find representatives in the target data, which can not only represent and associate data points belonging to each labeled category, but also discover novel categories in the target data, if any. To leverage labeled source data, we guide representative selection from labeled source to unlabeled target. We propose a joint optimization framework which alternately optimizes (1) representative selection in the target data and (2) discriminative feature learning from both the source and the target for better representative selection. Experiments on image and video datasets demonstrate that our proposed approach not only finds better representatives, but also can discover novel categories in the target data that are not in the source. Suchen Wang, Jingjing Meng, Junsong Yuan 0001, Yap-Peng Tan |
CVPR | 4 |
| 2019 | Scaling Object Detection by Transferring Classification WeightsabstractLarge scale object detection datasets are constantly increasing their size in terms of the number of classes and annotations count. Yet, the number of object-level categories annotated in detection datasets is an order of magnitude smaller than image-level classification labels. State-of-the art object detection models are trained in a supervised fashion and this limits the number of object classes they can detect. In this paper, we propose a novel weight transfer network (WTN) to effectively and efficiently transfer knowledge from classification network's weights to detection network's weights to allow detection of novel classes without box supervision. We first introduce input and feature normalization schemes to curb the under-fitting during training of a vanilla WTN. We then propose autoencoder-WTN (AE-WTN) which uses reconstruction loss to preserve classification network's information over all classes in the target latent space to ensure generalization to novel classes. Compared to vanilla WTN, AE-WTN obtains absolute performance gains of 6% on two Open Images evaluation sets with 500 seen and 57 novel classes respectively, and 25% on a Visual Genome evaluation set with 200 novel classes. Jason Kuen, Federico Perazzi, Zhe Lin 0001, Jianming Zhang 0001, Yap-Peng Tan |
ICCV | 5 |
| 2019 | Atrous convolutions spatial pyramid network for crowd counting and density estimation
Yap-Peng Tan |
Neurocomputing | 3 |
| 2019 | Boosting Positive and Unlabeled Learning for Anomaly Detection With Multi-FeaturesabstractOne of the key challenges of machine learning-based anomaly detection relies on the difficulty of obtaining anomaly data for training, which is usually rare, diversely distributed, and difficult to collect. To address this challenge, we formulate anomaly detection as a Positive and Unlabeled (PU) learning problem where only labeled positive (normal) data and unlabeled (normal and anomaly) data are required for learning an anomaly detector. As a semi-supervised learning method, it does not require providing labeled anomaly data for the training, thus it is easily deployed to various applications. As the unlabeled data can be extremely unbalanced, we introduce a novel PU learning method, which can tackle the situation where an unlabeled data set is mostly composed of positive instances. We start by using a linear model to extract the most reliable negative instances followed by a self-learning process to add reliable negative and positive instances with different speeds based on the estimated positive class prior. Furthermore, when feedback is available, we adopt boosting in the self-learning process to advantageously exploit the instability characteristic of PU learning. The classifiers in the self-learning process are weighted combined based on the estimated error rate to build the final classifier. Extensive experiments on six real datasets and one synthetic dataset show that our methods have better results under different conditions compared to existing methods. Jingjing Meng, Yap-Peng Tan, Junsong Yuan 0001 |
IEEE Trans. Multim. | 4 |
| 2018 | Motion-Guided Cascaded Refinement Network for Video Object SegmentationabstractDeep CNNs have achieved superior performance in many tasks of computer vision and image understanding. However, it is still difficult to effectively apply deep CNNs to video object segmentation(VOS) since treating video frames as separate and static will lose the information hidden in motion. To tackle this problem, we propose a Motion-guided Cascaded Refinement Network for VOS. By assuming the object motion is normally different from the background motion, for a video frame we first apply an active contour model on optical flow to coarsely segment objects of interest. Then, the proposed Cascaded Refinement Network(CRN) takes the coarse segmentation as guidance to generate an accurate segmentation of full resolution. In this way, the motion information and the deep CNNs can well complement each other to accurately segment objects from video frames. Furthermore, in CRN we introduce a Single-channel Residual Attention Module to incorporate the coarse segmentation map as attention, making our network effective and efficient in both training and testing. We perform experiments on the popular benchmarks and the results show that our method achieves state-of-the-art performance at a much faster speed. Ping Hu 0001, Gang Wang 0012, Xiangfei Kong, Jason Kuen, Yap-Peng Tan |
CVPR | 5 |
| 2018 | Stochastic Downsampling for Cost-Adjustable Inference and Improved Regularization in Convolutional NetworksabstractIt is desirable to train convolutional networks (CNNs) to run more efficiently during inference. In many cases however, the computational budget that the system has for inference cannot be known beforehand during training, or the inference budget is dependent on the changing real-time resource availability. Thus, it is inadequate to train just inference-efficient CNNs, whose inference costs are not adjustable and cannot adapt to varied inference budgets. We propose a novel approach for cost-adjustable inference in CNNs - Stochastic Downsampling Point (SDPoint). During training, SDPoint applies feature map downsampling to a random point in the layer hierarchy, with a random downsampling ratio. The different stochastic downsampling configurations known as SDPoint instances (of the same model) have computational costs different from each other, while being trained to minimize the same prediction loss. Sharing network parameters across different instances provides significant regularization boost. During inference, one may handpick a SDPoint instance that best fits the inference budget. The effectiveness of SDPoint, as both a cost-adjustable inference approach and a regularizer, is validated through extensive experiments on image classification. Jason Kuen, Xiangfei Kong, Zhe Lin 0001, Gang Wang 0012, Jianxiong Yin, Simon See, Yap-Peng Tan |
CVPR | 7 |
| 2018 | Multi-Label Deep Sparse HashingabstractIn this paper, we propose a multi-label deep sparse hashing (MDSH) to learn compact binary codes for efficient image retrieval. Unlike most existing supervised hashing methods which only exploit pairwise or triplet-wise similarity to learn binary codes, we perform deep network training such that optimal binary codes are obtained from a sparsity-based discriminative criterion. Specifically, we learn our hashing network by solving a multi-label classification problem with a sparse cross-entropy loss which ensures that sparse probabilities can be obtained while also learning the binary codes. By doing so, our network is able to scale well with ground truth labels which are generally sparse. Experimental results on two widely used multi-label image hashing datasets are presented to show the effectiveness of our proposed approach. Venice Erin Liong, Jiwen Lu, Yap-Peng Tan |
VCIP | 3 |
| 2018 | Sharable and Individual Multi-View Metric LearningabstractThis paper presents a sharable and individual multi-view metric learning (MvML) approach for visual recognition. Unlike conventional metric leaning methods which learn a distance metric on either a single type of feature representation or a concatenated representation of multiple types of features, the proposed MvML jointly learns an optimal combination of multiple distance metrics on multi-view representations, where not only it learns an individual distance metric for each view to retain its specific property but also a shared representation for different views in a unified latent subspace to preserve the common properties. The objective function of the MvML is formulated in the large margin learning framework via pairwise constraints, under which the distance of each similar pair is smaller than that of each dissimilar pair by a margin. Moreover, to exploit the nonlinear structure of data points, we extend MvML to a sharable and individual multi-view deep metric learning (MvDML) method by utilizing the neural network architecture to seek multiple nonlinear transformations. Experimental results on face verification, kinship verification, and person re-identification show the effectiveness of the proposed sharable and individual multi-view metric learning methods. Junlin Hu 0001, Jiwen Lu, Yap-Peng Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Nonlinear dictionary learning with application to image classification
Junlin Hu 0001, Yap-Peng Tan |
Pattern Recognit. | 2 |
| 2018 | Cross-Modal Discrete Hashing
Venice Erin Liong, Jiwen Lu, Yap-Peng Tan |
Pattern Recognit. | 3 |
| 2018 | Scenario-Based Insider Threat Detection From Cyber ActivitiesabstractAn insider threat scenario refers to the outcome of a set of malicious activities caused by intentional or unintentional misuse of the organization's systems, networks, data, and resources. Prevention of insider threat is difficult, since trusted partners of the organization are involved in it, who have authorized access to these confidential/sensitive resources. The state-of-the-art research on insider threat detection mostly focuses on developing unsupervised behavioral anomaly detection techniques with the objective of finding out anomalousness or abnormal changes in user behavior over time. However, an anomalous activity is not necessarily malicious that can lead to an insider threat scenario. As an improvement to the existing approaches, we propose a technique for insider threat detection from time-series classification of user activities. Initially, a set of single-day features is computed from the user activity logs. A time-series feature vector is next constructed from the statistics of each single-day feature over a period of time. The label of each time-series feature vector (whether malicious or nonmalicious) is extracted from the ground truth. To classify the imbalanced ground-truth insider threat data consisting of only a small number of malicious instances, we employ a cost-sensitive data adjustment technique that undersamples the nonmalicious class instances randomly. As a classifier, we employ a two-layered deep autoencoder neural network and compare its performance with other popularly used classifiers: random forest and multilayer perceptron. Encouraging results are obtained by evaluating our approach using the CMU Insider Threat Data, which is the only publicly available insider threat data set consisting of about 14-GB web-browsing logs, along with logon, device connection, file transfer, and e-mail log files. We observe that both deep autoencoder and random forest classifiers classify the dataadjusted time-series feature set with high precision, recall, and f-score. Although multilayer perceptron has a high recall, it suffers from a lower precision and f-score compared to the other two classifiers. Pratik Chattopadhyay, Lipo Wang 0001, Yap-Peng Tan |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2018 | Local Large-Margin Multi-Metric Learning for Face and Kinship VerificationabstractMetric learning has attracted wide attention in face and kinship verification, and a number of such algorithms have been presented over the past few years. However, most existing metric learning methods learn only one Mahalanobis distance metric from a single feature representation for each face image and cannot make use of multiple feature representations directly. In many face-related tasks, we can easily extract multiple features for a face image to extract more complementary information, and it is desirable to learn distance metrics from these multiple features, so that more discriminative information can be exploited than those learned from individual features. To achieve this, we present a large-margin multi-metric learning (LM3L) method for face and kinship verification, which jointly learns multiple global distance metrics under which the correlations of different feature representations of each sample are maximized, and the distance of each positive pair is less than a low threshold and that of each negative pair is greater than a high threshold. To better exploit the local structures of face images, we also propose a local metric learning and local LM3Lmethods to learn a set of local metrics. Experimental results on three face data sets show that the proposed methods achieve very competitive results compared with the state-of-the-art methods. Junlin Hu 0001, Jiwen Lu, Yap-Peng Tan, Junsong Yuan 0001, Jie Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Video Summarization Via Multiview Representative SelectionabstractVideo contents are inherently heterogeneous. To exploit different feature modalities in a diverse video collection for video summarization, we propose to formulate the task as a multiview representative selection problem. The goal is to select visual elements that are representative of a video consistently across different views (i.e., feature modalities). We present in this paper the multiview sparse dictionary selection with centroid co-regularization method, which optimizes the representative selection in each view, and enforces that the view-specific selections to be similar by regularizing them towards a consensus selection. We also introduce a diversity regularizer to favor a selection of diverse representatives. The problem can be efficiently solved by an alternating minimizing optimization with the fast iterative shrinkage thresholding algorithm. Experiments on synthetic data and benchmark video datasets validate the effectiveness of the proposed approach for video summarization, in comparison with other video summarization methods and representative selection methods such as K-medoids, sparse dictionary selection, and multiview clustering. Jingjing Meng, Suchen Wang, Hongxing Wang 0001, Junsong Yuan 0001, Yap-Peng Tan |
IEEE Trans. Image Process. | 5 |
| 2018 | Recurrent Spatial Pyramid CNN for Optical Flow EstimationabstractOptical flow estimation plays an important role in many multimedia and computer vision tasks. Although great progress has been made in applying convolutional neural networks (CNNs) to estimate optical flow in recent works, it is still difficult for CNNs to generate optical flow with the desired effectiveness and efficiency. Compared to CNN-based methods, conventional variational methods normally perform to optimize an energy function and produce optical flow with more precise details. Inspired by the effectiveness of variational methods and deep CNNs, we propose a recurrent spatial pyramid (RecSPy) network for optical flow estimation. To deal with large displacements and to decrease the number of parameters, we formulate the spatial pyramid as a recurrent process, and adopt a CNN to refine optical flow at each spatial scale. Furthermore, to improve the results with more precise details, we propose an energy function that encodes structure and constancy constraints to help refine the optical flow at each spatial scale. The combination of the proposed RecSPy network and the proposed energy-based refinement enables our system to estimate optical flow effectively and efficiently. Experimental results on the benchmarks validate the effectiveness and efficiency of the proposed method. Ping Hu 0001, Gang Wang 0012, Yap-Peng Tan |
IEEE Trans. Multim. | 3 |
| 2017 | Cross-Modal Deep Variational Hashing
Venice Erin Liong, Jiwen Lu, Yap-Peng Tan, Jie Zhou 0001 |
ICCV | 3 |
| 2017 | Learning a cross-modal hashing network for multimedia searchabstractIn this paper, we propose a cross-modal hashing network (CMHN) method to learn compact binary codes for cross-modality multimedia search. Unlike most existing cross-modal hashing methods which learn a single pair of projections to map each example into a binary vector, we design a deep neural network to learn multiple pairs of hierarchical non-linear transformations, under which the nonlinear characteristics of samples can be well exploited and the modality gap is well reduced. Our model is trained under an iterative optimization procedure which learns a (1) unified binary code discretely and discriminatively through a classification-based hinge-loss criterion, and (2) cross-modal hashing network, one deep network for each modality, through minimizing the quantization loss between real-valued neural code and binary code, and maximizing the variance of the learned neural codes. Experimental results on two benchmark datasets show the efficacy of the proposed approach. Venice Erin Liong, Jiwen Lu, Yap-Peng Tan |
ICIP | 3 |
| 2017 | Context-aware graph-based analysis for detecting anomalous activitiesabstractThis paper proposes a context-aware, graph-based approach for identifying anomalous user activities via user profile analysis, which obtains a group of users maximally similar among themselves as well as to the query during test time. The main challenges for the anomaly detection task are: (1) rare occurrences of anomalies making it difficult for exhaustive identification with reasonable false-alarm rate, and (2) continuously evolving new context-dependent anomaly types making it difficult to synthesize the activities apriori. Our proposed query-adaptive graph-based optimization approach, solvable using maximum flow algorithm, is designed to fully utilize both mutual similarities among the user models and their respective similarities with the query to shortlist the user profiles for a more reliable aggregated detection. Each user activity is represented using inputs from several multi-modal resources, which helps to localize anomalies from time-dependent data efficiently. Experiments on public datasets of insider threats and gesture recognition show impressive results. Sreyasee Das Bhattacharjee, Junsong Yuan 0001, Yap-Peng Tan |
ICME | 4 |
| 2017 | Positive and Unlabeled Learning for Anomaly Detection with Multi-featuresabstractAnomaly detection is of great interest to big data applications, and both supervised and unsupervised learning have been applied for anomaly detection. However, it still remains a challenging problem because: (1) for supervised learning, it is difficult to acquire training data for anomaly samples; while (2) for unsupervised learning, the performance may not be satisfactory due to the lack of training data. To address the limitations, we propose a hybrid solution by using both normal (positive) data and unlabeled data (could be positive or negative) for semi-supervised anomaly detection. Particularly, we introduce a new framework based on Positive and Unlabeled (PU) Learning using multi-features to detect anomalies. We extend previous PU learning methods to (1) better address unbalanced class problem which is typical for anomaly detection, and (2) handle multiple features for anomaly detection. An iterative algorithm is proposed to learn the anomaly classifier incrementally from the labeled normal data and also unlabeled data. Our proposed method is verified on three benchmark datasets and one synthetic dataset. Experimental results show that our method outperforms existing methods under different class priors and different proportions of given positive classes. Junsong Yuan 0001, Yap-Peng Tan |
ACM Multimedia | 4 |
| 2017 | Discriminative Deep Metric Learning for Face and Kinship VerificationabstractThis paper presents a new discriminative deep metric learning (DDML) method for face and kinship verification in wild conditions. While metric learning has achieved reasonably good performance in face and kinship verification, most existing metric learning methods aim to learn a single Mahalanobis distance metric to maximize the inter-class variations and minimize the intra-class variations, which cannot capture the nonlinear manifold where face images usually lie on. To address this, we propose a DDML method to train a deep neural network to learn a set of hierarchical nonlinear transformations to project face pairs into the same latent feature space, under which the distance of each positive pair is reduced and that of each negative pair is enlarged. To better use the commonality of multiple feature descriptors to make all the features more robust for face and kinship verification, we develop a discriminative deep multi-metric learning method to jointly learn multiple neural networks, under which the correlation of different features of each sample is maximized, and the distance of each positive pair is reduced and that of each negative pair is enlarged. Extensive experimental results show that our proposed methods achieve the acceptable results in both face and kinship verification. Jiwen Lu, Junlin Hu 0001, Yap-Peng Tan |
IEEE Trans. Image Process. | 3 |
| 2017 | Deep Video HashingabstractIn this work, we propose a deep video hashing (DVH) method for scalable video search. Unlike most existing video hashing methods that first extract features for each single frame and then use conventional image hashing techniques, our DVH learns binary codes for the entire video with a deep learning framework so that both the temporal and discriminative information can be well exploited. Specifically, we fuse the temporal information across different frames within each video to learn the feature representation under two criteria: the distance between a feature pair obtained at the top layer is small if they are from the same class, and large if they are from different classes; and the quantization loss between the real-valued features and the binary codes is minimized. We exploit different deep architectures to utilize spatial-temporal information in different manners and compare them with single-frame-based deep models and state-of-the-art image hashing methods. Experimental results demonstrate the effectiveness of our proposed method. Venice Erin Liong, Jiwen Lu, Yap-Peng Tan, Jie Zhou 0001 |
IEEE Trans. Multim. | 3 |
| 2017 | Deep Coupled Metric Learning for Cross-Modal MatchingabstractIn this paper, we propose a new deep coupled metric learning (DCML) method for cross-modal matching, which aims to match samples captured from two different modalities (e.g., texts versus images, visible versus near infrared images). Unlike existing cross-modal matching methods which learn a linear common space to reduce the modality gap, our DCML designs two feedforward neural networks which learn two sets of hierarchical nonlinear transformations (one set for each modality) to nonlinearly map samples from different modalities into a shared latent feature subspace, under which the intraclass variation is minimized and the interclass variation is maximized, and the difference of each data pair captured from two modalities of the same class is minimized, respectively. Experimental results on four different cross-modal matching datasets validate the efficacy of the proposed approach. Venice Erin Liong, Jiwen Lu, Yap-Peng Tan, Jie Zhou 0001 |
IEEE Trans. Multim. | 3 |
| 2016 | From Keyframes to Key Objects: Video Summarization by Representative Object Proposal SelectionabstractWe propose to summarize a video into a few key objects by selecting representative object proposals generated from video frames. This representative selection problem is formulated as a sparse dictionary selection problem, i.e., choosing a few representatives object proposals to reconstruct the whole proposal pool. Compared with existing sparse dictionary selection based representative selection methods, our new formulation can incorporate object proposal priors and locality prior in the feature space when selecting representatives. Consequently it can better locate key objects and suppress outlier proposals. We convert the optimization problem into a proximal gradient problem and solve it by the fast iterative shrinkage thresholding algorithm (FISTA). Experiments on synthetic data and real benchmark datasets show promising results of our key object summarization approach in video content mining and search. Comparisons with existing representative selection approaches such as K-mediod, sparse dictionary selection and density based selection validate that our formulation can better capture the key video objects despite appearance variations, cluttered backgrounds and camera motions. Jingjing Meng, Hongxing Wang 0001, Junsong Yuan 0001, Yap-Peng Tan |
CVPR | 4 |
| 2016 | Collaborative multi-view metric learning for visual classificationabstractMost of distance metric learning algorithms usually learn a single distance metric over the single-view data and cannot directly exploit multi-view data. In many visual classification applications, we have access to multi-view feature representations. To exploit more discriminative information for classification, it is desired to learn several distance metrics from multi-view data. To this aim, we propose a collaborative multi-view metric learning (CMML) method for visual classification. The proposed method jointly learns multiple distance metrics under which multiple feature representations are consistent across different views, i.e., the difference of the distance metrics learned in different views is enforced to be as small as possible. Experimental results on two visual classification tasks including face recognition and scene classification show the efficacy of the CMML method. Junlin Hu 0001, Jiwen Lu, Junsong Yuan 0001, Yap-Peng Tan |
ICME | 4 |
| 2016 | Nonlinear metric learning for visual trackingabstractWe propose a nonlinear metric learning (NML) method for visual tracking. Instead of utilizing the hand-crafted similarity measures, the NML tracker can automatically learn distance metrics from training data itself to categorize object and backgrounds in visual tracking. To exploit the nonlinear structures of samples, the NML tracker seeks several hierarchical nonlinear transformations by adopting the neural network architectures to map candidates and template into a latent subspace where the distance of each positive pair is smaller than that of each negative pair. In this learned metric space, the candidate that maintains the minimum distance to the template is treated as the final tracking result. Evaluation on 20 challenging videos shows the efficacy of the NML tracker. Jiwen Lu, Junlin Hu 0001, Yap-Peng Tan |
ICME | 3 |
| 2016 | Learning a Multi-class Discriminative Dictionary with Nonredundancy Constraints for Visual ClassificationabstractRecent studies have demonstrated advantages of sparse representation in providing an appealing paradigm for visual classification tasks. However, how to effectively learn a compact dictionary of superior reconstruction and discrimination power is still a challenging problem. In this paper, we concurrently exploit both the intra-class and the inter-class visual correlations to learn a multi-class discriminative dictionary. The intra-nonredundancy constraint prevents zero entities from appearing in the class-specific bases, thereby making the learned dictionary more stable. The inter-nonredundancy constraint effectively separates the common visual patterns from all the class-specific bases, yielding a more compact dictionary. Combining nonredundancy constraints with the reconstruction error and the classification error to form a unified objective function, our method can learn a superior dictionary and an optimal linear classifier simultaneously. Extensive experimental results demonstrate that the proposed algorithm achieves notable improvement over the state-of-the-art methods in image classification and visual tracking tasks. Yuwei Wu 0001, Junsong Yuan 0001, Yap-Peng Tan |
ACM Multimedia | 4 |
| 2016 | Deep Metric Learning for Visual TrackingabstractIn this paper, we propose a deep metric learning (DML) approach for robust visual tracking under the particle filter framework. Unlike most existing appearance-based visual trackers, which use hand-crafted similarity metrics, our DML tracker learns a nonlinear distance metric to classify the target object and background regions using a feed-forward neural network architecture. Since there are usually large variations in visual objects caused by varying deformations, illuminations, occlusions, motions, rotations, scales, and cluttered backgrounds, conventional linear similarity metrics cannot work well in such scenarios. To address this, our proposed DML tracker first learns a set of hierarchical nonlinear transformations in the feed-forward neural network to project both the template and particles into the same feature space where the intra-class variations of positive training pairs are minimized and the interclass variations of negative training pairs are maximized simultaneously. Then, the candidate that is most similar to the template in the learned deep network is identified as the true target. Experiments on the benchmark data set including 51 challenging videos show that our DML tracker achieves a very competitive performance with the state-of-the-art trackers. Junlin Hu 0001, Jiwen Lu, Yap-Peng Tan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2016 | Deep Transfer Metric LearningabstractConventional metric learning methods usually assume that the training and test samples are captured in similar scenarios so that their distributions are assumed to be the same. This assumption does not hold in many real visual recognition applications, especially when samples are captured across different data sets. In this paper, we propose a new deep transfer metric learning (DTML) method to learn a set of hierarchical nonlinear transformations for cross-domain visual recognition by transferring discriminative knowledge from the labeled source domain to the unlabeled target domain. Specifically, our DTML learns a deep metric network by maximizing the inter-class variations and minimizing the intra-class variations, and minimizing the distribution divergence between the source domain and the target domain at the top layer of the network. To better exploit the discriminative information from the source domain, we further develop a deeply supervised transfer metric learning (DSTML) method by including an additional objective on DTML, where the output of both the hidden layers and the top layer are optimized jointly. To preserve the local manifold of input data points in the metric space, we present two new methods, DTML with autoencoder regularization and DSTML with autoencoder regularization. Experimental results on face verification, person re-identification, and handwritten digit recognition validate the effectiveness of the proposed methods. Junlin Hu 0001, Jiwen Lu, Yap-Peng Tan, Jie Zhou 0001 |
IEEE Trans. Image Process. | 3 |
| 2016 | Robust Point Set Matching for Partial Face RecognitionabstractOver the past three decades, a number of face recognition methods have been proposed in computer vision, and most of them use holistic face images for person identification. In many real-world scenarios especially some unconstrained environments, human faces might be occluded by other objects, and it is difficult to obtain fully holistic face images for recognition. To address this, we propose a new partial face recognition approach to recognize persons of interest from their partial faces. Given a pair of gallery image and probe face patch, we first detect keypoints and extract their local textural features. Then, we propose a robust point set matching method to discriminatively match these two extracted local feature sets, where both the textural information and geometrical information of local features are explicitly used for matching simultaneously. Finally, the similarity of two faces is converted as the distance between these two aligned feature sets. Experimental results on four public face data sets show the effectiveness of the proposed approach. Renliang Weng, Jiwen Lu, Yap-Peng Tan |
IEEE Trans. Image Process. | 3 |
| 2016 | Query-Adaptive Small Object Search Using Object Proposals and Shape-Aware DescriptorsabstractWhile there has been a significant amount of work on object search and image retrieval, the focus has primarily been on establishing effective models for the whole images, scenes, and objects occupying a large portion of an image. In this paper, we propose to leverage object proposals to identify small and smooth-structured objects in a large image database. Unlike popular methods exploring a coarse image-level pairwise similarity, the search is designed to exploit the similarity measures at the proposal level. An effective graph-based query expansion strategy is designed to assess each of these better matched proposals against all its neighbors within the same image for a precise localization. Combined with a shape-aware feature descriptor EdgeBoW, a set of more insightful edge-weights and node-utility measures, the proposed search strategy can handle varying view angles, illumination conditions, deformation, and occlusion efficiently. Experiments performed on a number of other benchmark datasets show the powerful and superior generalization ability of this single integrated framework in dealing with both clutter-intensive real-life images and poor-quality binary document images at equal dexterity. Sreyasee Das Bhattacharjee, Junsong Yuan 0001, Yap-Peng Tan, Ling-Yu Duan |
IEEE Trans. Multim. | 3 |
| 2016 | Object Instance Search in Videos via Spatio-Temporal Trajectory DiscoveryabstractGiven a specific object as query, object instance search aims to not only retrieve the images or frames that contain the query, but also locate all its occurrences. In this work, we explore the use of spatio-temporal cues to improve the quality of object instance search from videos. To this end, we formulate this problem as the spatio-temporal trajectory search problem, where a trajectory is a sequence of bounding boxes that locate the object instance in each frame. The goal is to find the top- K trajectories that are likely to contain the target object. Despite the large number of trajectory candidates, we build on a recent spatio- temporal search algorithm for event detection to efficiently find the optimal spatio- temporal trajectories in large video volumes , with complexity linear to the video volume size. We solve the key bottleneck in applying this approach to object instance search by leveraging a randomized approach to enable fast scoring of any bounding boxes in the video volume. In addition , we present a new dataset for video object instance search. Experimental results on a 73-hour video dataset demonstrate that our approach improves the performance of video object instance search and localization over the state-of-the-art search and tracking methods. Jingjing Meng, Junsong Yuan 0001, Gang Wang 0012, Yap-Peng Tan |
IEEE Trans. Multim. | 5 |
| 2016 | Learning Cascaded Deep Auto-Encoder Networks for Face AlignmentabstractIn this paper, we propose a new cascaded deep auto-encoder networks (CDAN) approach for face alignment. Our framework consists of a global exemplar-based deep auto-encoder network (GEDAN) and a series of localized deep auto-encoder networks (LDAN) in a cascaded fashion. The global network takes a low-resolution holistic facial image as input and generates a preliminary facial landmark configuration. The following localized networks sample pose-indexed local features around current landmark positions, and refine the landmark positions with increasingly higher image resolutions. Our network architectures are designed to achieve greater robustness against pose variations as well as higher landmark estimation accuracy. Experimental results on three datasets show that the proposed approach achieves superior alignment accuracy with real-time speed. Renliang Weng, Jiwen Lu, Yap-Peng Tan, Jie Zhou 0001 |
IEEE Trans. Multim. | 3 |
| 2016 | A Survey on Visual Analytics of Social Media DataabstractThe unprecedented availability of social media data offers substantial opportunities for data owners, system operators, solution providers, and end users to explore and understand social dynamics. However, the exponential growth in the volume, velocity, and variability of social media data prevents people from fully utilizing such data. Visual analytics, which is an emerging research direction, has received considerable attention in recent years. Many visual analytics methods have been proposed across disciplines to understand large-scale structured and unstructured social media data. This objective, however, also poses significant challenges for researchers to obtain a comprehensive picture of the area, understand research challenges, and develop new techniques. In this paper, we present a comprehensive survey to characterize this fast-growing area and summarize the state-of-the-art techniques for analyzing social media data. In particular, we classify existing techniques into two categories: gathering information and understanding user behaviors. We aim to provide a clear overview of the research area through the established taxonomy. We then explore the design space and identify the research trends. Finally, we discuss challenges and open questions for future studies. Yingcai Wu, Nan Cao 0001, David Gotz, Yap-Peng Tan, Daniel A. Keim |
IEEE Trans. Multim. | 4 |
| 2015 | Deep transfer metric learningabstractConventional metric learning methods usually assume that the training and test samples are captured in similar scenarios so that their distributions are assumed to be the same. This assumption doesn't hold in many real visual recognition applications, especially when samples are captured across different datasets. In this paper, we propose a new deep transfer metric learning (DTML) method to learn a set of hierarchical nonlinear transformations for cross-domain visual recognition by transferring discriminative knowledge from the labeled source domain to the unlabeled target domain. Specifically, our DTML learns a deep metric network by maximizing the inter-class variations and minimizing the intra-class variations, and minimizing the distribution divergence between the source domain and the target domain at the top layer of the network. To better exploit the discriminative information from the source domain, we further develop a deeply supervised transfer metric learning (DSTML) method by including an additional objective on DTML where the output of both the hidden layers and the top layer are optimized jointly. Experimental results on cross-dataset face verification and person re-identification validate the effectiveness of the proposed methods. Junlin Hu 0001, Jiwen Lu, Yap-Peng Tan |
CVPR | 3 |
| 2015 | Fast object instance search in videos from one exampleabstractWe present an efficient approach to search for and locate all occurrences of a specific object in large video volumes, given a single query example. Locations of object occurrences are returned as spatio-temporal trajectories in the 3D video volume. Despite much work on object instance search in image datasets, these methods locate the object independently in each image, therefore do not preserve the spatio-temporal consistency in consecutive video frames. This results in sub-optimal performance if directly applied to videos, as will be shown in our experiments. We propose to locate the object jointly across video frames using spatio-temporal search. The efficiency and effectiveness of the proposed approach is demonstrated on a consumer video dataset consisting of crawled YouTube videos and mobile captured consumer clips. Our method significantly improves the localized search accuracy over the baseline, which treats each frame independently. Moreover, it is able to find the top 100 object trajectories in the 5.5-hour dataset within 30 seconds. Jingjing Meng, Junsong Yuan 0001, Yap-Peng Tan, Gang Wang 0012 |
ICIP | 3 |
| 2015 | Query-Adaptive Logo Search using Shape-Aware DescriptorsabstractWe propose a graph-based optimization framework to leverage category independent object proposals (candidate object regions) for logo search in a large scale image database. The proposed contour-based feature descriptor EdgeBoW is robust to view-angle changes, varying illumination conditions and can implicitly capture the significant object shape information. Having been equipped with a local descriptor, it can handle a fair amount of occlusion and deformation frequently present in a real-life scenario. Given a small set of initially retrieved candidate object proposals, a fast graph-based short-listing scheme is designed to exploit the mutual similarities among these proposals for eliminating outliers. In contrast to a coarse image-level pairwise similarity measure, this search focussed on a few specific image regions provides a more accurate method for matching. The proposed query expansion strategy aims to assess each of the remaining better matched proposals against all its neighbors within the same image for a precise localization. Combined with an efficient feature descriptor EdgeBoW, a set of more insightful edge-weights and node-utility measures can yield promising results, specially for object categories primarily defined by its shape. Extensive set of experiments performed on a number of benchmark datasets demonstrates its effectiveness and superior generalization ability in both clutter intensive real-life images and poor quality binary document images. Sreyasee Das Bhattacharjee, Junsong Yuan 0001, Yap-Peng Tan, Ling-Yu Duan |
ACM Multimedia | 3 |
| 2015 | Motion-compensated orthonormal expansion ℓ1-minimization for reference-driven MRI reconstruction using Augmented Lagrangian methods
Yap-Peng Tan |
J. Vis. Commun. Image Represent. | 2 |
| 2015 | QoE-aware video streaming for SVC over multiuser MIMO-OFDM systems
Maodong Li 0001, Peng Hui Tan, Sumei Sun, Yap-Peng Tan |
J. Vis. Commun. Image Represent. | 5 |
| 2015 | Rate and Power Allocation for Joint Coding and Transmission in Wireless Video Chat ApplicationsabstractWireless video chat is a power-consuming and bitrate-intensive application. Unlike video streaming, which is one-way traffic, video chat features distributed two-way traffic relayed via base stations, where resource allocation of a client affects the video quality seen by its communicating partner. In this paper, we study the mechanism design of this application via dynamic pricing, and seek efficiency and fairness of resource utilization. Specifically, we assume that the base station relays video bitstreams and charges a service price on the clients based on the transmission power consumption. Based on the price and a given power budget, the clients allocate bitrate and power for video coding and transmission such that the service price and the distortion seen by their partners are minimized. We study such network dynamics in Stackelberg game-theoretic framework. To solve the problem, we propose a complexity-scalable video encoding method and a power-rate-distortion (PRD) model for video chat. The model is more accurate in describing the PRD characteristics, yet of lower complexity in online updates of its coefficients. Based on the PRD model, we derive the distributed rate and power allocations for the clients. We show that a simple pricing update in the base stations is sufficient for optimal pricing. The proposed algorithms are optimal and converge to the Stackelberg equilibrium. Existing SNR- and power-based pricing schemes could not ensure fairness and efficiency simultaneously. We propose a hybrid pricing scheme that balances these conflicting criteria. Extensive simulations demonstrate superior performance of the proposed methods and solutions. Seong-Ping Chuah, Yap-Peng Tan |
IEEE Trans. Multim. | 2 |
| 2014 | Large Margin Multi-metric Learning for Face and Kinship Verification in the Wild
Junlin Hu 0001, Jiwen Lu, Junsong Yuan 0001, Yap-Peng Tan |
ACCV (3) | 4 |
| 2014 | Discriminative Deep Metric Learning for Face Verification in the WildabstractThis paper presents a new discriminative deep metric learning (DDML) method for face verification in the wild. Different from existing metric learning-based face verification methods which aim to learn a Mahalanobis distance metric to maximize the inter-class variations and minimize the intra-class variations, simultaneously, the proposed DDML trains a deep neural network which learns a set of hierarchical nonlinear transformations to project face pairs into the same feature subspace, under which the distance of each positive face pair is less than a smaller threshold and that of each negative pair is higher than a larger threshold, respectively, so that discriminative information can be exploited in the deep network. Our method achieves very competitive face verification performance on the widely used LFW and YouTube Faces (YTF) datasets. Junlin Hu 0001, Jiwen Lu, Yap-Peng Tan |
CVPR | 3 |
| 2014 | Efficient Sparsity Estimation via Marginal-Lasso Coding
Tzu-Yi Hung, Jiwen Lu, Yap-Peng Tan, Shenghua Gao |
ECCV (4) | 3 |
| 2014 | Cost optimal video transcoding in media cloud: Insights from user viewing patternabstractVideo transcoding has been touted as an enabling technology to support growing media consumption over heterogenous devices. However, on-line transcoding could incur tremendous, if not prohibitive, cost in deploying or renting resources. In this research, we leverage an insight into the viewing pattern of video consumers to reduce the operating cost of video transcoding services. Specifically, it has been reported that viewers tend to terminate their session before the whole video is watched. As such, it is not cost-efficient for service providers to store or transcode all segments of the videos. Built upon this insight, we propose a partial transcoding scheme for content management in a media cloud to reduce the operating cost. Particularly, each content is split into multiple segments and stored in different files of varying playback rates. Some of the segments are stored in cache, resulting in storage cost; while some are transcoded in real-time in case of cache miss, resulting in computing cost. We aim to minimize the long-term operational cost by determining the number of segments for each playback rate to be cached or transcoded in real-time. We formulate this partial transcoding scheme as a constrained integer optimization problem. Leveraging Lagrangian relaxation and a subgradient method, we obtain the approximate solution to the integer program. Numerical results indicate that our proposed partial transcoding scheme can save more than 30% of operational cost, compared with a brute-force scheme of caching all the segments. Guanyu Gao, Yonggang Wen 0001, Zhi Wang 0001, Wenwu Zhu 0001, Yap-Peng Tan |
ICME | 6 |
| 2014 | Complete discriminative feature learning: A new approach for heterogeneous face recognitionabstractIn this paper, we propose a new feature learning approach called complete discriminative feature learning (CDFL) for heterogeneous face recognition. Unlike most existing heterogeneous face recognition methods where hand-crafted feature descriptors are used for face representation, the proposed CD-FL aims to learn an optimal weighted discriminative image filter to improve learning discriminative filters, so that complete discriminative information is exploited and the feature difference between different modalities is effectively reduced, simultaneously. Experimental results shows that our approach consistently outperforms the state-of-the-art methods. Yi Jin 0001, Jiwen Lu, Qiuqi Ruan, Yap-Peng Tan |
ICME | 4 |
| 2014 | Multi-manifold metric learning for face recognition based on image sets
Likun Huang, Jiwen Lu, Yap-Peng Tan |
J. Vis. Commun. Image Represent. | 3 |
| 2014 | Neighborhood Repulsed Metric Learning for Kinship VerificationabstractKinship verification from facial images is an interesting and challenging problem in computer vision, and there are very limited attempts on tackle this problem in the literature. In this paper, we propose a new neighborhood repulsed metric learning (NRML) method for kinship verification. Motivated by the fact that interclass samples (without a kinship relation) with higher similarity usually lie in a neighborhood and are more easily misclassified than those with lower similarity, we aim to learn a distance metric under which the intraclass samples (with a kinship relation) are pulled as close as possible and interclass samples lying in a neighborhood are repulsed and pushed away as far as possible, simultaneously, such that more discriminative information can be exploited for verification. To make better use of multiple feature descriptors to extract complementary information, we further propose a multiview NRML (MNRML) method to seek a common distance metric to perform multiple feature fusion to improve the kinship verification performance. Experimental results are presented to demonstrate the efficacy of our proposed methods. Finally, we also test human ability in kinship verification from facial images and our experimental results show that our methods are comparable to that of human observers. Jiwen Lu, Xiuzhuang Zhou, Yap-Peng Tan, Jie Zhou 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2014 | Co-Learned Multi-View Spectral Clustering for Face Recognition Based on Image SetsabstractDifferent from the existing approaches that usually utilize single view information of image sets to recognize persons, multi-view information of image sets is exploited in this paper, where a novel method called Co-Learned Multi-View Spectral Clustering (CMSC) is proposed to recognize faces based on image sets. In order to make sure that a data point under different views is assigned to the same cluster, we propose an objective function that optimizes the approximations of the cluster indicator vectors for each view and meanwhile maximizes the correlations among different views. Instead of using an iterative method, we relax the constraints such that the objective function can be solved immediately. Experiments are conducted to demonstrate the efficiency and accuracy of the proposed CMSC method. Likun Huang, Jiwen Lu, Yap-Peng Tan |
IEEE Signal Process. Lett. | 3 |
| 2013 | Activity-based human identificationabstractWe investigate in this paper the problem of activity-based human identification. Different from most existing gait recognition methods where only human walking activity is considered and utilized for person identification, we aim to identify people from various activities such as eating, jumping, and weaving. For each video clip, we first extract binary human body masks by using background substraction, followed by computing the average energy image (AEI) features to represent each video clip. Then, a mapping is learned by applying an adaptive discriminant analysis (ADA) method to project AEI features into a low-dimensional subspace, such that the intra-class (activities performed by the same person) variations are minimized and the interclass (activities performed by different persons) are maximized, simultaneously. Moreover, interclass samples with large similarity difference are deemphasized and those with small difference are emphasized, such that more discriminative information can be used for recognition. Experimental results on three publicly available databases show the efficacy of our proposed approach. Tzu-Yi Hung, Jiwen Lu, Junlin Hu 0001, Yap-Peng Tan, Yongxin Ge |
ICASSP | 4 |
| 2013 | Robust Feature Set Matching for Partial Face RecognitionabstractOver the past two decades, a number of face recognition methods have been proposed in the literature. Most of them use holistic face images to recognize people. However, human faces are easily occluded by other objects in many real-world scenarios and we have to recognize the person of interest from his/her partial faces. In this paper, we propose a new partial face recognition approach by using feature set matching, which is able to align partial face patches to holistic gallery faces automatically and is robust to occlusions and illumination changes. Given each gallery image and probe face patch, we first detect key points and extract their local features. Then, we propose a Metric Learned Extended Robust Point Matching (MLERPM) method to discriminatively match local feature sets of a pair of gallery and probe samples. Lastly, the similarity of two faces is converted as the distance between two feature sets. Experimental results on three public face databases are presented to show the effectiveness of the proposed approach. Renliang Weng, Jiwen Lu, Junlin Hu 0001, Gao Yang 0001, Yap-Peng Tan |
ICCV | 5 |
| 2013 | Distributed rate and power allocation for wireless video chats via pricing schemesabstractVideo chat is a power and rate-intensive application which requires efficient resource utilization. Unlike video streaming which is generally one way, video chats characterize distributed two way traffics relayed via base stations. In this paper, we propose a distributed rate and power allocation framework for joint coding and transmission in wireless video chats. The base station imposes a service charge, which considers relay transmission power as a cost, for relaying video bitstreams. For clients, we derive the optimal rate and power allocation for video coding and transmission such that the network service charge and video distortion are minimized under a power constraint. For the base station, existing pricing schemes could not ensure fairness and efficiency simultaneously. We propose an optimal hybrid pricing scheme which allows balanced tradeoff between fairness and efficiency in network service. Network dynamics of video chats can be analyzed in the Stackelberg game framework, and shown to converge to the Stackelberg equilibrium. Extensive simulations confirm the performance analysis of the proposed solutions and the network dynamics. Seong-Ping Chuah, Yap-Peng Tan |
ICME | 3 |
| 2013 | Collaborative reconstruction-based manifold-manifold distance for face recognition with image setsabstractIn this paper, we propose a new collaborative reconstruction-based manifold-manifold distance (CRMMD) method for face recognition with image sets, where each gallery and probe sample is a set of face images captured from varying poses, illuminations and expressions. Given each face image set, we first model it as a nonlinear manifold and then the recognition task is converted as a manifold-manifold matching problem. For each manifold, we divide it into several clusters and describe each cluster by using a local model. Then, we use the local models from each gallery manifold to collaboratively reconstruct each local model of the testing manifold and the minimal reconstruction error is used for classification. Experimental results on three widely used face datasets are presented to show the effectiveness of the proposed method. Likun Huang, Jiwen Lu, Yap-Peng Tan |
ICME | 3 |
| 2013 | Graph-based sparse coding and embedding for activity-based human identificationabstractIn this paper, we propose a new graph-based sparse coding and embedding (GSCE) method for activity-based human identification. Different from human activity recognition which recognizes different types of human activities such as walking, running, eating, and drinking, in this study, we aim to identify persons from his/her activities. To our best knowledge, this problem has been seldom investigated in the literature. Given a training set of video clips, we first extract human body mask in each frame and learn a codebook to quantize these masks into a histogram feature by using a graphbased sparse coding technique to better preserve the similarity information of different frames within a same video clip. Moreover, we also learn a mapping to project each frame into a low-dimensional subspace to speed up the quantization procedure, such that more discriminative information can be further exploited for classification. Experimental results on three databases are presented to show the efficacy of the proposed method. Tzu-Yi Hung, Jiwen Lu, Yap-Peng Tan |
ICME | 3 |
| 2013 | Mobile media communication, processing, and analysis: A review of recent advancesabstractIn this paper, we review recent advances in mobile media communication, processing, and analysis. To identify the opportunities and challenges in fast growing mobile media computing, we discuss several emerging topics including mobile visual search, retargeting, mobile video streaming, and cloud based mobile media computing. According to the infrastructure of mobile devices vs. servers, we come up with essential concerns in mobile media computing such as wireless bandwidth consumption, mobile energy saving, media adaptation for better quality of services, the computational load shift from mobiles to servers, etc. With booming mobile Apps on diverse media consumption, it is envisioned that mobile media research and development is bringing about significant achievements in traditional topics of communication, processing, and analytics. Wen Gao 0001, Ling-Yu Duan, Jun Sun 0007, Junsong Yuan 0001, Yonggang Wen 0001, Yap-Peng Tan, Jianfei Cai 0001, Alex Chichung Kot |
ISCAS | 6 |
| 2013 | Cross-scene abnormal event detectionabstractThis paper presents an cross-scene abnormal event detection method by adopting Bag of Words (BoW) model with Spatial Pyramid Matching Kernel (SPM) cooperating with SIFT features and a SVM classifier. Different from existing abnormal event detection methods where abnormal events happened in a well-learned scene are considered and detected, we aim to detect concerned events in public where scenes can be unlearned before. Our method is motivated by the fact that the pattern of the notable events are similar and the learned models should be transferable to examine the events in other unlearned public scenes. To learn the patterns for an abnormal event, we divide the proposed method into two steps: feature coding and spatial pooling. For the feature coding step, the codebook is generated and the feature is quantized based on small patches. For the spatial pooling step, the patches are concatenating to exploit the spatial information of local regions. The intersection kernel is used to integrate with a SVM classifier. Experimental results on two benchmark databases demonstrate the efficacy of our proposed approach. Tzu-Yi Hung, Jiwen Lu, Yap-Peng Tan |
ISCAS | 3 |
| 2013 | A fast rate adaptation scheme for SVC based on the packet dependenciesabstractThe Scalable Video Coding (SVC) standard offers multiple scalabilities while maintaining high coding efficiency. However, the joint coding of multiple scalabilities complicates the rate adaptation as the SVC packets possess different priorities. To address this challenge, we propose in this paper a fast rate adaptation scheme for SVC in absence of the original video sequence. Specifically, we present an efficient algorithm to extract the SVC bitstream by a packet prioritization scheme, which is based on the analysis of packet dependencies (PD) in encoding of full scalability. Experimental results demonstrate that the proposed scheme achieves significant improvement in PSNR over the basic bit extraction scheme in SVC and attains comparable PSNR performance to the Quality Layer based rate adaptation scheme while significantly reducing the computational cost. Maodong Li 0001, Seong-Ping Chuah, Yap-Peng Tan |
ISCAS | 4 |
| 2013 | Complexity-scalable video coding and power-rate-distortion modeling forwireless video chat applicationsabstractWireless video chat is a power-consuming and of high bitrate application. To prolong the operational lifetime, optimal rate and power allocations in joint coding and transmission are necessary. We exploit low motion and high inter-frame correlation in video chats to determine a complexity-scalable video coding adaptation which is Pareto optimal. We propose a model that describes power-rate-distortion (PRD) characteristic of the complexity scalable video coding more accurately. As video contents are non-stationary, we formulate an online algorithm for the model's parameters updates. We demonstrate that the PRD model with online updates can be applied to solve the rate and power allocation problem optimally in joint coding and transmission. Simulation results confirm that the model describes PRD characteristics more accurately via online recursive updates. The resource allocation scheme which is based on the PRD model yields better video quality than recent methods in a resource-constrained wireless video chat application. Seong-Ping Chuah, Yap-Peng Tan |
VCIP | 3 |
| 2013 | Robust partial face recognition using instance-to-class distanceabstractWe present a new face recognition approach from partial face patches by using an instance-to-class distance. While numerous face recognition methods have been proposed over the past two decades, most of them recognize persons from whole face images. In many real world applications, partial faces usually occur in unconstrained scenarios such as visual surveillance systems. Hence, it is very important to recognize an arbitrary facial patch to enhance the intelligence of such systems. In this paper, we develop a robust partial face recognition approach based on local feature representation, where the similarity between each probe patch and gallery face is computed by using the instance-to-class distance with the sparse constraint. Experiments on two popular face datasets are presented to show the efficacy of our proposed method. Junlin Hu 0001, Jiwen Lu, Yap-Peng Tan |
VCIP | 3 |
| 2013 | On quality of experience of scalable video adaptation
Maodong Li 0001, Yap-Peng Tan |
J. Vis. Commun. Image Represent. | 3 |
| 2013 | Robust gait recognition via discriminative set matching
Nini Liu, Jiwen Lu, Gao Yang 0001, Yap-Peng Tan |
J. Vis. Commun. Image Represent. | 4 |
| 2013 | Discriminative Multimanifold Analysis for Face Recognition from a Single Training Sample per PersonabstractConventional appearance-based face recognition methods usually assume that there are multiple samples per person (MSPP) available for discriminative feature extraction during the training phase. In many practical face recognition applications such as law enhancement, e-passport, and ID card identification, this assumption, however, may not hold as there is only a single sample per person (SSPP) enrolled or recorded in these systems. Many popular face recognition methods fail to work well in this scenario because there are not enough samples for discriminant learning. To address this problem, we propose in this paper a novel discriminative multimanifold analysis (DMMA) method by learning discriminative features from image patches. First, we partition each enrolled face image into several nonoverlapping patches to form an image set for each sample per person. Then, we formulate the SSPP face recognition as a manifold-manifold matching problem and learn multiple DMMA feature spaces to maximize the manifold margins of different persons. Finally, we present a reconstruction-based manifold-manifold distance to identify the unlabeled subjects. Experimental results on three widely used face databases are presented to demonstrate the efficacy of the proposed approach. Jiwen Lu, Yap-Peng Tan, Gang Wang 0012 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | Hybrid Saliency Detection for ImagesabstractSaliency information interpreted from the visual stimuli can predict the attentional behaviour of human perception, thus playing a key role in visual signal processing. In this letter, we present a hybrid saliency detection method for images by which we automatically predict the saliency regions based on low-level and high-level cues. Unlike existing bottom-up and top-down attentional methods, we consider the high-level cue imposed by the photographer. Based on this assumption, we estimate the defocus map of the image and integrate it with other low-level features based on the Bayesian framework. We compare our algorithm to several state-of-the-art saliency detection methods based on the well-known 1000 image EPFL database, and demonstrate the superior performance of our proposed algorithm. Junsong Yuan 0001, Yap-Peng Tan |
IEEE Signal Process. Lett. | 3 |
| 2013 | Energy Minimization for Wireless Video Transmissions With Deadline and Reliability ConstraintsabstractIn wireless video transmissions, encoded video frames are often large in data load and truncated into many transport packets (TPs) for reliable transmissions. These TPs are to be delivered before a deadline at certain reliability that depends on the importance of the video frame. High power and bitrate transmission schemes are often deployed to ensure low loss rate, but at the cost of substantial energy consumption. This paper addresses the energy-minimizing transmission policy for highly reliable transmissions of a group of TP with a common deadline. We jointly adapt the transmission rate, transmission power, and retransmission limit to minimize the transmission energy while ensuring that the video frame is reliably delivered before a deadline. In a slow fading channel, we formulate a deterministic transmission policy that allocates a retransmission limit to the TPs and jointly optimizes with the transmission bitrate and power. Contrary to the intuition that avoids packet loss and retransmission to preserve energy, we demonstrate that allowing some retransmissions in the joint optimization consumes less energy, without compromising the target reliability. In a Rayleigh fading channel, a conventional approach adopts a supportable transmission rate and declares off-channel at deep fading. We generalize the approach to include various transmission schemes at different fading states and propose the probabilistic combination of these schemes. Given the fading statistics and a channel state, our policy determines to pause or to deploy a proper transmission scheme such that the video frame transmission is energy minimized, timely, and highly reliable. Extensive simulations confirm that the proposed transmission policies consume less energy than existing methods. In particular, when the deadline or the target reliability are tightened, the proposed policies yield even higher energy efficiency. Seong-Ping Chuah, Yap-Peng Tan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2013 | Image-to-Set Face Recognition Using Locality Repulsion Projections and Sparse Reconstruction-Based Similarity MeasureabstractFor many practical face recognition systems such as law enforcement, e-passport, and ID card identification, there is usually only a single sample per person (SSPP) enrolled in these systems, and many existing face recognition methods may fail to work well because there are not enough samples for discriminative feature extraction in this scenario. However, the probe samples of these face recognition systems are usually captured on the spot, and it is possible to collect multiple face images per person for on-location probing, which is potentially useful to improve the recognition performance. In this paper, we propose a method based on locality repulsion projections (LRP) and a sparse reconstruction-based similarity measure (SRSM) to address the problem of SSPP face recognition using multiple probe images. The LRP method is motivated by our observation that similar face images from different people may lie in a locality in the feature space and cause misclassifications. We design the method with the aim of separating the samples of different classes within a neighborhood through subspace projections for easier classification. To better characterize the similarity between each gallery face and the probe image set, we propose a SRSM method for assigning a label to each probe image set. Experimental results on five widely used face datasets are presented to demonstrate the effectiveness of the proposed approach. Jiwen Lu, Yap-Peng Tan, Gang Wang 0012, Gao Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2013 | Ordinary Preserving Manifold Analysis for Human Age and Head Pose EstimationabstractWe propose in this paper an ordinary preserving manifold analysis approach for human age and head pose estimation. While a large number of manifold learning algorithms have been proposed in the literature and some of them have been successfully applied to age/pose estimation, the ordinary characteristics of the age/pose information of samples have not been fully exploited to learn the low-dimensional discriminative features for these estimation tasks. To address this, we propose an ordinary preserving manifold analysis approach to seek a low-dimensional subspace such that the samples with similar label values (i.e., small age/pose difference) are projected to be as close as possible and those with dissimilar label values (i.e., large age/pose difference) as far as possible, simultaneously. Subsequently, we learn a multiple linear regression model to uncover the relation of these low-dimensional features and the ground-truth values of samples for age/pose estimation. Experimental results on facial age estimation, gait-based human age estimation, and head pose estimation are presented to demonstrate the efficacy of our proposed approach. Jiwen Lu, Yap-Peng Tan |
IEEE Trans. Hum. Mach. Syst. | 2 |
| 2013 | Cost-Sensitive Subspace Analysis and Extensions for Face RecognitionabstractConventional subspace-based face recognition methods seek low-dimensional feature subspaces to achieve high classification accuracy and assume the same loss from different types of misclassification. This assumption, however, may not hold in many practical face recognition systems as different types of misclassification could lead to different losses. Motivated by this concern, this paper proposes a cost-sensitive subspace analysis approach for face recognition. Our approach uses a cost matrix specifying different costs corresponding to different types of misclassifications, into two popular and widely used discriminative subspace analysis methods and devises the cost-sensitive linear discriminant analysis (CSLDA) and cost-sensitive marginal fisher analysis (CSMFA) methods, to achieve a minimum overall recognition loss by performing recognition in these learned low-dimensional subspaces. To better exploit the complementary information from multiple features for improved face recognition, we further propose a multiview cost-sensitive subspace analysis approach by seeking a common feature subspace to fuse multiple face features to improve the recognition performance. Extensive experimental results demonstrate the effectiveness of our proposed methods. Jiwen Lu, Yap-Peng Tan |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2013 | Scalable Resource Allocation for SVC Video Streaming Over Multiuser MIMO-OFDM NetworksabstractIn this paper, we propose a scalable resource allocation framework for streaming scalable videos over multiuser multiple-input multiple-output orthogonal frequency-division multiplexing (MIMO-OFDM) networks. We exploit the utilities of scalable videos produced by the scalable extension of H.264/AVC (SVC) and investigate the multidimensional diversities of the multiuser MIMO-OFDM wireless networks. First, we study the rate-utility relationship of SVC via a packet prioritization scheme. Based on the rate-utility analysis, a scalable resource-allocation framework is proposed to achieve differentiated service objectives for different scalable video layers. To provide users with fair opportunities to acquire basic viewing experience, a fair scheme is designed to guarantee that each user is entitled to a MAXMIN fairness to have their base layer video packets received. After all users have their base layer packets successfully scheduled, resources are distributed to exploit the network efficiency. The two schemes are integrated into a unified bit loading and power allocation solution to enhance the practicability of the scalable framework. Experiment results confirms that the proposed scheme handles fairness and efficiency better at different scenarios than the conventional schemes. Maodong Li 0001, Yap-Peng Tan |
IEEE Trans. Multim. | 3 |
| 2012 | Neighborhood repulsed metric learning for kinship verificationabstractKinship verification from facial images is a challenging problem in computer vision, and there is a very few attempts on tackling this problem in the literature. In this paper, we propose a new neighborhood repulsed metric learning (NRML) method for kinship verification. Motivated by the fact that interclass samples (without kinship relations) with higher similarity usually lie in a neighborhood and are more easily misclassified than those with lower similarity, we aim to learn a distance metric under which the intraclass samples (with kinship relations) are pushed as close as possible and interclass samples lying in a neighborhood are repulsed and pulled as far as possible, simultaneously, such that more discriminative information can be exploited for verification. Moreover, we propose a multiview NRM-L (MNRML) method to seek a common distance metric to make better use of multiple feature descriptors to further improve the verification performance. Experimental results are presented to demonstrate the efficacy of the proposed methods. Jiwen Lu, Junlin Hu 0001, Xiuzhuang Zhou, Yap-Peng Tan, Gang Wang 0012 |
CVPR | 5 |
| 2012 | Energy-minimized wireless video transmissions via sharing of retransmission limitsabstractIn wireless video streaming, a video frame is usually large in data load and is truncated into multiple transport packets for reliable transmissions. These transport packets should be transmitted successfully before a deadline for frame decoding. High bitrate and high power transmission schemes are often deployed to ensure reliable transmissions. Such a method however incurs substantial energy consumption. Unlike existing methods which assign the retransmission limit independently to each transport packet, we share the retransmission opportunity among transport packets of a video frame, and jointly optimize it with transmission rate and power such that the video frame delivery is energy efficient. Contrary to the common intuition that avoids packet loss and retransmission to preserve energy, our approach yields lower energy consumption by judiciously inducing loss and retransmission of the transport packets, without violating the deadline and reliability constraint. Numerical results reveal that the joint optimization improves energy efficiency in wireless video transmissions. The improvement is even more significant when the allocated transmission time is limited or the reliability constraint is tightened. Seong-Ping Chuah, Maodong Li 0001, Yap-Peng Tan |
ICIP | 4 |
| 2012 | Learning modality-invariant features for heterogeneous face recognition
Likun Huang, Jiwen Lu, Yap-Peng Tan |
ICPR | 3 |
| 2012 | Generalized subspace distance for set-to-set image classificationabstractRecent research in visual data classification often involves image sets and the measurement of dissimilarity between each pair of them. An effective solution is to model each image set using a subspace and compute the distance between these two subspaces as the dissimilarity between the sets. Several subspace similarity measures have been proposed in the literature. However, their relationships have not been well explored and most of them do not fully utilize the different importance of individual bases of each subspace. To consolidate this family of subspace-based measures, we propose a generalized subspace distance (GSD) framework and show that most existing subspace similarity measures can be considered as its special cases. To better utilize the different importance, we further propose a new fractional order weighted subspace distance (FOWSD) method within the GSD framework, by assigning different weights to the bases of each subspace and thus characterizing their different importance in similarity measurement. Experimental results on two image classification tasks including face recognition and object recognition are presented to show the effectiveness of the proposed method. Likun Huang, Jiwen Lu, Gao Yang 0001, Yap-Peng Tan |
ISCAS | 4 |
| 2012 | Video organization: Near-Duplicate Video clusteringabstractIt is not uncommon to see several videos of almost identical content on the internet. These near duplicates, coupled with the sheer number of videos, pose a big challenge to the effective organization of video clips online. We propose an adaptive classification approach to detect near-duplicate versions, and an integrated voting strategy to group clusters and to elect a representative for each cluster. Our proposed methods are based on our observation that near-duplicate videos usually span a small, albeit variable area in the feature space, while videos of different contents are scattered far apart. The classification method aims to select a suitable threshold by maximizing the margin for each video sequence in the similarity space, and the voting scheme focuses on merging subsets with mutual information based on neighbor information and inverted indices. Experimental results on an unconstrained web dataset including over 10000 videos demonstrate the efficacy of the proposed methods. Tzu-Yi Hung, Ce Zhu, Gao Yang 0001, Yap-Peng Tan |
ISCAS | 4 |
| 2012 | A scalable resource allocation framework for SVC video transmissions over downlink MIMO-OFDM networksabstractIn this paper, we propose a scalable resource allocation framework for transmission of multiple scalable videos over downlink MIMO-OFDM networks. First, we analyze the rate-utility relationship of scalable video by a packet prioritization scheme, which ensures an optimized video utility under the rate constraint. The scalable framework is then proposed to achieve differentiated service objectives for different video layers. To provide every user with a fair opportunity to receive video for basic viewing, a MAXMIN fairness is designed to have their base layer video packets received. When all users have their base layer packets successfully received, the resources are distributed to exploit network efficiency. Resource allocations are accomplished by the efficient bit loading and power allocation for different optimization goals. The performance of the proposed scheme is validated by comparing with conventional resource allocation schemes. Maodong Li 0001, Yap-Peng Tan |
ISCAS | 3 |
| 2012 | QoE-aware resource allocation for scalable video transmission over multiuser MIMO-OFDM systemsabstractWe investigate in this paper how to maximize multiuser Quality of Experience (QoE) when scalable videos are transmitted over Multiple-Input Multiple-Output (MIMO)-Orthogonal Frequency Division Multiplexing (OFDM) systems. We first study the QoE issues in Scalable Video Coding (SVC) adaptation by constructing a QoE assessment database. We derive the optimal scalability adaptation track for individual video and further summarize common scalability adaptation tracks for grouped videos. A rate-model is developed for SVC adaptation and is employed in designing an efficient resource allocation solution for SVC streaming over multiuser MIMO-OFDM systems. Specifically, time-frequency unit assignment, power allocation, and modulation selection are jointly optimized to maximize users' QoE. Experimental results show that the proposed QoE-aware scalability adaptation scheme significantly outperforms the conventional adaptation schemes, and the proposed QoE-aware resource allocation achieves better QoE performance when compared to existing resource allocation methods. Maodong Li 0001, Yap-Peng Tan |
VCIP | 3 |
| 2012 | QoE analysis for scalable video adaptationabstractQuality of Experience (QoE) serves as a key service goal in video applications. In this paper, we study the QoE issue in scalable video adaptation by constructing a subjective video quality assessment database based on the full scalability of SVC. We derive the optimal scalability adaptation track for individual video and further summarize common scalability adaptation tracks for grouped videos. The common track provides useful guidelines on how to adapt scalable video based on their content characteristics. A rate-QoE model is proposed accordingly for the SVC adaptation. Experimental analyses show that the novel QoE-aware scalability adaptation scheme significantly outperforms the existing ones. Maodong Li 0001, Yap-Peng Tan |
VCIP | 3 |
| 2012 | Cost-Sensitive Semi-Supervised Discriminant Analysis for Face RecognitionabstractThis paper presents a cost-sensitive semi-supervised discriminant analysis method for face recognition. While a number of semi-supervised dimensionality reduction algorithms have been proposed in the literature and successfully applied to face recognition in recent years, most of them aim to seek low-dimensional feature representations to achieve low classification errors and assume the same loss from all misclassifications in the feature representation/extraction phase. In many real-world face recognition applications, however, this assumption may not hold as different misclassifications could lead to different losses. For example, it may cause inconvenience to a gallery person who is misrecognized as an impostor and not allowed to enter the room by a face recognition-based door locker, but it could result in a serious loss or damage if an impostor is misrecognized as a gallery person and allowed to enter the room. Motivated by this concern, we propose in this paper a new method to learn a discriminative feature subspace by making use of both labeled and unlabeled samples and exploring different cost information of all the training samples simultaneously. Experimental results are presented to demonstrate the efficacy of the proposed method. Jiwen Lu, Xiuzhuang Zhou, Yap-Peng Tan, Jie Zhou 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2012 | Energy-Efficient Resource Allocation and Scheduling for Multicast of Scalable Video Over Wireless NetworksabstractIn this paper, we investigate optimal resource allocation and scheduling for scalable video multicast over wireless networks. The wireless video multicasting is a best-effort service which has limited transmission energy and channel access time. To cater for multi-resolution videos to heterogeneous clients and for channel adaptation, we adopt scalable video coding (SVC) with spatial, temporal and quality scalabilities. Our scalable video multicast system consists of a channel probing stage to gather the channel state information and a transmission stage to multicast videos to clients. We formulate the optimal resource allocation problem by maximizing the video quality of the clients subject to transmission energy and channel access constraints. We show that the problem is a joint optimization of the selection of modulation and coding scheme (MCS), and the transmission power allocation. By imposing a quality-of-service (QoS) constraint on the packet loss rate, we simplify the original problem to a binary knapsack problem which can be solved by a dynamic programming approach. Specifically, we first propose a multicast scheduling scheme based on the quality impact of each SVC layer. Guided by the content-aware multicast scheduling, we optimize the resource allocation for each SVC layer sequentially. Solution at each step takes into account of the channel condition, remaining resources, and client requirements. The proposed scheme is of linear complexity and leads to the maximized video quality for the admitted clients, while satisfying the energy budget and channel access constraints. Experiment results demonstrate that our scheme achieves notable video quality improvements for multicast clients, when compared to the state-of-the-art video multicast method. Seong-Ping Chuah, Yap-Peng Tan |
IEEE Trans. Multim. | 3 |
| 2011 | A novel study and analysis on segmental gait sequence recognitionabstractThis paper presents a novel study and analysis on two important problems in gait recognition: one is how to perform gait recognition with only segment of a complete gait cycle for test and the other is how static and dynamic information affect the recognition results in such situation. In conventional gait recognition research, the gait sequence of at least one cycle is usually needed to ensure the completeness of gait dynamics. And in this paper we will show that this is not a necessary condition and considerable recognition results could still be achieved with only part of the cycle based on appropriate similarity measure algorithms. Moreover, we will also show the different roles that static and dynamic information play in such situation through experimental study. To our knowledge, there is little work on these important problems, results from which can make gait recognition more applicable in frequently happened scenarios such as occlusion. Nini Liu, Yap-Peng Tan |
ICASSP | 2 |
| 2011 | Fusing shape and texture information for facial age estimationabstractThis paper presents a new human age estimation method by using multiple feature fusion via facial image analysis. Motivated by the fact that both shape and texture information of facial images can provide complementary information in characterizing human age, we propose fusing these two sources of information at the feature level by using canonical correlation analysis (CCA), a powerful and well-known tool that is well suitable for relating two sets of measurements, for enhanced facial age estimation. Then, we learn a multiple linear regression function to uncover the relation of the fused features and the ground-truth age values for age prediction. Experimental results are presented to demonstrate the efficacy of the pro posed method. Jiwen Lu, Yap-Peng Tan |
ICASSP | 2 |
| 2011 | Discriminative multi-manifold analysis for face recognition from a single training sample per personabstractConventional appearance-based face recognition methods usually assume there are multiple samples per person (MSPP) available during the training phase for discriminative feature extraction. In many practical face recognition applications such as law enhancement, e-passport and ID card identification, this assumption, however, may not hold as there is only a single sample per person (SSPP) enrolled or recorded in these systems. Many popular face recognition methods fail to work well in this scenario because there are not enough samples for discriminant learning. To address this problem, we propose in this paper a novel discriminative multi-manifold analysis (DMMA) method by learning discriminative features from image patches. First, we partition each enrolled image into several non-overlapping patches to form an image set for each sample per person. Then, we formulate the SSPP face recognition as a manifold-manifold matching problem and learn multiple DMMA feature spaces to maximize the manifold margins of different persons. Lastly, we propose a reconstruction-based manifold-manifold distance to identify the unlabeled subjects. Experimental results on three widely used face databases are presented to demonstrate the efficacy of the proposed approach. Jiwen Lu, Yap-Peng Tan, Gang Wang 0012 |
ICCV | 2 |
| 2011 | Combining Feature Context and Spatial Context for Image Pattern DiscoveryabstractOnce an image is decomposed into a number of visual primitives, e.g., local interest points or salient image regions, it is of great interests to discover meaningful visual patterns from them. Conventional clustering (e.g., k-means) of visual primitives, however, usually ignores the spatial dependency among them, thus cannot discover the high-level visual patterns of complex spatial structure. To overcome this problem, we propose to consider both spatial and feature contexts among visual primitives for pattern discovery. By discovering both spatial co-occurrence patterns among visual primitives and feature co-occurrence patterns among different types of features, our method can better handle the ambiguities of visual primitives, by leveraging these co-occurrences. We formulate the problem as a regularized k-means clustering, and propose an iterative bottom-up/top-down self-learning procedure to gradually refine the result until it converges. The experiments of image text on discovery and image region clustering convince that combining spatial and feature contexts can significantly improve the pattern discovery results. Hongxing Wang 0001, Junsong Yuan 0001, Yap-Peng Tan |
ICDM | 3 |
| 2011 | Joint power allocation and bit loading for enhanced SVC video downlink transmissions over SDMA/OFDMA networksabstractWe address in this paper the utility maximization of scalable video transmission under limited radio resource in a multi-user, multi-antenna downlink. The proposed utility maximization framework comprises two parts: a utility analysis of scalable video based on approximated rate-distortion function and a resource allocation scheme by joint power allocation and bit loading. Scalable video packets are entitled with differentiated priority due to the inherent dependency relationship in a layer-based encoding scheme. First, we analyze the rate-utility of scalable video by packet prioritization, which plays a key role in video utility maximization under a fixed rate constraint. Then to cater for diverse rate-utility character and time-varying network condition for multiple scalable video downlink over SDMA/OFDMA system, a joint power allocation and bit loading scheme is proposed to increase the overall video utility under limited radio resource. The resource allocation is accomplished by three important procedures, user grouping, bit loading and power adaptation. The performance of the proposed scheme is validated by comparing with conventional radio resource allocation schemes. Maodong Li 0001, Yap-Peng Tan |
ICME | 3 |
| 2011 | Set-to-set gait recognition across varying views and walking conditionsabstractThis paper examines the multiview gait recognition problem in which human gait sequences are collected from several different views simultaneously. Motivated by the fact that set-based feature representation can handle certain intra-subject variations, we propose a new Multiview Subspace Representation (MSR) method for gait recognition across varying views and walking conditions. It takes samples collected from different views of the same subject as a feature set and uses a subspace to represent such information. Then, the similarity of two subjects is measured by the distance between two subspaces and a simple yet effective Weighted Subspace Distance (WSD) algorithm is applied to calculate the similarity. There are two notable advantages of our proposed method: 1) we need not know the exact view of the test gait sequence in advance, and 2) some extent of intra-subject variations can be effectively handled. Experimental results on two benchmark multi-view gait databases are presented to demonstrate the effectiveness of the proposed method. Nini Liu, Jiwen Lu, Yap-Peng Tan, Maodong Li 0001 |
ICME | 3 |
| 2011 | Locality repulsion projections for image-to-set face recognitionabstractConventional face recognition usually assumes that both the training and test phases employ the same form of data in a face recognition system. In many real world face recognition applications such as e-passport and ID card identification, it is common to have only a single sample per person enrolled or recorded in these systems because it is generally difficult to collect additional samples for training. In the testing phase, however, the probe samples are usually captured on spot and it is possible to collect a number of face samples for recognition. This problem is defined as image-to-set face recognition in this paper, which is essentially different from most existing image-to-image or set-to-set face recognition problems. Specifically, we propose in this paper a new locality repulsion projections (LRP) method to address this problem. Motivated by the fact that interclass face samples with higher similarity usually lie in a locality and are more easily misclassified than those with lower similarity, we aim to learn a mapping to project the original face samples onto a low-dimensional feature subspace such that samples lying in a locality are repulsed and more discriminative information can be exploited for recognition. To better characterize the similarity between a gallery sample and a testing set, we further propose a reconstruction-based point-to-set similarity measure to identify the unlabeled subjects. Experimental results on two widely used face databases are presented to demonstrate the efficacy of the proposed method. Jiwen Lu, Yap-Peng Tan |
ICME | 2 |
| 2011 | Adaptive maximum margin criterion for image classificationabstractWe propose in this paper a novel adaptive maximum margin criterion (AMMC) method for image classification. While a large number of discriminant analysis algorithms have been proposed in recent years, most of them consider an equal importance of each training sample and ignore the different contributions of these samples to learn the discriminative feature subspace for classification. Motivated by the fact that some training samples are more effectual in learning the low-dimensional feature space than other samples, we propose using different weights to characterize the different contributions of the training samples and incorporate such weighting information into the popular maximum margin criterion algorithm to devise the corresponding AMMC for image classification. Moreover, we extend the proposed MMC algorithm to the semi-supervised case, namely, semi-supervised adaptive maximum margin criterion (SAMMC), by making use of both labeled and unlabeled samples to further improve the classification performance. Experimental results are presented to demonstrate the efficacy of the proposed methods. Jiwen Lu, Yap-Peng Tan |
ICME | 2 |
| 2011 | Frame-level quantization control for perceptual quality constrained H.264/AVC video codingabstractAchieving the desired visual quality is an important objective of video compression. In this paper, we present a frame-level quantization control approach for perceptual quality constrained video coding which aims to compress video at a certain perceptual quality level. We develop a general learning- based framework in which different video quality measures can be adopted. We model the rate-quality characteristic of the video using v-support vector regression with a Gaussian radial basis function as the kernel. Such a rate-quality modeling is useful for practical video coding with perceptual quality constraint. Specifically, we show how to minimize the bit rate cost and satisfy the quality constraint by exploiting the relationship between the quantization and advanced video quality measure. The advantage of our proposed approach in practical H.264/AVC video coding is demonstrated through simulations and evidenced by favorable experimental results. Yap-Peng Tan |
ISCAS | 2 |
| 2011 | A MAXMIN resource allocation approach for scalable video delivery over multiuser MIMO-OFDM systemsabstractIn this paper, we propose a novel approach for scalable video delivery over multiuser Multiple Input Multiple Output-Orthogonal Frequency Division Multiplexing (MIMO-OFDM) systems with MAXMIN fairness ensured. Scalable Video Coding (SVC) is efficient for rate adaptation and unequal protection in wireless networks where heterogeneous mobile clients have different hardware specifications and time-varying channel conditions. Using the scalability characteristics of SVC, we propose a general approach to ensure fairness for multiple scalable video downlink from a multiple-antenna base station. The fairness is achieved by packet priority analysis referring to layer dependency from application layer and a MAXMIN resource allocation at the physical layer. Based on packet priorities, time-frequency resource, power and modulation schemes are adaptively selected for transmission. This scheme significantly improves the overall system performance and guarantee fairness among users, as demonstrated by experimental comparisons with conventional radio resource allocation schemes. Maodong Li 0001, Yap-Peng Tan |
ISCAS | 3 |
| 2011 | Blind PSNR estimation using shifted blocks for JPEG imagesabstractWe propose a no-reference method for estimating the peak signal-to-noise ratio (PSNR) of images subject to quantization noise in the block-based discrete cosine transform (bDCT) domain. The proposed method uses bDCT coefficients with shifted block boundary positions from the compressed image to estimate the Laplace distribution parameter of the bDCT coefficients of the original image. The resulting distributions are used to estimate the quantization noise power and the PSNR of the compressed image. Gao Yang 0001, Yap-Peng Tan |
ISCAS | 2 |
| 2011 | Estimating relative objective quality among images compressed from the same originalabstractThis paper presents a new method for estimating the relative compression distortion among images encoded from the same original. Increasingly, many compressed versions of the same image are available over the Internet, but users may not have access to the original to measure their objective quality. We propose to cross-check each image against the quantization constraints of another version and obtain a conservative estimation of the difference in mean squared error (ΔMSE), indicating the relative objective quality between these images. Our proposed method is able to judge the relative quality correctly among compressed images with PSNR difference of as little as 0.5 dB, and it is applicable to images in mixed compression formats, such as JPEG and JPEG2000. Gao Yang 0001, Ci Wang, Yap-Peng Tan |
ISCAS | 3 |
| 2011 | Face recognition using an enhanced age simulation methodabstractWe propose in this paper an enhanced age simulation method for face recognition across age differences. Since the within-class variations caused by age differences are generally much larger than the between-class variations caused by different identities at similar ages, we propose an age simulation method to reduce the facial appearance difference of the same person across age variations. Since it's very difficult to obtain a sequence of facial images of the person at continuous ages, we first implement a filling process to enlarge our limited training database and classify it into different age groups. Then, we represent each age group by a mean feature vector. With these mean feature vectors, typical vectors for different ages are generated. Lastly, we synthesize a new virtual facial image at a target age by combining the training samples and typical feature vectors. Experimental results are presented to demonstrate the effectiveness of the proposed method. Jiwen Lu, Yap-Peng Tan |
VCIP | 3 |
| 2011 | Improved discriminant locality preserving projections for face and palmprint recognition
Jiwen Lu, Yap-Peng Tan |
Neurocomputing | 2 |
| 2011 | Packet scheduling with playout adaptation for scalable video delivery over wireless networks
Tzu-Yi Hung, Yap-Peng Tan |
J. Vis. Commun. Image Represent. | 3 |
| 2011 | Cross-layer optimization for SVC video delivery over the IEEE 802.11e wireless networks
Maodong Li 0001, Yap-Peng Tan |
J. Vis. Commun. Image Represent. | 3 |
| 2011 | Channel Access Allocation for Scalable Video Transmission Over Contention-Based Wireless NetworksabstractContention-based wireless networks make use of random access delay in the medium access control (MAC) layer to share the channel efficiently among users. This channel access scheme poses challenges to video traffic which is delay-sensitive. When channel access is limited, how to optimally allocate channel access and judiciously drop video packets remains an open issue. In this paper, we investigate the optimal channel access allocation and preemptive packet drop strategy for scalable video packets using contention-based MAC. We first establish a relationship to associate the characteristics of video packets, the cumulative distribution function of channel access delays and the expected video quality at the receiver. Based on the relationship, we define a video utility function that explicitly takes into account cumulative distribution of channel access delay over contention-based wireless links, video packet length and truncation, packet loss rate and quality impacts of SVC layers. As the analytical model of the video utility function is rather complex for optimization, we further propose an approximation model for the video utility function to transform the optimization problem into a convex problem which is a solvable by using the dual decomposition method. Simulation results confirm that our proposed method achieves better PSNR performance over the existing approach, especially when the total channel access for the video traffic is limited. Seong-Ping Chuah, Yap-Peng Tan |
IEEE Signal Process. Lett. | 3 |
| 2011 | Joint Subspace Learning for View-Invariant Gait RecognitionabstractWe propose in this paper a novel joint subspace learning (JSL) method for view-invariant gait recognition. Inspired by the finding that if a 3-D object can be well represented by the weighted sum of a sufficiently small number of prototypes in the same view, then the representation coefficients are generally consistent across different views, we propose to use these coefficients as view-invariant features for gait recognition. Firstly, we conduct JSL to obtain the prototypes of different views. Then, we represent each sample in both the gallery set and probe set acquired from different views as a linear combination of these prototypes in the corresponding views, and extract the coefficients for feature representation. Lastly, we perform recognition by using a simple nearest neighbor rule. Experimental results on the widely used CASIA-B gait database demonstrate the effectiveness of the proposed method. Nini Liu, Jiwen Lu, Yap-Peng Tan |
IEEE Signal Process. Lett. | 3 |
| 2011 | Nearest Feature Space Analysis for ClassificationabstractWe propose in this letter two new subspace learning methods, called nearest feature space analysis (NFSA) and discriminant nearest feature space analysis (DNFSA), for pattern classification. While many subspace learning algorithms have been proposed in recent years, most of them apply the conventional nearest neighbor (NN) metric to derive the subspace and may not effectively characterize the geometrical information of the samples, especially when the number of training samples per class is limited. In this paper, we propose using the nearest feature space (NFS) metric to seek a NFSA subspace to improve the discriminating power of the subspace for classification. To further enhance the discriminative power of NFSA, we also propose a new DNFSA method to minimize the within-class feature space (FS) distances and maximize the between-class FS distances simultaneously in the derived subspace. Experimental results on face and facial expression recognition are presented to demonstrate the efficacy of the proposed methods. Jiwen Lu, Yap-Peng Tan |
IEEE Signal Process. Lett. | 2 |
| 2011 | Binocular Just-Noticeable-Difference Model for Stereoscopic ImagesabstractConventional 2-D Just-Noticeable-Difference (JND) models measure the perceptible distortion of visual signal based on monocular vision properties by presenting a single image for both eyes. However, they are not applicable for stereoscopic displays in which a pair of stereoscopic images is presented to a viewer's left and right eyes, respectively. Some unique binocular vision properties, e.g., binocular combination and rivalry, need to be considered in the development of a JND model for stereoscopic images. In this letter, we propose a binocular JND (BJND) model based on psychophysical experiments which are conducted to model the basic binocular vision properties in response to asymmetric noises in a pair of stereoscopic images. The first experiment exploits the joint visibility thresholds according to the luminance masking effect and the binocular combination of noises. The second experiment examines the reduction of visual sensitivity in binocular vision due to the contrast masking effect. Based on these experiments, the developed BJND model measures the perceptible distortion of binocular vision for stereoscopic images. Subjective evaluations on stereoscopic images validate of the proposed BJND model. Yin Zhao, Ce Zhu, Yap-Peng Tan, Lu Yu 0003 |
IEEE Signal Process. Lett. | 4 |
| 2011 | Real-Time, Adaptive, and Locality-Based Graph Partitioning Method for Video Scene ClusteringabstractWe propose in this paper an efficient, adaptive, and locality-based graph partitioning method for video scene clustering. First, a graph partitioning method is proposed to group video shots into scenes, and a peer-group filtering (PGF) scheme is used to identify all the shots similar to each particular shot based on Fisher's discriminant analysis. To work with computable shot similarity measures that have only limited discriminating power, we develop a graph partitioning scheme to cluster the shots by maximizing the likeness of shots within the same cluster and minimizing that between different clusters. Second, considering that video data are normally obtained and viewed sequentially, we propose to perform a locality-based PGF and graph partitioning on video segments with 50 shots, 100 shots, and so on. This proposed locality-based method has the advantage that the number of scene clusters is not required to be known a priori, and it can achieve performance comparable to that processing on the whole video sequence. Experimental results are presented to demonstrate the effectiveness and efficiency of the proposed method. Hong Lu 0001, Yap-Peng Tan, Xiangyang Xue 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2011 | Introduction to the ICME2010 Special IssueabstractThe 15 papers in this special issue are extended versions of papers presented at the 2010 IEEE International Conference on Multimedia and Expo (ICME), held in Singapore on July 19-23, 2010. These papers cover a wide range of topics in multimedia including user interface, content understanding, mobility, 3-D processing, storage, and forensics. Zicheng Liu 0001, Ming-Ting Sun, Chia-Wen Lin, Zhengyou Zhang, Zhu Liu 0001, Homer H. Chen, Yap-Peng Tan, Oscar C. Au |
IEEE Trans. Multim. | 7 |
| 2010 | Cost-sensitive subspace learning for face recognitionabstractConventional subspace learning-based face recognition aims to attain low recognition errors and assumes same loss from all misclassifications. In many real-world face recognition applications, however, this assumption may not hold as different misclassifications could lead to different losses. For example, it may cause inconvenience to a gallery person who is mis-recognized as an impostor and not allowed to enter the room by a face recognition-based door-locker, but it could result in a serious loss or damage if an impostor is mis-recognized as a gallery person and allowed to enter the room. Motivated by this concern, we propose in this paper a cost-sensitive subspace learning approach for face recognition. Our approach incorporates a cost matrix, which specifies the different costs associated with misclassifications of subjects, into three popular subspace learning algorithms and devise the corresponding cost-sensitive methods, namely, cost-sensitive principal component analysis (CSPCA), cost-sensitive linear discriminant analysis (CSLDA), and cost-sensitive locality preserving projections (CSLPP), to achieve a minimum overall recognition loss by performing recognition in the low-dimensional subspaces derived. Experimental results are presented to demonstrate the efficacy of the proposed approach. Jiwen Lu, Yap-Peng Tan |
CVPR | 2 |
| 2010 | Facial part displacement effect on template-based gender and ethnicity classificationabstractVisual information such as gender, age and ethnicity play critical roles in human identification. Most of gender and ethnicity recognition research works use the full face considering equal discriminant capability for different face parts. In this paper, we improve the gender and ethnicity recognition, by employing the optimum decision making rule on the confidence level of automatically separated face regions using the modified Golden ratio mask. Faces are preprocessed with multiple base point photometric normalization to prevent facial parts displacement in the noted mask, due to different facial parts' distances of people. SVM is employed as the classifier on the extracted Gabor features of each patch to get its confidence level. The final classification results are obtained based on the output of each patch decision using the optimum decision making rule. Finally, using the most accurate normalization approach for each patch, we could achieve 94% and 98% for gender and ethnicity respectively on a dataset composed of FERET and PEAL frontal face images. Fahimeh Saei Manesh, Mohammad Ghahramani, Yap-Peng Tan |
ICARCV | 3 |
| 2010 | A comparative study of age-invariant face recognition with different feature representationsabstractAge invariant face recognition is an important yet less investigated problem in the face recognition community. In this paper, we empirically evaluate state-of-the-art facial feature representations for age-invariant face recognition. Three representative features including local binary pattern (LBP), Gabor wavelets and gradient orientation pyramid (GOP) were applied, followed by a principal component analysis (PCA) to reduce the dimensions of the extracted features. Experimental results on the MORPH database, one of the largest publicly available face dataset containing thousands of longitudinal images are presented. Experimental results show that Gabor wavelets feature with five scales and eight orientations is the optimal feature representation method for age-invariant face recognition. Cui Meng, Jiwen Lu, Yap-Peng Tan |
ICARCV | 3 |
| 2010 | Effects of facial alignment for age estimationabstractAge estimation is an important enabling capability for the near future, especially in applications related to Human Computer Interaction. Perhaps due to technical difficulties, age estimation has only recently begun to receive more attention. One of the more important pre-processing steps before age estimation is facial alignment, which spatially transforms a face image to align certain facial features, in order to maximize classification accuracy. However the capability for facial alignment in automated age estimation literature is commonly assumed, and the effects of facial alignment is not addressed, even though it is an important issue especially for live deployment. In this paper, we present, to our knowledge, the results of the first systematic investigation on the effects of facial alignment on age estimation accuracy, and conclude with some directions on further topics of investigation in this area. Hee Lin Wang, Jian-Gang Wang 0001, Weiyun Yau, Xing Lun Chua, Yap-Peng Tan |
ICARCV | 5 |
| 2010 | Joint packet prioritization and QoS mapping for SVC over wlansabstractThe emerging H.264/AVC extension, SVC encoding standard facilitates the truncation of bitstreams at certain points to fit in with wireless network variations. In this paper, we propose a joint packet prioritization and QoS mapping strategy based on priority analysis of SVC packets and the exploration of the service differentiations among IEEE 802.11e EDCA access categories. The proposed scheme enables the interaction among different layers, providing differentiated services for scalable video packets. The cross-layer optimization is performed based on SVC packet information at application layer, differentiated access categories at MAC layer and interface queue (IFQ) control at link layer. Based on these cross-layer information, the proposed joint packet prioritization and QoS mapping strategy minimizes the packet loss impact of visual quality. The proposed approach shows significant performance improvement when compared to conventional schemes for SVC streaming over 802.11e wireless networks. Maodong Li 0001, Yap-Peng Tan |
ICASSP | 3 |
| 2010 | View invariant gait recognitionabstractIn this paper, we attempt to enhance the overall recognition rate for view-invariant gait recognition. We propose a simple but efficient framework for this task with training gait sequences from multiple views. A most important problem in the framework is about the optimal choice for the training views, that is, how many views are enough to ensure a satisfying overall performance and how to combine these views to achieve the optimal performance. To solve this problem, we execute intensive experiments and give reasonable optimal choices based on the experimental results. Besides, the gait feature descriptor and the fusion method we develop for the framework also contribute to the promising results. We propose to use mean of Radon transforms of the silhouettes as the descriptor which is very competent for view-invariant application. Moreover, the combination of class correlation and view correlation is applied to score level fusion of results from different views. The CASIA database B which contains gait data from 11 views distributed uniformly in range of [0°, 180°] is chosen in our experiments. Nini Liu, Yap-Peng Tan |
ICASSP | 2 |
| 2010 | Gait-based human age estimationabstractWe investigate in this paper the problem of estimating human ages from gait signatures. To our knowledge, this problem has not been formally addressed in the literature. Estimating human ages at a distance has a number of potential applications, including visual surveillance and monitoring in such public places as airports, railway stations, shopping malls, and various building entrances. Motivated by the fact that human gait appearances vary between males and females even within the same age group, we learn a multi-label-guided (MLG) subspace to better characterize and correlate the age and gender information of a person for estimating his/her age. As human ages assume only nonnegative values and existing multi-label learning techniques mainly deal with ensembles of different binary classes, we devise an effective label encoding scheme to convert each age value to a binary sequence, making conventional multi-label learning suitable for our task. Our experimental results clearly demonstrate the feasibility of using gait signatures to estimate human age and the efficacy of our proposed method. Jiwen Lu, Yap-Peng Tan |
ICASSP | 2 |
| 2010 | An efficient multicast algorithm for the scalable extension of H.264/AVC over IEEE 802.11 WLANsabstractWe present an efficient resource allocation algorithm for scalable video multicast over wireless links with limited channel access time and transmit energy budget. We first introduce a multicast architecture for the scalable extension of H.264/AVC based video delivery over IEEE 802.11 WLANs. We show that the resource allocation problem is a mixed-integer program for which we propose an efficient solution. Our algorithm explicitly takes into account of relevant factors, such as quality demands of multicast clients, wireless channel conditions, packet length, and different priorities of SVC layers. We maximize the utility of video streaming whilst satisfying channel access and energy constraints. The proposed solution jointly considers PHY mode selection and SNR allocation. Our simulation results demonstrate that the proposed algorithm achieves significant performance improvements for multicast clients in resource-constrained wireless networks. Seong-Ping Chuah, Yap-Peng Tan |
ICIP | 3 |
| 2010 | Cost-sensitive subspace learning for human age estimationabstractThis paper presents a novel cost-sensitive subspace learning approach for human age estimation using face and gait signatures. Motivated by the fact that mis-estimating the age information of a person from a facial image or gait sequence could lead to different errors, we propose in this paper two new cost-sensitive subspace learning methods for human age estimation. Our approach incorporates a cost matrix, which specifies the different error associated with mis-estimating each sample, into two popular subspace learning algorithms and devise the corresponding cost-sensitive methods, namely, cost-sensitive principal component analysis (CSPCA), and cost-sensitive locality preserving projections (CSLPP), to project high-dimensional face and gait samples into the low-dimensional subspaces derived. To uncover the relation of the projected features and the ground-truth age values, we learn a multiple linear regression function with a quadratic model for age estimation. Experimental results on the MORPH face database and the USF gait database are presented to demonstrate the efficacy of our proposed methods. Jiwen Lu, Yap-Peng Tan |
ICIP | 2 |
| 2010 | View recognition of human gait sequences in videosabstractWe investigate in this paper the problem of view recognition of human gait sequences in videos. To our knowledge, the problem has not been formally addressed in the literature. Recognizing the views of human gait sequences has a number of potential applications, including visual surveillance and view-invariant human gait recognition. Motivated by the fact that human gait sequences collected from two views with small differences are more easily mis-recognized than those with large differences, we propose a new adaptive discriminant analysis (ADA) method by imposing large penalties on interclass samples with small differences and small penalties on those samples with large differences simultaneously, such that the discriminating power of the extracted features can be boosted for view recognition. Experimental results are presented to demonstrate the efficacy of the proposed approach. Jiwen Lu, Yap-Peng Tan |
ICIP | 2 |
| 2010 | Enhancing incremental learning/recognition via efficient neighborhood estimationabstractOne of problems with incremental learning approaches is that the quality of learning samples cannot be controlled since they are provided on-line. Current incremental learning algorithms ignore the issue by accepting each sample unconditionally and assimilating it into the existing recognition system. However, improper sample, caused by a partial occlusion, badly illumination condition and unusual pose, deteriorates the overall performance of the recognition system significantly. In order to overcome this issue, in this paper, based on hypothesis and verification paradigm, we propose a criterion for sample selection, and present a novel system for automatically evaluating the qualities of new added samples during the incremental learning procedure. Following this, a subspace is learned by using only the samples that are considered to be good after verification. The high computation cost, incurred by inclusion of sample verification into incremental learning framework, is greatly reduced using relative distance filtering. The experimental results demonstrated the superiority of our approach, in both recognition accuracy and robustness. Likun Huang, Jianchao Yao, Yap-Peng Tan |
ICME | 3 |
| 2010 | Playout adaptation based packet scheduling for scalable video delivery over wireless linksabstractIn this paper, we propose an efficient packet scheduling algorithm based on a novel adaptive playout for scalable video delivery over wireless links. The proposed playout adaptation algorithm monitors the playout buffer status to adjust the playout speed by active and passive ways to prevent playout buffer from underflow and avoid annoying playout interruptions. The active playout adaptation proportionally adjusts the playout speed based on the buffer fullness. The passive playout adaptation is enabled when the network is seriously congested and adjusts the playout speed to the lowest one directly. In addition to the adaptive playout deadline, the packet priority and channel conditions are incorporated in the proposed packet scheduling algorithm. Packets are selected for transmission by maximizing the quality of the received video. The playout-deadline aware packet retransmission is also performed. The simulation results show that the proposed approach can efficiently reduce the playout latency and minimize the video quality degradation. Tzu-Yi Hung, Yap-Peng Tan |
ICME | 3 |
| 2010 | Media-rich interactive mobile learning assistantabstractThe proposed system aims to provide useful tool kits to create more effective and interactive learning environment through the use of mobile devices. This system facilitates a flexible way in delivering the course contents and interaction among the instructor and students. It allows students to take notes or to submit answers and feedbacks during a lecture in media formats customized to the users' needs and preferences. Efficient delivery schemes are proposed for media content delivery to optimize system and network resources. Furthermore, useful statistics together with location-based analysis are provided to aid the instructor in accessing the current level of understanding or attention of the students. Yii Leong Ling, Xuan Jing, Koh Kok Sun, Yap-Peng Tan |
ICME | 5 |
| 2010 | Efficient packet scheduling for scalable video delivery to mobile clientsabstractScalable Video Coding (SVC) provides an efficient solution for video adaptation to satisfy different requirements from heterogeneous mobile clients due to their display sizes and channel conditions. Based on the encoding structure and dependency relationship, SVC packets are entitled with different priorities in presenting quality of video sequences. In this paper, a functional model is derived to calculate packet priority index for multiple scalable video streaming to heterogeneous mobile clients. SVC packet layer ID information is utilized for packet prioritization which provides a fast and efficient implementation in packet scheduling. The accuracy and efficacy of the model are validated by comparisons with the maximum performance achieved from the distortion based prioritization. The generality is considered under different coding structures and the utility is convinced through comparisons with conventional streaming schemes in which different resolution videos are transmitted through IEEE 802.11e wireless networks. Maodong Li 0001, Seong-Ping Chuah, Yap-Peng Tan |
ISCAS | 4 |
| 2010 | Perceptually optimized error resilient transcoding using attention-based intra refreshabstractWhile deployment of wireless channels has become widespread and fast-growing for mobile applications, transmitting data over these existing error-prone networks can be very unreliable and challenging due to time-varying interference and channel errors. Many error-resilient algorithms have been proposed to provide adequate resilient features in order to protect video data from channel errors. However, these algorithms often aim to achieve the optimal decoded video quality in terms of mean square error without any consideration for the visual quality. In this paper, we present a perceptually error-resilient method for video transcoding based on the attention-based intra refresh technique and the characteristics of the human visual system to enhance the perceptual performance of the transcoded video. Specifically, the foveated just noticeable distortion and visual attention models are employed to estimate the perceptual loss impact due to error propagation for allocating intra-refreshed macroblocks in the transcoded video. Experimental results show that the proposed method can achieve a much better performance than the existing methods in terms of both the visual quality and perceptual quality measure. Yap-Peng Tan |
ISCAS | 3 |
| 2010 | Pattern Space Maintenance for Data Updates and Interactive MiningabstractThis article addresses the incremental and decremental maintenance of the frequent pattern space. We conduct an in‐depth investigation on how the frequent pattern space evolves under both incremental and decremental updates. Based on the evolution analysis, a new data structure, Generator‐Enumeration Tree (GE‐tree), is developed to facilitate the maintenance of the frequent pattern space. With the concept of GE‐tree, we propose two novel algorithms, Pattern Space Maintainer+ (PSM+) and Pattern Space Maintainer− (PSM−), for the incremental and decremental maintenance of frequent patterns. Experimental results demonstrate that the proposed algorithms, on average, outperform the representative state‐of‐the‐art methods by an order of magnitude. Mengling Feng, Guozhu Dong, Jinyan Li 0001, Yap-Peng Tan, Limsoon Wong |
Comput. Intell. | 4 |
| 2010 | Uncorrelated discriminant simplex analysis for view-invariant gait signal computing
Jiwen Lu, Yap-Peng Tan |
Pattern Recognit. Lett. | 2 |
| 2010 | Perception-Aware Multiple Scalable Video Streaming Over WLANsabstractIn this letter, we consider how to efficiently transmit multiple video programs over IEEE 802.11e WLANs using scalable video coding technique. Scalable video offers flexibilities and functionalities for video adaptation according to the time-varying wireless channel conditions. We examine the perceptible quality impact of scalable video packets and maximize the minimum perceptible video utility of each scalable video stream. The quality of service is optimized by QoS mapping such that scalable video packets with higher impact on perceptible quality are better protected by the enhanced distributed channel access mechanism (EDCA). Using a MAXMIN strategy, we achieve fair distribution of videos to end users based on the spatial, temporal, and quality scalabilities offered by scalable video coding. Our simulation results show the efficacy and performance of the proposed approach. Maodong Li 0001, Yap-Peng Tan |
IEEE Signal Process. Lett. | 3 |
| 2010 | Uncorrelated Discriminant Nearest Feature Line Analysis for Face RecognitionabstractWe propose in this letter a new subspace learning method, called uncorrelated discriminant nearest feature line analysis (UDNFLA), for face recognition. Motivated by the fact that existing nearest feature line (NFL) can effectively characterize the geometrical information of face samples, and uncorrelated features are desirable for many pattern analysis applications, we propose using the NFL metric to seek a feature subspace such that the within-class feature line (FL) distances are minimized and between-class FL distances are maximized simultaneously in the reduced subspace, and impose an uncorrelated constraint to make the extracted features statistically uncorrelated. Experimental results on two widely used face databases demonstrate the efficacy of the proposed method. Jiwen Lu, Yap-Peng Tan |
IEEE Signal Process. Lett. | 2 |
| 2010 | Frame Rate Up-Conversion Using Trilateral FilteringabstractFrame rate up-conversion (FRUC) can enhance the visual quality of low frame rate video presented on liquid crystal display. To minimize the difference between a reference block and an interpolated block, an effective FRUC algorithm partitions a large block into several sub-blocks of smaller size and estimates their motions. Motion estimation searches for the block which has the minimum difference (cost) with the processed block in terms of some block matching distortion and motion discontinuity. As convexity and convergence of the cost function are not guaranteed, the computational cost for such motion estimation is usually extensive or unpredictable. In our proposed FRUC method, the two predictions of a frame to be interpolated are generated through shifting its nearest neighbor frames in the previous and following directions with the motion vectors estimated between them. The initial interpolated frame and its pixel's reliability are subsequently estimated from these two predictions. We then apply a trilateral filter on the initial prediction to correct the unreliable pixels and to restore the missing pixels. Our proposed method not only reduces the computation for refining motion vectors, but also suppresses the interpolation noises and misregistration errors. We have conducted extensive experiments and the results show that the proposed algorithm outperforms the existing methods with better objective and subjective visual quality, and achieves about 3 dB on-average peak signal-to-noise ratio improvement. Ci Wang, Lei Zhang 0006, Yuwen He, Yap-Peng Tan |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2010 | A doubly weighted approach for appearance-based subspace learning methodsabstractWe propose in this paper a doubly weighted subspace learning approach for face representation and recognition. Motivated by the fact that some face samples and parts are more effectual in characterizing and recognizing faces, we construct two weighting matrices based on pairwise similarity of face samples within a same class and discriminant score of each pixel within a face sample to duly emphasize both the between-sample and within-sample features. We then incorporate these two weighting matrices into three popular subspace learning methods, namely principal component analysis, linear discriminant analysis, and nonnegative matrix factorization, to obtain the discriminative features of faces for recognition. Moreover, the proposed doubly weighted technique can be readily extended to other newly proposed subspace learning algorithms to improve their performance. Experimental results show that the proposed approach can effectively enhance the discriminant power of the extracted face features and outperform existing, nonweighted subspace learning algorithms. The performance gain is even more apparent for cases with imbalanced training samples. Jiwen Lu, Yap-Peng Tan |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2010 | Gait-Based Human Age EstimationabstractWe investigate in this paper the problem of estimating human ages from gait signatures. To our knowledge, this problem has not been formally addressed in the literature. Estimating human ages at a distance has a number of potential applications, including visual surveillance and monitoring in such public places as airports, railway stations, shopping malls, and various building entrances. Motivated by the fact that human gait appearances vary between males and females even within the same age group, we learn a multilabel-guided subspace to better characterize and correlate the age and gender information of a person for estimating human age. As human ages assume only nonnegative values and existing multilabel learning techniques mainly deal with ensembles of binary classes, we devise an effective label encoding scheme to convert each age value to a binary sequence, making conventional multilabel learning suitable for our task. To better characterize human gait appearance and enhance the robustness of the proposed age estimation method, we extract a set of over-complete Gabor features including both Gabor magnitude and Gabor phase information of a gait sequence and perform multiple feature fusion to enhance the age estimation performance. Our experimental results clearly demonstrate the feasibility of using gait signatures to estimate human age and the efficacy of our proposed method. Jiwen Lu, Yap-Peng Tan |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2010 | Regularized Locality Preserving Projections and Its Extensions for Face RecognitionabstractWe propose in this paper a parametric regularized locality preserving projections (LPP) method for face recognition. Our objective is to regulate the LPP space in a parametric manner and extract useful discriminant information from the whole feature space rather than a reduced projection subspace of principal component analysis. This results in better locality preserving power and higher recognition accuracy than the original LPP method. Moreover, the proposed regularization method can easily be extended to other manifold learning algorithms and to effectively address the small sample size problem. Experimental results on two widely used face databases demonstrate the efficacy of the proposed method. Jiwen Lu, Yap-Peng Tan |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2009 | Doubly weighted nonnegative matrix factorization for imbalanced face recognitionabstractWe propose in this paper a novel doubly weighted nonnegative matrix factorization (DWNMF) method for imbalanced face recognition. Motivated by the fact that some face samples and certain parts of each face sample are more useful for recognition, we construct two weighted matrices based on the pairwise similarity of face samples in the same class and the discriminant score of each face pixel. Compared with the existing NMF algorithm, the proposed DWNMF method can more effectively exploit the discriminative and geometrical information of face samples, and it is especially suitable for imbalanced face recognition. Experimental results are presented to demonstrate the efficacy of the proposed method. Jiwen Lu, Yap-Peng Tan |
ICASSP | 2 |
| 2009 | Two-directional two-dimensional discriminant locality preserving projections for image recognitionabstractWe propose in this paper an improved manifold learning method called two-directional two-dimensional discriminant locality preserving projections, (2D)2-DLPP, for efficient image recognition. As the existing method of two-dimensional discriminant locality preserving projections (2D-DLPP) mainly relies upon the local structure information in the rows of images, we first derive an alternative 2D-DLPP algorithm that makes use of the information in the columns. Exploiting the local structure and discriminant information in both the rows and the columns, we develop the (2D)2-DLPP method for efficient image feature extraction and dimensionality reduction. Experimental results on two benchmark image datasets show the effectiveness of the proposed method. Jiwen Lu, Yap-Peng Tan |
ICASSP | 2 |
| 2009 | Enhanced gait recognition based on weighted dynamic featureabstractGait Energy Image (GEI) has been shown to be a robust gait descriptor for gait recognition, and many algorithms based on GEI have been proposed. We propose in this paper an improved algorithm to exploit the discriminative information of GEI in identifying walking people based on gait sequences. Specifically, we first obtain the discriminative power of each pixel in the GEI, referred to as feature weight or feature score, through statistic learning from the whole gallery set. We then generate a binary mask for each frame in a gait sequence according to the intensity value of the GEI to separate the dynamic part from static part of GEI. Combining the feature score and the binary mask, we arrive at a new feature for every GEI for discriminative representation and effective recognition. Experimental results on both NLPR and USF databases show the effectiveness of our proposed algorithm in terms of gait recognition rate. Nini Liu, Jiwen Lu, Yap-Peng Tan |
ICIP | 3 |
| 2009 | On the method of multicopy video enhancement in transform domainabstractIncreasingly, we can obtain more than one compressed copy of the same video content with different levels of visual quality over the Internet. As the original source video is not always available, how to choose or derive a video of the best quality from these copies becomes a challenging and interesting problem. In this paper, we address this new research problem by blindly enhancing the quality of the video reconstructed from multiple compressed copies of the same video content. The aim is to reconstruct a video that achieves better quality than any of the available copies. Specifically, we propose to reconstruct each coefficient of the video in the transform domain by using a narrow quantization constraint set derived from the multiple compressed copies together with the Cauchy distribution model for each AC transform coefficient to minimize the distortion. Analytical and experimental results show the effectiveness of the proposed method. Yap-Peng Tan |
ICIP | 2 |
| 2009 | Throughput Adaptation for Scalable Video Multicast in Wireless NetworksabstractWe present a novel method of cross-layer adaptation for scalable video coding (SVC) extension of H.264/AVC multicast in wireless network. Transmission strategy for the single-hop multicast adapts according to wireless channel condition and SVC layer. Efficient transmit power and unequal FEC bit budget allocations are proposed while existing adaptive technique for throughput optimization is applied. We formulated a convex optimization problem to obtain the optimal time-division scheduler for all multicasts such that the wireless network delivers all SVC layers in noisy wireless channels and severe network capacity constraints. The cross-layer adaptation technique is verified by simulation results. Seong-Ping Chuah, Tianxiao Ye, Yap-Peng Tan, Hock Chuan Chua |
ISCAS | 4 |
| 2009 | Uncorrelated Multilinear Geometry Preserving Projections for Multimodal Biometrics RecognitionabstractWe propose in this paper a novel supervised manifold learning algorithm, called uncorrelated multilinear geometry preserving projections (UMGPP), incorporating both the Fisher criterion and manifold criterion to learn multiple interrelated subspaces in an iterative manner for efficient multimodal biometric recognition. In contrast to the existing GPP algorithm, UMGPP learns multiple feature subspaces directly in higher order tensor space to preserve the structural information of original biometrics datum and obtains an increased number of uncorrelated projection directions, which enable UMGPP to out-perform GPP for multimodal biometrics recognition. Compared with other conventional information fusion-based multimodal recognition methods, UMGPP well exploits the relationship of different modality of the same individual and learns more efficient subspaces for feature extraction. Experimental results are presented to demonstrate the efficacy of the proposed method. Jiwen Lu, Yap-Peng Tan |
ISCAS | 2 |
| 2009 | Frame Rate Up-conversion with Edge-weighted Motion Estimation and Trilateral InterpolationabstractWe propose in this paper a novel video frame-rate up-conversion method using edge-weighted motion estimation and trilateral filtering. First, the method produces two frames of the intermediate frame from its previous and following frames with estimated motion vectors. It then obtains an initial estimate from the two frames and calculates its pixel reliability. Finally, a trilateral filter is applied on the initial estimate to correct the unreliable pixels and missing pixels. Compared with the existing methods, the proposed method not only reduces the computation cost on motion vector estimation, but also suppresses the interpolation noises and motion compensation errors. Experimental results show that the proposed method provides better subjective and objective quality, and obtains up to 4-dB PSNR improvement. Lei Zhang 0006, Ci Wang, Wenjun Zhang 0001, Yap-Peng Tan |
ISCAS | 4 |
| 2009 | Reconstructing Videos From Multiple Compressed CopiesabstractA single source video may be compressed using different encoders with different settings. In the context of online video sharing, many such video copies exist, and an end user may have access to a few of them and would like to reconstruct a video sequence with quality superior to all available copies. In this paper, we propose a scheme to improve the video quality by projecting the reconstructed video onto the quantization constraint sets defined by multiple video copies. Experimental results show that the proposed method is capable of improving video quality both subjectively and objectively. Ci Wang, Gao Yang 0001, Yap-Peng Tan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2008 | Negative Generator Border for Effective Pattern Maintenance
Mengling Feng, Jinyan Li 0001, Limsoon Wong, Yap-Peng Tan |
ADMA | 4 |
| 2008 | Enhanced face recognition using tensor neighborhood preserving discriminant projectionsabstractWe propose in this paper a novel subspace learning method called tensor neighborhood preserving discriminant projections (TNPDP) for face recognition. Compared with the conventional appearance-based face recognitionmethod, the proposed TNPDP does not need to perform image-to-vector conversion and can well preserve the structure of the original image. Different from the existing tensor-based recognition approaches such as tensor subspace analysis (TSA) and discriminant analysis with tensor representation (DATER), TNPDP considers locality and discriminative information simultaneously and can find the optimal tensor subspace that best maintains locality neighborhood manifold and discriminates different classes by maximizing the between-class scatter while minimizing the within-class scatter. Experimental results on two benchmark face databases demonstrate the effectiveness of the proposed method and indicate that TNPDP is better than TSA and DATER, as well as other popular face recognition methods such as principal component analysis (PCA) and linear discrimination analysis (LDA). Jiwen Lu, Yap-Peng Tan |
ICIP | 2 |
| 2008 | Efficient clustering of face sequences with application to character-based movie browsingabstractHuman face has emerged as a useful feature in recent advances in semantic-based video analysis. Due to the large variation of facial poses in videos, conventional face recognition methods for authentication and identification generally experience difficulties in this particular application domain. Based on the recently proposed affinity propagation algorithm, we present in this paper a new method to match and cluster human faces with large pose variations. We also devise a novel approach to exploring the storyline of character(s) of particular interest for efficient movie navigation and browsing. Experimental results are reported to show the efficacy of our method. Ji Tao, Yap-Peng Tan |
ICIP | 2 |
| 2008 | Regularized dequantization for image copies compressed with different quantization parametersabstractA single source image can be compressed into many different copies with different quantization parameters. An end user may have access to a few of such copies (e.g., over the Internet), but would like to obtain one that has superior quality to any available copy. We propose in this paper to improve the reconstruction quality by using regularization and refining the narrowed quantization constraint set (NQCS) from such image copies. Since the size of NQCS is generally large, we also discuss the estimation of initial images for regularization by applying proper probability models. Experimental results show that the proposed scheme can consistently improve the quality of the reconstructed images in both subjective and objective terms. Gao Yang 0001, Ci Wang, Yap-Peng Tan |
ICIP | 3 |
| 2008 | Adaptive downsampling/upsampling for better video compression at low bit rateabstractTo transmit video contents over limited bandwidth network, video bitstreams may need to reduce the bit rate by encoding with coarse quantization parameters at the expense of degrading quality. At low bit rates, better coding quality can be achieved by downsampling the video prior to compression and upsampling later after decompression. In this paper, we present an adaptive downsampling/upsampling video coding scheme in order to achieve better video quality at low bit rates in terms of both measure and visual quality. In particular, appropriate downsampling directions/ratios and quantization step sizes are adaptively decided for encoding different regions of video frame with the consideration of local contents. Experimental results have shown the better performance of the proposed scheme over the regular coding and downsampling-based coding scheme with fixed downscaling ratio. In addition, the proposed scheme significantly raises the critical bit rate below which a downsampling-based coding scheme outperforms the regular coding. Yap-Peng Tan, Weisi Lin |
ISCAS | 2 |
| 2008 | Face clustering in videos using constraint propagationabstractIn this paper, we propose a novel approach to automatic detection and clustering of human faces presented in videos. In each video shot, continuously appearing human faces are firstly associated to form face sequences. Instead of matching the face sequences directly, we partition them into subsequences consisting of similar poses for the ease of comparison. Face subsequences can then be clustered by graph partitioning with the computed affinity matrix. Prior to that, however, a set of constraints need to be formulated so as to incorporate domain knowledge into the graph. Moreover, we propose a constraint propagation algorithm to fully exploit the space-level implications of these constraints. Experimental results demonstrate the effectiveness of our approach in identifying the main cast in movie clips. Ji Tao, Yap-Peng Tan |
ISCAS | 2 |
| 2008 | Numerical error analysis for super-resolution reconstructionabstractSuperresolution reconstruction (SR) is a technique for reconstructing a high resolution image from several low resolution ones. Recently, some researchers study the relationship of SR error and magnification factor, and give out the efficient and forbidden magnification range in their published literature. Based on their work, we describe SR error bound in more explicit way and reveal the essential factor affecting SR performance, if infinite low resolution (LR) observations are given. For SR implementation, smooth regularization serves as an efficient tool for producing stable SR solution, but this operator degrades SR performance and produces overly smooth SR image at certain magnifications. To show its mechanism, we augment SR model by integrating regularization operator into it, and then present the curve of SR error bound vs magnification factors. It is proven that the forbidden magnification range no longer exists when the regularization is involved. Ci Wang, Yap-Peng Tan, Kap Luk Chan |
ISCAS | 2 |
| 2008 | Efficient all-zero block detection algorithm for H.264 integer transformabstractA maximum magnitude probability (MMP) model is introduced in this paper to study contributions of individual DCT coefficient to the all-zero block (AZB) detection. Based on the MMP model, an efficient AZB detection algorithm is proposed to reduce the overall computation complexity. The new method is further combined with a recently published multi-threshold method [6] to achieve an adaptive detection of all-zero DCT sub-blocks. Experiment results show that the proposed algorithm can achieve a remarkable reduction of computational complexity from the original H.264 encoder with minor PSNR loss. Some comparison with similar lossy encoding schemes is also presented. Tianxiao Ye, Yap-Peng Tan, Ping Xue 0001 |
ISCAS | 2 |
| 2007 | Detection and Removal of Rainboweffect ArtifactsabstractDue to the imperfect separation of luminance and chroma signals in receiver's demodulation, composite video signals suffer from rainbow effect artifacts, which present themselves as interlaced color stripes in regions of high luminance frequency and high luminance intensity. In this paper, we derive the formulas governing the rainbow effects. Based on these formulas, we propose a novel method to detect and remove these annoying artifacts. Experimental results on both captured video frames and simulated frames show that our method can remove the rainbow-effect artifacts effectively and improve the image quality notably. Lanlan Chang, Yap-Peng Tan, Hock Chuan Chua |
ICIP (1) | 2 |
| 2007 | An Efficient and Effective Color Filter Array Demosaicking MethodabstractTo reduce the cost and size, most digital still cameras (DSCs) capture only one color value at each pixel, and the results - color filter array samples - are then interpolated by a demosaicking method to construct a full-color image. Many advanced demosaicking methods have been proposed recently. However, the high complexity of these methods could prevent them from being used in DSCs. In this paper we propose an efficient and effective demosaicking method, which substitutes high-frequency component of color values in the spatial rather than frequency domain. We also propose a simple ternary, anisotropic interpolation scheme to obtain an initial full-color image required in the spatial-domain high-frequency substitution. Experimental results show that the proposed method can outperform recent state-of-the-art methods in terms of both PSNR performance and perceptual results, at the same time reducing the computational cost substantially. Naixiang Lian, Yap-Peng Tan |
ICIP (4) | 2 |
| 2007 | Retinal Vessel Detection using Self-Matched FilteringabstractAutomated analysis of retinal images usually requires estimating the positions of blood vessels, which contain important features for image alignment and abnormality detection. Matched filtering can produce the best results but is difficult to implement because the vessel orientations and widths are unknown beforehand. Many researchers use Hessian filtering, which provides an estimate for vessel orientation through the use of three orientation templates. We propose a novel filtering approach, called self-matched filtering, which is based on the 180deg rotated version of the noisy vessel signal in the local neighborhood. We show that even though the proposed filter achieves half the signal-to-noise ratio of a matched filter, it does not require the estimation of the vessel scale and orientation, and can outperform Hessian filtering by up to a factor of two in terms of miss detection error. Naixiang Lian, Vitali Zagorodnov, Yap-Peng Tan |
ICIP (6) | 3 |
| 2007 | Evolution and Maintenance of Frequent Pattern Space When Transactions Are Removed
Mengling Feng, Guozhu Dong, Jinyan Li 0001, Yap-Peng Tan, Limsoon Wong |
PAKDD | 4 |
| 2007 | Adaptive Filtering for Color Filter Array DemosaickingabstractMost digital still cameras acquire imagery with a color filter array (CFA), sampling only one color value for each pixel and interpolating the other two color values afterwards. The interpolation process is commonly known as demosaicking. In general, a good demosaicking method should preserve the high-frequency information of imagery as much as possible, since such information is essential for image visual quality. We discuss in this paper two key observations for preserving high-frequency information in CFA demosaicking: (1) the high frequencies are similar across three color components, and (2) the high frequencies along the horizontal and vertical axes are essential for image quality. Our frequency analysis of CFA samples indicates that filtering a CFA image can better preserve high frequencies than filtering each color component separately. This motivates us to design an efficient filter for estimating the luminance at green pixels of the CFA image and devise an adaptive filtering approach to estimating the luminance at red and blue pixels. Experimental results on simulated CFA images, as well as raw CFA data, verify that the proposed method outperforms the existing state-of-the-art methods both visually and in terms of peak signal-to-noise ratio, at a notably lower computational cost. Naixiang Lian, Lanlan Chang, Yap-Peng Tan, Vitali Zagorodnov |
IEEE Trans. Image Process. | 3 |
| 2006 | People Counting by Video Segmentation and TrackingabstractIn this paper, we present a novel approach to counting number of people that pass the view of an overhead mounted camera. Moving people are first detected as blobs and represented by binary masks, based on which possible multi-person blobs are further segmented into isolated persons according to their areas and locations. Each single person is tracked through consecutive frames using a correlation-based algorithm and a state diagram is proposed to count people entering and leaving the scene. Experimental results show that our approach is able to achieve promising results Hartono Septian, Ji Tao, Yap-Peng Tan |
ICARCV | 3 |
| 2006 | An Efficient Early Termination Algorithm of Intra Prediction for H.264/AVCabstractWe propose in this paper an efficient early termination algorithm of intra prediction for H.264/AVC. It uses the spatial correlation after 16times16 inter prediction to make the judgement on whether to discard intra prediction or not. Experimental results demonstrate that the proposed algorithm can save the encoding time of H.264/AVC (JM98) between 25-45% with negligible degradation in the quality Hong Lu 0001, Xiangyang Xue 0001, Yap-Peng Tan |
ICARCV | 4 |
| 2006 | Image Denoising Using Optimal Color Space ProjectionabstractDenoising of color images can be improved by exploiting strong correlation between high-frequency content of different color components. We show that for typical color images high correlation also means similarity, and propose to exploit this property using an optimal luminance/color-difference space projection. Experimental results confirm that denoising in the proposed color space yields superior performance, both in PSNR and visual quality sense, compared to that of existing solutions. Naixiang Lian, Vitali Zagorodnov, Yap-Peng Tan |
ICASSP (2) | 3 |
| 2006 | Reversing Demosaicking and Compression in Color Filter Array ImagesabstractAn alternative processing chain of digital still camera (DSC) has been proposed by moving the compression process before demosaicking to avoid additional data redundancy introduced by the demosaicking process. Recent empirical studies have shown that the alternative processing chain can actually outperform the conventional one in terms of image quality at low compression ratios. To provide theoretically sound basis for such conclusion, we propose analytical models for the reconstruction errors of the two processing chains. The models developed confirm the results of existing empirical studies and also allows performance predictions for more advanced compression and demosaicking methods, thus providing important cues for development in this area. Naixiang Lian, Lanlan Chang, Vitali Zagorodnov, Yap-Peng Tan |
ICIP | 4 |
| 2006 | Error Inhomogeneity of Wavelet Image CompressionabstractDespite the popularity of wavelet-based image compression, its error inhomogeneity-the error that is different for even and odd pixel locations, has not been previously analyzed and formally addressed. The difference on PSNR performance can be substantial, up to 3.4 dB for some images and compression ratios. In this paper, we show that the error inhomogeneity is caused by asymmetrical filtering of quantization errors after the upsampling step in wavelet synthesis process. We also develop a model that also allows predicting the amount of inhomogeneity for a given wavelet. Furthermore, we show how to redesign wavelet filters to reduce the error inhomogeneity. Naixiang Lian, Vitali Zagorodnov, Yap-Peng Tan |
ICIP | 3 |
| 2006 | Content-based video copy detection with video signatureabstractIn this paper, an efficient content-based video copy detection method is proposed by using video signatures. Consecutive frame features are mapped into hashes with a time constraint. We employ a winnowing scheme to generate video signatures, which are a subset of the hashes of a video, resulting in notably reduced computing complexity. Experimental results on test videos show that the proposed method is promising. Zhenyan Li, Yap-Peng Tan |
ISCAS | 2 |
| 2006 | Video denoising using vector estimation of wavelet coefficientsabstractWavelet-based image denoising can be extended to a video by applying it to each video frame independently. The denoising performance can be improved by exploiting inter-frame correlations, for example, using appropriate temporal filtering. However, fixed temporal filters might not perform sufficiently well due to their inability to cope with the variability of inter-frame correlations across the video. While many adaptive temporal filtering approaches for denoising in spatial domain have been proposed, they do not straightforwardly extend to wavelet-based denoising. We propose a vector extension of popular hidden Markov tree modeling that flexibly exploits the color and frame dependency of wavelet coefficients. Experimental results confirm that the vector estimator of wavelet coefficients yields denoising performance superior to that of existing solutions, both in CPSNR and visual quality sense. Naixiang Lian, Vitali Zagorodnov, Yap-Peng Tan |
ISCAS | 3 |
| 2006 | Quickest change detection for health-care video surveillanceabstractDetecting changes in video scenes is of fundamental importance for various video surveillance tasks. Of particular interest are abnormal changes of foreground human behaviors/activities that could pose damages or dangers to human properties and lives. In this paper, we propose a unified sequential approach to detecting, as soon as possible, human fall incidents for health-care purpose. Specifically, aspect ratio of human body is extracted as the representative feature, based on which an event-inference module parses observed feature sequences for possible falling behavioral signs. Experimental results are reported to show the efficacy of the proposed approach. Ji Tao, M. Turjo, Yap-Peng Tan |
ISCAS | 3 |
| 2006 | A low complexity H.263 to H.264 transcoderabstractA low complex H.263 to H.264 P-frame transcoder is presented in this paper. The new scheme uses an early-stop skip-mode process, a motion vector refinement algorithm and a simplified inter-mode decision module. Experiment results show that the proposed transcoding scheme can achieve a remarkable reduction of computational complexity at the cost of a negligible rate-distortion loss (under 0.6 dB) as compared with the full transcoder. Tianxiao Ye, Yap-Peng Tan, Ping Xue 0001 |
ISCAS | 2 |
| 2006 | Efficient block-matching motion estimation based on Integral frame attributesabstractBlock-based motion estimation is widely used in video compression for exploiting video temporal redundancy. Although effective, the process is arguably the most computationally intensive part of a typical video encoder. To speed up the process, a large number of fast block-matching algorithms (BMAs) have been proposed for motion estimation by limiting the number of search locations or simplifying the measure of match between the two blocks under comparison. In this paper we propose a new BMA that measures the match by using such features as block sum and block variance, which can be easily computed using integral frame attributes. Experimental results show that the proposed BMA can reduce the computational complexity notably and achieve compression performance very close to that of existing BMAs using conventional block-matching measures. Yap-Peng Tan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2006 | Reversing Demosaicking and Compression in Color Filter Array Image Processing: Performance Analysis and ModelingabstractIn the conventional processing chain of single-sensor digital still cameras (DSCs), the images are captured with color filter arrays (CFAs) and the CFA samples are demosaicked into a full color image before compression. To avoid additional data redundancy created by the demosaicking process, an alternative processing chain has been proposed to move the compression process before the demosaicking. Recent empirical studies have shown that the alternative chain can outperform the conventional one in terms of image quality at low compression ratios. To provide a theoretically sound basis for such conclusion, we propose analytical models for the reconstruction errors of the two processing chains. The models developed confirm the results of existing empirical studies and provide better understanding of DSC processing chains. The modeling also allows performance predictions for more advanced compression and demosaicking methods, thus providing important cues for future development in this area. Naixiang Lian, Lanlan Chang, Vitali Zagorodnov, Yap-Peng Tan |
IEEE Trans. Image Process. | 4 |
| 2006 | Edge-preserving image denoising via optimal color space projectionabstractDenoising of color images can be done on each color component independently. Recent work has shown that exploiting strong correlation between high-frequency content of different color components can improve the denoising performance. We show that for typical color images high correlation also means similarity, and propose to exploit this strong intercolor dependency using an optimal luminance/color-difference space projection. Experimental results confirm that performing denoising on the projected color components yields superior denoising performance, both in peak signal-to-noise ratio and visual quality sense, compared to that of existing solutions. We also develop a novel approach to estimate directly from the noisy image data the image and noise statistics, which are required to determine the optimal projection. Naixiang Lian, Vitali Zagorodnov, Yap-Peng Tan |
IEEE Trans. Image Process. | 3 |
| 2005 | Improved color filter array demosaicking by accurate luminance estimationabstractLuminance information plays an important role in dictating the quality of a color image, and some existing color filter array (CFA) demosaicking methods achieve superiority by first reconstructing a satisfactory luminance plane. It is however difficult to accurately estimate the luminance from CFA samples since their color spectra are generally aliased. Extending a state-of-the-art luminance-based demosaicking scheme, in this paper we propose an improved demosaicking method using an efficient filter to estimate the luminance at green pixels and employing an effective edge-adaptive interpolation scheme to obtain the luminance at red and blue pixels. Experimental results demonstrate that not only is the proposed method less complex, it also performs noticeably better, both visually and in terms of peak signal-to-noise ratio, comparing with several recent methods. Naixiang Lian, Lanlan Chang, Yap-Peng Tan |
ICIP (1) | 3 |
| 2005 | Efficient H.263 to H.264/AVC video transcoding using enhanced rate controlabstractA new video coding standard H.264/AVC has been recently developed and standardized, which represents a number of advances in standard video coding technology and is expected to replace the existing standards such as H.263 and MPEG-1/2/4. In this paper, we propose an enhanced rate control method for H.263 to H.264/AVC transcoding. Specifically, we develop a model to approximate the relationship between the total number of coding bits and quantization step sizes between the preceded and transcoded videos for selecting quantization parameters at the sequence and frame levels. In addition, a new frame-layer bit allocation scheme based on the frame complexity obtained from the preceded video is employed in order to achieve more accurate bit rate and constant visual quality. The experimental results show the accuracy of the model and the effectiveness of the proposed method. Yap-Peng Tan |
ICIP (3) | 2 |
| 2005 | Accurate face localization in videos using effective information propagationabstractIn this paper, we present a novel approach to accurate detection and tracking of human faces in videos. The idea is to propagate the information of a group of seed faces detected off-line with high certainties to recover the faces undetected or detected with only low confidence. Specifically, our approach first estimates from the color of the seed faces a person-specific skin color model, and based on which a particle filtering is performed for sequential face tracking. Then, a backward propagation scheme is devised to optimize the overall localization results in terms of smoothness. Experimental results demonstrate the efficacy of our approach. Ji Tao, Yap-Peng Tan |
ICIP (3) | 2 |
| 2005 | Efficient Video Clip Retrieval Using Index StructureabstractRetrieving similar video clips from large video database requires high query efficiency, precision and recall, which remains a challenging problem since the traditional query algorithms are inefficient and time-consuming. In this paper, we adopt the high-dimensional index structure vector-approximation file (VA-file) to organize the video database, and propose a new similarity measure which takes the temporal order among the video representations into account to improve the accuracy of query. Based on the VA-file and similarity measure, a new video clip retrieval algorithm is proposed in our method to achieve high query efficiency by using restricted sliding window to construct candidate video clips. Experimental results show that the proposed video retrieval method is efficient and effective Linjun Yang, Hong Lu 0001, Xiangyang Xue 0001, Yap-Peng Tan |
MMSP | 5 |
| 2005 | Relative risk and odds ratio: a data mining perspectiveabstractWe are often interested to test whether a given cause has a given effect. If we cannot specify the nature of the factors involved, such tests are called model-free studies. There are two major strategies to demonstrate associations between risk factors (ie. patterns) and outcome phenotypes (ie. class labels). The first is that of prospective study designs, and the analysis is based on the concept of "relative risk": What fraction of the exposed (ie. has the pattern) or unexposed (ie. lacks the pattern) individuals have the phenotype (ie. the class label)? The second is that of retrospective designs, and the analysis is based on the concept of "odds ratio": The odds that a case has been exposed to a risk factor is compared to the odds for a case that has not been exposed. The efficient extraction of patterns that have good relative risk and/or odds ratio has not been previously studied in the data mining context. In this paper, we investigate such patterns. We show that this pattern space can be systematically stratified into plateaus of convex spaces based on their support levels. Exploiting convexity, we formulate a number of sound and complete algorithms to extract the most general and the most specific of such patterns at each support level. We compare these algorithms. We further demonstrate that the most efficient among these algorithms is able to mine these sophisticated patterns at a speed comparable to that of mining frequent closed patterns, which are patterns that satisfy considerably simpler conditions. Haiquan Li, Jinyan Li 0001, Limsoon Wong, Mengling Feng, Yap-Peng Tan |
PODS | 5 |
| 2005 | Color image denoising using wavelets and minimum cut analysisabstractWavelet thresholding has proven to be an efficient edge-preserving denoising method for grayscale images, especially when it exploits the interscale correlations of wavelet coefficients. Intrascale correlations can further improve the denoising performance, but the gain for grayscale images is generally small. In this letter, we demonstrate that the gain can become substantial in color image denoising, especially for smooth image color-difference components. We then propose a new denoising method, based on the minimum cut algorithm, to exploit both the interscale and intrascale correlations of wavelet coefficients. The proposed method achieves up to 5-dB gain in peak signal-to-noise ratio for color-difference images and leads to fewer visual color artifacts. Naixiang Lian, Vitali Zagorodnov, Yap-Peng Tan |
IEEE Signal Process. Lett. | 3 |
| 2005 | An effective post-refinement method for shot boundary detectionabstractIn content-based video analysis, shot boundary detection (SBD) is a common first step which segments video data into elementary shots, each comprising a sequence of consecutive frames recording a video event or scene continuous in time and space. Many SBD methods have been proposed in the literature, and experimental results show that the existing methods work reasonably well for abrupt shot boundaries, but less effectively for gradual shot boundaries. In this paper, we propose an effective post-refinement method for identifying actual shot boundaries from the results obtained by existing SBD methods. The proposed method formulates the SBD problem as sequential detection of changes in the underlying feature distributions whose parameters are estimated from existing video shots. Specifically, the proposed post-refinement method enhances the performance of SBD by identifying as many false positives (false detections) and false negatives (miss detections) as possible. Experiments conducted on a large set of test videos, whose initial shot boundaries are obtained by four existing SBD methods, show that the proposed post-refinement method can improve markedly the detection recall and precision and is rather insensitive to the thresholds used by the existing methods in detecting the initial shot boundaries. Hong Lu 0001, Yap-Peng Tan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2004 | Improved shot boundary detection method based on text edgesabstractShot boundary detection is a pre-requisite technique for video indexing and retrieval. To avoid the influence of flashlight on abrupt shot detection, many edge-based techniques are studied thoroughly. However, these techniques are still susceptible to miss and mistake detecting the abrupt changes. Our observation shows that one of the reasons for these errors is the existence of superimposed text which has rich edges and is ever presented in video frames. To provide a solution, we present a novel method that utilizes the edge type, text edge (edge in text area) or non-text-edge (edge in other text area), reducing erroneous detection with the appearance of video text. Compared to other edge-based detection techniques, experimental results show that our proposed method achieves preferable performance. Liuhong Liang, Yang Liu 0246, Xiangyang Xue 0001, Hong Lu 0001, Yap-Peng Tan |
ICARCV | 5 |
| 2004 | Effective video text detection using line featuresabstractText superimposed on video frames provides synoptic or supplemental information on video semantics. In this paper, we propose a novel method to detect superimposed text effectively. First, we detect edges by an improved Canny edge detector. Then, a line-feature vector graph is generated based on the edge map and the stroke information is extracted. Finally text regions are generated and filtered according to line features. Experimental results show that, without much increasing the computational cost, our proposed method could suppress the false alarms notably. Furthermore, our method can be easily customized to applications with different tradeoffs in recall and precision. Yang Liu 0246, Hong Lu 0001, Xiangyang Xue 0001, Yap-Peng Tan |
ICARCV | 4 |
| 2004 | People monitoring using face recognition with observation constraints
Ji Tao, Yap-Peng Tan |
ICIP | 2 |
| 2004 | Combined use of spatial and spectral correlations for enhanced color filter array demosaickingabstractSingle-sensor digital still cameras capture imagery with a color filter array in a way that each sensor pixel samples only one of the primary color values. To render a full-color image from CFA samples, the missing color values need to be interpolated from the neighboring samples. This process is commonly referred to as demosaicking. In this paper, we present two main contributions to demosaicking. First, we stress the importance and necessity of well exploiting both image spatial and spectral correlations, and characterize the demosaicking artifacts due to inadequate use of either correlation. Second, based on the insights gained from our study, we propose several effective schemes to enhance two existing state-of-the-art demosaicking methods. Experimental results demonstrate the improvement of the enhanced methods, both visually and quantitatively. Lanlan Chang, Yap-Peng Tan |
ICME | 2 |
| 2004 | Adaptive binarization method for document image analysisabstractThis paper proposes an adaptive binarization method, based on the criterion of maximizing local contrast, for document image analysis. The proposed method has overcome, to a large extent, the general problems of poor quality document images, such as non-uniform illumination, undesirable shadows and random noise. It was tested against a variety of challenging images, and the experimental results are presented to show the effectiveness and superiority of the proposed method. Mengling Feng, Yap-Peng Tan |
ICME | 2 |
| 2004 | Video segmentation based on sequential change detectionabstractIn content-based video analysis, substantial research efforts have been focused on developing techniques to detect the boundaries between two successive shots, each comprising consecutive frames filmed with a single camera act. However, there is still room for further improvement in the detection performance. With the use of sequential change detection and the help of nonparametric density estimation principles, we propose A new shot boundary detection method that can maintain not only satisfactory detection accuracy, but also consistent detection performance based on the results of various test videos. Zhenyan Li, Hong Lu 0001, Yap-Peng Tan |
ICME | 3 |
| 2004 | On the methods and performances of rational downsizing video transcoding
Yap-Peng Tan, Yongqing Liang 0001, Haiwei Sun |
Signal Process. Image Commun. | 1 |
| 2004 | A probabilistic approach to incorporating domain knowledge for closed-room people monitoring
Ji Tao, Yap-Peng Tan |
Signal Process. Image Commun. | 2 |
| 2004 | Fast block-based motion estimation using integral framesabstractBlock-based motion estimation is widely used in video compression for exploitation of video temporal redundancy. Although effective, the process is arguably the most computationally intensive part of a typical video encoder. To speed up the process, a large number of fast block-matching algorithms (BMAs) have been proposed for motion estimation by limiting the number of search locations or simplifying the measure of match between two blocks under comparison. In this paper we propose a new BMA that measures the match by using the block sum of pixel values, a quantity that can be easily computed from an integral frame. Experimental results show that the proposed BMA can reduce notably the computational load and achieve compression performance very close to that of existing BMAs using conventional block-matching measures. Yap-Peng Tan |
IEEE Signal Process. Lett. | 2 |
| 2004 | A vision-based approach to early detection of drowning incidents in swimming poolsabstractWe present in this paper a vision-based approach to detection of drowning incidents in swimming pools at the earliest possible stage. The proposed approach consists of two main parts: a vision component which can reliably detect and track swimmers in spite of large scene variations of monitored pool areas, and an event-inference module which parses observation sequences of swimmer features for possible drowning behavioral signs. The vision component employs a model-based approach to represent and differentiate the background pool areas and foreground swimmers. The event-inference module is constructed based on a finite state machine, which integrates several reasoning rules formulated from universal motion characteristics of drowning swimmers. Possible drowning incidents are quickly detected using a sequential change detection algorithm. We have applied the proposed approach to a number of video clips of simulated drowning and obtained promising results as reported in this paper. Wenmiao Lu, Yap-Peng Tan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2003 | An effective post-refinement method for shot boundary detectionabstractIn content-based video analysis, shot boundary detection is a common first step to segment video data into fundamental units of shots, each composing consecutive frames filmed with a single camera act. Many methods have been proposed in the literature for detection of shot boundaries. In this paper, we propose a new and effective post-refinement method on the detected shot boundaries by performing sequential detection of abrupt change in two underlying distributions. Experimental results show that the proposed method can eliminate most false detections and also recover many missed detections from the original detected shot boundaries, attaining better detection performance. Hong Lu 0001, Yap-Peng Tan |
ICIP (2) | 2 |
| 2003 | Arbitrary downsizing video transcoding using H.26L standardabstractDownsizing video transcoding is a useful technique for adapting the bit rate or frame size of a precoded video to suit better the constraints of different transmission networks and receiving devices. To avoid frame skipping and maintain good quality at low bit rates, we have proposed H.263-based video transcoding techniques for downszing a precoded video by an arbitrary factor. In this paper, we extend the proposed techniques by using the newly developed H.26L standard in order to fully exploit the advantages of arbitrary downsizing video transcoding. Experimental results show that the proposed method can achieve much better video quality than that using H.263 standard. Haiwei Sun, Yap-Peng Tan |
ICIP (1) | 2 |
| 2003 | Image retrieval with SVM active learning embedding Euclidean searchabstractImage retrieval with relevance feedback suffers from the small sample problem. Recently, SVM active learning has been proposed to tackle this problem, showing promising results. However, a small but sufficient number of initially labelled samples are still required to ensure subsequent efficient active learning and good retrieval performance. In the existing method, the user is asked to label more images before active learning starts. In this paper, a method of embedding Euclidean search into SVM active learning is proposed. With the help of Euclidean search, the adverse effect on retrieval performance due to lack of initially labelled samples can be reduced. Experimental results demonstrate the improvement by the proposed method, especially when the number of initially labelled samples is small. Lei Wang 0001, Kap Luk Chan, Yap-Peng Tan |
ICIP (1) | 3 |
| 2003 | Unsupervised clustering of dominant scenes in sports video
Hong Lu 0001, Yap-Peng Tan |
Pattern Recognit. Lett. | 2 |
| 2003 | Color filter array demosaicking: new method and performance measuresabstractSingle-sensor digital cameras capture imagery by covering the sensor surface with a color filter array (CFA) such that each sensor pixel only samples one of three primary color values. To render a full-color image, an interpolation process, commonly referred to as CFA demosaicking, is required to estimate the other two missing color values at each pixel. In this paper, we present two contributions to the CFA demosaicking: a new and improved CFA demosaicking method for producing high quality color images and new image measures for quantifying the performance of demosaicking methods. The proposed demosaicking method consists of two successive steps: an interpolation step that estimates missing color values by exploiting spatial and spectral correlations among neighboring pixels, and a post-processing step that suppresses noticeable demosaicking artifacts by adaptive median filtering. Moreover, in recognition of the limitations of current image measures, we propose two types of image measures to quantify the performance of different demosaicking methods; the first type evaluates the fidelity of demosaicked images by computing the peak signal-to-noise ratio and CIELAB DeltaE(*)(ab) for edge and smooth regions separately, and the second type accounts for one major demosaicking artifact-zipper effect. We gauge the proposed demosaicking method and image measures using several existing methods as benchmarks, and demonstrate their efficacy using a variety of test images. Wenmiao Lu, Yap-Peng Tan |
IEEE Trans. Image Process. | 2 |
| 2002 | Content-based sports video analysis and modelingabstractWe propose in this paper some new methods for analyzing and modeling sports based on their low-level visual are first automatically identified by using the color features derived from its video are first automatically identified by using the color features derived from its video shots. To improve the performance on the identification of dominant scenes and reduce the dependency on a proper threshold, a comparison on different forms of shot color features, including shot color histograms, their principal components, and subspace linear discriminant representations, is performed. Second, the content compactness and motion attributes of clustered video shots are analyzed to differentiate dominant scene types. A scene transition diagram is then constructed to form a structural descriptor for sports video contents. Third, the video shots belonging to each dominant scene are processed, using customized schemes and domain specific knowledge, to identify interesting play events in the sports video. Experimental results on identification of dominant scenes, structural desprictors and high-level ball videos, are presented to demonstrate the possible applications of the proposed methods. Hong Lu 0001, Yap-Peng Tan |
ICARCV | 2 |
| 2002 | A novel multi-scale spatial-color descriptor for content-based image retrievalabstractWe propose in this paper a new approach to combine color and spatial features for content-based image retrieval applications. The color feature is derived using a fuzzy c-means clustering algorithm in a perceptually uniform color space. To further improve the image discrimination power, the spatial information of each image is measured by a multi-scale color histogram set derived from a set of hierarchical image partitions. The proposed method is capable of representing each image compactly and retrieving similar images effectively. Experimental results suggest that the proposed method can achieve consistently better performance compared to those using the popular conventional color histogram and generalized color histogram. Anning Ouyang, Yap-Peng Tan |
ICARCV | 2 |
| 2002 | A camera-based system for early detection of drowning incidentsabstractWe present in this paper a camera-based system for detecting drowning incidents in a swimming pool at the earliest possible stage. The system consists of two main parts: a vision component which can reliably detect and track swimmers in spite of large scene variations of monitored pool areas, and an event-inference module which parses observation sequences of swimmer features for possible drowning behavioral signs. The vision component employs a model-based approach to represent and differentiate background pool areas and foreground swimmers. The event-inference module is constructed based on a finite state machine, which integrates several reasoning rules formulated from universal motion characteristics of drowning swimmers. Possible drowning incidents are quickly detected using a sequential change detection algorithm. The proposed system has been applied to a number of video clips of simulated drowning, and promising results have been obtained. Wenmiao Lu, Yap-Peng Tan |
ICIP (3) | 2 |
| 2002 | Model-based clustering and analysis of video scenesabstractWe make two contributions. First, we develop an unsupervised method to discover clusters of video scenes and summarize them with a concise Gaussian mixture model. To search for the best possible model, an effective procedure is devised to compare among models with different dimensions (i.e., numbers of mixture components) and, for a given dimension, among models with different parameters. Second, we propose a scene interference measure to characterize the interaction among different scenes of a video sequence. When applied to the clustered video scenes, the measure can reveal the dominant video segments of a class of videos without requiring much domain-specific knowledge. The proposed methods have been tested with a large number of sports videos and promising results are reported. Yap-Peng Tan, Hong Lu 0001 |
ICIP (1) | 1 |
| 2002 | Video transcoding for fast forward/reverse video playbackabstractFast forward and fast reverse playbacks are two common video browsing functions provided in many analog and digital video players. They help users quickly find and access video segments of interest by scanning through the content of a video at a faster than normal playback speed. We propose a video transcoding approach to realizing fast forward and reverse video playbacks by generating a new compressed video from a pre-coded video. To reduce the computational requirements, we design and compare several fast algorithms for estimating the motion vectors required in transcoded video. To accommodate changes due to frame skipping for fast video playback, we also alter the group-of-pictures structure of transcoded video. In addition, subjective tests are conducted to assess the minimum video peak-signal-to-noise-ratio degradation that is perceptible to viewers at different fast playback speeds. To this end, we obtain an adaptive video transcoding method, which combines intra-coding and inter-coding with a fast motion vector reestimation method to strike a good balance between computational complexity and transcoded video quality. Experimental results are reported to show the efficacy of the proposed method. Yap-Peng Tan, Yongqing Liang 0002 |
ICIP (1) | 1 |
| 2002 | On model-based clustering of video scenes using sceneletsabstractWe propose in this paper a model-based approach to clustering video scenes based on scenelets. We define a video scenelet as a short consecutive sample of frames of a video sequence. The approach makes use of an unsupervised method to represent scenelets of a video with a concise Gaussian mixture model and cluster them into different video scenes according to their visual similarities. In particular the expectation-maximization algorithm is employed to estimate the unknown model parameters, and Bayesian information criterion is used to determine the optimal number and model of scene clusters in a principled manner. This approach is fundamentally different from many existing video clustering methods, as it does not require explicit knowledge of shot boundaries. Instead, the shot boundaries can also be obtained as a by-product of the scene clustering process. The proposed methods have been tested with various types of sports videos and promising results are reported in this paper. Hong Lu 0001, Yap-Peng Tan |
ICME (1) | 2 |
| 2002 | On the methods and applications of arbitrarily downsizing video transcodingabstractVideo transcoding is a common technique for adapting the bitrate or spatial/temporal resolution of a compressed video to suit different transmission bandwidths or receiving devices. To reduce the computational complexity, many fast methods have been proposed to estimate the motion vectors required for downsizing a pre-coded video by an integer factor. We develop and compare several fast video transcoding methods for downsizing a pre-coded video by an arbitrary factor Methods which out-perform others under different conditions are identified and discussed. To exploit fully the advantages of arbitrarily downsizing video transcoding, we also design a scheme to determine the reduced frame size that can sustain the best possible video quality for a given target bitrate. Experimental results are presented to show the performance of the proposed video transcoding methods. Yap-Peng Tan, Haiwei Sun, Yongqing Liang 0002 |
ICME (1) | 1 |
| 2002 | Arbitrary downsizing video transcoding using fast motion vector reestimationabstractWe propose a video transcoding method for arbitrarily downsizing a precoded video by reestimating from its original motion vectors the new motion vectors required to code the downsized video. Compared with the existing methods for downsizing a precoded video by an integral factor, the main advantage of the proposed method is that the spatial resolution of the precoded video can be freely adjusted to meet different bandwidth and device requirements. Experimental results show that the proposed method can obtain arbitrarily downsized video with good perceptual quality while reducing the computational complexity of the process. Yongqing Liang 0002, Lap-Pui Chau, Yap-Peng Tan |
IEEE Signal Process. Lett. | 3 |
| 2001 | A new content-based hybrid video transcoding methodabstractMany video transcoding architectures have been proposed to reduce the bitrates of compressed video bitstreams. Requantization of transformed coefficients (such as DCT coefficients), spatial resolution down-sampling, and temporal resolution down-sampling are the three transcoding tools which are commonly used to reduce compressed video bitrates. We propose a new video transcoding method that combines the use of different transcoding tools to transcode compressed video bit streams to various target bitrates while maintaining good subjective/objective video quality. The transcoding tools are selected based on the video content, which is characterized by two video content descriptors inferring the video spatial and motion activities, respectively. These two content descriptors can be readily computed from the compressed video bit streams without full decompression. Experimental results show that the proposed method can provide better video quality than other existing transcoders that make use of a single transcoding tool. Yongqing Liang 0002, Yap-Peng Tan |
ICIP (1) | 2 |
| 2001 | Layering-based color filter array interpolationabstractThis paper presents a new method to interpolate the color filter array (CFA) pattern that is commonly used in a single-sensor digital camera. The proposed method involves two main steps: slicing the color planes into layers to identify smooth local image regions; estimating the missing color values based on the spectral correlation between color planes. The performance of the proposed algorithm is evaluated by comparing the results with those generated by other existing methods. It shows that the proposed method outperforms other methods in terms of both subjective and objective image quality. Wenmiao Lu, Yap-Peng Tan |
ICIP (3) | 2 |
| 2001 | Sports video analysis and structuringabstractWe propose a new method for structuring and analyzing sports video through clustering of video shots based on low-level visual content. The dominant scenes of a sports video are first automatically extracted by using the color features derived from its video shots. The video shots belonging to each dominant scene are then processed to identify interesting play events in the sports video. To reduce the influence of threshold selection on the results, different forms of shot color features, including shot color histogram and its principal component and subspace linear discriminant representations, are also examined. Experimental results on various kinds of sports videos, such as tennis, volleyball, basketball and football videos, are presented to demonstrate the effectiveness of the proposed method. Hong Lu 0001, Yap-Peng Tan |
MMSP | 2 |
| 2000 | Rapid estimation of camera motion from compressed video with application to video annotationabstractAs digital video becomes more pervasive, efficient ways of searching and annotating video according to content will be increasingly important. Such tasks arise, for example, in the management of digital video libraries for content-based retrieval and browsing. We develop tools based on camera motion for analyzing and annotating a class of structured video using the low-level information available directly from MPEG-compressed video. In particular, we show that in certain structured settings, it is possible to obtain reliable estimates of camera motion by directly processing data easily obtained from the MPEG format. Working directly with the compressed video greatly reduces the processing time and enhances storage efficiency. As an illustration of this idea, we have developed a simple basketball annotation system which combines the low-level information extracted from an MPEG stream with the prior knowledge of basketball structure to provide high-level content analysis, annotation, and browsing for events such as wide-angle and close-up views, fast breaks, probable shots at the basket, etc. The methods used in this example should also be useful in the analysis of high-level content of structured video in other domains. Yap-Peng Tan, Drew D. Saur, Sanjeev R. Kulkarni, Peter J. Ramadge |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1999 | A robust sequential approach for the detection of defective pixels in an image sensorabstractLarge image sensors usually contain same defects. Defects are pixels with abnormal photo-responsibility. As a result they often generate outputs different from their adjacent pixel outputs and seriously degrade the visual quality of the captured images. However, it is not economically feasible to produce sensors with no defects for rendering images. A limited number of defects are usually allowed in an image sensor as long as the defective outputs can be corrected with post signal processing techniques. In this paper we present a robust sequential approach for detecting sensor defects from a sequence of images captured by the sensor. With this approach no extra non-volatile memory is required in the sensor device to store the locations of sensor defects. In addition, the detection and correction of image defective outputs can be performed efficiently in a computer host. Experimental results of this approach are reported in the paper. Yap-Peng Tan, Tinku Acharya |
ICASSP | 1 |
| 1999 | A Framework for Measuring Video Similarity and Its Application to Video Query by ExampleabstractThe usefulness of a video database relies on whether the video of interest can be easily located. To allow exploring, browsing, and retrieving videos according to their visual content, efficient techniques for evaluating the visual similarity between different video clips are necessary. We present a framework for measuring video similarity across different resolutions-both spatial and temporal. In particular, the video clips to be compared can be properly aligned through the use of suitable weighting functions and alignment constraints. Dynamic programming techniques are employed to obtain the video similarity measure with a reasonable computational cost. An application to searching MPEG compressed video by example is presented to demonstrate the potential use of the proposed video similarity measure. Yap-Peng Tan, Sanjeev R. Kulkarni, Peter J. Ramadge |
ICIP (2) | 1 |
| 1996 | Video shot classification using human facesabstractPeople usually make up a lot of the information content in videos. The abilities to answer queries and facilitate browsing related to people in videos are crucial. In a single video sequence, a particular person may appear multiple number of times. We propose a scheme to automatically detect the repeated occurrences of the same people to enable fast people related searching. In particular, we propose a video shot classification scheme using human faces, regardless of scale and background. Video shots are classified by clustering facial features extracted from these shots. Potential applications include video indexing and browsing. Employing unsupervised clustering algorithms, this scheme requires no human intervention. Experimental results on a 4-minute news sequence show that it achieves encouraging results. Yin Chan, Shang-Hung Lin, Yap-Peng Tan, Sun-Yuan Kung |
ICIP (3) | 3 |
| 1996 | Extracting good features for motion estimationabstractSelecting image features whose correspondences can be accurately established between images is a key step in many image processing problems, such as camera and object motion estimation, 3D structure reconstruction, and image registration. In this paper, we present a new method of selecting good features for estimating motion from images. Our approach is different from other existing approaches in that we formulate feature tracking as a signal parameter estimation problem, give a quantitative measure of feature quality in terms of how accurately the feature can be tracked, and can adaptively select features with different shapes and sizes which depend on the local variations of the images. Through the analysis of this feature quality measure, we can characterize the basic properties that allow a feature to be well tracked. Some experimental results are shown to demonstrate the advantages and robustness of the proposed method. Yap-Peng Tan, Sanjeev R. Kulkarni, Peter J. Ramadge |
ICIP (1) | 1 |
| 1995 | A new method for camera motion parameter estimationabstractWe derive a six parameter system to estimate and compensate the effects of camera motion-zoom, pan, tilt and swing. As compared to other existing methods, this model describes more precisely the effect of different kinds of camera motions. A recursive least-squares estimator has been used to solve for the motion parameters. Experiments suggest that our algorithm converges to satisfactory results when about 10 pairs of corresponding pairs between two image frames are available. Yap-Peng Tan, Sanjeev R. Kulkarni, Peter J. Ramadge |
ICIP | 1 |