Pong C. Yuen

dblp:y/PongChiYuen · also Pong Chi Yuen · DBLP profile ↗
← Back
195ranked-venue papers
10as first author
52since 2021 · last 2026
0000-0002-9343-2202ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 121 · 8 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 89 · 2 first-author · 25 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 14 since 2021Security and privacy · 14 · 5 since 2021Human-computer interaction and ubiquitous computing · 14 · 1 first-authorDatabases, data management, data science and information retrieval · 6 · 3 since 2021Computer networks · 3 · 3 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Federated Single-Positive Multi-Label Learning
abstract
Single-positive multi-label learning (SPMLL) aims to train a multi-label classifier from data with single-positive label, to predict all applicable labels during testing. However, existing SPMLL methods are tailored for centralized datasets, which fail to be directly deployed to distributed setting like federated learning. In this paper, we start the first attempt to study federated single-positive multi-label learning (FedSPMLL), aiming to collaboratively train a SPMLL model from distributed data. To achieve this, we need to address challenges caused by label incompleteness: limited generalization ability of local model and overweighting contribution of client with local dataset suffering from severe label incompleteness. To this end, we propose a novelFedLOGmethod, guidingFedSPMLL with predicateLOGic-modeled label correlation. Enabling the informative knowledge extraction from limited data, we propose to model label correlation within local dataset using predicate logic. To alleviate false negative label issue, we propose to transfer confident label correlation knowledge to local model by self-distillation. To downweight the contribution of unreliable client owning dataset with severe label incompleteness, we propose a new measurement of label incompleteness to adjust client contribution for a fair aggregation. We establish a comprehensive FedSPMLL benchmark. And extensive experiments demonstrate the superiority of our FedLOG method.
Mang Ye, Andy Jinhua Ma, Pong C. Yuen
IEEE Trans. Inf. Forensics Secur.4
2026 Laboratory Test-Guided Medical Image Generation for Multi-Modal Disease Prediction
abstract
The integration of laboratory tests and medical images is crucial in making accurate disease prediction. However, imaging data exhibits temporal sparsity, compared to frequently collected laboratory tests. This temporal sparsity limits effective multi-modal interaction, which in turn degrades the prediction accuracy. We address this issue by generating additional medical images at more time points, conditioned on the laboratory tests. Inspired by the pivotal role of organs in mediating laboratory tests and imaging abnormalities, we propose an Organ-Centric Modal-Shared Image Generator. It converts laboratory tests into imaging abnormalities through two key components: 1) Organ-Centric Graph: It positions organs as central nodes connecting laboratory tests and imaging abnormalities; and 2) Knowledge-Guided Modal-Shared Trajectory Module: It binds multi-modal features across time into a unified organ state trajectory. Experimental results demonstrate that our method improves multi-modal prediction performance across various diseases. Code is available at https://github.com/LyapunovStability/Lab_Guide_Med_Image _Gen.
Jingwen Xu 0002, Fei Lyu 0004, Pong C. Yuen
IEEE Trans. Medical Imaging4
2025 Beyond Myopia: Enhancing Few-Shot Open-Set Recognition via Hyperopia Distillation
abstract
Existing few-shot open-set recognition (FSOR) methods primarily employ the meta-learning mechanism, in which each meta-task randomly selects a small subset of base classes as knowns, and samples an equal number of classes from the remaining base classes as pseudo-unknowns. While effective, these methods potentially face two critical weaknesses: i) Class-identity overlapping: The same classes are designated as knowns in one meta-task but may be considered as pseudo-unknowns in another, leading to conflicts across meta-tasks and consequently degrading the model’s performance; ii) Narrow pseudo-unknown utilization: Each meta-task selects only a limited number of base classes as pseudo-unknowns rather than providing a broader view of more available base classes. Fundamentally, these issues arise from the myopia of the meta-learners in existing methods, as they lack a broad view on all available base classes. To this end, we strategically propose a novel Hyperopia Distillation Enhancement framework (HDE) for FSOR, which encourages the meta-learner to observe a broader view of pseudo-unknown classes without too worrying about the class-imbalance issue, while effectively mitigating the class-identity overlapping problem. The key to HDE lies in its dual hyperopia distillation mechanism, which enhances the meta-learner by hyperopically distilling available full-view inter-class relationships and more pseudo-unknown knowledge. Extensive experiments verify the effectiveness of our HDE.
Chuanxing Geng, Xiangshu Ding, Songcan Chen, Pong C. Yuen
ECAI4
2025 LoD: Loss-difference OOD Detection by Intentionally Label-Noisifying Unlabeled Wild Data
abstract
Using unlabeled wild data containing both in-distribution (ID) and out-of-distribution (OOD) data to improve the safety and reliability of models has recently received increasing attention. Existing methods either design customized losses for labeled ID and unlabeled wild data then perform joint optimization, or first filter out OOD data from the latter then learn an OOD detector. While achieving varying degrees of success, two potential issues remain: (i) Labeled ID data typically dominates the learning of models, inevitably making models tend to fit OOD data as IDs; (ii) The selection of thresholds for identifying OOD data in unlabeled wild data usually faces dilemma due to the unavailability of pure OOD samples. To address these issues, we propose a novel loss-difference OOD detection framework (LoD) by intentionally label-noisifying unlabeled wild data. Such operations not only enable labeled ID data and OOD data in unlabeled wild data to jointly dominate the models' learning but also ensure the distinguishability of the losses between ID and OOD samples in unlabeled wild data, allowing the classic clustering technique (e.g., K-means) to filter these OOD samples without requiring thresholds any longer. We also provide theoretical foundation for LoD's viability, and extensive experiments verify its superiority.
Chuanxing Geng, Xinrui Wang 0003, Dong Liang 0008, Songcan Chen, Pong C. Yuen
IJCAI6
2025 Test-Time Training with Diversified Local Aggregation Consistency for Mortality Prediction using Clinical Time Series
abstract
Mortality prediction is necessary for patients in the Intensive Care Unit (ICU). Clinical time series provide essential insights for making accurate predictions. However, existing prediction models often struggle with domain shifts when applied across different domains. Privacy concern hinders model calibration due to the forbidden data sharing across domains. Test-Time Training (TTT) has been increasingly researched to tackle the above issues by updating a source model to each single target sample before inference. While massive vision-based TTT methods are proposed, deploying TTT in clinical time series still faces the unique challenge of temporal imbalance: Time points tend to cluster around specific periods in some patients. Neglecting the temporal imbalance in TTT can make the model biased toward dense local pattern, resulting in unsatisfactory prediction. To overcome this challenge, we propose a novel Test-Time Training method with Diversified Local Aggregation Consistency (DLAC-TTT). During the test-time update, DLAC-TTT focuses on the distinct temporal distributions within each patient, enforcing their local diversity and global consistency through aggregation. In this way, it can mitigate the over-reliance on specific local patterns and well integrate diverse local patterns for global learning. Extensive experiments show that DLAC-TTT can boost the generalization performance across real-world clinical datasets from different medical institutes.
Jingwen Xu 0002, Fei Lyu 0004, Pong C. Yuen
KDD (2)3
2025 Camera-Based Bi-Modal PPG-SCG: Sleep Privacy-Protected Contactless Vital Signs Monitoring
abstract
The monitoring of respiratory rate (RR), heart rate (HR), HR variability (HRV), and blood pressure (BP) during sleep allows for a comprehensive evaluation of sleep quality, facilitating the understanding and improvement of a person’s sleep health. Contactless physiological monitoring using cameras has gained popularity recently due to its convenient, infection-free, continuous, and versatile nature. However, the privacy concerns limit the application of camera-based solutions in sleep monitoring setups. This study proposes a novel hybrid setup that integrates camera-based seismocardiography (CamSCG) and photoplethysmography (CamPPG) for contactless measurement of RR, HR, and HRV during sleep while simultaneously estimating BP. For the proximal SCG, we employed camera-based laser speckle vibrometry to measure cardiac motions from the chest, and benchmarked it with a millimeter-wave radar (RFSCG). For the distal photoplethysmographic (PPG), a defocused camera was utilized to measure pulse signals from the facial skin while protecting privacy. In this setup, we analyzed the single-modality in measuring RR, HR, and HRV, and established two bi-modalities (CamSCG-CamPPG and RFSCG-CamPPG) to measure pulse transit time (PTT) features for BP calibration. The benchmark involving 19 subjects highlights the potential of camera-based bi-modal SCG-PPG for privacy-protected vital signs monitoring during sleep.
Yingen Zhu, Yao Ge 0002, Dongmin Huang, Pong C. Yuen, Fu Xiao 0001, Wenjin Wang 0002
IEEE Internet Things J.6
2025 Build Yourself Before Collaboration: Vertical Federated Learning With Limited Aligned Samples
abstract
Vertical Federated Learning (VFL) has emerged as a crucial privacy-preserving learning paradigm that involves training models using distributed features from shared samples. However, the performance of VFL can be hindered when the number of shared or aligned samples is limited, a common issue in mobile environments where user data are diverse and unaligned across multiple devices. Existing approaches use feature generation and pseudo-label estimation for unaligned samples to address this issue, unavoidably introducing noise during the generation process. In this work, we propose Local Enhanced Effective Vertical Federated Learning (LEEF-VFL), which fully utilizes unaligned samples in the local learning before collaboration. Unlike previous methods that overlook private labels owned by each client, we leverage these private labels to learn from all local samples, constructing robust local models to serve as solid foundations for collaborative learning. Additionally, we reveal that the limited number of aligned samples introduces distribution bias from global data distribution. In this case, we propose to minimize the distribution discrepancies between the aligned samples and the global data distribution to enhance collaboration. Extensive experiments demonstrate the effectiveness of LEEF-VFL in addressing the challenges of limited aligned samples, making it suitable for VFL in mobile computing environments.
Wei Shen 0006, Mang Ye, Wei Yu 0009, Pong C. Yuen
IEEE Trans. Mob. Comput.4
2025 Polarity Prompting Vision Foundation Models for Pathology Image Analysis
abstract
The sharp rise in non-alcoholic fatty liver disease (NAFLD) cases has become a major health concern in recent years. Accurately identifying tissue alteration regions is crucial for NAFLD diagnosis but challenging with small-scale pathology datasets. Recently, prompt tuning has emerged as an effective strategy for adapting vision models to small-scale data analysis. However, current prompting techniques, designed primarily for general image classification, use generic cues that are inadequate when dealing with the intricacies of pathological tissue analysis. To solve this problem, we introduce Quantitative Attribute-based Polarity Visual Prompting (Q-PoVP), a new prompting method for pathology image analysis. Q-PoVP introduces two types of measurable attributes: K-function-based spatial attributes and histogram-based morphological attributes. Both help to measure tissue conditions quantitatively. We develop a quantitative attribute-based polarity visual prompt generator that converts quantitative visual attributes into positive and negative visual prompts, facilitating a more comprehensive and nuanced interpretation of pathological images. To enhance feature discrimination, we introduce a novel orthogonal-based polarity visual prompt tuning technique that disentangles and amplifies positive visual attributes while suppressing negative ones. We extensively tested our method on three different tasks. Our task-specific prompting demonstrates superior performance in both diagnostic accuracy and interpretability compared to existing methods. This dual advantage makes it particularly valuable for clinical settings, where healthcare providers require not only reliable results but also transparent reasoning to support informed patient care decisions. Code is available at https://github.com/7LFB/Q-PoVP.
Chong Yin, Si-Qi Liu 0003, Kaiyang Zhou, Vincent Wai-Sun Wong, Pong C. Yuen
IEEE Trans. Medical Imaging5
2025 Allosteric Feature Collaboration for Model-Heterogeneous Federated Learning
abstract
Although federated learning (FL) has achieved outstanding results in privacy-preserved distributed learning, the setting of model homogeneity among clients restricts its wide application in practice. This article investigates a more general case, namely, model-heterogeneous FL (M-hete FL), where client models are independently designed and can be structurally heterogeneous. M-hete FL faces new challenges in collaborative learning because the parameters of heterogeneous models could not be directly aggregated. In this article, we propose a novel allosteric feature collaboration (AlFeCo) method, which interchanges knowledge across clients and collaboratively updates heterogeneous models on the server. Specifically, an allosteric feature generator is developed to reveal task-relevant information from multiple client models. The revealed information is stored in the client-shared and client-specific codes. We exchange client-specific codes across clients to facilitate knowledge interchange and generate allosteric features that are dimensionally variable for model updates. To promote information communication between different clients, a dual-path (model-model and model-prediction) communication mechanism is designed to supervise the collaborative model updates using the allosteric features. Client models are fully communicated through the knowledge interchange between models and between models and predictions. We further provide theoretical evidence and convergence analysis to support the effectiveness of AlFeCo in M-hete FL. The experimental results show that the proposed AlFeCo method not only performs well on classical FL benchmarks but also is effective in model-heterogeneous federated antispoofing. Our codes are publicly available at https://github.com/ybaoyao/AlFeCo.
Baoyao Yang, Pong C. Yuen, Yiqun Zhang 0006, An Zeng
IEEE Trans. Neural Networks Learn. Syst.2
2024 XFibrosis: Explicit Vessel-Fiber Modeling for Fibrosis Staging from Liver Pathology Images
abstract
The increasing prevalence of non-alcoholic fatty liver disease (NAFLD) has caused public concern in recent years. The high prevalence and risk of severe complications make monitoring NAFLD progression a public health priority. Fibrosis staging from liver biopsy images plays a key role in demonstrating the histological progression of NAFLD. Fibrosis mainly involves the deposition of fibers around vessels. Current deep learning-based fi-brosis staging methods learn spatial relationships between tissue patches but do not explicitly consider the relation-ships between vessels and fibers, leading to limited performance and poor interpretability. In this paper, we propose an eXplicit vessel-fiber modeling method for Fibrosis staging from liver biopsy images, namely XFibrosis. Specifically, we transform vessels and fibers into graph-structured representations, where their micro-structures are depicted by vessel-induced primal graphs andfiber-induced dual graphs, respectively. Moreover, the fiber-induced dual graphs also represent the connectivity information between vessels caused by fiber deposition. A primal-dual graph convolution module is designed to facilitate the learning of spatial relationships between vessels and fibers, allowing for the joint exploration and interaction of their micro-structures. Experiments conducted on two datasets have shown that explicitly modeling the relationship between vessels and fibers leads to improved fibrosis staging and en-hanced interpretability.
Chong Yin, Si-Qi Liu 0003, Fei Lyu 0004, Sune Darkner, Vincent Wai-Sun Wong, Pong C. Yuen
CVPR7
2024 Prompting Vision Foundation Models for Pathology Image Analysis
abstract
The rapid increase in cases of non-alcoholic fatty liver disease (NAFLD) in recent years has raised significant public concern. Accurately identifying tissue alteration regions is crucial for the diagnosis of NAFLD, but this task presents challenges in pathology image analysis, particularly with small-scale datasets. Recently, the paradigm shift from full fine-tuning to prompting in adapting vision foundation models has offered a new perspective for small-scale data analysis. However, existing prompting methods based on task-agnostic prompts are mainly developed for generic image recognition, which fall short in providing instructive cues for complex pathology images. In this paper, we propose Quantitative Attribute-based Prompting (QAP), a novel prompting method specifically for liver pathology image analysis. QAP is based on two quantitative attributes, namely K-function-based spatial attributes and histogram-based morphological attributes, which are aimed for quantitative assessment of tissue states. Moreover, a conditional prompt generator is designed to turn these instance-specific attributes into visual prompts. Extensive experiments on three diverse tasks demonstrate that our task-specific prompting method achieves better diagnostic performance as well as better interpretability. Code is available at https://github.com/7LFBIQAP.
Chong Yin, Si-Qi Liu 0003, Kaiyang Zhou, Vincent Wai-Sun Wong, Pong C. Yuen
CVPR5
2024 Bottom-Up Domain Prompt Tuning for Generalized Face Anti-spoofing
Si-Qi Liu 0003, Pong C. Yuen
ECCV (70)3
2024 Multi-scale Value-Density Transformer with Medical Semantic Guidance for Disease Risk Prediction Based on Clinical Time Series
Jingwen Xu 0002, Xiaoge Wei, Pong C. Yuen
ICPR (23)3
2024 Dynamic against Dynamic: An Open-Set Self-Learning Framework
Chuanxing Geng, Pong C. Yuen, Songcan Chen
IJCAI3
2024 Superpixel-Guided Segment Anything Model for Liver Tumor Segmentation with Couinaud Segment Prompt
Fei Lyu 0004, Jingwen Xu 0002, Grace Lai-Hung Wong, Pong C. Yuen
MICCAI (8)5
2024 Temporal Neighboring Multi-modal Transformer with Missingness-Aware Prompt for Hepatocellular Carcinoma Prediction
Jingwen Xu 0002, Fei Lyu 0004, Grace Lai-Hung Wong, Pong C. Yuen
MICCAI (1)5
2024 HistoSyn: Histomorphology-Focused Pathology Image Synthesis
Chong Yin, Si-Qi Liu 0003, Vincent Wai-Sun Wong, Pong C. Yuen
MICCAI (4)4
2024 Symptom Disentanglement in Chest X-Ray Images for Fine-Grained Progression Learning
Jingwen Xu 0002, Fei Lyu 0004, Pong C. Yuen
MICCAI (1)4
2024 BindingSiteDTI: differential-scale binding site modelling for drug-target interaction prediction
abstract
MOTIVATION: Enhanced by contemporary computational advances, the prediction of drug-target interactions (DTIs) has become crucial in developing de novo and effective drugs. Existing deep learning approaches to DTI prediction are frequently beleaguered by a tendency to overfit specific molecular representations, which significantly impedes their predictive reliability and utility in novel drug discovery contexts. Furthermore, existing DTI networks often disregard the molecular size variance between macro molecules (targets) and micro molecules (drugs) by treating them at an equivalent scale that undermines the accurate elucidation of their interaction. RESULTS: We propose a novel DTI network with a differential-scale scheme to model the binding site for enhancing DTI prediction, which is named as BindingSiteDTI. It explicitly extracts multiscale substructures from targets with different scales of molecular size and fixed-scale substructures from drugs, facilitating the identification of structurally similar substructural tokens, and models the concealed relationships at the substructural level to construct interaction feature. Experiments conducted on popular benchmarks, including DUD-E, human, and BindingDB, shown that BindingSiteDTI contains significant improvements compared with recent DTI prediction methods. AVAILABILITY AND IMPLEMENTATION: The source code of BindingSiteDTI can be accessed at https://github.com/MagicPF/BindingSiteDTI.
Chong Yin, Si-Qi Liu 0003, Zhaoxiang Bian, Pong C. Yuen
Bioinform.6
2024 Special section: Best papers of the international conference on pattern recognition and artificial intelligence (ICPRAI) 2022
Mounim A. El-Yacoubi, Umapada Pal 0001, Eric Granger, Pong C. Yuen
Pattern Recognit. Lett.4
2024 Robust Remote Photoplethysmography Estimation With Environmental Noise Disentanglement
abstract
Remote Photoplethysmography (rPPG) has been attracting increasing attention due to its potential in a wide range of application scenarios such as physical training, clinical monitoring, and face anti-spoofing. On top of conventional solutions, deep-learning approach starts to dominate in rPPG estimation and achieves top-level performance. However, most of them try to integrate preprocessing steps such as the ROI selection into an end-to-end network, which may diverge the attention and also limit the generalization in other scenarios with different input skin regions. In this work, we focus on learning the intrinsic rPPG feature and design a lightweight but effective rPPG estimation network based on spatiotemporal convolution. To further improve the robustness, on top of the basic design we propose the Noise-Disentangled DeeprPPG (ND-DeeprPPG) by disentangling the environmental noise from the raw rPPG feature with an adversarial canonical correlation analysis learning strategy. Background regions are employed as a reference to guide the noise disentangling in a self-supervised manner. Extensive experiments show that our ND-DeeprPPG not only outperforms the state-of-the-arts on heart rate estimation but also exhibits promising robustness in cross-skin-region, cross-dataset scenarios and other rPPG-based tasks.
Si-Qi Liu 0003, Pong C. Yuen
IEEE Trans. Image Process.2
2024 Local Style Transfer via Latent Space Manipulation for Cross-Disease Lesion Segmentation
abstract
Automatic lesion segmentation is important for assisting doctors in the diagnostic process. Recent deep learning approaches heavily rely on large-scale datasets, which are difficult to obtain in many clinical applications. Leveraging external labelled datasets is an effective solution to tackle the problem of insufficient training data. In this paper, we propose a new framework, namely LatenTrans, to utilize existing datasets for boosting the performance of lesion segmentation in extremely low data regimes. LatenTrans translates non-target lesions into target-like lesions and expands the training dataset with target-like data for better performance. Images are first projected to the latent space via aligned style-based generative models, and rich lesion semantics are encoded using the latent codes. A novel consistency-aware latent code manipulation module is proposed to enable high-quality local style transfer from non-target lesions to target-like lesions while preserving other parts. Moreover, we propose a new metric, Normalized Latent Distance, to solve the question of how to select an adequate one from various existing datasets for knowledge transfer. Extensive experiments are conducted on segmenting lung and brain lesions, and the experimental results demonstrate that our proposed LatenTrans is superior to existing methods for cross-disease lesion segmentation.
Fei Lyu 0004, Mang Ye, Terry Cheuk-Fung Yip, Grace Lai-Hung Wong, Pong C. Yuen
IEEE J. Biomed. Health Informatics5
2024 Federated Generalized Face Presentation Attack Detection
abstract
Face presentation attack detection (fPAD) plays a critical role in the modern face recognition pipeline. An fPAD model with good generalization can be obtained when it is trained with face images from different input distributions and different types of spoof attacks. In reality, training data (both real face images and spoof images) are not directly shared between data owners due to legal and privacy issues. In this article, with the motivation of circumventing this challenge, we propose a federated face presentation attack detection (FedPAD) framework that simultaneously takes advantage of rich fPAD information available at different data owners while preserving data privacy. In the proposed framework, each data owner (referred to as data centers) locally trains its own fPAD model. A server learns a global fPAD model by iteratively aggregating model updates from all data centers without accessing private data in each of them. Once the learned global model converges, it is used for fPAD inference. To equip the aggregated fPAD model in the server with better generalization ability to unseen attacks from users, following the basic idea of FedPAD, we further propose a federated generalized face presentation attack detection (FedGPAD) framework. A federated domain disentanglement strategy is introduced in FedGPAD, which treats each data center as one domain and decomposes the fPAD model into domain-invariant and domain-specific parts in each data center. Two parts disentangle the domain-invariant and domain-specific features from images in each local data center. A server learns a global fPAD model by only aggregating domain-invariant parts of the fPAD models from data centers, and thus, a more generalized fPAD model can be aggregated in server. We introduce the experimental setting to evaluate the proposed FedPAD and FedGPAD frameworks and carry out extensive experiments to provide various insights about federated learning for fPAD.
Rui Shao 0001, Pramuditha Perera, Pong C. Yuen, Vishal M. Patel
IEEE Trans. Neural Networks Learn. Syst.3
2023 Density-Aware Temporal Attentive Step-wise Diffusion Model For Medical Time Series Imputation
abstract
Medical time series have been widely employed for disease prediction. Missing data hinders accurate prediction. While existing imputation methods partially solve the problem, there are two challenges for medical time series: (1) High dimensionality: Existing imputation methods existing methods suffer from the trade-off between accuracy and computational efficiency. (2) Irregularity: Medical time series exhibit the dynamic temporal relationship that changes over varying sampling densities. However, existing methods mainly take the stationary mechanism, which struggles with capturing the dynamic temporal relationships. To overcome the above deficiencies, we propose a Density-Aware Temporal Attentive Step-wise Diffusion Model (DA-TASWDM), which imputes each time step based on a non-iterative diffusion model and captures inter-step dependency with the density-aware time similarity. Specifically, DA-TASWDM exploits two novel modules: (1) Density-Aware Temporal Attention (DA-TA): It correlates inter-step values from the time embedding similarity adjusted with varying sampling densities. (2) Non-Iterative Step-wise Diffusion Imputer (NI-SWDI): It directly recovers the missing values at each time step from noise without diffusion iteration. Compared with the existing methods, DA-TASWDM can achieve promising accuracy without sacrificing computational efficiency. Extensive experimental results on three real-world datasets demonstrate that our method can significantly outperform state-of-the-art methods in both imputation and post-imputation performance.
Jingwen Xu 0002, Fei Lyu 0004, Pong C. Yuen
CIKM3
2023 Dual-bridging with Adversarial Noise Generation for Domain Adaptive rPPG Estimation
abstract
The remote photoplethysmography (rPPG) technique can estimate pulse-related metrics (e.g. heart rate and respiratory rate) from facial videos and has a high potential for health monitoring. The latest deep rPPG methods can model in-distribution noise due to head motion, video compression, etc., and estimate high-quality rPPG signals under similar scenarios. However, deep rPPG models may not generalize well to the target test domain with unseen noise and distortions. In this paper, to improve the generalization ability of rPPG models, we propose a dual-bridging network to reduce the domain discrepancy by aligning intermediate domains and synthesizing the target noise in the source domain for better noise reduction. To comprehensively explore the target domain noise, we propose a novel adversarial noise generation in which the noise generator indirectly competes with the noise reducer. To further improve the robustness of the noise reducer, we propose hard noise pattern mining to encourage the generator to learn hard noise patterns contained in the target domain features. We evaluated the proposed method on three public datasets with different types of interferences. Under different crossdomain scenarios, the comprehensive results show the effectiveness of our method.
Jingda Du, Si-Qi Liu 0003, Bochao Zhang, Pong C. Yuen
CVPR4
2023 Tackling Model Mismatch with Mixup Regulated Test-Time Training
abstract
Test-time training (TTT) is an emerging approach for addressing the problem of domain shift. In its framework, a test-time training phase is inserted between the training phase and the test phase. During the test-time training phase, the representation layers are adapted using an auxiliary task. Then the updated model will be used in the test phase. Although the idea is very intuitive, TTT does not demonstrate competitive performance compared with some other domain adaption methods. In this paper, we present both theoretical and empirical analyses to explain the subpar performance of TTT. In particular, we point out that TTT causes a new kind of problem, which we term as Model Mismatch. To address this problem of Model Mismatch, we analyse a simple yet effective method inspired by the idea of mixup in robust training. Such effectiveness is shown in the experimental results.
Bochao Zhang, Rui Shao 0001, Jingda Du, Pong C. Yuen, Wei Luo 0001
DSAA4
2023 E2: Entropy Discrimination and Energy Optimization for Source-free Universal Domain Adaptation
abstract
Universal domain adaptation (UniDA) transfers knowledge under both distribution and category shifts. Most UniDA methods accessible to source-domain data during model adaptation may result in privacy policy violation and source-data transfer inefficiency. To address this issue, we propose a novel source-free UniDA method coupling confidence-guided entropy discrimination and likelihood-induced energy optimization. The entropy-based separation of target-known and unknown classes is too conservative for known-class prediction. Thus, we derive the confidence-guided entropy by scaling the normalized prediction score with the known-class confidence, that more known-class samples are correctly predicted. Due to difficult estimation of the marginal distribution without source-domain data, we constrain the target-domain marginal distribution by maximizing (minimizing) the known (unknown)-class likelihood, which equals free energy optimization. Theoretically, the overall optimization amounts to decreasing and increasing internal energy of known and unknown classes in physics, respectively. Extensive experiments demonstrate the superiority of the proposed method.
Andy Jinhua Ma, Pong C. Yuen
ICME3
2023 Pseudo-Label Guided Image Synthesis for Semi-Supervised COVID-19 Pneumonia Infection Segmentation
abstract
Coronavirus disease 2019 (COVID-19) has become a severe global pandemic. Accurate pneumonia infection segmentation is important for assisting doctors in diagnosing COVID-19. Deep learning-based methods can be developed for automatic segmentation, but the lack of large-scale well-annotated COVID-19 training datasets may hinder their performance. Semi-supervised segmentation is a promising solution which explores large amounts of unlabelled data, while most existing methods focus on pseudo-label refinement. In this paper, we propose a new perspective on semi-supervised learning for COVID-19 pneumonia infection segmentation, namely pseudo-label guided image synthesis. The main idea is to keep the pseudo-labels and synthesize new images to match them. The synthetic image has the same COVID-19 infected regions as indicated in the pseudo-label, and the reference style extracted from the style code pool is added to make it more realistic. We introduce two representative methods by incorporating the synthetic images into model training, including single-stage Synthesis-Assisted Cross Pseudo Supervision (SA-CPS) and multi-stage Synthesis-Assisted Self-Training (SA-ST), which can work individually as well as cooperatively. Synthesis-assisted methods expand the training data with high-quality synthetic data, thus improving the segmentation performance. Extensive experiments on two COVID-19 CT datasets for segmenting the infections demonstrate our method is superior to existing schemes for semi-supervised segmentation, and achieves the state-of-the-art performance on both datasets. Code is available at: https://github.com/FeiLyu/SASSL.
Fei Lyu 0004, Mang Ye, Jonathan Frederik Carlsen, Kenny Erleben, Sune Darkner, Pong C. Yuen
IEEE Trans. Medical Imaging6
2022 Anatomical prior-inspired label refinement for weakly supervised liver tumor segmentation with volume-level labels
Fei Lyu 0004, Andy Jinhua Ma, Pong C. Yuen
BMVC3
2022 Learning Sparse Interpretable Features For NAS Scoring From Liver Biopsy Images
abstract
Liver biopsy images play a key role in the diagnosis of global non-alcoholic fatty liver disease (NAFLD). The NAFLD activity score (NAS) on liver biopsy images grades the amount of histological findings that reflect the progression of NAFLD. However, liver biopsy image analysis remains a challenging task due to its complex tissue structures and sparse distribution of histological findings. In this paper, we propose a sparse interpretable feature learning method (SparseX) to efficiently estimate NAS scores. First, we introduce an interpretable spatial sampling strategy based on histological features to effectively select informative tissue regions containing tissue alterations. Then, SparseX formulates the feature learning as a low-rank decomposition problem. Non-negative matrix factorization (NMF)-based attributes learning is embedded into a deep network to compress and select sparse features for a small portion of tissue alterations contributing to diagnosis. Experiments conducted on the internal Liver-NAS and public SteatosisRaw datasets show the effectiveness of the proposed method in terms of classification performance and interpretability. regions containing tissue alterations. Then, SparseX formulates the feature learning as a low-rank decomposition problem. Non-negative matrix factorization (NMF)-based attributes learning is embedded into a deep network to compress and select sparse features for a small portion of tissue alterations contributing to diagnosis. Experiments conducted on the internal Liver-NAS and public SteatosisRaw datasets show the effectiveness of the proposed method in terms of classification performance and interpretability.
Chong Yin, Si-Qi Liu 0003, Vincent Wai-Sun Wong, Pong C. Yuen
IJCAI4
2022 Open-Set Adversarial Defense with Clean-Adversarial Mutual Learning
Rui Shao 0001, Pramuditha Perera, Pong C. Yuen, Vishal M. Patel
Int. J. Comput. Vis.3
2022 Augmentation Invariant and Instance Spreading Feature for Softmax Embedding
abstract
Deep embedding learning plays a key role in learning discriminative feature representations, where the visually similar samples are pulled closer and dissimilar samples are pushed away in the low-dimensional embedding space. This paper studies the unsupervised embedding learning problem by learning such a representation without using any category labels. This task faces two primary challenges: mining reliable positive supervision from highly similar fine-grained classes, and generalizing to unseen testing categories. To approximate the positive concentration and negative separation properties in category-wise supervised learning, we introduce a data augmentation invariant and instance spreading feature using the instance-wise supervision. We also design two novel domain-agnostic augmentation strategies to further extend the supervision in feature space, which simulates the large batch training using a small batch size and the augmented features. To learn such a representation, we propose a novel instance-wise softmax embedding, which directly perform the optimization over the augmented instance features with the binary discrmination softmax encoding. It significantly accelerates the learning speed with much higher accuracy than existing methods, under both seen and unseen testing categories. The unsupervised embedding performs well even without pre-trained network over samples from fine-grained categories. We also develop a variant using category-wise supervision, namely category-wise softmax embedding, which achieves competitive performance over the state-of-of-the-arts, without using any auxiliary information or restrict sample mining.
Mang Ye, Jianbing Shen, Xu Zhang 0022, Pong C. Yuen, Shih-Fu Chang
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Cross-Domain Missingness-Aware Time-Series Adaptation With Similarity Distillation in Medical Applications
abstract
Medical time series of laboratory tests has been collected in electronic health records (EHRs) in many countries. Machine-learning algorithms have been proposed to analyze the condition of patients using these medical records. However, medical time series may be recorded using different laboratory parameters in different datasets. This results in the failure of applying a pretrained model on a test dataset containing a time series of different laboratory parameters. This article proposes to solve this problem with an unsupervised time-series adaptation method that generates time series across laboratory parameters. Specifically, a medical time-series generation network with similarity distillation is developed to reduce the domain gap caused by the difference in laboratory parameters. The relations of different laboratory parameters are analyzed, and the similarity information is distilled to guide the generation of target-domain specific laboratory parameters. To further improve the performance in cross-domain medical applications, a missingness-aware feature extraction network is proposed, where the missingness patterns reflect the health conditions and, thus, serve as auxiliary features for medical analysis. In addition, we also introduce domain-adversarial networks in both feature level and time-series level to enhance the adaptation across domains. Experimental results show that the proposed method achieves good performance on both private and publicly available medical datasets. Ablation studies and distribution visualization are provided to further analyze the properties of the proposed method.
Baoyao Yang, Mang Ye, Qingxiong Tan, Pong C. Yuen
IEEE Trans. Cybern.4
2022 Learning Temporal Similarity of Remote Photoplethysmography for Fast 3D Mask Face Presentation Attack Detection
abstract
To detect 3D mask face presentation attack, remote Photoplethysmography (rPPG), a biomedical technique that measures the heartbeat signal remotely with a normal RGB camera, is adopted as a robust liveness cue. Although existing rPPG-based solutions exhibit strong performance in experiments, the required observation time is too long (10-12 seconds) to be user-friendly in real applications such as E-payment and smartphone unlock. To shorten the observation time (within 1-second), we propose a fast rPPG-based 3D mask presentation attack detection (PAD) method by analyzing the similarity of rPPG signals in the time domain. In particular, based on facial and background local rPPG signals, we design a set of temporal similarity features to investigate the robust properties of rPPG shape and phase. Following the same direction, we refine the traditional rPPG extractor into a learnable network to cooperate with our TSrPPG feature for better robustness. An effective but lightweight spatiotemporal convolution network is constructed with a self-supervised learning strategy, aiming at enhancing the consistency of genuine facial rPPG signals and reducing the correlation of rPPG signals on masked faces. Extensive experiments are conducted on 3DMAD, HKBU-MARs V1+ and V2+, and CSMAD, which totally involve 18772 short-term video slots with a large number of real-world variations, in terms of mask type, mask transmittance, lighting condition, recording device, resolution of facial region, and compression configuration. Our proposed method persists the good performance of rPPG-based solution with only 1-second observation and outperforms the state-of-the-art competitors on discriminability and generalizability. Evaluations on prints attack, display attack, and disguise attacks with transparent masks, make-up and tattoo further exhibit its potential on handling a wider variety of attacks. To our best knowledge, this is the first work that addresses the length of observation time issue of rPPG-based 3D mask PAD.
Si-Qi Liu 0003, Xiangyuan Lan, Pong C. Yuen
IEEE Trans. Inf. Forensics Secur.3
2022 Revealing Task-Relevant Model Memorization for Source-Protected Unsupervised Domain Adaptation
abstract
Source-data-free unsupervised domain adaptation (SF-UDA) is an approach to improve model performance in the target domain without accessing the source data. Some SF-UDA methods have been proposed and achieved promising results using the information from source-model parameters. However, current research on information security confirms the ability of a well-trained model to memorize its training data. Therefore, SF-UDA methods that access model parameters remain at risk of privacy disclosure. This paper introduces a new topic of source-protected UDA (SP-UDA) that adapts the source model to the target domain while protecting the source-domain data and model privacy. In SP-UDA, only a black-box source model and a set of unlabeled target data are available for domain adaptation. We consider SP-UDA from a new perspective of model memorization revelation. A Source-Protected Generative Model (SPGM) is developed to reveal task-relevant memorization from the source model. SPGM directly distills the inverse process of the source model without access to source-model parameters to meet the privacy protection objective in SP-UDA. The SPGM is learned under the supervision of a newly designed metric named privacy-protected transfer (PPT). The PPT metric measures the transferability and desensitization of the generated data to encourage the SPGM to extract task-relevant information rather than the unintended memorization. A set of desensitized pseudo data is then generated as substitutes for the real source data in UDA. The performance of the proposed method has been validated in four cross-dataset recognition applications with encouraging results.
Baoyao Yang, Andy Jinhua Ma, Pong C. Yuen
IEEE Trans. Inf. Forensics Secur.3
2022 Model-Induced Generalization Error Bound for Information-Theoretic Representation Learning in Source-Data-Free Unsupervised Domain Adaptation
abstract
Many unsupervised domain adaptation (UDA) methods have been developed and have achieved promising results in various pattern recognition tasks. However, most existing methods assume that raw source data are available in the target domain when transferring knowledge from the source to the target domain. Due to the emerging regulations on data privacy, the availability of source data cannot be guaranteed when applying UDA methods in a new domain. The lack of source data makes UDA more challenging, and most existing methods are no longer applicable. To handle this issue, this paper analyzes the cross-domain representations in source-data-free unsupervised domain adaptation (SF-UDA). A new theorem is derived to bound the target-domain prediction error using the trained source model instead of the source data. On the basis of the proposed theorem, information bottleneck theory is introduced to minimize the generalization upper bound of the target-domain prediction error, thereby achieving domain adaptation. The minimization is implemented in a variational inference framework using a newly developed latent alignment variational autoencoder (LA-VAE). The experimental results show good performance of the proposed method in several cross-dataset classification tasks without using source data. Ablation studies and feature visualization also validate the effectiveness of our method in SF-UDA.
Baoyao Yang, Hao-Wei Yeh, Tatsuya Harada, Pong C. Yuen
IEEE Trans. Image Process.4
2022 Weakly Supervised Liver Tumor Segmentation Using Couinaud Segment Annotation
abstract
Automatic liver tumor segmentation is of great importance for assisting doctors in liver cancer diagnosis and treatment planning. Recently, deep learning approaches trained with pixel-level annotations have contributed many breakthroughs in image segmentation. However, acquiring such accurate dense annotations is time-consuming and labor-intensive, which limits the performance of deep neural networks for medical image segmentation. We note that Couinaud segment is widely used by radiologists when recording liver cancer-related findings in the reports, since it is well-suited for describing the localization of tumors. In this paper, we propose a novel approach to train convolutional networks for liver tumor segmentation using Couinaud segment annotations. Couinaud segment annotations are image-level labels with values ranging from 1 to 8, indicating a specific region of the liver. Our proposed model, namely CouinaudNet, can estimate pseudo tumor masks from the Couinaud segment annotations as pixel-wise supervision for training a fully supervised tumor segmentation model, and it is composed of two components: 1) an inpainting network with Couinaud segment masks which can effectively remove tumors for pathological images by filling the tumor regions with plausible healthy-looking intensities; 2) a difference spotting network for segmenting the tumors, which is trained with healthy-pathological pairs generated by an effective tumor synthesis strategy. The proposed method is extensively evaluated on two liver tumor segmentation datasets. The experimental results demonstrate that our method can achieve competitive performance compared to the fully supervised counterpart and the state-of-the-art methods while requiring significantly less annotation effort.
Fei Lyu 0004, Andy Jinhua Ma, Terry Cheuk-Fung Yip, Grace Lai-Hung Wong, Pong C. Yuen
IEEE Trans. Medical Imaging5
2022 Learning From Synthetic CT Images via Test-Time Training for Liver Tumor Segmentation
abstract
Automatic liver tumor segmentation could offer assistance to radiologists in liver tumor diagnosis, and its performance has been significantly improved by recent deep learning based methods. These methods rely on large-scale well-annotated training datasets, but collecting such datasets is time-consuming and labor-intensive, which could hinder their performance in practical situations. Learning from synthetic data is an encouraging solution to address this problem. In our task, synthetic tumors can be injected to healthy images to form training pairs. However, directly applying the model trained using the synthetic tumor images on real test images performs poorly due to the domain shift problem. In this paper, we propose a novel approach, namely Synthetic-to-Real Test-Time Training (SR-TTT), to reduce the domain gap between synthetic training images and real test images. Specifically, we add a self-supervised auxiliary task, i.e., two-step reconstruction, which takes the output of the main segmentation task as its input to build an explicit connection between these two tasks. Moreover, we design a scheduled mixture strategy to avoid error accumulation and bias explosion in the training process. During test time, we adapt the segmentation model to each test image with self-supervision from the auxiliary task so as to improve the inference performance. The proposed method is extensively evaluated on two public datasets for liver tumor segmentation. The experimental results demonstrate that our proposed SR-TTT can effectively mitigate the synthetic-to-real domain shift problem in the liver tumor segmentation task, and is superior to existing state-of-the-art approaches.
Fei Lyu 0004, Mang Ye, Andy Jinhua Ma, Terry Cheuk-Fung Yip, Grace Lai-Hung Wong, Pong C. Yuen
IEEE Trans. Medical Imaging6
2021 Hierarchical Information Passing Based Noise-Tolerant Hybrid Learning for Semi-Supervised Human Parsing
abstract
Deep learning based human parsing methods usually require a large amount of training data to reach high performance. However, it is costly and time-consuming to obtain manually annotated high quality labels for a large scale dataset. To alleviate annotation efforts, we propose a new semi-supervised human parsing method for which we only need a small number of labels for training. First, we generate high quality pseudo labels on unlabeled images using a hierarchical information passing network (HIPN), which reasons human part segmentation in a coarse to fine manner. Furthermore, we develop a noise-tolerant hybrid learning method, which takes advantage of positive and negative learning to better handle noisy pseudo labels. When evaluated on standard human parsing benchmarks, our HIPN achieves a new state-of-the-art performance. Moreover, our noise-tolerant hybrid learning method further improves the performance and outperforms the state-of-the-art semi-supervised method (i.e. GRN) by 4.47 points w.r.t mIoU on the LIP dataset.
Yunan Liu 0001, Shanshan Zhang 0001, Jian Yang 0003, Pong C. Yuen
AAAI4
2021 Federated Test-Time Adaptive Face Presentation Attack Detection with Dual-Phase Privacy Preservation
abstract
Face presentation attack detection (fPAD) plays a critical role in the modern face recognition pipeline. The generalization ability of face presentation attack detection models to unseen attacks has become a key issue for real-world deployment, which can be improved when models are trained with face images from different input distributions and different types of spoof attacks. In reality, due to legal and privacy issues, training data (both real face images and spoof images) are not allowed to be directly shared between different data sources. In this paper, to circumvent this challenge, we propose a Federated Test-Time Adaptive Face Presentation Attack Detection with Dual-Phase Privacy Preservation framework, with the aim of enhancing the generalization ability of fPAD models in both training and testing phase while preserving data privacy. In the training phase, the proposed framework exploits the federated learning technique, which simultaneously takes advantage of rich fPAD information available at different data sources by aggregating model updates from them without accessing their private data. To further boost the generalization ability, in the testing phase, we explore test-time adaptation by minimizing the entropy of fPAD model prediction on the testing data, which alleviates the domain gap between training and testing data and thus reduces the generalization error of a fPAD model. We introduce the experimental setting to evaluate the proposed framework and carry out extensive experiments to provide various insights about the proposed method for fPAD.
Rui Shao 0001, Bochao Zhang, Pong C. Yuen, Vishal M. Patel
FG3
2021 Adaptive Multi-Domain Learning for Outdoor 3d Human Pose and Shape Estimation
abstract
It is an extremely challenging task to estimate 3D human pose and shape in outdoor scenes for which we can hardly obtain precise ground truth data for training. Previous methods usually use multiple datasets collected at different scenes to train their models, including those collected in laboratories with precise ground truth and those collected at outdoor scenes with estimated or even no ground truth. Since data from different scenes are included in training, it is necessary to handle the domain difference problem, which unfortunately has never been considered by previous works. In this paper, we first point out this problem and then address it via a novel cascade multi-domain learning module (CMDL), where multiple adapters are employed to extract more discriminative features for different domains. We show that our method with CMDL outperforms previous methods in outdoor scenes. In principle, the proposed CMDL module can be easily applied on top of any arbitrary 3D human pose and shape approach.
Zhaoyang Gui, Shanshan Zhang 0001, Kangkan Wang, Jian Yang 0003, Pong C. Yuen
ICASSP5
2021 Cooperative Joint Attentive Network for Patient Outcome Prediction on Irregular Multi-Rate Multivariate Health Data
abstract
Due to the dynamic health status of patients and discrepant stability of physiological variables, health data often presents as irregular multi-rate multivariate time series (IMR-MTS) with significantly varying sampling rates. Existing methods mainly study changes of IMR-MTS values in the time domain, without considering their different dominant frequencies and varying data quality. Hence, we propose a novel Cooperative Joint Attentive Network (CJANet) to analyze IMR-MTS in frequency domain, which adaptively handling discrepant dominant frequencies while tackling diverse data qualities caused by irregular sampling. In particular, novel dual-channel joint attention is designed to jointly identify important magnitude and phase signals while detecting their dominant frequencies, automatically enlarging the positive influence of key variables and frequencies. Furthermore, a new cooperative learning module is introduced to enhance information exchange between magnitude and phase channels, effectively integrating global signals to optimize the network. A frequency-aware fusion strategy is finally designed to aggregate the learned features. Extensive experimental results on real-world medical datasets indicate that CJANet significantly outperforms existing methods and provides highly interpretable results.
Qingxiong Tan, Mang Ye, Grace Lai-Hung Wong, Pong C. Yuen
IJCAI4
2021 A Segmentation-Assisted Model for Universal Lesion Detection with Partial Labels
Fei Lyu 0004, Baoyao Yang, Andy Jinhua Ma, Pong C. Yuen
MICCAI (5)4
2021 Focusing on Clinically Interpretable Features: Selective Attention Regularization for Liver Biopsy Image Classification
Chong Yin, Si-Qi Liu 0003, Rui Shao 0001, Pong C. Yuen
MICCAI (5)4
2021 SoFA: Source-data-free Feature Alignment for Unsupervised Domain Adaptation
abstract
Applying a trained model on a new scenario may suffer from domain shift. Unsupervised domain adaptation (UDA) has been proven to be an effective approach to solve the problem of domain shift by leveraging both data from the scenario that the model was trained on (source) and the new scenario (target). Although the source data are available for training the source model, there is no guarantee that the source data will still be available when applying UDA in the future due to emerging regulations on privacy of data. This results in the in-applicability of most existing UDA methods in the absence of source data. This paper proposes a source-data-free feature alignment (SoFA) method to address this problem by only using the trained source model and unlabeled target data. The source model is used to predict the labels for target data, and we model the generation process from predicted classes to input data to infer the latent features for alignment. Specifically, a mixture of Gaussian distributions is induced from the predicted classes as the reference distribution. The encoded target features are then aligned to the reference distribution via variational inference to extract class semantics without accessing source data. Relationship of the proposed method and the theory of domain adaptation is provided to verify the performance. Experimental results show the proposed method achieves higher or comparable accuracy compared to the existing methods in several cross-dataset classification tasks. Ablation studies are also conducted to confirm the importance of latent feature alignment to adaptation performance.
Hao-Wei Yeh, Baoyao Yang, Pong C. Yuen, Tatsuya Harada
WACV3
2021 Importance-aware personalized learning for early risk prediction using static and dynamic health data
abstract
OBJECTIVE: Accurate risk prediction is important for evaluating early medical treatment effects and improving health care quality. Existing methods are usually designed for dynamic medical data, which require long-term observations. Meanwhile, important personalized static information is ignored due to the underlying uncertainty and unquantifiable ambiguity. It is urgent to develop an early risk prediction method that can adaptively integrate both static and dynamic health data. MATERIALS AND METHODS: Data were from 6367 patients with Peptic Ulcer Bleeding between 2007 and 2016. This article develops a novel End-to-end Importance-Aware Personalized Deep Learning Approach (eiPDLA) to achieve accurate early clinical risk prediction. Specifically, eiPDLA introduces a long short-term memory with temporal attention to learn sequential dependencies from time-stamped records and simultaneously incorporating a residual network with correlation attention to capture their influencing relationship with static medical data. Furthermore, a new multi-residual multi-scale network with the importance-aware mechanism is designed to adaptively fuse the learned multisource features, automatically assigning larger weights to important features while weakening the influence of less important features. RESULTS: Extensive experimental results on a real-world dataset illustrate that our method significantly outperforms the state-of-the-arts for early risk prediction under various settings (eg, achieving an AUC score of 0.944 at 1 year ahead of risk prediction). Case studies indicate that the achieved prediction results are highly interpretable. CONCLUSION: These results reflect the importance of combining static and dynamic health data, mining their influencing relationship, and incorporating the importance-aware mechanism to automatically identify important features. The achieved accurate early risk prediction results save precious time for doctors to timely design effective treatments and improve clinical outcomes.
Qingxiong Tan, Mang Ye, Andy Jinhua Ma, Terry Cheuk-Fung Yip, Grace Lai-Hung Wong, Pong C. Yuen
J. Am. Medical Informatics Assoc.6
2021 Learning adaptive geometry for unsupervised domain adaptation
Baoyao Yang, Pong C. Yuen
Pattern Recognit.2
2021 Self-Fusion Convolutional Neural Networks
Shenjian Gong, Shanshan Zhang 0001, Jian Yang 0003, Pong C. Yuen
Pattern Recognit. Lett.4
2021 Multi-Channel Remote Photoplethysmography Correspondence Feature for 3D Mask Face Presentation Attack Detection
abstract
With the advancement of 3D printing technologies, 3D mask presentation attack becomes a critical challenge in face recognition. To tackle the 3D mask presentation attack detection (PAD), remote Photoplethysmography (rPPG) is employed as an intrinsic detection cue which is independent of the mask material and appearance quality. Although the effectiveness of existing rPPG-based methods has been verified, they may not be robust enough when rPPG signals are contaminated by noise. To identify the heartbeat information from the noisy raw rPPG signals, we propose a new 3D mask PAD feature, multi-channel rPPG correspondence feature (MCCFrPPG) with the global noise-aware template learning and verification framework. To further boost the discriminability, temporal variation of the rPPG signal is considered and extracted through the multi-channel time-frequency analysis scheme. This paper also extends HKBU-MARs V2 dataset with more customized high-quality masks and increases the number of videos by two times. Comprehensive experiments were performed on existing 3D mask datasets and the extended HKBU-MARs V2+, which totally covers 3 types of masks, 12 different light settings and 6 cameras. The results not only justify the effectiveness and robustness of the proposed MCCFrPPG on 3D mask attacks but also indicate its potential on handling the replay attack with camera motion and dim light.
Si-Qi Liu 0003, Xiangyuan Lan, Pong C. Yuen
IEEE Trans. Inf. Forensics Secur.3
2021 SecureFace: Face Template Protection
abstract
It has been shown that face images can be reconstructed from their representations (templates). We propose a randomized CNN to generate protected face biometric templates given the input face image and a user-specific key. The use of user-specific keys introduces randomness to the secure template and hence strengthens the template security. To further enhance the security of the templates, instead of storing the key, we store a secure sketch that can be decoded to generate the key with genuine queries submitted to the system. We have evaluated the proposed protected template generation method using three benchmarking datasets for the face (FRGC v2.0, CFP, and IJB-A). The experimental results justify that the protected template generated by the proposed method are non-invertible and cancellable, while preserving the verification performance.
Guangcan Mai, Kai Cao 0001, Xiangyuan Lan, Pong C. Yuen
IEEE Trans. Inf. Forensics Secur.4
2021 Explainable Uncertainty-Aware Convolutional Recurrent Neural Network for Irregular Medical Time Series
abstract
Influenced by the dynamic changes in the severity of illness, patients usually take examinations in hospitals irregularly, producing a large volume of irregular medical time-series data. Performing diagnosis prediction from the irregular medical time series is challenging because the intervals between consecutive records significantly vary along time. Existing methods often handle this problem by generating regular time series from the irregular medical records without considering the uncertainty in the generated data, induced by the varying intervals. Thus, a novel Uncertainty-Aware Convolutional Recurrent Neural Network (UA-CRNN) is proposed in this article, which introduces the uncertainty information in the generated data to boost the risk prediction. To tackle the complex medical time series with subseries of different frequencies, the uncertainty information is further incorporated into the subseries level rather than the whole sequence to seamlessly adjust different time intervals. Specifically, a hierarchical uncertainty-aware decomposition layer (UADL) is designed to adaptively decompose time series into different subseries and assign them proper weights in accordance with their reliabilities. Meanwhile, an Explainable UA-CRNN (eUA-CRNN) is proposed to exploit filters with different passbands to ensure the unity of components in each subseries and the diversity of components in different subseries. Furthermore, eUA-CRNN incorporates with an uncertainty-aware attention module to learn attention weights from the uncertainty information, providing the explainable prediction results. The extensive experimental results on three real-world medical data sets illustrate the superiority of the proposed method compared with the state-of-the-art methods.
Qingxiong Tan, Mang Ye, Andy Jinhua Ma, Baoyao Yang, Terry Cheuk-Fung Yip, Grace Lai-Hung Wong, Pong C. Yuen
IEEE Trans. Neural Networks Learn. Syst.7
2021 Spatial-temporal Regularized Multi-modality Correlation Filters for Tracking with Re-detection
abstract
The development of multi-spectrum image sensing technology has brought great interest in exploiting the information of multiple modalities (e.g., RGB and infrared modalities) for solving computer vision problems. In this article, we investigate how to exploit information from RGB and infrared modalities to address two important issues in visual tracking: robustness and object re-detection. Although various algorithms that attempt to exploit multi-modality information in appearance modeling have been developed, they still face challenges that mainly come from the following aspects: (1) the lack of robustness to deal with large appearance changes and dynamic background, (2) failure in re-capturing the object when tracking loss happens, and (3) difficulty in determining the reliability of different modalities. To address these issues and perform effective integration of multiple modalities, we propose a new tracking-by-detection algorithm called Adaptive Spatial-temporal Regulated Multi-Modality Correlation Filter. Particularly, an adaptive spatial-temporal regularization is imposed into the correlation filter framework in which the spatial regularization can help to suppress effect from the cluttered background while the temporal regularization enables the adaptive incorporation of historical appearance cues to deal with appearance changes. In addition, a dynamic modality weight learning algorithm is integrated into the correlation filter training, which ensures that more reliable modalities gain more importance in target tracking. Experimental results demonstrate the effectiveness of the proposed method.
Xiangyuan Lan, Zifei Yang, Wei Zhang 0021, Pong C. Yuen
ACM Trans. Multim. Comput. Commun. Appl.4
2020 Regularized Fine-Grained Meta Face Anti-Spoofing
abstract
Face presentation attacks have become an increasingly critical concern when face recognition is widely applied. Many face anti-spoofing methods have been proposed, but most of them ignore the generalization ability to unseen attacks. To overcome the limitation, this work casts face anti-spoofing as a domain generalization (DG) problem, and attempts to address this problem by developing a new meta-learning framework called Regularized Fine-grained Meta-learning. To let our face anti-spoofing model generalize well to unseen attacks, the proposed framework trains our model to perform well in the simulated domain shift scenarios, which is achieved by finding generalized learning directions in the meta-learning process. Specifically, the proposed framework incorporates the domain knowledge of face anti-spoofing as the regularization so that meta-learning is conducted in the feature space regularized by the supervision of domain knowledge. This enables our model more likely to find generalized learning directions with the regularized meta-learning for face anti-spoofing task. Besides, to further enhance the generalization ability of our model, the proposed framework adopts a fine-grained learning strategy that simultaneously conducts meta-learning in a variety of domain shift scenarios in each iteration. Extensive experiments on four public datasets validate the effectiveness of the proposed method.
Rui Shao 0001, Xiangyuan Lan, Pong C. Yuen
AAAI3
2020 DATA-GRU: Dual-Attention Time-Aware Gated Recurrent Unit for Irregular Multivariate Time Series
abstract
Due to the discrepancy of diseases and symptoms, patients usually visit hospitals irregularly and different physiological variables are examined at each visit, producing large amounts of irregular multivariate time series (IMTS) data with missing values and varying intervals. Existing methods process IMTS into regular data so that standard machine learning models can be employed. However, time intervals are usually determined by the status of patients, while missing values are caused by changes in symptoms. Therefore, we propose a novel end-to-end Dual-Attention Time-Aware Gated Recurrent Unit (DATA-GRU) for IMTS to predict the mortality risk of patients. In particular, DATA-GRU is able to: 1) preserve the informative varying intervals by introducing a time-aware structure to directly adjust the influence of the previous status in coordination with the elapsed time, and 2) tackle missing values by proposing a novel dual-attention structure to jointly consider data-quality and medical-knowledge. A novel unreliability-aware attention mechanism is designed to handle the diversity in the reliability of different data, while a new symptom-aware attention mechanism is proposed to extract medical reasons from original clinical records. Extensive experimental results on two real-world datasets demonstrate that DATA-GRU can significantly outperform state-of-the-art methods and provide meaningful clinical interpretation.
Qingxiong Tan, Mang Ye, Baoyao Yang, Si-Qi Liu 0003, Andy Jinhua Ma, Terry Cheuk-Fung Yip, Grace Lai-Hung Wong, Pong C. Yuen
AAAI8
2020 Open-Set Adversarial Defense
Rui Shao 0001, Pramuditha Perera, Pong C. Yuen, Vishal M. Patel
ECCV (17)3
2020 A General Remote Photoplethysmography Estimator with Spatiotemporal Convolutional Network
abstract
Remote PPG (rPPG) has been attracting increasing attention due to its potential in a wide range of application scenarios such as clinical monitoring, physical training and face presentation attack detection. On top of manually designed solutions, the deep-learning approach appears in rPPG estimation and achieves top-level performance. However, most of them try to integrate the steps of both preprocessing and ROI selection into an end-to-end network, which limits the generalization in other applications that use different input skin regions. The ROI selection model learned on face videos may not adapt to skin regions from other body parts with different appearance and size. In this paper, we leave the preprocessing apart and propose a lightweight rPPG estimation network-DeeprPPG for general use. DeeprPPG is based on spatiotemporal convolutions and can be used as a well-defined module in wider application scenarios with different types of input skin. To further boost the robustness, a spatiotemporal rPPG aggregation strategy is designed to adaptively aggregate rPPG signals from multiple skin regions into the final one. Extensive experiments are conducted and the results illustrate its robustness when facing unseen skin regions and unseen scenarios.
Si-Qi Liu 0003, Pong C. Yuen
FG2
2020 Temporal Similarity Analysis of Remote Photoplethysmography for Fast 3D Mask Face Presentation Attack Detection
abstract
To tackle the 3D mask face presentation attack, remote Photoplethysmography (rPPG), a biomedical technique that can detect heartbeat signal remotely, is employed as an intrinsic liveness cue. Although existing rPPG-based methods exhibit encouraging results, they require long observation time (10-12 seconds) to identify the heartbeat information, which limits their employment in real applications such as smartphone unlock and e-payment. To shorten the observation time (within 1-second) while keeping the performance, we propose a fast rPPG-based 3D mask presentation attack detection (PAD) method by analyzing the similarity of local facial rPPG signals in the time domain. In particular, a set of temporal similarity features of facial and background local rPPG signals are designed and fused to adapt the real world variations based on rPPG shape and phase properties. For better evaluation under practical variations, we build the HKBU-MARsV2+ dataset that includes 16 masks from 2 types and 6 lighting conditions. Finally, extensive experiments are conducted on 11092 shortterm video slots from 4 datasets with a large number of real- world variations, in terms of mask type, lighting condition, camera, resolution of face region, and compression setting. Results show that the proposed TSrPPG outperforms the state-of-the-art competitors dramatically on discriminabil- ity and generalizability. To our best knowledge, this is the first work that addresses the length of observation time issue of rPPG-based 3D mask PAD.
Si-Qi Liu 0003, Xiangyuan Lan, Pong C. Yuen
WACV3
2020 Temporal matrix completion with locally linear latent factors for medical applications
Andy Jinhua Ma, Jacky C. P. Chan, Frodo Kin-Sun Chan, Pong C. Yuen, Terry Cheuk-Fung Yip, Yee-Kit Tse, Vincent Wai-Sun Wong, Grace Lai-Hung Wong
Artif. Intell. Medicine4
2020 Invariant subspace learning for time series data based on dynamic time warping distance
Huiqi Deng, Weifu Chen, Andy Jinhua Ma, Pong C. Yuen, Guo-Can Feng
Pattern Recognit.5
2020 Modality-correlation-aware sparse representation for RGB-infrared object tracking
Xiangyuan Lan, Mang Ye, Shengping Zhang, Huiyu Zhou 0001, Pong C. Yuen
Pattern Recognit. Lett.5
2020 Bi-Directional Center-Constrained Top-Ranking for Visible Thermal Person Re-Identification
abstract
Visible thermal person re-identification (VT-REID) is a task of matching person images captured by thermal and visible cameras, which is an extremely important issue in night-time surveillance applications. Existing cross-modality recognition works mainly focus on learning sharable feature representations to handle the cross-modality discrepancies. However, apart from the cross-modality discrepancy caused by different camera spectrums, VT-REID also suffers from large cross-modality and intra-modality variations caused by different camera environments and human poses, and so on. In this paper, we propose a dual-path network with a novel bi-directional dual-constrained top-ranking (BDTR) loss to learn discriminative feature representations. It is featured in two aspects: 1) end-to-end learning without extra metric learning step and 2) the dual-constraint simultaneously handles the cross-modality and intra-modality variations to ensure the feature discriminability. Meanwhile, a bi-directional center-constrained top-ranking (eBDTR) is proposed to incorporate the previous two constraints into a single formula, which preserves the properties to handle both cross-modality and intra-modality variations. The extensive experiments on two cross-modality re-ID datasets demonstrate the superiority of the proposed method compared to the state-of-the-arts.
Mang Ye, Xiangyuan Lan, Zheng Wang 0007, Pong C. Yuen
IEEE Trans. Inf. Forensics Secur.4
2020 PurifyNet: A Robust Person Re-Identification Model With Noisy Labels
abstract
Person re-identification (Re-ID) has been widely studied by learning a discriminative feature representation with a set of well-annotated training data. Existing models usually assume that all the training samples are correctly annotated. However, label noise is unavoidable due to false annotations in large-scale industrial applications. Different from the label noise problem in image classification with abundant samples, the person Re-ID task with label noise usually has very limited annotated samples for each identity. In this paper, we propose a robust deep model, namely PurifyNet, to address this issue. PurifyNet is featured in two aspects: 1) it jointly refines the annotated labels and optimizes the neural networks by progressively adjusting the predicted logits, which reuses the wrong labels rather than simply filtering them; 2) it can simultaneously reduce the negative impact of noisy labels and pay more attention to hard samples with correct labels by developing a hard-aware instance re-weighting strategy. With limited annotated samples for each identity, we demonstrate that hard sample mining is crucial for label corrupted Re-ID task, while it is usually ignored in existing robust deep learning methods. Extensive experiments on three datasets demonstrate the robustness of PurifyNet over the competing methods under various settings. Meanwhile, we show that it consistently improves the unsupervised/video-based Re-ID methods. Code is available at: https://github.com/mangye16/ReID-Label-Noise.
Mang Ye, Pong C. Yuen
IEEE Trans. Inf. Forensics Secur.2
2019 Cross-Domain Visual Representations via Unsupervised Graph Alignment
abstract
In unsupervised domain adaptation, distributions of visual representations are mismatched across domains, which leads to the performance drop of a source model in the target domain. Therefore, distribution alignment methods have been proposed to explore cross-domain visual representations. However, most alignment methods have not considered the difference in distribution structures across domains, and the adaptation would subject to the insufficient aligned cross-domain representations. To avoid the misclassification/misidentification due to the difference in distribution structures, this paper proposes a novel unsupervised graph alignment method that aligns both data representations and distribution structures across the source and target domains. An adversarial network is developed for unsupervised graph alignment, which maps both source and target data to a feature space where data are distributed with unified structure criteria. Experimental results show that the graph-aligned visual representations achieve good performance on both crossdataset recognition and cross-modal re-identification.
Baoyao Yang, Pong C. Yuen
AAAI2
2019 UA-CRNN: Uncertainty-Aware Convolutional Recurrent Neural Network for Mortality Risk Prediction
abstract
Accurate prediction of mortality risk is important for evaluating early treatments, detecting high-risk patients and improving healthcare outcomes. Predicting mortality risk from the irregular clinical time series data is challenging due to the varying time intervals in the consecutive records. Existing methods usually solve this issue by generating regular time series data from the original irregular data without considering the uncertainty in the generated data, caused by varying time intervals. In this paper, we propose a novel Uncertainty-Aware Convolutional Recurrent Neural Network (UA-CRNN), which incorporates the uncertainty information in the generated data to improve the mortality risk prediction performance. To handle the complex clinical time series data with sub-series of different frequencies, we propose to incorporate the uncertainty information into the sub-series level rather than the whole time series data. Specifically, we design a novel hierarchical uncertainty-aware decomposition layer (UADL) to adaptively decompose time series into different sub-series and assign them proper weights according to their reliabilities. Experimental results on two real-world clinical datasets demonstrate that the proposed UA-CRNN method significantly outperforms state-of-the-art methods in both short-term and long-term mortality risk predictions.
Qingxiong Tan, Andy Jinhua Ma, Mang Ye, Baoyao Yang, Huiqi Deng, Vincent Wai-Sun Wong, Yee-Kit Tse, Terry Cheuk-Fung Yip, Grace Lai-Hung Wong, Jessica Yuet-Ling Ching, Francis Ka-Leung Chan, Pong C. Yuen
CIKM12
2019 Multi-Adversarial Discriminative Deep Domain Generalization for Face Presentation Attack Detection
abstract
Face presentation attacks have become an increasingly critical issue in the face recognition community. Many face anti-spoofing methods have been proposed, but they cannot generalize well on "unseen" attacks. This work focuses on improving the generalization ability of face anti-spoofing methods from the perspective of the domain generalization. We propose to learn a generalized feature space via a novel multi-adversarial discriminative deep domain generalization framework. In this framework, a multi-adversarial deep domain generalization is performed under a dual-force triplet-mining constraint. This ensures that the learned feature space is discriminative and shared by multiple source domains, and thus is more generalized to new face presentation attacks. An auxiliary face depth supervision is incorporated to further enhance the generalization ability. Extensive experiments on four public datasets validate the effectiveness of the proposed method.
Rui Shao 0001, Xiangyuan Lan, Jiawei Li 0003, Pong C. Yuen
CVPR4
2019 Unsupervised Embedding Learning via Invariant and Spreading Instance Feature
abstract
This paper studies the unsupervised embedding learning problem, which requires an effective similarity measurement between samples in low-dimensional embedding space. Motivated by the positive concentrated and negative separated properties observed from category-wise supervised learning, we propose to utilize the instance-wise supervision to approximate these properties, which aims at learning data augmentation invariant and instance spread-out features. To achieve this goal, we propose a novel instance based softmax embedding method, which directly optimizes the `real' instance features on top of the softmax function. It achieves significantly faster learning speed and higher accuracy than all existing methods. The proposed method performs well for both seen and unseen testing categories with cosine similarity. It also achieves competitive performance even without pre-trained network over samples from fine-grained categories.
Mang Ye, Xu Zhang 0022, Pong C. Yuen, Shih-Fu Chang
CVPR3
2019 Variation Generalized Feature Learning via Intra-view Variation Adaptation
abstract
This paper addresses the variation generalized feature learning problem in unsupervised video-based person re-identification (re-ID). With advanced tracking and detection algorithms, large-scale intra-view positive samples can be easily collected by assuming that the image frames within the tracking sequence belong to the same person. Existing methods either directly use the intra-view positives to model cross-view variations or simply minimize the intra-view variations to capture the invariant component with some discriminative information loss. In this paper, we propose a Variation Generalized Feature Learning (VGFL) method to learn adaptable feature representation with intra-view positives. The proposed method can learn a discriminative re-ID model without any manually annotated cross-view positive sample pairs. It could address the unseen testing variations with a novel variation generalized feature learning algorithm. In addition, an Adaptability-Discriminability (AD) fusion method is introduced to learn adaptable video-level features. Extensive experiments on different datasets demonstrate the effectiveness of the proposed method.
Jiawei Li 0003, Mang Ye, Andy Jinhua Ma, Pong C. Yuen
IJCAI4
2019 On the Reconstruction of Face Images from Deep Face Templates
abstract
State-of-the-art face recognition systems are based on deep (convolutional) neural networks. Therefore, it is imperative to determine to what extent face templates derived from deep networks can be inverted to obtain the original face image. In this paper, we study the vulnerabilities of a state-of-the-art face recognition system based on template reconstruction attack. We propose a neighborly de-convolutional neural network (NbNet) to reconstruct face images from their deep templates. In our experiments, we assumed that no knowledge about the target subject and the deep network are available. To train the NbNet reconstruction models, we augmented two benchmark face datasets (VGG-Face and Multi-PIE) with a large collection of images synthesized using a face generator. The proposed reconstruction was evaluated using type-I (comparing the reconstructed images against the original face images used to generate the deep template) and type-II (comparing the reconstructed images against a different face image of the same subject) attacks. Given the images reconstructed from NbNets, we show that for verification, we achieve TAR of 95.20 percent (58.05 percent) on LFW under type-I (type-II) attacks @ FAR of 0.1 percent. Besides, 96.58 percent (92.84 percent) of the images reconstructed from templates of partition fa (fb) can be identified from partition fa in color FERET. Our study demonstrates the need to secure deep templates in face recognition systems.
Guangcan Mai, Kai Cao 0001, Pong C. Yuen, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2019 Body Parts Synthesis for Cross-Quality Pose Estimation
abstract
Although encouraging results have been obtained in human pose estimation in recent years, the performance may degrade dramatically when the image quality differs between training and testing data sets. This paper addresses problems in cross-image-quality human pose estimation. To achieve this, we follow an unsupervised domain adaptation approach, in which labels in the target domain are unavailable. Unlike existing unsupervised domain adaptation methods that find label information from unlabeled data, the target pose information (label) is instead generated by synthesizing body parts with similar image-quality of the target domain. A translative dictionary is learned to associate the source and target domains, and a cross-quality adaptation model is developed to refine the source pose estimator using the synthesized target body parts. We perform cross-quality experiments on three data sets with different image quality by using two state-of-the-art pose estimators, and compare the proposed method with five unsupervised domain adaptation methods. Our experimental results show that the proposed method outperforms not only the source pose estimators, but also other unsupervised domain adaptation methods.
Baoyao Yang, Andy Jinhua Ma, Pong C. Yuen
IEEE Trans. Circuits Syst. Video Technol.3
2019 Joint Discriminative Learning of Deep Dynamic Textures for 3D Mask Face Anti-Spoofing
abstract
Three-dimensional mask spoofing attacks have been one of the main challenges in face recognition. Compared with a 3D mask, a real face displays different facial motion patterns that are reflected by different facial dynamic textures. However, a large portion of these facial motion differences is subtle. We find that the subtle facial motion can be fully captured by multiple deep dynamic textures from a convolutional layer of a convolutional neural network, but not all deep dynamic textures from different spatial regions and different channels of a convolutional layer are useful for differentiation of subtle motions between real faces and 3D masks. In this paper, we propose a novel feature learning model to learn discriminative deep dynamic textures for 3D mask face anti-spoofing. A novel joint discriminative learning strategy is further incorporated in the learning model to jointly learn the spatial- and channel-discriminability of the deep dynamic textures. The proposed joint discriminative learning strategy can be used to adaptively weight the discriminability of the learned feature from different spatial regions or channels, which ensures that more discriminative deep dynamic textures play more important roles in face/mask classification. Experiments on several publicly available data sets validate that the proposed method achieves promising results in intra- and cross-data set scenarios.
Rui Shao 0001, Xiangyuan Lan, Pong C. Yuen
IEEE Trans. Inf. Forensics Secur.3
2019 Dynamic Graph Co-Matching for Unsupervised Video-Based Person Re-Identification
abstract
Cross-camera label estimation from a set of unlabelled training data is an extremely important component in unsupervised person re-identification (re-ID) systems. With the estimated labels, existing advanced supervised learning methods can be leveraged to learn discriminative re-ID models. In this paper, we utilize the graph matching technique for accurate label estimation due to its advantages in optimal global matching and intra-camera relationship mining. However, the graph structure constructed with non-learnt similarity measurement cannot handle the large cross-camera variations, which leads to noisy and inaccurate label outputs. This paper designs a Dynamic Graph Matching (DGM) framework, which improves the label estimation process by iteratively refining the graph structure with better similarity measurement learnt from intermediate estimated labels. In addition, we design a positive re-weighting strategy to refine the intermediate labels, which enhances the robustness against inaccurate matching output and noisy initial training data. To fully utilize the abundant video information and reduce false matchings, a co-matching strategy is further incorporated into the framework. Comprehensive experiments conducted on three video benchmarks demonstrate that DGM outperforms state-of-the-art unsupervised re-ID methods and yields competitive performance to fully supervised upper bounds.
Mang Ye, Jiawei Li 0003, Andy Jinhua Ma, Liang Zheng 0001, Pong C. Yuen
IEEE Trans. Image Process.5
2018 Robust Collaborative Discriminative Learning for RGB-Infrared Tracking
abstract
Tracking target of interests is an important step for motion perception in intelligent video surveillance systems. While most recently developed tracking algorithms are grounded in RGB image sequences, it should be noted that information from RGB modality is not always reliable (e.g. in a dark environment with poor lighting condition), which urges the need to integrate information from infrared modality for effective tracking because of the insensitivity to illumination condition of infrared thermal camera. However, several issues encountered during the tracking process limit the fusing performance of these heterogeneous modalities: 1) the cross-modality discrepancy of visual and motion characteristics, 2) the uncertainty of degree of reliability in different modalities, and 3) large target appearance variations and background distractions within each modality. To address these issues, this paper proposes a novel and optimal discriminative learning framework for multi-modality tracking. In particular, the proposed discriminative learning framework is able to: 1) jointly eliminate outlier samples caused by large variations and learn discriminability-consistent features from heterogeneous modalities, and 2) collaboratively perform modality reliability measurement and target-background separation. Extensive experiments on RGB-infrared image sequences demonstrate the effectiveness of the proposed method.
Xiangyuan Lan, Mang Ye, Shengping Zhang, Pong C. Yuen
AAAI4
2018 Domain-Shared Group-Sparse Dictionary Learning for Unsupervised Domain Adaptation
abstract
Unsupervised domain adaptation has been proved to be a promising approach to solve the problem of dataset bias. To employ source labels in the target domain, it is required to align the joint distributions of source and target data. To do this, the key research problem is to align conditional distributions across domains without target labels. In this paper, we propose a new criterion of domain-shared group-sparsity that is an equivalent condition for conditional distribution alignment. To solve the problem in joint distribution alignment, a domain-shared group-sparse dictionary learning method is developed towards joint alignment of conditional and marginal distributions. A classifier for target domain is trained using the domain-shared group-sparse coefficients and the target-specific information from the target data. Experimental results on cross-domain face and object recognition show that the proposed method outperforms eight state-of-the-art unsupervised domain adaptation algorithms.
Baoyao Yang, Andy Jinhua Ma, Pong C. Yuen
AAAI3
2018 Hierarchical Discriminative Learning for Visible Thermal Person Re-Identification
abstract
Person re-identification is widely studied in visible spectrum, where all the person images are captured by visible cameras. However, visible cameras may not capture valid appearance information under poor illumination conditions, e.g, at night. In this case, thermal camera is superior since it is less dependent on the lighting by using infrared light to capture the human body. To this end, this paper investigates a cross-modal re-identification problem, namely visible-thermal person re-identification (VT-REID). Existing cross-modal matching methods mainly focus on modeling the cross-modality discrepancy, while VT-REID also suffers from cross-view variations caused by different camera views. Therefore, we propose a hierarchical cross-modality matching model by jointly optimizing the modality-specific and modality-shared metrics. The modality-specific metrics transform two heterogenous modalities into a consistent space that modality-shared metric can be subsequently learnt. Meanwhile, the modality-specific metric compacts features of the same person within each modality to handle the large intra-modality intra-person variations (e.g. viewpoints, pose). Additionally, an improved two-stream CNN network is presented to learn the multi-modality sharable feature representations. Identity loss and contrastive loss are integrated to enhance the discriminability and modality-invariance with partially shared layer parameters. Extensive experiments illustrate the effectiveness and robustness of the proposed method.
Mang Ye, Xiangyuan Lan, Jiawei Li 0003, Pong C. Yuen
AAAI4
2018 A Hybrid Residual Network and Long Short-Term Memory Method for Peptic Ulcer Bleeding Mortality Prediction
Qingxiong Tan, Andy Jinhua Ma, Huiqi Deng, Vincent Wai-Sun Wong, Yee-Kit Tse, Terry Cheuk-Fung Yip, Grace Lai-Hung Wong, Jessica Yuet-Ling Ching, Francis Ka-Leung Chan, Pong C. Yuen
AMIA10
2018 Remote Photoplethysmography Correspondence Feature for 3D Mask Face Presentation Attack Detection
Si-Qi Liu 0003, Xiangyuan Lan, Pong C. Yuen
ECCV (16)3
2018 Robust Anchor Embedding for Unsupervised Video Person re-IDentification in the Wild
Mang Ye, Xiangyuan Lan, Pong C. Yuen
ECCV (7)3
2018 Visible Thermal Person Re-Identification via Dual-Constrained Top-Ranking
abstract
Cross-modality person re-identification between the thermal and visible domains is extremely important for night-time surveillance applications. Existing works in this filed mainly focus on learning sharable feature representations to handle the cross-modality discrepancies. However, besides the cross-modality discrepancy caused by different camera spectrums, visible thermal person re-identification also suffers from large cross-modality and intra-modality variations caused by different camera views and human poses. In this paper, we propose a dual-path network with a novel bi-directional dual-constrained top-ranking loss to learn discriminative feature representations. It is advantageous in two aspects: 1) end-to-end feature learning directly from the data without extra metric learning steps, 2) it simultaneously handles the cross-modality and intra-modality variations to ensure the discriminability of the learnt representations. Meanwhile, identity loss is further incorporated to model the identity-specific information to handle large intra-class variations. Extensive experiments on two datasets demonstrate the superior performance compared to the state-of-the-arts.
Mang Ye, Zheng Wang 0007, Xiangyuan Lan, Pong C. Yuen
IJCAI4
2018 Feature Constrained by Pixel: Hierarchical Adversarial Deep Domain Adaptation
abstract
In multimedia analysis, one objective of unsupervised visual domain adaptation is to train a classifier that works well on a target domain given labeled source samples and unlabeled target samples. Feature alignment of two domains is the key issue which should be addressed to achieve this objective. Inspired by the recent study of Generative Adversarial Networks (GAN) in domain adaptation, this paper proposes a new model based on Generative Adversarial Network, named Hierarchical Adversarial Deep Network (HADN), which jointly optimizes the feature-level and pixel-level adversarial adaptation within a hierarchical network structure. Specifically, the hierarchical network structure ensures that the knowledge from pixel-level adversarial adaptation can be back propagated to facilitate the feature-level adaptation, which achieves a better feature alignment under the constraint of pixel-level adversarial adaptation. Extensive experiments on various visual recognition tasks show that the proposed method performs favorably against or better than competitive state-of-the-art methods.
Rui Shao 0001, Xiangyuan Lan, Pong C. Yuen
ACM Multimedia3
2018 Robust Shapelets Learning: Transform-Invariant Prototypes
Huiqi Deng, Weifu Chen, Andy Jinhua Ma, Pong C. Yuen, Guo-Can Feng
PRCV (3)5
2018 Semi-supervised Region Metric Learning for Person Re-identification
Jiawei Li 0003, Andy Jinhua Ma, Pong C. Yuen
Int. J. Comput. Vis.3
2018 Inference-Based Similarity Search in Randomized Montgomery Domains for Privacy-Preserving Biometric Identification
abstract
Similarity search is essential to many important applications and often involves searching at scale on high-dimensional data based on their similarity to a query. In biometric applications, recent vulnerability studies have shown that adversarial machine learning can compromise biometric recognition systems by exploiting the biometric similarity information. Existing methods for biometric privacy protection are in general based on pairwise matching of secured biometric templates and have inherent limitations in search efficiency and scalability. In this paper, we propose an inference-based framework for privacy-preserving similarity search in Hamming space. Our approach builds on an obfuscated distance measure that can conceal Hamming distance in a dynamic interval. Such a mechanism enables us to systematically design statistically reliable methods for retrieving most likely candidates without knowing the exact distance values. We further propose to apply Montgomery multiplication for generating search indexes that can withstand adversarial similarity analysis, and show that information leakage in randomized Montgomery domains can be made negligibly small. Our experiments on public biometric datasets demonstrate that the inference-based approach can achieve a search accuracy close to the best performance possible with secure computation methods, but the associated cost is reduced by orders of magnitude compared to cryptographic primitives.
Yi Wang 0017, Jianwu Wan, Jun Guo 0001, Yiu-Ming Cheung, Pong C. Yuen
IEEE Trans. Pattern Anal. Mach. Intell.5
2018 Learning domain-shared group-sparse representation for unsupervised domain adaptation
Baoyao Yang, Andy Jinhua Ma, Pong C. Yuen
Pattern Recognit.3
2018 Learning Common and Feature-Specific Patterns: A Novel Multiple-Sparse-Representation-Based Tracker
abstract
The use of multiple features has been shown to be an effective strategy for visual tracking because of their complementary contributions to appearance modeling. The key problem is how to learn a fused representation from multiple features for appearance modeling. Different features extracted from the same object should share some commonalities in their representations while each feature should also have some feature-specific representation patterns which reflect its complementarity in appearance modeling. Different from existing multi-feature sparse trackers which only consider the commonalities among the sparsity patterns of multiple features, this paper proposes a novel multiple sparse representation framework for visual tracking which jointly exploits the shared and feature-specific properties of different features by decomposing multiple sparsity patterns. Moreover, we introduce a novel online multiple metric learning to efficiently and adaptively incorporate the appearance proximity constraint, which ensures that the learned commonalities of multiple features are more representative. Experimental results on tracking benchmark videos and other challenging videos demonstrate the effectiveness of the proposed tracker.
Xiangyuan Lan, Shengping Zhang, Pong C. Yuen, Rama Chellappa
IEEE Trans. Image Process.3
2018 Point-to-Set Distance Metric Learning on Deep Representations for Visual Tracking
abstract
For autonomous driving application, a car shall be able to track objects in the scene in order to estimate where and how they will move such that the tracker embedded in the car can efficiently alert the car for effective collision-avoidance. Traditional discriminative object tracking methods usually train a binary classifier via a support vector machine (SVM) scheme to distinguish the target from its background. Despite demonstrated success, the performance of the SVM-based trackers is limited because the classification is carried out only depending on support vectors (SVs) but the target's dynamic appearance may look similar to the training samples that have not been selected as SVs, especially when the training samples are not linearly classifiable. In such cases, the tracker may drift to the background and fail to track the target eventually. To address this problem, in this paper, we propose to integrate the point-to-set/image-to-imageSet distance metric learning (DML) into visual tracking tasks and take full advantage of all the training samples when determining the best target candidate. The point-to-set DML is conducted on convolutional neural network features of the training data extracted from the starting frames. When a new frame comes, target candidates are first projected to the common subspace using the learned mapping functions, and then the candidate having the minimal distance to the target template sets is selected as the tracking result. Extensive experimental results show that even without model update the proposed method is able to achieve favorable performance on challenging image sequences compared with several state-of-the-art trackers.
Shengping Zhang, Yuankai Qi, Feng Jiang 0001, Xiangyuan Lan, Pong C. Yuen, Huiyu Zhou 0001
IEEE Trans. Intell. Transp. Syst.5
2017 Robust MIL-Based Feature Template Learning for Object Tracking
abstract
Because of appearance variations, training samples of the tracked targets collected by the online tracker are required for updating the tracking model. However, this often leads to tracking drift problem because of potentially corrupted samples: 1) contaminated/outlier samples resulting from large variations (e.g. occlusion, illumination), and 2) misaligned samples caused by tracking inaccuracy. Therefore, in order to reduce the tracking drift while maintaining the adaptability of a visual tracker, how to alleviate these two issues via an effective model learning (updating) strategy is a key problem to be solved. To address these issues, this paper proposes a novel and optimal model learning (updating) scheme which aims to simultaneously eliminate the negative effects from these two issues mentioned above in a unified robust feature template learning framework. Particularly, the proposed feature template learning framework is capable of: 1) adaptively learning uncontaminated feature templates by separating out contaminated samples, and 2) resolving label ambiguities caused by misaligned samples via a probabilistic multiple instance learning (MIL) model. Experiments on challenging video sequences show that the proposed tracker performs favourably against several state-of-the-art trackers.
Xiangyuan Lan, Pong C. Yuen, Rama Chellappa
AAAI2
2017 A competition on generalized software-based face presentation attack detection in mobile scenarios
abstract
In recent years, software-based face presentation attack detection (PAD) methods have seen a great progress. However, most existing schemes are not able to generalize well in more realistic conditions. The objective of this competition is to evaluate and compare the generalization performances of mobile face PAD techniques under some real-world variations, including unseen input sensors, presentation attack instruments (PAI) and illumination conditions, on a larger scale OULU-NPU dataset using its standard evaluation protocols and metrics. Thirteen teams from academic and industrial institutions across the world participated in this competition. This time typical liveness detection based on physiological signs of life was totally discarded. Instead, every submitted system relies practically on some sort of feature representation extracted from the face and/or background regions using hand-crafted, learned or hybrid descriptors. Interesting results and findings are presented and discussed in this paper.
Zinelabidine Boulkenafet, Jukka Komulainen, Zahid Akhtar, Azeddine Benlamoudi, Djamel Samai, Salah Eddine Bekhouche, Abdelkrim Ouafi, Fadi Dornaika, Abdelmalik Taleb-Ahmed, Fei Peng 0001, L. B. Zhang, Min Long 0003, Shruti Bhilare, Vivek Kanhangad, Artur Costa-Pazo, Esteban Vázquez-Fernández, Daniel Pérez-Cabo, J. J. Moreira-Perez, Daniel González-Jiménez, Amir Mohammadi, Sushil Bhattacharjee, Sébastien Marcel, Svetlana Volkova, N. Abe, X. Feng, Z. Xia, Rui Shao 0001, Pong C. Yuen, Waldir R. de Almeida, Fernanda A. Andaló, Rafael Padilha, Gabriel Bertocco, William Dias, Jacques Wainer, Ricardo da Silva Torres, Anderson Rocha 0001, Marcus A. Angeloni, Guilherme Folego, Alan Godoy, Abdenour Hadid
IJCB33
2017 On the guessability of binary biometric templates: A practical guessing entropy based approach
abstract
A security index for biometric systems is essential because biometrics have been widely adopted as a secure authentication component in critical systems. Most of bio-metric systems secured by template protection schemes are based on binary templates. To adopt popular template protection schemes such as fuzzy commitment and fuzzy extractor that can be applied on binary templates only, non-binary templates (e.g., real-valued, point-set based) need to be converted to binary. However, existing security measurements for binary template based biometric systems either cannot reflect the actual attack difficulties or are too computationally expensive to be practical. This paper presents an acceleration of the guessing entropy which reflects the expected number of guessing trials in attacking the binary template based biometric systems. The acceleration benefits from computation reuse and pruning. Experimental results on two datasets show that the acceleration has more than 6x, 20x, and 200x speed up without losing the estimation accuracy in different system settings.
Guangcan Mai, Meng-Hui Lim, Pong C. Yuen
IJCB3
2017 Deep convolutional dynamic texture learning with adaptive channel-discriminability for 3D mask face anti-spoofing
abstract
3D mask spoofing attack has been one of the main challenges in face recognition. A real face displays a different motion behaviour compared to a 3D mask spoof attempt, which is reflected by different facial dynamic textures. However, the different dynamic information usually exists in the subtle texture level, which cannot be fully differentiated by traditional hand-crafted texture-based methods. In this paper, we propose a novel method for 3D mask face anti-spoofing, namely deep convolutional dynamic texture learning, which learns robust dynamic texture information from fine-grained deep convolutional features. Moreover, channel-discriminability constraint is adaptively incorporated to weight the discriminability of feature channels in the learning process. Experiments on both public datasets validate that the proposed method achieves promising results under intra and cross dataset scenario.
Rui Shao 0001, Xiangyuan Lan, Pong C. Yuen
IJCB3
2017 Dynamic Label Graph Matching for Unsupervised Video Re-identification
abstract
Label estimation is an important component in an unsupervised person re-identification (re-ID) system. This paper focuses on cross-camera label estimation, which can be subsequently used in feature learning to learn robust re-ID models. Specifically, we propose to construct a graph for samples in each camera, and then graph matching scheme is introduced for cross-camera labeling association. While labels directly output from existing graph matching methods may be noisy and inaccurate due to significant cross-camera variations, this paper propose a dynamic graph matching (DGM) method. DGM iteratively updates the image graph and the label estimation process by learning a better feature space with intermediate estimated labels. DGM is advantageous in two aspects: 1) the accuracy of estimated labels is improved significantly with the iterations; 2) DGM is robust to noisy initial training data. Extensive experiments conducted on three benchmarks including the large-scale MARS dataset show that DGM yields competitive performance to fully supervised baselines, and outperforms competing unsupervised learning methods.1
Mang Ye, Andy Jinhua Ma, Liang Zheng 0001, Jiawei Li 0003, Pong C. Yuen
ICCV5
2017 Binary feature fusion for discriminative and secure multi-biometric cryptosystems
Guangcan Mai, Meng-Hui Lim, Pong C. Yuen
Image Vis. Comput.3
2017 Plant identification via multipath sparse coding
Heyan Zhu, Shengping Zhang, Pong C. Yuen
Multim. Tools Appl.4
2017 An Asymmetric Distance Model for Cross-View Feature Mapping in Person Reidentification
abstract
Person reidentification, which matches person images of the same identity across nonoverlapping camera views, becomes an important component for cross-camera-view activity analysis. Most (if not all) person reidentification algorithms are designed based on appearance features. However, appearance features are not stable across nonoverlapping camera views under dramatic lighting change, and those algorithms assume that two cross-view images of the same person can be well represented either by exploring robust and invariant features or by learning matching distance. Such an assumption ignores the nature that images are captured under different camera views with different camera characteristics and environments, and thus, mostly there exists large discrepancy between the extracted features under different views. To solve this problem, we formulate an asymmetric distance model for learning camera-specific projections to transform the unmatched features of each view into a common space where discriminative features across view space are extracted. A cross-view consistency regularization is further introduced to model the correlation between view-specific feature transformations of different camera views, which reflects their nature relations and plays a significant role in avoiding overfitting. A kernel cross-view discriminant component analysis is also presented. Extensive experiments have been conducted to show that asymmetric distance modeling is important for person reidentification, which matches the concerns on cross-disjoint-view matching, reporting superior performance compared with related distance learning methods on six publically available data sets.
Ying-Cong Chen, Wei-Shi Zheng 0001, Jian-Huang Lai, Pong C. Yuen
IEEE Trans. Circuits Syst. Video Technol.4
2017 Robust Visual Tracking via Basis Matching
abstract
Most existing tracking approaches are based on either the tracking by detection framework or the tracking by matching framework. The former needs to learn a discriminative classifier using positive and negative samples, which will cause tracking drift due to unreliable samples. The latter usually performs tracking by matching local interest points between a target candidate and the tracked target, which is not robust to target appearance changes over time. In this paper, we propose a novel tracking by matching framework for robust tracking based on basis matching rather than point matching. In particular, we learn the target model from target images using a set of Gabor basis functions, which have large responses on the corresponding spatial positions after a max pooling. During tracking, a target candidate is evaluated by computing the responses of the Gabor basis functions on their corresponding spatial positions. The experimental results on a set of challenging sequences validate that the performance of the proposed tracking method outperforms those of several state-of-the-art methods.
Shengping Zhang, Xiangyuan Lan, Yuankai Qi, Pong C. Yuen
IEEE Trans. Circuits Syst. Video Technol.4
2016 3D Mask Face Anti-spoofing with Remote Photoplethysmography
Si-Qi Liu 0003, Pong C. Yuen, Shengping Zhang, Guoying Zhao 0001
ECCV (7)2
2016 Generalized face anti-spoofing by detecting pulse from face videos
abstract
Face biometric systems are vulnerable to spoofing attacks. Such attacks can be performed in many ways, including presenting a falsified image, video or 3D mask of a valid user. A widely used approach for differentiating genuine faces from fake ones has been to capture their inherent differences in (2D or 3D) texture using local descriptors. One limitation of these methods is that they may fail if an unseen attack type, e.g. a highly realistic 3D mask which resembles real skin texture, is used in spoofing. Here we propose a robust anti-spoofing method by detecting pulse from face videos. Based on the fact that a pulse signal exists in a real living face but not in any mask or print material, the method could be a generalized solution for face liveness detection. The proposed method is evaluated first on a 3D mask spoofing database 3DMAD to demonstrate its effectiveness in detecting 3D mask attacks. More importantly, our cross-database experiment with high quality REAL-F masks shows that the pulse based method is able to detect even the previously unseen mask type whereas texture based methods fail to generalize beyond the development data. Finally, we propose a robust cascade system combining two complementary attack-specific spoof detectors, i.e. utilize pulse detection against print attacks and color texture analysis against video attacks.
Jukka Komulainen, Guoying Zhao 0001, Pong C. Yuen, Matti Pietikäinen
ICPR4
2016 Robust Joint Discriminative Feature Learning for Visual Tracking
Xiangyuan Lan, Shengping Zhang, Pong C. Yuen
IJCAI3
2016 Improving posture classification accuracy for depth sensor-based human activity monitoring in smart environments
abstract
Smart environments and monitoring systems are popular research areas nowadays due to its potential to enhance the quality of life. Applications such as human behavior analysis and workspace ergonomics monitoring are automated, thereby improving well-being of individuals with minimal running cost. The central problem of smart environments is to understand what the user is doing in order to provide the appropriate support. While it is difficult to obtain information of full body movement in the past, depth camera based motion sensing technology such as Kinect has made it possible to obtain 3D posture without complex setup. This has fused a large number of research projects to apply Kinect in smart environments. The common bottleneck of these researches is the high amount of errors in the detected joint positions, which would result in inaccurate analysis and false alarms. In this paper, we propose a framework that accurately classifies the nature of the 3D postures obtained by Kinect using a max-margin classifier. Different from previous work in the area, we integrate the information about the reliability of the tracked joints in order to enhance the accuracy and robustness of our framework. As a result, apart from general classifying activity of different movement context, our proposed method can classify the subtle differences between correctly performed and incorrectly performed movement in the same context. We demonstrate how our framework can be applied to evaluate the user’s posture and identify the postures that may result in musculoskeletal disorders. Such a system can be used in workplace such as offices and factories to reduce risk of injury. Experimental results have shown that our method consistently outperforms existing algorithms in both activity classification and posture healthiness classification. Due to the low cost and the easy deployment process of depth camera based motion sensors, our framework can be applied widely in home and office to facilitate smart environments.
Edmond S. L. Ho, Jacky C. P. Chan, Donald C. K. Chan, Hubert P. H. Shum, Yiu-Ming Cheung, Pong C. Yuen
Comput. Vis. Image Underst.6
2016 Combination of spatio-temporal and transform domain for sparse occlusion estimation by optical flow
Pengguang Chen, Xingming Zhang 0001, Pong C. Yuen, Aihua Mao
Neurocomputing3
2016 Ensembling over-segmentations: From weak evidence to strong segmentation
Dong Huang 0001, Jian-Huang Lai, Chang-Dong Wang 0001, Pong C. Yuen
Neurocomputing4
2016 Learning discriminability-preserving histogram representation from unordered features for multibiometric feature-fused-template protection
Meng-Hui Lim, Sunny Verma, Guangcan Mai, Pong C. Yuen
Pattern Recognit.4
2016 Entropy Measurement for Biometric Verification Systems
abstract
Biometric verification systems are designed to accept multiple similar biometric measurements per user due to inherent intrauser variations in the biometric data. This is important to preserve reasonable acceptance rate of genuine queries and the overall feasibility of the recognition system. However, such acceptance of multiple similar measurements decreases the imposter's difficulty of obtaining a system-acceptable measurement, thus resulting in a degraded security level. This deteriorated security needs to be measurable to provide truthful security assurance to the users. Entropy is a standard measure of security. However, the entropy formula is applicable only when there is a single acceptable possibility. In this paper, we develop an entropy-measuring model for biometric systems that accepts multiple similar measurements per user. Based on the idea of guessing entropy, the proposed model quantifies biometric system security in terms of adversarial guessing effort for two practical attacks. Excellent agreement between analytic and experimental simulation-based measurement results on a synthetic and a benchmark face dataset justify the correctness of our model and thus the feasibility of the proposed entropy-measuring approach.
Meng-Hui Lim, Pong C. Yuen
IEEE Trans. Cybern.2
2015 Online Dictionary Learning on Symmetric Positive Definite Manifolds with Vision Applications
abstract
Symmetric Positive Definite (SPD) matrices in the form of region covariances are considered rich descriptors for images and videos. Recent studies suggest that exploiting the Riemannian geometry of the SPD manifolds could lead to improved performances for vision applications. For tasks involving processing large-scale and dynamic data in computer vision, the underlying model is required to progressively and efficiently adapt itself to the new and unseen observations. Motivated by these requirements, this paper studies the problem of online dictionary learning on the SPD manifolds. We make use of the Stein divergence to recast the problem of online dictionary learning on the manifolds to a problem in Reproducing Kernel Hilbert Spaces, for which, we develop efficient algorithms by taking into account the geometric structure of the SPD manifolds. To our best knowledge, our work is the first study that provides a solution for online dictionary learning on the SPD manifolds. Empirical results on both large-scale image classification task and dynamic video processing tasks validate the superior performance of our approach as compared to several state-of-the-art algorithms.
Shengping Zhang, Shiva Prasad Kasiviswanathan, Pong C. Yuen, Mehrtash Harandi
AAAI3
2015 Modeling spatial relations of human body parts for indexing and retrieving close character interactions
abstract
Retrieving pre-captured human motion for analyzing and synthesizing virtual character movement have been widely used in Virtual Reality (VR) and interactive computer graphics applications. In this paper, we propose a new human pose representation, called Spatial Relations of Human Body Parts (SRBP), to represent spatial relations between body parts of the subject(s), which intuitively describes how much the body parts are interacting with each other. Since SRBP is computed from the local structure (i.e. multiple body parts in proximity) of the pose instead of the information from individual or pairwise joints as in previous approaches, the new representation is robust to minor variations of individual joint location. Experimental results show that SRBP outperforms the existing skeleton-based motion retrieval and classification approaches on benchmark databases.
Edmond S. L. Ho, Jacky C. P. Chan, Yiu-Ming Cheung, Pong C. Yuen
VRST4
2015 Text string detection for loosely constructed characters with arbitrary orientations
Jian-Huang Lai, Pong C. Yuen
Neurocomputing3
2015 Learning Compact Binary Codes for Hash-Based Fingerprint Indexing
abstract
Compact binary codes can in general improve the speed of searches in large-scale applications. Although fingerprint retrieval was studied extensively with real-valued features, only few strategies are available for search in Hamming space. In this paper, we propose a theoretical framework for systematically learning compact binary hash codes and develop an integrative approach to hash-based fingerprint indexing. Specifically, we build on the popular minutiae cylinder code (MCC) and are inspired by observing that the MCC bit-based representation is bit-correlated. Accordingly, we apply the theory of Markov random field to model bit correlations in MCC. This enables us to learn hash bits from a generalized linear model whose maximum likelihood estimates can be conveniently obtained using established algorithms. We further design a hierarchical fingerprint indexing scheme for binary hash codes. Under the new framework, the code length can be significantly reduced from 384 to 24 bits for each minutiae representation. Statistical experiments on public fingerprint databases demonstrate that our proposed approach can significantly improve the search accuracy of the benchmark MCC-based indexing scheme. The binary hash codes can achieve a significant search speedup compared with the MCC bit-based representation.
Yi Wang 0017, Yiu-Ming Cheung, Pong C. Yuen
IEEE Trans. Inf. Forensics Secur.4
2015 Joint Sparse Representation and Robust Feature-Level Fusion for Multi-Cue Visual Tracking
abstract
Visual tracking using multiple features has been proved as a robust approach because features could complement each other. Since different types of variations such as illumination, occlusion, and pose may occur in a video sequence, especially long sequence videos, how to properly select and fuse appropriate features has become one of the key problems in this approach. To address this issue, this paper proposes a new joint sparse representation model for robust feature-level fusion. The proposed method dynamically removes unreliable features to be fused for tracking by using the advantages of sparse representation. In order to capture the non-linear similarity of features, we extend the proposed method into a general kernelized framework, which is able to perform feature fusion on various kernel spaces. As a result, robust tracking performance is obtained. Both the qualitative and quantitative experimental results on publicly available videos show that the proposed method outperforms both sparse representation-based and fusion based-trackers.
Xiangyuan Lan, Andy Jinhua Ma, Pong C. Yuen, Rama Chellappa
IEEE Trans. Image Process.3
2015 Cross-Domain Person Reidentification Using Domain Adaptation Ranking SVMs
abstract
This paper addresses a new person reidentification problem without label information of persons under nonoverlapping target cameras. Given the matched (positive) and unmatched (negative) image pairs from source domain cameras, as well as unmatched (negative) and unlabeled image pairs from target domain cameras, we propose an adaptive ranking support vector machines (AdaRSVMs) method for reidentification under target domain cameras without person labels. To overcome the problems introduced due to the absence of matched (positive) image pairs in the target domain, we relax the discriminative constraint to a necessary condition only relying on the positive mean in the target domain. To estimate the target positive mean, we make use of all the available data from source and target domains as well as constraints in person reidentification. Inspired by adaptive learning methods, a new discriminative model with high confidence in target positive mean and low confidence in target negative image pairs is developed by refining the distance model learnt from the source domain. Experimental results show that the proposed AdaRSVM outperforms existing supervised or unsupervised, learning or non-learning reidentification methods without using label information in target cameras. Moreover, our method achieves better reidentification performance than existing domain adaptation methods derived under equal conditional probability assumption.
Andy Jinhua Ma, Jiawei Li 0003, Pong C. Yuen, Ping Li 0001
IEEE Trans. Image Process.3
2014 Multi-cue Visual Tracking Using Robust Feature-Level Fusion Based on Joint Sparse Representation
abstract
The use of multiple features for tracking has been proved as an effective approach because limitation of each feature could be compensated. Since different types of variations such as illumination, occlusion and pose may happen in a video sequence, especially long sequence videos, how to dynamically select the appropriate features is one of the key problems in this approach. To address this issue in multi-cue visual tracking, this paper proposes a new joint sparse representation model for robust feature-level fusion. The proposed method dynamically removes unreliable features to be fused for tracking by using the advantages of sparse representation. As a result, robust tracking performance is obtained. Experimental results on publicly available videos show that the proposed method outperforms both existing sparse representation based and fusion-based trackers.
Xiangyuan Lan, Andy Jinhua Ma, Pong C. Yuen
CVPR3
2014 Fingerprint Geometric Hashing Based on Binary Minutiae Cylinder Codes
abstract
Identity management has become increasingly more difficult with biometric big data. Hash-based indexing methods are promising for efficient searches in the high-dimensional space. Geometric hashing is one of the popular methods and has seen many of its variants proposed in the literature for fingerprint identification. Most of them use the same real-valued measures of local geometric invariants for both index creation and feature comparison. In this paper, we propose to build a 3D geometric hash table for storing binary minutiae cylinder codes with access keys that collectively describe the global geometric configuration. The proposed scheme is more robust against sample noise and distortion, and its most computation intensive part can be done efficiently in Hamming space. We perform fingerprint indexing experiments on the public benchmark databases of FVC2002 DB1 and NIST DB14. The results show that the performance of our approach can converge faster to high hit rates with lower penetration rates compared to other hash-based fingerprint indexing methods.
Yi Wang 0017, Yiu-Ming Cheung, Pong C. Yuen
ICPR4
2014 Reduced Analytic Dependency Modeling: Robust Fusion for Visual Recognition
Andy Jinhua Ma, Pong C. Yuen
Int. J. Comput. Vis.2
2014 Masquerade attack on transform-based binary-template protection based on perceptron learning
Yi C. Feng, Meng-Hui Lim, Pong C. Yuen
Pattern Recognit.3
2014 Face hallucination with imprecise-alignment using iterative sparse representation
Jian-Huang Lai, Pong C. Yuen, Wilman W. W. Zou
Pattern Recognit.3
2013 Domain Transfer Support Vector Ranking for Person Re-identification without Target Camera Label Information
abstract
This paper addresses a new person re-identification problem without the label information of persons under non-overlapping target cameras. Given the matched (positive) and unmatched (negative) image pairs from source domain cameras, as well as unmatched (negative) image pairs which can be easily generated from target domain cameras, we propose a Domain Transfer Ranked Support Vector Machines (DTRSVM) method for re-identification under target domain cameras. To overcome the problems introduced due to the absence of matched (positive) image pairs in target domain, we relax the discriminative constraint to a necessary condition only relying on the positive mean in target domain. By estimating the target positive mean using source and target domain data, a new discriminative model with high confidence in target positive mean and low confidence in target negative image pairs is developed. Since the necessary condition may not truly preserve the discriminability, multi-task support vector ranking is proposed to incorporate the training data from source domain with label information. Experimental results show that the proposed DTRSVM outperforms existing methods without using label information in target cameras. And the top 30 rank accuracy can be improved by the proposed method upto 9.40% on publicly available person re-identification datasets.
Andy Jinhua Ma, Pong C. Yuen, Jiawei Li 0003
ICCV2
2013 Topology Aware Data-Driven Inverse Kinematics
abstract
Abstract Creating realistic human movement is a time consuming and labour intensive task. The major difficulty is that the user has to edit individual joints while maintaining an overall realistic and collision free posture. Previous research suggests the use of data‐driven inverse kinematics, such that one can focus on the control of a few joints, while the system automatically composes a natural posture. However, as a common problem of kinematics synthesis, penetration of body parts is difficult to avoid in complex movements. In this paper, we propose a new data‐driven inverse kinematics framework that conserves the topology of the synthesizing postures. Our system monitors and regulates the topology changes using the Gauss Linking Integral (GUI), such that penetration can be efficiently prevented. As a result, complex motions with tight body movements, as well as those involving interaction with external objects, can be simulated with minimal manual intervention. Experimental results show that using our system, the user can create high quality human motion in real‐time by controlling a few joints using a mouse or a multi‐touch screen. The movement generated is both realistic and penetration free. Our system is best applied for interactive motion design in computer animations and games.
Edmond S. L. Ho, Hubert P. H. Shum, Yiu-Ming Cheung, Pong C. Yuen
Comput. Graph. Forum4
2013 Linear Dependency Modeling for Classifier Fusion and Feature Combination
abstract
This paper addresses the independent assumption issue in fusion process. In the last decade, dependency modeling techniques were developed under a specific distribution of classifiers or by estimating the joint distribution of the posteriors. This paper proposes a new framework to model the dependency between features without any assumption on feature/classifier distribution, and overcomes the difficulty in estimating the high-dimensional joint density. In this paper, we prove that feature dependency can be modeled by a linear combination of the posterior probabilities under some mild assumptions. Based on the linear combination property, two methods, namely, Linear Classifier Dependency Modeling (LCDM) and Linear Feature Dependency Modeling (LFDM), are derived and developed for dependency modeling in classifier level and feature level, respectively. The optimal models for LCDM and LFDM are learned by maximizing the margin between the genuine and imposter posterior probabilities. Both synthetic data and real datasets are used for experiments. Experimental results show that LCDM and LFDM with dependency modeling outperform existing classifier level and feature level combination methods under nonnormal distributions and on four real databases, respectively. Comparing the classifier level and feature level fusion methods, LFDM gives the best performance.
Andy Jinhua Ma, Pong C. Yuen, Jian-Huang Lai
IEEE Trans. Pattern Anal. Mach. Intell.2
2013 Semi-supervised metric learning via topology preserving multiple semi-supervised assumptions
Qianying Wang 0001, Pong C. Yuen, Guo-Can Feng
Pattern Recognit.2
2013 Supervised Spatio-Temporal Neighborhood Topology Learning for Action Recognition
abstract
Supervised manifold learning has been successfully applied to action recognition, in which class label information could improve the recognition performance. However, the learned manifold may not be able to well preserve both the local structure and global constraint of temporal labels in action sequences. To overcome this problem, this paper proposes a new supervised manifold learning algorithm called supervised spatio-temporal neighborhood topology learning (SSTNTL) for action recognition. By analyzing the topological characteristics in the context of action recognition, we propose to construct the neighborhood topology using both supervised spatial and temporal pose correspondence information. Employing the property in locality preserving projection (LPP), SSTNTL solves the generalized eigenvalue problem to obtain the best projections that not only separates data points from different classes, but also preserves local structures and temporal pose correspondence of sequences from the same class. Experimental results demonstrate that SSTNTL outperforms the manifold embedding methods with other topologies or local discriminant information. Moreover, compared with state-of-the-art action recognition algorithms, SSTNTL gives convincing performance for both human and gesture action recognition.
Andy Jinhua Ma, Pong C. Yuen, Wilman W. W. Zou, Jian-Huang Lai
IEEE Trans. Circuits Syst. Video Technol.2
2013 Low-Resolution Face Tracker Robust to Illumination Variations
abstract
In many practical video surveillance applications, the faces acquired by outdoor cameras are of low resolution and are affected by uncontrolled illumination. Although significant efforts have been made to facilitate face tracking or illumination normalization in unconstrained videos, the approaches developed may not be effective in video surveillance applications. This is because: 1) a low-resolution face contains limited information, and 2) major changes in illumination on a small region of the face make the tracking ineffective. To overcome this problem, this paper proposes to perform tracking in an illumination-insensitive feature space, called the gradient logarithm field (GLF) feature space. The GLF feature mainly depends on the intrinsic characteristics of a face and is only marginally affected by the lighting source. In addition, the GLF feature is a global feature and does not depend on a specific face model, and thus is effective in tracking low-resolution faces. Experimental results show that the proposed GLF-based tracker works well under significant illumination changes and outperforms many state-of-the-art tracking algorithms.
Wilman W. W. Zou, Pong C. Yuen, Rama Chellappa
IEEE Trans. Image Process.2
2012 Reduced Analytical Dependency Modeling for Classifier Fusion
Andy Jinhua Ma, Pong C. Yuen
ECCV (3)2
2012 Ulcer detection in wireless capsule endoscopy images
Lecheng Yu, Pong C. Yuen, Jian-Huang Lai
ICPR2
2012 Similarity Learning Based on Semi-Supervised Graph for Classification
abstract
Similarity measurement is crucial for classification. Based on the manifold assumption, many graph-based algorithms were developed. Almost all methods follow the k-rule or ε-rule to construct a graph, and then focus on the algorithms based on the graph. However, the graph may not represent the local structure well, and it does not fully utilize the label information yet. The local structure can be presented by the local density and the distance between the samples and their neighbors. And the graph constructed by the guidance of label information will be better approximate of the relationship of the input data. In this paper, we propose an adaptive semi-supervised graph constructing method. The similarity is learned when constructing the graph. The advantages of the similarity learned by our method include: (1) The similarity is measured along the manifold by constructing a graph; (2) nearby points and points in the same cluster share high similarity; (3) samples from the same class have higher similarity than samples from different classes. Experimental results show that using the proposed similarity for classification task could get better recognition accuracy.
Qianying Wang 0001, Pong C. Yuen, Guo-Can Feng, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.2
2012 Image super-resolution by textural context constrained visual vocabulary
Pong C. Yuen, Jian-Huang Lai
Signal Process. Image Commun.2
2012 Binary Discriminant Analysis for Generating Binary Face Template
abstract
Although biometrics is more reliable, robust and convenient than traditional methods, security and privacy concerns are growing. Biometric templates stored in databases are vulnerable to attacks if they are not protected. To solve this problem, a biometric cryptosystem approach that combines cryptography and biometrics has been proposed. Under this approach, helper data is stored in a database rather than the original reference biometric templates. The helper data is generated from the original reference biometric templates and a cryptographic key with error-correcting coding schemes. During decoding, the same cryptographic key can be released from the helper data if and only if the input query data is close enough to the reference. It is assumed that the helper data does not reveal any information about the original reference biometric templates. Thus, the biometric cryptosystem approach can protect the original reference templates. However, error-correcting coding algorithms (e.g., the fuzzy commitment scheme and fuzzy vault) normally require finite input. As most face templates are real-valued templates, a binarization scheme transforming the original real-valued face templates into binary templates is required. Most existing binarization schemes are performed in an ad hoc manner and do not consider the discriminability of the binary template. The recognition accuracy based on the binary templates is thus degraded. In view of this limitation, we propose a new binarization scheme by optimizing binary template discriminability. A novel binary discriminant analysis is developed to transform a real-valued template into a binary template. Differentiation is hard to perform in binary space and direct optimization is difficult. To solve this problem, we construct a continuous function based on the perceptron to optimize binary template discriminability. Our experimental results show that the proposed algorithm improves binary template discriminability.
Yi C. Feng, Pong C. Yuen
IEEE Trans. Inf. Forensics Secur.2
2012 Very Low Resolution Face Recognition Problem
abstract
This paper addresses the very low resolution (VLR) problem in face recognition in which the resolution of the face image to be recognized is lower than 16 × 16. With the increasing demand of surveillance camera-based applications, the VLR problem happens in many face application systems. Existing face recognition algorithms are not able to give satisfactory performance on the VLR face image. While face super-resolution (SR) methods can be employed to enhance the resolution of the images, the existing learning-based face SR methods do not perform well on such a VLR face image. To overcome this problem, this paper proposes a novel approach to learn the relationship between the high-resolution image space and the VLR image space for face SR. Based on this new approach, two constraints, namely, new data and discriminative constraints, are designed for good visuality and face recognition applications under the VLR problem, respectively. Experimental results show that the proposed SR algorithm based on relationship learning outperforms the existing algorithms in public face databases.
Wilman W. W. Zou, Pong C. Yuen
IEEE Trans. Image Process.2
2011 Similarity learning for semi-supervised multi-class boosting
abstract
In semi-supervised classification boosting, a similarity measure is demanded in order to measure the distance between samples (both labeled and unlabeled). However, most of the existing methods employed a simple metric, such as Euclidian distance, which may not be able to truly reflect the actual similarity/distance. This paper presents a novel similarity learning method based on the geodesic distance. It incorporates the manifold, margin and the density information of the data which is important in semi-supervised classification. The proposed similarity measure is then applied to a semi-supervised multi-class boosting (SSMB) algorithm. In turn, the three semi-supervised assumptions, namely smoothness, low density separation and manifold assumption, are all satisfied. We evaluate the proposed method on UCI databases. Experimental results show that the SSMB algorithm with proposed similarity measure outperforms the SSMB algorithm with Euclidian distance.
Qianying Wang 0001, Pong C. Yuen, Guo-Can Feng
ICASSP2
2011 Linear dependency modeling for feature fusion
abstract
This paper addresses the independent assumption issue in fusion process. In the last decade, dependency modeling techniques were developed under a specific distribution of classifiers. This paper proposes a new framework to model the dependency between features without any assumption on feature/classifier distribution. In this paper, we prove that feature dependency can be modeled by a linear combination of the posterior probabilities under some mild assumptions. Based on the linear combination property, two methods, namely Linear Classifier Dependency Modeling (LCDM) and Linear Feature Dependency Modeling (LFDM), are derived and developed for dependency modeling in classifier level and feature level, respectively. The optimal models for LCDM and LFDM are learned by maximizing the margin between the genuine and imposter posterior probabilities. Both synthetic data and real datasets are used for experiments. Experimental results show that LFDM outperforms all existing combination methods.
Andy Jinhua Ma, Pong C. Yuen
ICCV2
2011 Face tracking in low resolution videos under illumination variations
abstract
In practical face tracking applications, the face region is often small and affected by illumination variations. We address this problem by using a new feature, namely the Gradient-Logarithmic Field (GLF) feature, in the particle filter framework. The GLF feature is robust under illumination variations and the GLF-based tracker does not assume any model for the face being tracked and is effective in low-resolution video. Experimental results show that the proposed GFL-based tracker works well under significant illumination changes and outperforms some of the state-of-the-art algorithms.
Wilman W. W. Zou, Rama Chellappa, Pong C. Yuen
ICIP3
2011 Kernel machine-based rank-lifting regularized discriminant analysis method for face recognition
Pong C. Yuen, Xuehui Xie
Neurocomputing2
2011 Learning low-rank Mercer kernels with fast-decaying spectrum
Binbin Pan, Jian-Huang Lai, Pong C. Yuen
Neurocomputing3
2011 A Boosted Co-Training Algorithm for Human Action Recognition
abstract
This paper proposes a boosted co-training algorithm for human action recognition. To address the view-sufficiency and view-dependency issues in co-training, two new confidence measures, namely, inter-view confidence and intra-view confidence, are proposed. They are dynamically fused into a semi-supervised learning process. Mutual information is employed to quantify the inter-view uncertainty and measure the independence among respective views. Intra-view confidence is estimated from boosted hypotheses to measure the total data inconsistency of labeled data and unlabeled data. Given a small set of labeled videos and a large set of unlabeled videos, the proposed semi-supervised learning algorithm trains a classifier by maximizing the inter-view confidence and intra-view confidence, and dynamically incorporating unlabeled data into the labeled data set. To evaluate the proposed boosted co-training algorithm, eigen-action and information saliency feature vectors are employed as two input views. The KTH and Weizmann human action databases are used for experiments, average recognition accuracy of 93.2% and 99.6% are obtained, respectively.
Chang Liu 0108, Pong C. Yuen
IEEE Trans. Circuits Syst. Video Technol.2
2011 Normalization of Face Illumination Based on Large-and Small-Scale Features
abstract
A face image can be represented by a combination of large-and small-scale features. It is well-known that the variations of illumination mainly affect the large-scale features (low-frequency components), and not so much the small-scale features. Therefore, in relevant existing methods only the small-scale features are extracted as illumination-invariant features for face recognition, while the large-scale intrinsic features are always ignored. In this paper, we argue that both large-and small-scale features of a face image are important for face restoration and recognition. Moreover, we suggest that illumination normalization should be performed mainly on the large-scale features of a face image rather than on the original face image. A novel method of normalizing both the Small-and Large-scale (S&L) features of a face image is proposed. In this method, a single face image is first decomposed into large-and small-scale features. After that, illumination normalization is mainly performed on the large-scale features, and only a minor correction is made on the small-scale features. Finally, a normalized face image is generated by combining the processed large-and small-scale features. In addition, an optional visual compensation step is suggested for improving the visual quality of the normalized image. Experiments on CMU-PIE, Extended Yale B, and FRGC 2.0 face databases show that by using the proposed method significantly better recognition performance and visual results can be obtained as compared to related state-of-the-art methods.
Xiaohua Xie, Wei-Shi Zheng 0001, Jian-Huang Lai, Pong C. Yuen, Ching Y. Suen
IEEE Trans. Image Process.4
2010 Binary Discriminant Analysis for Face Template Protection
abstract
Biometric cryptosystem (BC) is a very secure approach for template protection because the stored template is encrypted. The key issues in BC approach include(i) limited capability in handling intra-class variations and (ii) binary input is required. To overcome these problems, this paper adopts the concept of discriminative analysis and develops a new binary discriminant analysis (BDA) method to convert a real valued template to a binary template. Experimental results on CMU-PIE and FRGC face databases show that the proposed BDA method outperforms existing template binarization schemes.
Yi C. Feng, Pong C. Yuen
ICPR2
2010 Learning the Relationship Between High and Low Resolution Images in Kernel Space for Face Super Resolution
abstract
This paper proposes a new nonlinear face super resolution algorithm to address an important issue in face recognition from surveillance video namely, recognition of low resolution face image with nonlinear variations. The proposed method learns the nonlinear relationship between low resolution face image and high resolution face image in (nonlinear) kerkernel feature spacenel feature space. Moreover, the discriminative term can be easily included in the proposed framework. Experimental results on CMU-PIE and FRGC v2.0 databases show that proposed method outperforms existing methods as well as the recognition based on high resolution images.
Wilman W. W. Zou, Pong C. Yuen
ICPR2
2010 Video-Shot Transition Detection: a Spatio-Temporal Saliency Approach
abstract
This paper proposes a novel approach for video-shot transition detection using spatio-temporal saliency. Both temporal and spatial information are combined to generate a saliency map, and features are available based on the change of saliency. Considering the context of shot changes, a statistical detector is constructed to determine all types of shot transitions by the minimization of the detection-error probability simultaneously under the same framework. The evaluation performed on videos of various content types demonstrates that the proposed approach outperforms a more recent method and two publicly available systems, namely VideoAnnex and VCM.
Xian Wu 0001, Jian-Huang Lai, Pong C. Yuen
Int. J. Pattern Recognit. Artif. Intell.3
2010 Human action recognition using boosted EigenActions
Chang Liu 0108, Pong C. Yuen
Image Vis. Comput.2
2010 Interactive imaging and vision - Ideas, algorithms and applications
Guoping Qiu, Pong C. Yuen
Pattern Recognit.2
2010 Discriminability and reliability indexes: Two new measures to enhance multi-image face recognition
Weiwen Zou, Pong C. Yuen
Pattern Recognit.2
2010 A hybrid approach for generating secure and discriminating face template
abstract
Biometric template protection is one of the most important issues in deploying a practical biometric system. To tackle this problem, many algorithms, that do not store the template in its original form, have been reported in recent years. They can be categorized into two approaches, namely biometric cryptosystem and transform-based. However, most (if not all) algorithms in both approaches offer a trade-off between the template security and matching performance. Moreover, we believe that no single template protection method is capable of satisfying the security and performance simultaneously. In this paper, we propose a hybrid approach which takes advantage of both the biometric cryptosystem approach and the transform-based approach. A three-step hybrid algorithm is designed and developed based on random projection, discriminability-preserving (DP) transform, and fuzzy commitment scheme. The proposed algorithm not only provides good security, but also enhances the performance through the DP transform. Three publicly available face databases, namely FERET, CMU-PIE, and FRGC, are used for evaluation. The security strength of the binary templates generated from FERET, CMU-PIE, and FRGC databases are 206.3, 203.5, and 347.3 bits, respectively. Moreover, noninvertibility analysis and discussion on data leakage of the proposed hybrid algorithm are also reported. Experimental results show that, using Fisherface to construct the input facial feature vector (face template), the proposed hybrid method can improve the recognition accuracy by 4%, 11%, and 15% on the FERET, CMU-PIE, and FRGC databases, respectively. A comparison with the recently developed random multispace quantization biohashing algorithm is also reported.
Yi C. Feng, Pong C. Yuen, Anil K. Jain 0001
IEEE Trans. Inf. Forensics Secur.2
2010 Penalized preimage learning in kernel principal component analysis
abstract
Finding the preimage of a feature vector in kernel principal component analysis (KPCA) is of crucial importance when KPCA is applied in some applications such as image preprocessing. Since the exact preimage of a feature vector in the kernel feature space, normally, does not exist in the input data space, an approximate preimage is learned and encouraging results have been reported in the last few years. However, it is still difficult to find a "good" estimation of preimage. As estimation of preimage in kernel methods is ill-posed, how to guide the preimage learning for a better estimation is important and still an open problem. To address this problem, a penalized strategy is developed in this paper, where some penalization terms are used to guide the preimage learning process. To develop an efficient penalized technique, we first propose a two-step general framework, in which a preimage is directly modeled by weighted combination of the observed samples and the weights are learned by some optimization function subject to certain constraints. Compared to existing techniques, this would also give advantages in directly turning preimage learning into the optimization of the combination weights. Under this framework, a penalized methodology is developed by integrating two types of penalizations. First, to ensure learning a well-defined preimage, of which each entry is not out of data range, convexity constraint is imposed for learning the combination weights. More insight effects of the convexity constraint are also explored. Second, a penalized function is integrated as part of the optimization function to guide the preimage learning process. Particularly, the weakly supervised penalty is proposed, discussed, and extensively evaluated along with Laplacian penalty and ridge penalty. It could be further interpreted that the learned preimage can preserve some kind of pointwise conditional mutual information. Finally, KPCA with preimage learning is applied on face image data sets in the aspects of facial expression normalization, face image denoising, recovery of missing parts from occlusion, and illumination normalization. Experimental results show that the proposed preimage learning algorithm obtains lower mean square error (MSE) and better visual quality of reconstructed images.
Wei-Shi Zheng 0001, Jian-Huang Lai, Pong C. Yuen
IEEE Trans. Neural Networks3
2009 Interpolatory Mercer kernel construction for kernel direct LDA on face recognition
abstract
This paper proposes a novel methodology on Mercer kernel construction using interpolatory strategy. Based on a given symmetric and positive semi-definite matrix (Gram matrix) and Cholesky decomposition, it first constructs a nonlinear mapping Φ, which is well-defined on the training data. This mapping is then extended to the whole input feature space by utilizing Lagrange interpolatory basis functions. The kernel function constructed by inner product is proven to be a Mercer kernel function. The self-constructed interpolatory Mercer (IM) kernel keeps the Gram matrix unchanged on the training samples. To evaluate the performance of the proposed IM kernel, a popular kernel direct linear discriminant analysis (KDDA) method for face recognition is selected. Comparing with RBF kernel based KDDA method on two face databases, namely FERET and CMU PIE databases, the IM kernel based KDDA approach could increase the performance by around 20% on CMU PIE database.
Pong C. Yuen
ICASSP2
2009 Object motion detection using information theoretic spatio-temporal saliency
Chang Liu 0108, Pong C. Yuen, Guoping Qiu
Pattern Recognit.2
2009 Perturbation LDA: Learning the difference between the class empirical mean and its expectation
Wei-Shi Zheng 0001, Jian-Huang Lai, Pong C. Yuen, Stan Z. Li
Pattern Recognit.3
2009 Learning Kernel in Kernel-Based LDA for Face Recognition Under Illumination Variations
abstract
Kernel-based methods have been proved to be an effective approach for face recognition in dealing with complex and nonlinear face image variations. While many encouraging results have been reported, the selection of kernel is ratheradhoc. This letter proposes a systematic method to construct a new kernel for kernel discriminant analysis, which is good for handling illumination problem. The proposed method first learns a kernel matrix by maximizing the difference between inter-class and intra-class similarities under the Lambertian model, and then generalizes the kernel matrix to our proposed ILLUM kernel using the scattered data interpolation technique. Experiments on the Yale-B and the CMU PIE face databases show that, the proposed kernel outperforms the popular Gaussian kernel in Kernel Discriminant Analysis and the recognition rate can be improved around 10%.
Xiaozhang Liu, Pong C. Yuen, Guo-Can Feng
IEEE Signal Process. Lett.3
2008 Face illumination normalization on large and small scale features
abstract
It is well known that the effect of illumination is mainly on the large-scale features (low-frequency components) of a face image. In solving the illumination problem for face recognition, most (if not all) existing methods either only use extracted small-scale features while discard large-scale features, or perform normalization on the whole image. In the latter case, small-scale features may be distorted when the large-scale features are modified. In this paper, we argue that large-scale features of face image are important and contain useful information for face recognition as well as visual quality of normalized image. Moreover, this paper suggests that illumination normalization should mainly perform on large-scale features of face image rather than the whole face image. Along this line, a novel framework for face illumination normalization is proposed. In this framework, a single face image is first decomposed into large- and small- scale feature images using logarithmic total variation (LTV) model. After that, illumination normalization is performed on large-scale feature image while small-scale feature image is smoothed. Finally, a normalized face image is generated by combination of the normalized large-scale feature image and smoothed small-scale feature image. CMU PIE and (Extended) YaleB face databases with different illumination variations are used for evaluation and the experimental results show that the proposed method outperforms existing methods.
Xiaohua Xie, Wei-Shi Zheng 0001, Jian-Huang Lai, Pong C. Yuen
CVPR4
2008 Boosting EigenActions: A new algorithm for human action categorization
abstract
This paper proposes a boosting EigenActions algorithm for human action categorization. In determining the EigenActions, a spatio-temporal information saliency is first calculated from the video sequence by estimating pixel density function. Since human action can be approximated as a periodic motion, salient action unit, which is one cycle of the motion, is extracted and EigenActions are determined using principle component analysis. A human action classifier is developed by multi-class Adaboost algorithm. Weizmann human action database with ninety different human actions is used to evaluate our proposed algorithm. The recognition accuracy is 98.3%. A comparison with two latest methods on human action recognition is also reported.
Chang Liu 0108, Pong C. Yuen
FG2
2008 Learning a discriminative sparse tri-value transform
abstract
Simple binary patterns have been successfully used for extracting feature representations for visual object classification. In this paper, we present a method to learn a set of discriminative tri-value patterns for projecting high dimensional raw visual inputs into a low dimensional subspace for tasks such as face detection. Unlike previous methods that use predefined simple transform bases to generate tens of thousands features first and then use machine learning to select the most useful features, our method attempts to learn discriminative transform bases directly. Since it would be extremely hard to develop analytical solutions, we define an objective function that can be solved using simulated annealing. To reduce the search space, we impose sparseness and smoothness constraints on the transform bases. Experimental results demonstrate that our method is effective and provides an alternative approach to effective visual object classification.
Zhenhua Qu, Guoping Qiu, Pong C. Yuen
ICPR3
2008 Kernel parameter optimization of Kernel-based LDA methods
abstract
Kernel approach has been employed to solve classification problem with complex distribution by mapping the input space to higher dimensional feature space. However, one of the crucial factors in the kernel approach is the choosing of kernel parameters which highly affect the performance and stability of the kernel-based learning methods. In view of this limitation, this paper adopts the eigenvalue stability bounded margin maximization (ESBMM) algorithm to automatically tune the multiple kernel parameters for kernel-based LDA methods. To demonstrate its effectiveness, the ESBMM algorithm has been extended and applied on two existing kernel-based LDA methods. Experimental results show that after applying the ESBMM algorithm, the performance of these two methods are both improved.
Jian Huang 0009, Pong C. Yuen, Jun Zhang 0003, Jian-Huang Lai
IJCNN3
2008 Incremental Linear Discriminant Analysis for Face Recognition
abstract
Dimensionality reduction methods have been successfully employed for face recognition. Among the various dimensionality reduction algorithms, linear (Fisher) discriminant analysis (LDA) is one of the popular supervised dimensionality reduction methods, and many LDA-based face recognition algorithms/systems have been reported in the last decade. However, the LDA-based face recognition systems suffer from the scalability problem. To overcome this limitation, an incremental approach is a natural solution. The main difficulty in developing the incremental LDA (ILDA) is to handle the inverse of the within-class scatter matrix. In this paper, based on the generalized singular value decomposition LDA (LDA/GSVD), we develop a new ILDA algorithm called GSVD-ILDA. Different from the existing techniques in which the new projection matrix is found in a restricted subspace, the proposed GSVD-ILDA determines the projection matrix in full space. Extensive experiments are performed to compare the proposed GSVD-ILDA with the LDA/GSVD as well as the existing ILDA methods using the face recognition technology face database and the Carneggie Mellon University Pose, Illumination, and Expression face database. Experimental results show that the proposed GSVD-ILDA algorithm gives the same performance as the LDA/GSVD with much smaller computational complexity. The experimental results also show that the proposed GSVD-ILDA gives better classification performance than the other recently proposed ILDA algorithms.
Haitao Zhao 0002, Pong C. Yuen
IEEE Trans. Syst. Man Cybern. Part B2
2007 Class-Distribution Preserving Transform for Face Biometric Data Security
abstract
This paper addresses the face data variations problem in biometric cryptosystems in which the cryptographic technique is applied to biometric system. To overcome the limitation, this paper introduce a new class-distribution preserving transform to biometric cryptosystems. The basic idea is to transform a real value face feature vector to a binary feature vector using a random points set. The proposed transform is integrated into a BCH coding technique. Fisherface algorithm is used for feature extraction and ORL face database is selected for experiments. It is shown that only around 0.8% accuracy is degraded in comparing with the original Fisherface algorithm while the system security can be enhanced by 126 bits.
Yi C. Feng, Pong C. Yuen
ICASSP (2)2
2007 Face Recognition by Regularized Discriminant Analysis
abstract
When the feature dimension is larger than the number of samples the small sample-size problem occurs. There is great concern about it within the face recognition community. We point out that optimizing the Fisher index in linear discriminant analysis does not necessarily give the best performance for a face recognition system. We propose a new regularization scheme. The proposed method is evaluated using the Olivetti Research Laboratory database, the Yale database, and the Feret database.
Dao-Qing Dai, Pong C. Yuen
IEEE Trans. Syst. Man Cybern. Part B2
2007 Choosing Parameters of Kernel Subspace LDA for Recognition of Face Images Under Pose and Illumination Variations
abstract
This paper addresses the problem of automatically tuning multiple kernel parameters for the kernel-based linear discriminant analysis (LDA) method. The kernel approach has been proposed to solve face recognition problems under complex distribution by mapping the input space to a high-dimensional feature space. Some recognition algorithms such as the kernel principal components analysis, kernel Fisher discriminant, generalized discriminant analysis, and kernel direct LDA have been developed in the last five years. The experimental results show that the kernel-based method is a good and feasible approach to tackle the pose and illumination variations. One of the crucial factors in the kernel approach is the selection of kernel parameters, which highly affects the generalization capability and stability of the kernel-based learning methods. In view of this, we propose an eigenvalue-stability-bounded margin maximization (ESBMM) algorithm to automatically tune the multiple parameters of the Gaussian radial basis function kernel for the kernel subspace LDA (KSLDA) method, which is developed based on our previously developed subspace LDA method. The ESBMM algorithm improves the generalization capability of the kernel-based LDA method by maximizing the margin maximization criterion while maintaining the eigenvalue stability of the kernel-based LDA method. An in-depth investigation on the generalization performance on pose and illumination dimensions is performed using the YaleB and CMU PIE databases. The FERET database is also used for benchmark evaluation. Compared with the existing PCA-based and LDA-based methods, our proposed KSLDA method, with the ESBMM kernel parameter estimation algorithm, gives superior performance.
Jian Huang 0009, Pong C. Yuen, Jian-Huang Lai
IEEE Trans. Syst. Man Cybern. Part B2
2007 Human Face Image Searching System Using Sketches
abstract
This paper reports a human face image searching system using sketches. A two-phase method, namely, sketch-to-mug-shot matching and human face image searching using relevance feedback, is designed and developed. In the sketch-to-mug-shot matching phase, we have developed a facial feature matching algorithm using local and global features. A point distribution model is employed to represent local facial features while the global feature consists of a set of the geometrical relationship between facial features. It is found that the performance of the sketch-to-mug-shot matching is good if the sketch image looks like the mug shot image in the database. However, in some situations, it is hard to construct a sketch that looks like the photograph. To overcome this limitation, this paper makes use of the concept of ldquohuman-in-the-looprdquo and proposes a human face image searching algorithm using relevance feedback in the second phase. Positive and negative samples will be collected from the user. A feedback algorithm that employs subspace linear discriminant analysis for online learning of the optimal projection for face representation is then designed and developed. The proposed system has been evaluated using the FERET database and a Japanese database with hundreds of individuals. The results are encouraging.
Pong C. Yuen, C. H. Man
IEEE Trans. Syst. Man Cybern. Part A1
2006 Two-step single parameter regularization fisher discriminant method for face recognition
abstract
In face recognition tasks, Fisher discriminant analysis (FDA) is one of the promising methods for dimensionality reduction and discriminant feature extraction. The objective of FDA is to find an optimal projection matrix, which maximizes the between-class-distance and simultaneously minimizes within-class-distance. The main limitation of traditional FDA is the so-called Small Sample Size (3S) problem. It induces that the within-class scatter matrix is singular and then the traditional FDA fails to perform directly for pattern classification. To overcome 3S problem, this paper proposes a novel two-step single parameter regularization Fisher discriminant (2SRFD) algorithm for face recognition. The first semi-regularized step is based on a rank lifting theorem. This step adjusts both the projection directions and their corresponding weights. Our previous three-to-one parameter regularized technique is exploited in the second stage, which just changes the weights of projection directions. It is shown that the final regularized within-class scatter matrix approaches the original within-class scatter matrix as the single parameter tends to zero. Also, our method has good computational complexity. The proposed method has been tested and evaluated with three public available databases, namely ORL, CMU PIE and FERET face databases. Comparing with existing state-of-the-art FDA-based methods in solving the S3 problem, the proposed 2SRFD approach gives the best performance.
Pong C. Yuen, Jian Huang 0009, Bin Fang 0001
Int. J. Pattern Recognit. Artif. Intell.2
2006 Face and Eye Detection from Head and Shoulder Image on Mobile Devices
abstract
With the advance of semiconductor technology, the current mobile devices support multimodal input and multimedia output. In turn, human computer communication applications can be developed in mobile devices such as mobile phone and PDA. This paper addresses the research issues of face and eye detection on mobile devices. The major obstacles that we need to overcome are the relatively low processor speed, low storage memory and low image (CMOS senor) quality. To solve these problems, this paper proposes a novel and efficient method for face and eye detection. The proposed method is based on color information because the computation time is small. However, the color information is sensitive to the illumination changes. In view of this limitation, this paper proposes an adaptive Illumination Insensitive (AI2) Algorithm, which dynamically calculates the skin color region based on an image color distribution. Moreover, to solve the strong sunlight effect, which turns the skin color pixel into saturation, a dual-color-space model is also developed. Based on AI2algorithm and face boundary information, face region is located. The eye detection method is based on an average integral of density, projection techniques and Gabor filters. To quantitatively evaluate the performance of the face and eye detection, a new metric is proposed. 2158 head & shoulder images captured under uncontrolled indoor and outdoor lighting conditions are used for evaluation. The accuracy in face detection and eye detection are 98% and 97% respectively. Moreover, the average computation time of one image using Matlab code in Pentium III 700MHz computer is less than 15 seconds. The computational time will be reduced to tens hundreds of millisecond (ms) if low level programming language is used for implementation. The results are encouraging and show that the proposed method is suitable for mobile devices.
Jian-Huang Lai, Pong C. Yuen
Int. J. Pattern Recognit. Artif. Intell.2
2006 A novel incremental principal component analysis and its application for face recognition
abstract
Principal component analysis (PCA) has been proven to be an efficient method in pattern recognition and image analysis. Recently, PCA has been extensively employed for face-recognition algorithms, such as eigenface and fisherface. The encouraging results have been reported and discussed in the literature. Many PCA-based face-recognition systems have also been developed in the last decade. However, existing PCA-based face-recognition systems are hard to scale up because of the computational cost and memory-requirement burden. To overcome this limitation, an incremental approach is usually adopted. Incremental PCA (IPCA) methods have been studied for many years in the machine-learning community. The major limitation of existing IPCA methods is that there is no guarantee on the approximation error. In view of this limitation, this paper proposes a new IPCA method based on the idea of a singular value decomposition (SVD) updating algorithm, namely an SVD updating-based IPCA (SVDU-IPCA) algorithm. In the proposed SVDU-IPCA algorithm, we have mathematically proved that the approximation error is bounded. A complexity analysis on the proposed method is also presented. Another characteristic of the proposed SVDU-IPCA algorithm is that it can be easily extended to a kernel version. The proposed method has been evaluated using available public databases, namely FERET, AR, and Yale B, and applied to existing face-recognition algorithms. Experimental results show that the difference of the average recognition accuracy between the proposed incremental method and the batch-mode method is less than 1%. This implies that the proposed SVDU-IPCA method gives a close approximation to the batch-mode PCA method.
Haitao Zhao 0002, Pong C. Yuen, James T. Kwok
IEEE Trans. Syst. Man Cybern. Part B2
2005 A Novel Regularized Fisher Discriminant Method for Face Recognition Based on Subspace and Rank Lifting Scheme
Pong C. Yuen, Jian Huang 0009, Jian-Huang Lai, Jianliang Tang
ACII2
2005 A new regularized linear discriminant analysis method to solve small sample size problems
abstract
This paper presents a new regularization technique to deal with the small sample size (S3) problem in linear discriminant analysis (LDA) based face recognition. Regularization on the within-class scatter matrix Sw has been shown to be a good direction for solving the S3 problem because the solution is found in full space instead of a subspace. The main limitation in regularization is that a very high computation is required to determine the optimal parameters. In view of this limitation, this paper re-defines the three-parameter regularization on the within-class scatter matrix [Formula: see text], which is suitable for parameter reduction. Based on the new definition of [Formula: see text], we derive a single parameter (t) explicit expression formula for determining the three parameters and develop a one-parameter regularization on the within-class scatter matrix. A simple and efficient method is developed to determine the value of t. It is also proven that the new regularized within-class scatter matrix [Formula: see text] approaches the original within-class scatter matrix Sw as the single parameter tends to zero. A novel one-parameter regularization linear discriminant analysis (1PRLDA) algorithm is then developed. The proposed 1PRLDA method for face recognition has been evaluated with two public available databases, namely ORL and FERET databases. The average recognition accuracies of 50 runs for ORL and FERET databases are 96.65% and 94.00%, respectively. Comparing with existing LDA-based methods in solving the S3 problem, the proposed 1PRLDA method gives the best performance.
Pong C. Yuen, Jian Huang 0009
Int. J. Pattern Recognit. Artif. Intell.2
2005 Universal writing model for recovery of writing sequence of static handwriting images
abstract
Online features have been proven to be more robust information for handwriting recognition than an offline static image due to dynamic aspects, such as the writing sequence of strokes. The estimation of temporal information from a static image becomes an important issue. This paper presents a new statistical method to reconstruct the writing order of a handwritten signature from a two-dimensional static image. The reconstruction process consists of two phases, namely the training phase and the testing phase. In the training phase, the writing order with other attributes, such as length and direction, are extracted and analyzed from a set of training online handwritten signatures. A Universal Writing Model (UWM), which consists of a set of distribution functions, is then constructed. In the testing phase, the UWM is applied to reconstruct the writing order of an offline signature. 300 offline signatures with ground truth are used for evaluation. Experimental results show that about one-eighth of the reconstructed writing sequences are the same as the actual writing sequences.
Kai Kwong Lau, Pong C. Yuen, Yuan Yan Tang
Int. J. Pattern Recognit. Artif. Intell.2
2005 Optimal Subspace Analysis for Face Recognition
abstract
Fisher Linear Discriminant Analysis (LDA) has been successfully used as a data discriminantion technique for face recognition. This paper has developed a novel subspace approach in determining the optimal projection. This algorithm effectively solves the small sample size problem and eliminates the possibility of losing discriminative information. Through the theoretical derivation, we compared our method with the typical PCA-based LDA methods, and also showed the relationship between our new method and perturbation-based method. The feasibility of the new algorithm has been demonstrated by comprehensive evaluation and comparison experiments with existing LDA-based methods.
Haitao Zhao 0002, Pong C. Yuen, Jing-Yu Yang 0001
Int. J. Pattern Recognit. Artif. Intell.2
2005 Directed connection measurement for evaluating reconstructed stroke sequence in handwriting images
Kai Kwong Lau, Pong C. Yuen, Yuan Yan Tang
Pattern Recognit.2
2005 Kernel machine-based one-parameter regularized Fisher discriminant method for face recognition
abstract
This paper addresses two problems in linear discriminant analysis (LDA) of face recognition. The first one is the problem of recognition of human faces under pose and illumination variations. It is well known that the distribution of face images with different pose, illumination, and face expression is complex and nonlinear. The traditional linear methods, such as LDA, will not give a satisfactory performance. The second problem is the small sample size (S3) problem. This problem occurs when the number of training samples is smaller than the dimensionality of feature vector. In turn, the within-class scatter matrix will become singular. To overcome these limitations, this paper proposes a new kernel machine-based one-parameter regularized Fisher discriminant (K1PRFD) technique. K1PRFD is developed based on our previously developed one-parameter regularized discriminant analysis method and the well-known kernel approach. Therefore, K1PRFD consists of two parameters, namely the regularization parameter and kernel parameter. This paper further proposes a new method to determine the optimal kernel parameter in RBF kernel and regularized parameter in within-class scatter matrix simultaneously based on the conjugate gradient method. Three databases, namely FERET, Yale Group B, and CMU PIE, are selected for evaluation. The results are encouraging. Comparing with the existing LDA-based methods, the proposed method gives superior results.
Pong C. Yuen, Jian Huang 0009, Dao-Qing Dai
IEEE Trans. Syst. Man Cybern. Part B2
2005 GA-fisher: a new LDA-based face recognition algorithm with selection of principal components
abstract
This paper addresses the dimension reduction problem in Fisherface for face recognition. When the number of training samples is less than the image dimension (total number of pixels), the within-class scatter matrix (Sw) in Linear Discriminant Analysis (LDA) is singular, and Principal Component Analysis (PCA) is suggested to employ in Fisherface for dimension reduction of Sw so that it becomes nonsingular. The popular method is to select the largest nonzero eigenvalues and the corresponding eigenvectors for LDA. To attenuate the illumination effect, some researchers suggested removing the three eigenvectors with the largest eigenvalues and the performance is improved. However, as far as we know, there is no systematic way to determine which eigenvalues should be used. Along this line, this paper proposes a theorem to interpret why PCA can be used in LDA and an automatic and systematic method to select the eigenvectors to be used in LDA using a Genetic Algorithm (GA). A GA-PCA is then developed. It is found that some small eigenvectors should also be used as part of the basis for dimension reduction. Using the GA-PCA to reduce the dimension, a GA-Fisher method is designed and developed. Comparing with the traditional Fisherface method, the proposed GA-Fisher offers two additional advantages. First, optimal bases for dimensionality reduction are derived from GA-PCA. Second, the computational efficiency of LDA is improved by adding a whitening procedure after dimension reduction. The Face Recognition Technology (FERET) and Carnegie Mellon University Pose, Illumination, and Expression (CMU PIE) databases are used for evaluation. Experimental results show that almost 5 % improvement compared with Fisherface can be obtained, and the results are encouraging.
Wei-Shi Zheng 0001, Jian-Huang Lai, Pong C. Yuen
IEEE Trans. Syst. Man Cybern. Part B3
2004 Incremental PCA based face recognition
abstract
In the real world, learning is often expected to be a continuous process, which is capable of incorporating new facts into the past experience. However, currently many typical face recognition methods, such as eigenface and Fisherface, have only focused on non-incremental learning tasks, where the learning stops once the training set has been duly processed. In this paper, we present a PCA-based algorithm for face recognition, which takes the incremental learning in account. This method can update the principal subspace without simply re-computing the eigen decomposition from scratch.
Haitao Zhao 0002, Pong C. Yuen, James T. Kwok
ICARCV2
2004 Generalized Spectroface For Face Recognition
abstract
Spectroface is a face representation method using wavelet transform and Fourier transform and has been proved to be invariant to translation, on-the-plane rotation and scale. Two types of spectrofaces, namely first order and second order spectrofaces, have been proposed and successfully applied for face recognition. This paper reports a generalized spectroface, in which, we prove that any continuous and compact supported low-pass filters can be used to replace the wavelet transform. This feature provides a high flexibility of the spectroface representation. It is also proved that the generalized spectroface is translation invariant. A simple but effective feature selection algorithm is also proposed for the generalized spectroface in which the recognition is further increased. Three standard databases from Yale University, Olivette Research Laboratory and MIT, are used to evaluate the proposed method. The recognition accuracy is as high as 99.29%. If we consider the top three matches, the accuracy increases to 99.64%.
Jian-Huang Lai, Pong C. Yuen, Dong-Gao Deng
Int. J. Pattern Recognit. Artif. Intell.2
2003 Recovery of Writing Sequence of Static Images of Handwriting using UWM
abstract
It is generally agreed that an on-line recognition system is always reliable than an off-line one. It is due to the availability of the dynamic information, especially the writing sequence of the strokes. This paper presents a new statistical method to reconstruct the writing order of a handwritten script from a two-dimensional static image. The reconstruction process consists of two phases, named the training phase and the testing phase. In the training phase, the writing order with other attributes, such as length and direction, are extracted from a set of training on-line handwritten scripts statistically to form a universal writing model (UWM). In the testing phase, UWM is applied to reconstruct the drawing order of offline handwritten scripts by finding the highest total probability. 300 off-line signatures with ground truth are used for evaluation. Experimental results show that the reconstructed writing sequence using UWM is close to the actual writing sequence. 1.
Kai Kwong Lau, Pong C. Yuen, Yuan Yan Tang
ICDAR2
2003 Regularized discriminant analysis and its application to face recognition
Dao-Qing Dai, Pong C. Yuen
Pattern Recognit.2
2002 Tongue image matching using color content
Chun-hung Li, Pong C. Yuen
Pattern Recognit.2
2002 Face representation using independent component analysis
Pong C. Yuen, Jian-Huang Lai
Pattern Recognit.1
2001 Transductive Learning: Learning Iris Data with Two Labeled Data
Chun-hung Li, Pong C. Yuen
ICANN2
2001 Semi-supervised Learning in Medical Image Database
Chun-hung Li, Pong C. Yuen
PAKDD2
2001 Multi-cues eye detection on gray intensity image
Guo-Can Feng, Pong C. Yuen
Pattern Recognit.2
2001 Face recognition using holistic Fourier invariant features
Jian-Huang Lai, Pong C. Yuen, Guo-Can Feng
Pattern Recognit.2
2000 Virtual View Face Image Synthesis Using 3D Spring-Based Face Model from a Single Image
abstract
It is known that 2D views of a person can be synthesised if the face 3D model of that person is available. This paper proposes a new method, called 3D spring-based face model (SBFM), to determine the precise face model of a person with different poses and facial expressions from a single image. The SBFM combines the concepts of generic 3D face model in computer graphics and deformable template in computer vision. Face image databases from MIT AI laboratory and Yale University are used to test our proposed method and the results are encouraging.
Guo-Can Feng, Pong C. Yuen, Jian-Huang Lai
FG2
2000 Face Recognition Based on Local Fisher Features
Dao-Qing Dai, Guo-Can Feng, Jian-Huang Lai, Pong C. Yuen
ICMI4
2000 EDT Based Tracing Maximum Thinning Algorithm on Grey Scale Images
abstract
Most of the thinning algorithm nowadays are based on bilevel images. In recognition of hand-written words, such as signatures, the use of grey-scale image is better because much data is available from the image. In this article, we propose an efficient thinning algorithm based on Euclidean distance transformation (EDT) on grey level image. The output of our algorithm consists of the skeleton of the object as well as the intersection points. Moreover, the proposed algorithm is efficient and accurate in finding the skeleton.
Kai Kwong Lau, Pong C. Yuen, Yuan Yan Tang
ICPR2
2000 View Synthesis Under Perspective Projection
abstract
This paper addresses the issue of generating a 2D view of a 3D object from its other 2D views. Linear Combination method is the typical approach to this problem. However, a 2D view cannot be represented by a linear combination of other 2D views under perspective projection. Instead, we have presented a solution under perspective projection. The proposed method also applies to the construction of virtual frontal view face image and the results are encouraging.
Guo-Can Feng, Jian-Huang Lai, Pong C. Yuen
Int. J. Pattern Recognit. Artif. Intell.3
2000 Regularized color clustering for medical image database
abstract
A regularized color clustering algorithm is proposed to solve the color clustering problem in medical image database. By incorporating both measures of cluster separability and cluster compactness, regularized color clustering allows the automatic extraction of significant color groups with varying populations. Experimental results in different color spaces show that the regularized color clustering gives superior results in extracting significant distinct/abnormal color clusters without significant increases in cluster compactness. Furthermore, results of color clustering in different color spaces show that the LUV color space is more suitable for color clustering. Methods for selecting the regularization constants have also been suggested.
Chun-hung Li, Pong C. Yuen
IEEE Trans. Medical Imaging2
2000 Recognition of head-and-shoulder face image using virtual frontal-view image
abstract
This paper addresses the problem of face recognition under varying poses. To recognize a face under different poses, one approach is to use a human face 3D model. This approach is flexible but the equipment for acquiring the 3D face image is very expensive. The second approach is view-based. However, the complexity of the system is very high, as it requires constructing a representation for each view. For a 3D rotation, construction of dozens of representations may be required. This paper proposes a new idea to transform the face with unknown pose into frontal view for recognition. To construct the virtual frontal view image, we have developed an algorithm for detecting facial landmarks, which are then used to estimate the orientation of the face. A generic 3D spring-based face model is developed to transform the unknown face image into virtual frontal-view image. Finally, a spectroface method, which is based on wavelet transform and Fourier transform, is developed to recognize the virtual frontal face image. The proposed method has been tested by 1145 face images from 85 persons with different poses, facial expressions and small occlusions. The recognition accuracy for the best match is 84.7%. If we consider the top three matches, the accuracy increases to 92.9%.
Guo-Can Feng, Pong C. Yuen
IEEE Trans. Syst. Man Cybern. Part A2
1999 A contour detection method: Initialization and contour model
Pong C. Yuen, Guo-Can Feng, J. P. Zhou
Pattern Recognit. Lett.1
1998 Human face image retrieval system for large database
abstract
Addresses the speed problem in a human face image retrieval system from a large database. A novel method based on the wavelet transform and principal component analysis (PCA) is developed and presented. The computational load of the proposed method is greatly reduced compared with the original PCA based method. Moreover, the accuracy of the proposed method is improved.
Pong C. Yuen, Guo-Can Feng, Dao-Qing Dai
ICPR1
1998 Printed Chinese Character Similarity Measurement Using Ring Projection and Distance Transform
abstract
This paper presents a new Chinese character similarity measurement method based on the ring projection algorithm and distance transform. The ring projection algorithm is used to transform a character image with two independent variables into a function of one independent variable in the ring projection space. This representation of character in the ring projection space has been proved to be in orientation and scale invariant. However, this representation will be distorted nonlinearly in the presence of noise. Therefore, common linear metrics such as Euclidean distance, cannot be applied to measure distance. To solve the nonlinear distortion problem, distance transform is proposed as a nonlinear metric. The similarity measurement is performed using the distance transformed image in the ring projection space. A number of Chinese characters are selected to evaluate the capability of the proposed measurement scheme and the results are encouraging.
Pong C. Yuen, Guo-Can Feng, Yuan Yan Tang
Int. J. Pattern Recognit. Artif. Intell.1
1998 Contour length terminating criterion for snake model
Y. Y. Wong, Pong C. Yuen, Chong Sze Tong
Pattern Recognit.2
1998 Segmented snake for contour detection
Y. Y. Wong, Pong C. Yuen, Chong Sze Tong
Pattern Recognit.2
1998 Variance projection function and its application to eye detection for human face recognition
Guo-Can Feng, Pong C. Yuen
Pattern Recognit. Lett.2
1996 A novel method for parameter estimation of digital arc
Pong C. Yuen, Guo-Can Feng
Pattern Recognit. Lett.1
1996 Brick-wall structured segmentation for interpolative vector quantization of images
W. F. Lee, Pong C. Yuen, C. K. Chan
Signal Process. Image Commun.2
1995 Local feature detector using 3-point matching and dynamic programming
Pong C. Yuen, Peter Wai-Ming Tsang
Image Vis. Comput.1
1995 Localization of dominant points for image coding
Pong C. Yuen, W. F. Lee, Peter Wai-Ming Tsang, F. K. Lam
Pattern Recognit. Lett.1
1995 Localization of dominant points for object recognition: A scale-space approach
Pong C. Yuen, Peter Wai-Ming Tsang
Signal Process.1
1994 Detection of dominant points on an object boundary: a discontinuity approach
Peter Wai-Ming Tsang, Pong C. Yuen, Kai Kwong Lam
Image Vis. Comput.2
1994 Classification of partially occluded objects using 3-point matching and distance transformation
Peter Wai-Ming Tsang, Pong C. Yuen, Kai Kwong Lam
Pattern Recognit.2
1994 Robust matching process: a dominant point approach
Pong C. Yuen, Peter Wai-Ming Tsang, Kai Kwong Lam
Pattern Recognit. Lett.1
1993 Recognition of partially occluded objects
abstract
A computer vision system for the recognition of real world image is developed and reported. The system is capable of identifying multiple overlapped objects in a scene without stringent restrictions on their size, shape and orientation. An object shape is identified by the system through the detection of selected discrete feature segments in the contour code instead of attempting to search for a complete boundary. Consequently, an object that is partially occluded can still be recognized with its remaining unmasked portion. Extraction of salient features from an unknown geometry is performed using the nonlinear elastic matching technique. This algorithm is insensitive to sizing and distortions of the feature segments, hence reducing the problems caused by the error imposed during the image capturing process. A multilayer artificial neural network is used to provide the final identification of an unknown object based on the extracted features. A case study on the recognition of handtools with different surface reflectiveness is presented as an example. Possible improvements in the performance of the system are discussed.>
Peter Wai-Ming Tsang, Pong C. Yuen
IEEE Trans. Syst. Man Cybern.2
1992 Recognition of occluded objects
Peter Wai-Ming Tsang, Pong C. Yuen, Kai Kwong Lam
Pattern Recognit.2