EDBT 2026 Demo / reviewers in the wild / expert
Gee-Sern Hsu
dblp:64/3643 · also Gee-Sern Jison Hsu, Jison Hsu
· DBLP profile ↗
50ranked-venue papers
26as first author
16since 2021 · last 2026
0000-0003-2631-0448ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 41 · 23 first-author · 13 since 2021Artificial intelligence and machine learning · 20 · 11 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 7 · 3 first-author · 2 since 2021Security and privacy · 3 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Diffusion-to-detection (D2D): A training-free framework for zero-shot anomaly detection
Jun-Yi Chen, Gee-Sern Hsu, Yu-Hsuan Chiu 0001 |
Neurocomputing | 2 |
| 2025 | Attribute-Specified Generation And Style-Transfer Diffusion For Face Recognition EnhancementabstractWe propose a face synthesis approach to augment the training data of a face recognition model for performance improvement. Our approach consists of two major modules, the Source Face Generator (SFG) and the Style-Transfer Diffusion (STD) module. The SFG, composed of a prompt maker and a Multi-Modal Diffusion Transformer (DiT) model, aims to generate diversified and unique frontal faces, each tailored to a specific ethnicity, gender, age, and country group. The STD is designed to duplicate styles from a given set of reference faces onto the source faces generated by SFG while meticulously preserving the source facial identity. Our approach not only preserves the source facial identity within specific demographic groups but also enhances the diversity of the generated faces, effectively mitigating the attribute imbalance issue associated with real-world data. Experiments show that the training data augmented by the synthetic data made by our approach leads to clear performance improvement on state-of-the-art face recognition models. Ching-Hsun Chang, Ming-Hao Lee, Yu-Hsuan Chiu 0001, Gee-Sern Hsu, Sheng-Luen Chung |
ICIP | 4 |
| 2025 | Anchor-Guided Contrastive Learning for User Identification Based On Video PreferencesabstractUser identification through aesthetic preferences has gained attention as a promising direction in social behavioral biometrics, offering a non-invasive and privacy-conscious alternative to traditional physiological identifiers. Unlike still images or single-modal inputs, video-based aesthetic preferences provide richer, temporally-aware insights into user behavior, enabling a deeper understanding of personal taste and style. Despite this potential, existing approaches often treat preference items independently and fail to capture the interrelationships within a user’s preference set. To address these limitations, this paper introduces Preference-Aware Set Encoding with Contrastive Personalization (PASE), a novel deep learning framework designed to model user identity based on structured preferred video sets. The proposed method integrates a Cross-Video Attention Encoder to learn co-preference patterns across videos, a User-Anchor Contrastive Loss to align personalized embeddings, and a Cross-Set Mixup Regularization technique to improve generalization by simulating diverse preference scenarios. Evaluation on a curated aesthetic dataset demonstrates that PASE achieves 98.38% identification accuracy, outperforming unimodal and multi-modal baselines. These findings highlight the unique advantages of leveraging video-based aesthetic information for biometric identification, particularly in applications demanding both accuracy and user-friendly privacy safeguards. Fariha Iffath, Gee-Sern Hsu, Marina L. Gavrilova |
SMC | 2 |
| 2025 | Style-Preserving Generator for Synthetic License Plate RecognitionabstractWe propose the Style-Preserving Generator (SPG) to generate synthetic license plate data to train License Plate Recognition (LPR) models, and compare the performance with the same models trained on real-world data. The proposed SPG can edit the characters on real-world license plates while maintaining their original styles, allowing synthetic license plate data to be generated with user-specified characters. We can therefore synthesize license plates with desired characters to effectively alleviate the data attribute imbalance and privacy issues associated with real-world license plates. To the best of our knowledge, this work is the first study to present the making of synthetic LP data by proposing a novel text-editing approach tailor-made for LP data, that is the proposed SPG. The SPG consists of a transformer, a source encoder, a source style encoder, a character mask decoder, a target generator, and a target discriminator. Given a source license plate image and a specified text as input, these components collaborate to compute the self- and cross-attention embeddings, predict character masks, and generate a synthetic license plate in the source style but with source characters replaced by the specified characters. We adopt a two-phase training scheme. Phase 1 training uses synthetic data only, but Phase 2 training uses synthetic and real-life data. To showcase the effectiveness of the SPG, we introduce a new benchmark dataset, the LP-2025 (License Plate 2025), which alleviates the limitations of existing datasets and presents new challenges for license plate recognition and generative models. We validate SPG performance on the LP-2025 dataset and other benchmark datasets and compare it against state-of-the-art text-editing approaches. Gee-Sern Hsu, Wei-Jun Lin, Wei-Chun Hsieh, Wei-Zhe Jian, Sheng-Luen Chung, Marina L. Gavrilova |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Pose Adapted Shape Learning for Large-Pose Face ReenactmentabstractWe propose the Pose Adapted Shape Learning (PASL) for large-pose face reenactment. The PASL framework consists of three modules, namely the Pose-Adapted face Encoder (PAE), the Cycle-consistent Shape Generator (CSG), and the Attention-Embedded Generator (AEG). Different from previous approaches that use a single face encoder for identity preservation, we propose multiple Pose-Adapted face Encodes (PAEs) to better preserve facial identity across large poses. Given a source face and a reference face, the CSG generates a recomposed shape that fuses the source identity and reference action in the shape space and meets the cycle consistency requirement. Taking the shape code and the source as inputs, the AEG learns the attention within the shape code and between the shape code and source style to enhance the generation of the desired target face. As existing benchmark datasets are inappropriate for evaluating large-pose face reenactment, we propose a scheme to compose large-pose face pairs and introduce the MPIE-LP (Large Pose) and VoxCeleb2-LP datasets as the new large-pose benchmarks. We compared our approach with state-of-the-art methods on MPIE-LP and VoxCeleb2-LP for large-pose performance and on VoxCelebl for the common scope of pose variation. Gee-Sern Hsu, Jie-Ying Zhang, Huang Yu Hsiang, Wei-Jie Hong |
CVPR | 1 |
| 2024 | Contour-Guided Context Learning for Scene Text Recognition
Wei-Chun Hsieh, Gee-Sern Hsu, Jun-Yi Chen, Moi Hoon Yap, Zi-Chun Chao |
ICPR (20) | 2 |
| 2024 | 3D Pose-Based Evaluation of the Risk of Sarcopenia
Bo-Cheng Liao, Jie-Syuan Wu, Chen-Lung Tang, Gee-Sern Hsu, Jiunn-Horng Kang |
ICPR (16) | 4 |
| 2024 | A Novel Laguerre Voronoi Diagram Token Filtering Strategy for Computer Vision TransformersabstractTransformer architectures emerged as frontrunning approaches for computer vision classification tasks. Although transformers are faster than recurrent networks, they can become time and memory-inefficient. This paper proposes a novel filtering strategy based on Laguerre Voronoi diagram-based feature similarity measure. The approach allows to choose the most influential tokens and thus to reduce redundant information processing and extra computational costs. The filtering strategy can significantly reduce computation and memory overhead of vision transformers, depending on the number of patches to process. By comparing the proposed strategy with existing token filtering strategies, we report that the proposed architecture performs better than the comparators while reducing the computational overhead. Abu Quwsar Ohi, Gee-Sern Hsu, Marina L. Gavrilova |
IV | 2 |
| 2022 | Dual-Generator Face ReenactmentabstractWe propose the Dual-Generator (DG) network for largepose face reenactment. Given a source face and a reference face as inputs, the DG network can generate an output face that has the same pose and expression as the reference face, and has the same identity as the source face. As most approaches do not particularly consider large-pose reenactment, the proposed approach addresses this issue by incorporating a 3D landmark detector into the framework and considering a loss function to capture visible local shape variation across large pose. The DG network consists of two modules, the ID-preserving Shape Generator (IDSG) and the Reenacted Face Generator (RFG). The IDSG encodes the 3D landmarks of the reference face into a reference landmark code, and encodes the source face into a source face code. The reference landmark code and the source face code are concatenated and decoded to a set of target landmarks that exhibits the pose and expression of the reference face and preserves the identity of the source face. The RFG is partially built on the StarGAN2 generator with modifications on the input and layer settings, and with a facial style encoder added in. Given the target landmarks made by the IDSG and the source face as inputs, the RFG generates the target face with the desired identity, pose and expression. We evaluate our approach on the RaFD, MPIE, VoxCeleb1, and VoxCeleb2 benchmarks and compare with state-of-the-art methods. Gee-Sern Hsu, Chun-Hung Tsai, Hung-Yi Wu |
CVPR | 1 |
| 2022 | AgeTransGAN for Facial Age Transformation with Rectified Performance Metrics
Gee-Sern Hsu, Rui-Cang Xie, Zhi-Ting Chen, Yu-Hong Lin |
ECCV (12) | 1 |
| 2022 | Graph Refinement with Regression Prior for 3D Face ReconstructionabstractWe propose the Graph Refinement with Regression Prior (GRRP) framework for 3D face reconstruction. The GRRP framework is composed of two component networks, namely the 3D Face Regressor and the Graph-Convolutional Refining (GCR) Network. The 3D Face Regressor is made of a face reconstruction pre-trained model, which is not updated during training and offers a pseudo target to guide the training of the GCR. The GCR proposed to design the graph-refining architecture with appropriate graph convolutional settings for enhancing 3D reconstruction accuracy. Apart from the network design, we also highlight the influences of two objective functions, the adaptive vertex loss and surface normal loss, and point out the difficulty of generating 3D faces of local details. The adaptive vertex loss is designed to adaptively optimize the local facial meshes so that the local undesired vertices can be rectified while the global mesh can be kept, and the surface normal loss is designed to eliminate the flying vertices problem of the generated shape surface. The proposed framework is verified through the experiments on the AFLW2000-3D and MICC Florence datasets and compared with contemporary approaches for evaluation. Chia-Hao Tang, Gee-Sern Hsu, Ting-Yu Tai |
ICIP | 2 |
| 2021 | Age-Style and Alignment Augmentation for Facial Age Estimation
Yu-Hong Lin, Chia-Hao Tang, Zhi-Ting Chen, Gee-Sern Hsu, Md. Shopon, Marina L. Gavrilova |
CAIP (2) | 4 |
| 2021 | Pose-Guided And Style-Transferred Face ReenactmentabstractUnlike most face reenactment that deals with minor pose variation, we propose the Pose-Guided Reenactment (PGR) framework for large-pose face reenactment. The proposed PGR is composed of a landmark detector, a landmark encoder, a face encoder, a landmark decoder, a style encoder, a face generator and two discriminators. The training of these components is divided into two phases. In Phase I, the landmark encoder is trained to encode the landmarks of the reference face to a reference landmark code, and the face encoder is trained to encode the source face to a source face code. The reference landmark code and the source face code are entered to the landmark decoder, which generates a target landmark set. In Phase II, the face generator is trained to generate the target face with the desired identity, pose and expression, given the target landmark set and the source face as input. To handle large pose, we include the large pose data in the training set and propose the pose-guided landmark switch to control the change of the facial landmarks during pose variation. Experiments on MPIE and VoxCeleb1 benchmark databases show that the proposed PGR can effectively reenact the faces with large pose, and delivers a state-of-the-art overall performance compared with other approaches. Gee-Sern Hsu, Hung-Yi Wu |
ICIP | 1 |
| 2021 | Multi-View Normalization For Face RecognitionabstractUnlike the majority of face normalization that focuses on single-view frontalization, we propose the Multi-View Normalization (MVN) framework to normalize an arbitrary face to multiple desired poses with balanced illumination and neutral expression. Taking the advantages of generative and adversarial learning, the proposed MVN transforms a face into a set of multi-view faces with facial identity well preserved, offering a better representation to handle face recognition. The MVN is designed to learn the transformation from a input set to seven output sets. The input set contains faces collected in the wild with arbitrary poses, lighting conditions and expressions. The seven output sets include seven poses from $0^{\circ}\sim 90^{\circ}$ in yaw with $15^{\mathrm{\circ}}$ interval with balanced illumination and neutral expression. The MVN is composed of one face encoder, seven pose-specific generators and seven sets of discriminators. The encoder is made of a face recognition expert network, which is not updated during training and acts as a facial feature extractor. The generators are trained to transform the input set to the seven output sets. The discriminators are trained to not only ensure the photo-realistic quality of the generated faces, but also force the poses of the generated faces to the desired poses with corresponding facial appearances. Experiments on several benchmark datasets show that the proposed MVN demonstrates a competitive performance to state-of-the-art approaches. Chia-Hao Tang, Yi-Mei Chou, Gee-Sern Hsu |
ICIP | 3 |
| 2021 | R-MNet: A Perceptual Adversarial Network for Image InpaintingabstractFacial image inpainting is a problem that is widely studied, and in recent years the introduction of Generative Adversarial Networks, has led to improvements in the field. Unfortunately some issues persists, in particular when blending the missing pixels with the visible ones. We address the problem by proposing a Wasserstein GAN combined with a new reverse mask operator, namely Reverse Masking Network (R-MNet), a perceptual adversarial network for image inpainting. The reverse mask operator transfers the reverse masked image to the end of the encoder-decoder network leaving only valid pixels to be inpainted. Additionally, we propose a new loss function computed in feature space to target only valid pixels combined with adversarial training. These then capture data distributions and generate images similar to those in the training data with achieved realism (realistic and coherent) on the output images. We evaluate our method on publicly available dataset, and compare with state-of-the-art methods. We show that our method is able to generalize to high-resolution inpainting task, and further show more realistic outputs that are plausible to the human visual system when compared with the state-of-the-art methods. https://github.com/Jireh-Jam/R-MNet-Inpainting-keras Jireh Jam, Connah Kendrick, Vincent Drouard, Kevin Walker, Gee-Sern Hsu, Moi Hoon Yap |
WACV | 5 |
| 2021 | A comprehensive review of past and present image inpainting methods
Jireh Jam, Connah Kendrick, Kevin Walker, Vincent Drouard, Gee-Sern Hsu, Moi Hoon Yap |
Comput. Vis. Image Underst. | 5 |
| 2020 | Contrastive Data Learning for Facial Pose and Illumination NormalizationabstractFace normalization can be a crucial step when handling generic face recognition. We propose the Pose and Illumination Normalization (PIN) framework with contrast data learning for face normalization. The PIN framework is designed to learn the transformation from a source set to a target set. The source set and the target set compose a contrastive data set for learning. The source set contains faces collected in the wild and thus covers a wide range of variation across illumination, pose, expression and other variables. The target set contains face images taken under controlled conditions and all faces are in frontal pose and balanced in illumination. The PIN framework is composed of an encoder, a decoder and two discriminators. The encoder is made of a state-of-the-art face recognition network and acts as a facial feature extractor, which is not updated during training. The decoder is trained on both the source and target sets, and aims to learn the transformation from the source set to the target set; and therefore, it can transform an arbitrary face into a illumination and pose normalized face. The discriminators are trained to ensure the photo-realistic quality of the normalized face images generated by the decoder. The loss functions employed in the decoder and discriminators are appropriately designed and weighted for yielding better normalization outcomes and recognition performance. We verify the performance of the propose framework on several benchmark databases, and compare with state-of-the-art approaches. Gee-Sern Hsu, Chia-Hao Tang, Svetlana N. Yanushkevich, Marina L. Gavrilova |
ICPR | 1 |
| 2020 | A deep learning framework for heart rate estimation from facial videos
Gee-Sern Hsu, Rui-Cang Xie, Arul-Murugan Ambikapathi, Kae-Jy Chou |
Neurocomputing | 1 |
| 2019 | Identification of Partially Occluded Pharmaceutical Blister PackagesabstractMedical dispensing refers to the in-office preparation and delivery of prescription drugs, which is mostly dispensed by the units of blister packages. The objective of the study is to design an image-based blister package identification solution, which is capable of identifying a fetched drug based on a pair of the two opposite camera images of the hand-held drug. To this aim, this paper proposes a deep learning based Hand-held Blister Identification network (HBIN) to identify partially occluded blister packages present in arbitrary positions and orientation with possibly cluttered backgrounds. The proposed HBIN is a two-stage network that contains Blister cropping network (BCN) followed by RTT identification network (RIN). The BCN subnetwork, an image to image translation deep learning network, is to crop both side contours of the hand-held drug, before the pair of cropped contours can be juxtaposed as a fixed sized and fixed orientation RTT (rectified two-sides template) for final identification in the RIN sub-network. A blister package dataset containing a total of 30,394 images based on 230 types, typically found in hospital dispensing stations, have been collected and labeled. With extensive test, the accuracy of the primitive primitive HBIN attains an F-score of more than 94.33% for testing data from similar backgrounds and an F-score of 79.80% for dissimilar backgrounds. Although still a prototype, the preliminary results show the feasibility of identifying blister packages during retrieval process without resorting to bar codes nor RFID tags. Sheng-Luen Chung, Chih-Fang Chen, Gee-Sern Hsu, Shen-Te Wu |
AVSS | 3 |
| 2019 | Face Recognition with Disentangled Facial Representation Learning and Data AugmentationabstractWe address two issues for tackling face recognition across pose, one is disentangled representation learning and the other is training data augmentation. To have better training properties, we propose the Representation-Learning Wasserstein-GAN (RL-WGAN) with three component networks for learning the disentangled facial representation. As the learning based on imbalanced data often leads to biased estimation, we proposed a data augmentation scheme that exploits the 3D Morphable Model (3DMM) for generating faces of desired poses. The RL-WGAN and the data augmentation are verified in the experiments with benchmark databases, and compared with contemporary approaches for performance evaluation. Chia-Hao Tang, Gee-Sern Hsu, Moi Hoon Yap |
ICIP | 2 |
| 2018 | Deep Hybrid Network for Automatic Quantitative Analysis of Facial ParalysisabstractWe propose the Deep Hybrid Network (DHN) for the analysis of facial paralysis syndrome. This is a pioneering work that explores the deep-learning for the facial paralysis study. The proposed DHN consists of three component networks, the first detects the subject's face, the second detects the landmarks and the edges on the detected faces, and the third detects the local paralysis regions. One novelty of this research is the exploration of facial edge features in the analysis of facial paralysis. Additionally, we introduce the first public database for facial paralysis study, as the previous studies were all evaluated on proprietary databases, making the comparison with other methods difficult. Our database includes 32 videos of 21 patients collected from YouTube. To enhance the robustness against expression variations, we include the CK+ facial expression database in the training and testing phases. We show that the proposed DHN does not just detect the local paralysis regions, but also captures the intensity of the syndrome over time, enabling the quantitative description of the syndrome. Experiments show that the proposed approach offers an accurate and efficient solution for facial paralysis analysis. Gee-Sern Hsu, Min-Hsiang Chang |
AVSS | 1 |
| 2018 | Edge-Coupled and Multi-Dropout Face AlignmentabstractWe propose the Edge-Coupled Multi-Dropout (ERN) framework for face alignment. Two features make the ERN framework an effective solution, one is the coupling of edges in the input and the other is the multiple dropout implemented at the convolution layers of the network. The core part of ERN consists of two component networks, the Edge Detection Network (EDN) and Mutiple Dropout Network (CD-VGG). Given a face, the EDN Detects the edges around the face and facial components where the facial landmarks are most likely located. The EDN output the detected facial edge and the face are then entered as input to the CD-VGG for locating the landmarks. The ERN framework also embeds a pose regressor following a face detector, making the collaboration of the EDN and CD-VGG a pose-oriented task. The ERN is tested on several benchmark databases, particularly on those with large poses, to emphasize the effectiveness for handling difficult cases. Gee-Sern Hsu, Kang-Chi Ho |
ICIP | 1 |
| 2018 | CoVieW'18: The 1st Workshop and Challenge on Comprehensive Video Understanding in the WildabstractThe 1st Workshop and Challenge on Comprehensive Video Understanding in the Wild, dubbed CoVieW'18, is held in Seoul, Korea on October 22, 2018, in conjuction with ACM Multimedia 2018. The workshop aims to solve the joint and comprehensive understanding problem in untrimmed videos with a particular emphasis on joint action and scene recognition. The workshop encourages researchers to participate in joint action and scene recognition challenge in untrimmed videos and to report their results. The workshop program includes 1 keynote speech, 2 invited speakers, 6 regular and challenge papers. The developments made in the workshop will deliver a step change in a variety of video applications. Kwanghoon Sohn, Ming-Hsuan Yang 0001, Hyeran Byun, Jongwoo Lim, Gee-Sern Hsu, Stephen Lin 0001, Euntai Kim, Seungryong Kim |
ACM Multimedia | 5 |
| 2018 | Hybrid Ageing Patterns for face age estimation
Choon-Ching Ng, Moi Hoon Yap, Yi-Tseng Cheng, Gee-Sern Hsu |
Image Vis. Comput. | 4 |
| 2018 | Robust cross-pose face recognition using landmark oriented depth warping
Gee-Sern Hsu, Arul-Murugan Ambikapathi, Sheng-Luen Chung, Hung-Cheng Shie |
J. Vis. Commun. Image Represent. | 1 |
| 2018 | Fast Landmark Localization With 3D Component Reconstruction and CNN for Cross-Pose RecognitionabstractTwo approaches are proposed for cross-pose face recognition, one is built on the handcrafted features extracted from the 3D reconstruction of facial components and the other is built on the learned features from a deep convolutional neural network (CNN). As both approaches rely on facial landmarks for alignment across large poses, we propose the Fast Hierarchical Model (FHM) for locating cross-pose facial landmarks in real time. Unlike most 3D approaches that consider holistic faces, the first proposed approach considers 3D facial components. It segments each 2D face in the gallery into components, reconstructs the 3D surface for each component, and recognizes a query face by component features. The core part of the CNN-based approach is a modified VGG network. We study the performance with different settings on the training set, including the synthesized data from 3D reconstruction, the real-life data from an in-the-wild database, and both types of data combined. The two recognition approaches and the FHM are evaluated in extensive experiments and compared with state-of-the-art methods to demonstrate their efficacy. Gee-Sern Hsu, Hung-Cheng Shie, Cheng-Hua Hsieh, Jui-Shan Chan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | Robust license plate detection in the wildabstractLicense Plate Detection (LPD) is the pivotal step for License Plate Recognition. In this work, we explore and customize state-of-the-art detection approaches for exclusively handling the LPD in the wild. In-the-wild LPD considers license plates captured in challenging conditions caused by bad weathers, lighting, traffics, and other factors. As conventional methods failed to handle these inevitable conditions, we explore the latest deep learning based detectors, namely YOLO (You-Only-Look-Once) and its variant YOLO-9000 (referred here as YOLO-2), and customize them for effectively handling the LPD. The prime customizations include modification of the grid size and of the bounding box parameter estimation, and the composition of a more challenging AOLPE (Application-Oriented License Plate Extended) database for performance evaluation. The AOLPE database is an extended version of the AOLP database [1] with additional images taken under extreme but frequently-encountered conditions. As the original YOLO and YOLO-2 are not designed for the LPD, they failed to handle the LPD on the AOLPE without the customizations. This study can be one of the pioneering works that revise state-of-the-art real-time deep networks for handling the LPD. It also serves as a case study for those who wish to customize existing deep networks for detecting specific objects. In addition to a pioneering explorations of deep networks for handling the in-the-wild LPD, our contribution also includes the release of the AOLPE database and evaluation protocol for a novel benchmark for the LPD. Gee-Sern Hsu, Arul-Murugan Ambikapathi, Sheng-Luen Chung, Cheng-Po Su |
AVSS | 1 |
| 2017 | Deep learning with time-frequency representation for pulse estimation from facial videosabstractAccurate pulse estimation is of pivotal importance in acquiring the critical physical conditions of human subjects under test, and facial video based pulse estimation approaches recently gained attention owing to their simplicity. In this work, we have endeavored to develop a novel deep learning approach as the core part for pulse (heart rate) estimation by using a common RGB camera. Our approach consists of four steps. We first begin by detecting the face and its landmarks, and thereby locate the required facial ROI. In Step 2, we extract the sample mean sequences of the R, G, and B channels from the facial ROI, and explore three processing schemes for noise removal and signal enhancement. In Step 3, the Short-Time Fourier Transform (STFT) is employed to build the 2D Time-Frequency Representations (TFRs) of the sequences. The 2D TFR enables the formulation of the pulse estimation as an image-based classification problem, which can be solved in Step 4 by a deep Con-volutional Neural Network (CNN). Our approach is one of the pioneering works for attempting real-time pulse estimation using a deep learning framework. We have developed a pulse database, called the Pulse from Face (PFF), and used it to train the CNN. The PFF database will be made publicly available to advance related research. When compared to state-of-the-art pulse estimation approaches on the standard MAHNOB-HCI database, the proposed approach has exhibited superior performance. Gee-Sern Hsu, Arul-Murugan Ambikapathi, Ming-Shiang Chen |
IJCB | 1 |
| 2017 | Cross-pose landmark localization using multi-dropout frameworkabstractWe propose the Multiple Dropout Framework (MDF) for facial landmark localization across large poses. Unlike most landmark detectors only work for poses less than 45 degree in yaw, the proposed MDF works for pose as large as 90 degree, i.e., full profile. In the proposed MDF, the Single Shot Multibox Detector (SSD) [10] is tailored for fast and precise face detection. Given an SSD detected face, a Multiple Dropout Network (MDN) is proposed to classify the face into either frontal or profile pose, and for each pose another MDN is configured for detecting pose-oriented landmarks. As the MDF framework contains one MDN (pose) classifier and two MDN (landmark) regressors, this study aims to determine the MDN structures and settings appropriate for handling classification and regression tasks. The MDN framework demonstrates the following advantages and observations. (1) Landmark detection across poses can be better approached by incorporating a pose classifier with pose-oriented landmark regressors. (2) Multiple dropouts are required for stabilizing the training of regressor networks. (3) Additional hand-crafted features, such as the Local Binary Pattern (LBP), can improve the accuracy of landmark localization. (4) Face profiling is a powerful tool for offering a large cross-pose training set. A comparison study on benchmark databases shows that the MDN delivers a competitive performance to the state-of-the-art approaches for face alignment across large poses. Gee-Sern Hsu, Cheng-Hua Hsieh |
IJCB | 1 |
| 2017 | Face recognition by facial attribute assisted networkabstractWe propose the Facial Attribute Assistant Network (FAAN) for face recognition in the wild. The FAAN is developed to imitate some general human description of a face using facial attributes, such as gender, ethnicity, hair color and with or without eyeglasses. Given a face as the input, the FAAN renders an identity descriptor (I-descriptor) and an attribute descriptor (A-descriptor) in the output. These descriptors are exploited with the Minimum Distance Subset scheme, proposed in this study, for handling image-set based recognition. When searching for a match in a gallery set, only the subjects with similar attributes as of the probe, measured by the A-descriptors, are selected for matching by comparing the I-descriptors. The FAAN is built on the Residual Network (ResNet), which is first trained for face identification and then fine tuned for attribute identification. This study shows that attributes can effectively reduce the search space and make the search more efficient and accurate. Compared with other state-of-the-art methods, the FAAN demonstrates a competitive performance on the MPIE and IJB-A benchmarks. Jui-Shan Chan, Gee-Sern Hsu, Hung-Cheng Shie, Yan-Xiang Chen |
ICIP | 2 |
| 2017 | Multi-dropout regression for wide-angle landmark localizationabstractWe propose the Multi-Dropout Regression Network (MDRN) for real-time facial landmark localization across extreme poses. Different from most landmark localization methods only work for -45° to 45° in yaw, the proposed MDRN works for the full coverage of -90° to 90°. It employs the Single Shot Multibox Detector (SSD) [1] as a preprocessor for fast and accurate face detection. Given an SSD detected face, the MDRN locates the landmarks. The MDRN is composed of 2 double-layer convolution blocks, 1 triple-layer convolution block and 3 fully-connected layers. Unlike most networks with only one dropout layer connected to the last convolution layer, in the MDRN each convolution block is followed by a max-pooling layer and a dropout layer ahead of connecting to the next processing layer. Experiments reveal that multiple dropouts better stabilize the regression and improve the accuracy of landmark localization. To locate the landmarks on profile faces and other extreme poses, the MDRN is trained on an augmented database composed of imagery of synthesized poses. A comparison study shows that the proposed solution delivers a comparable performance to the state of the art for wide-angle landmark localization. Gee-Sern Hsu, Cheng-Hua Hsieh |
ICIP | 1 |
| 2015 | Regressive Tree Structured Model for Facial Landmark LocalizationabstractAlthough the Tree Structured Model (TSM) is proven effective for solving face detection, pose estimation and landmark localization in an unified model, its sluggish run time makes it unfavorable in practical applications, especially when dealing with cases of multiple faces. We propose the Regressive Tree Structure Model (RTSM) to improve the run-time speed and localization accuracy. The RTSM is composed of two component TSMs, the coarse TSM (c-TSM) and the refined TSM (r-TSM), and a Bilateral Support Vector Regressor (BSVR). The c-TSM is built on the low-resolution octaves of samples so that it provides coarse but fast face detection. The r-TSM is built on the mid-resolution octaves so that it can locate the landmarks on the face candidates given by the c-TSM and improve precision. The r-TSM based landmarks are used in the forward BSVR as references to locate the dense set of landmarks, which are then used in the backward BSVR to relocate the landmarks with large localization errors. The forward and backward regression goes on iteratively until convergence. The performance of the RTSM is validated on three benchmark databases, the Multi-PIE, LFPW and AFW, and compared with the latest TSM to demonstrate its efficacy. Gee-Sern Hsu, Kai-Hsiang Chang, Shih-Chieh Huang |
ICCV | 1 |
| 2015 | Face detection and landmark localization using Bilayer Tree Structured ModelabstractAlthough the Tree Structured Model (TSM) is proven effective for face detection, pose estimation and landmark localization, its sluggish runtime makes it unfavorable in practical applications. We propose the Bilayer Tree Structure Model (BTSM) to improve the run-time speed while keeping the performance unchanged or slightly better. The BTSM is composed of two component TSMs, the coarse c-TSM and the refined r-TSM. The c-TSM is trained on low-resolution samples so that it can provide coarse but fast detection, The r-TSM is trained on mid-resolution samples so that it can locate precise part locations. The performance of the BTSM is validated on three benchmark databases, the Multi-PIE, LFPW and AFW, and compared with the latest TSM to demonstrate its efficacy. Gee-Sern Hsu, Kai-Hsiang Chang, Shih-Chieh Huang, Sheng-Luen Chung |
ICIP | 1 |
| 2014 | RGB-D Based Face Reconstruction and RecognitionabstractMost RGB-D based research focuses on gesture analysis, scene reconstruction and SLAM, but only few study its impacts on face recognition. A common yet challenging scenario considered in face recognition across pose takes a single 2D face of frontal pose as the galley and other poses as the probe set. We consider a similar scenario but with a RGB-D image pair taken at frontal pose in the gallery, only 2D images with a large scope of poses in the probe set, and study the advantage of the additional depth map on top of the regular RGB image. We formulate the 3D face reconstruction using the RGB-D image as a constrained optimization, and compare the results with different reconstruction settings. The reconstructed 3D face allows the generation of 2D face with specific poses, which can be matched with the probes. Experiments on the Biwi Kinect Head Pose Database and Eurecom Database show that the additional depth map substantially improves the cross-pose recognition performance, and the depth-based component selection also improves the recognition under occlusion and expression variation. Gee-Sern Hsu, Yu-Lun Liu 0002, Hsiao-Chia Peng, Sheng-Luen Chung |
ICPR | 1 |
| 2014 | A Framework for Making Face Detection Benchmark DatabasesabstractThe images in face detection benchmark databases are mostly taken by consumer cameras, and thus are constrained by popular preferences, including a frontal pose and balanced lighting conditions. A good face detector should consider beyond such constraints and work well for other types of images, for example, those captured by a surveillance camera. To overcome such constraints, a framework is proposed to transform a mother database, originally made for benchmarking face recognition, to daughter datasets that are good for benchmarking face detection. The daughter datasets can be customized to meet the requirements of various performance criteria; therefore, a face detector can be better evaluated on desired datasets. The framework is composed of two phases: 1) intrinsic parametrization and 2) extrinsic parametrization. The former parametrizes the intrinsic variables that affect the appearance of a face, and the latter parametrizes the extrinsic variables that determine how faces appear on an image. Experiments reveal that the proposed framework can generate not just data that are similar to those available from popular benchmark databases, but also those that are hardly available from existing databases. The datasets generated by the proposed framework offer the following advantages: 1) they can define the performance specification of a face detector in terms of the detection rates on variables with different variation scopes; 2) they can benchmark the performance on one single or multiple variables, which can be difficult to collect; and 3) their ground truth is available when the datasets are generated, avoiding the time-consuming manual annotation. Gee-Sern Hsu, Tsu-Ying Chu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2014 | RGB-D-Based Face Reconstruction and RecognitionabstractMost RGB-D-based research focuses on scene reconstruction, gesture analysis, and simultaneous localization and mapping, but only a few study its impacts on face recognition. A common yet challenging scenario considered in face recognition takes a single 2D face of frontal pose as the gallery and other poses as the probe set. We consider a similar scenario but with an RGB-D image pair taken at frontal pose for each subject in the gallery, only 2D images with a large scope of pose variations in the probe set, and study the advantage of the additional depth map on top of the regular RGB image. To tackle the cases with depth map corrupted by quantization noise, which are often encountered when the face is not close enough to the RGB-D camera, we propose a resurfacing approach as a preprocessing phase. We formulate the 3D face reconstruction using the RGB-D image as a constrained optimization and compare the results with different reconstruction settings. The reconstructed 3D face allows the generation of 2D face with specific poses, which can be matched against the probes. To deal with occlusion and expression variations, an automatic landmark detection algorithm is exploited to identify the parts on a given probe that are good for recognition. Experiments on benchmark databases show that the additional depth map substantially improves the cross-pose recognition performance, and the landmark-based component selection also improves the recognition under occlusion and expression variation. The performance comparison with other contemporary approaches also shows the effectiveness of the proposed approach. Gee-Sern Hsu, Yu-Lun Liu 0002, Hsiao-Chia Peng, Po-Xun Wu |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2013 | Heterogeneous feature code for expression recognitionabstractThe Heterogeneous Feature Code (HFC), a coding scheme based on both human and machine selected local features, is proposed for expression recognition. The HFC consists of two component codes, the Human Observable Code (HOC) and Boost Feature Code (BFC). The HOC is developed to capture the local deformation patches observable to humans when the face is showing an expression. Different expressions appear with a specific set of such patches with different deformation patterns at different locations, which are considered in the configuration of the HOC codewords. The BFC is built upon the mutually connected Haar-like features selected by a set of Adaboost classifiers followed by a multi-class SVM classifier. Unlike the HOC features, the BFC features can hardly be selected by human eyes. The HFC is probably the first code that combines human selected features and machine selected features, and proven effective for expression recognition. Performance evaluation on the Cohn-Kanade extension (CK+) database and the Japanese Female Facial Expression (JAFFE) shows that the HFC outperforms either HOC or BFC component code alone, and is competitive to the state-of-the-art. Gee-Sern Hsu, Shang-Min Yeh |
ICIP | 1 |
| 2013 | Facial Trait CodeabstractWe propose a facial trait code (FTC) to encode human facial images, and apply it to face recognition. Extracted from an exhaustive set of local patches cropped from a large stack of faces, the facial traits and the associated trait patterns can accurately capture the appearance of a given face. The extraction has two phases. The first phase is composed of clustering and boosting upon a training set of faces with neutral expression, even illumination, and frontal pose. The second phase focuses on the extraction of the facial trait patterns from the set of faces with variations in expression, illumination, and poses. To apply the FTC to face recognition, two types of codewords, hard and probabilistic, with different metrics for characterizing the facial trait patterns are proposed. The hard codeword offers a concise representation of a face, while the probabilistic codeword enables matching with better accuracy. Our experiments compare the proposed FTC to other algorithms on several public datasets, all showing promising results. Ping-Han Lee, Gee-Sern Hsu, Tsuhan Chen, Yi-Ping Hung |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2012 | A comparison study on appearance-based object recognition
Gee-Sern Hsu, Truong Tan Loc, Sheng-Luen Chung |
ICPR | 1 |
| 2012 | Pedestrian tracking in low contrast regions using aggregated background model and Silhouette Components
Gee-Sern Hsu, Hong Phuoc Nguyen, Chien-Hung Wu, Sheng-Luen Chung |
ICPR | 1 |
| 2012 | Color and illumination invariant dice recognitionabstractA system is proposed for automatic reading of the number of dots on dice in general table game settings. Different from previous dice recognition systems which recognize dice of a specific color using a single top-view camera in an enclosure with controlled settings, the proposed one uses multiple cameras to recognize dice of various colors posed in a wide range of viewing angle and under uncontrolled conditions. It is composed of three modules. Module-1 locates the dice using the gradient-conditioned color segmentation (GCCS), proposed in this paper, to segment dice of arbitrary colors from the background. Module-2 exploits the local invariant features good for building homographies across multiple views and lighting conditions. The homographies are used to enhance coplanar features and weaken non-coplanar features, giving a solution to segment the top faces of the dice and make up the features ruined by possible specular reflection. To identify the dots on the segmented top faces, an MSER detector is embedded in Module-3 for its consistency in locating the dot regions regardless of illumination and viewpoint variations. Experiments show that the proposed system performs satisfactorily in various test conditions. Gee-Sern Hsu, Hsiao-Chia Peng, Shang-Min Yeh |
SMC | 1 |
| 2012 | Subject-Specific and Pose-Oriented Facial Features for Face Recognition Across PosesabstractMost face recognition scenarios assume that frontal faces or mug shots are available for enrollment to the database, faces of other poses are collected in the probe set. Given a face from the probe set, one needs to determine whether a match in the database exists. This is under the assumption that in forensic applications, most suspects have their mug shots available in the database, and face recognition aims at recognizing the suspects when their faces of various poses are captured by a surveillance camera. This paper considers a different scenario: given a face with multiple poses available, which may or may not include a mug shot, develop a method to recognize the face with poses different from those captured. That is, given two disjoint sets of poses of a face, one for enrollment and the other for recognition, this paper reports a method best for handling such cases. The proposed method includes feature extraction and classification. For feature extraction, we first cluster the poses of each subject's face in the enrollment set into a few pose classes and then decompose the appearance of the face in each pose class using Embedded Hidden Markov Model, which allows us to define a set of subject-specific and pose-priented (SSPO) facial components for each subject. For classification, an Adaboost weighting scheme is used to fuse the component classifiers with SSPO component features. The proposed method is proven to outperform other approaches, including a component-based classifier with local facial features cropped manually, in an extensive performance evaluation study. Ping-Han Lee, Gee-Sern Hsu, Yun-Wen Wang, Yi-Ping Hung |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2011 | Dice Recognition in Uncontrolled Illumination Conditions by Local Invariant Features
Gee-Sern Hsu, Hsiao-Chia Peng, Chyi-Yeu Lin, Pendry Alexandra |
CAIP (2) | 1 |
| 2010 | Local Empirical Templates and Density Ratios for People Counting
Dao Huu Hung, Sheng-Luen Chung, Gee-Sern Hsu |
ACCV (4) | 3 |
| 2010 | Robust Face Recognition Using Probabilistic Facial Trait Code
Ping-Han Lee, Gee-Sern Hsu, Szu-Wei Wu, Yi-Ping Hung |
ECCV (1) | 2 |
| 2010 | Benchmark face detection using a face recognition databaseabstractA framework is proposed to generate datasets good for benchmarking face detection using database meant for benchmarking face recognition. Instead of the common way of collecting images manually, the datasets from the proposed framework are made by a synthesis process with two phases: intrinsic parameterization and extrinsic parameterization. The former parameterizes the intrinsic variables that affect the appearance of a face, while the latter parameterizes the extrinsic variables that dominate how faces appear on background images as required by a test criterion. Experiments reveal that the proposed framework can generate test samples similar to those available from a popular face detection database, and also samples unavailable from existing face databases. Gee-Sern Hsu, Thu Ha Tran, Sheng-Luen Chung |
ICIP | 1 |
| 2010 | Probabilistic Facial Trait Code for face recognitionabstractRecently, a new facial encoding scheme, namely Facial Trait Code (FTC), was proposed. FTC encoded human faces into a series of integers. Distances between codewords of different people were maximized during the code construction. FTC was applied to solve face recognition, and was reported with promising verification rates. However, due to several simplifications in the FTC encoding, its performance degraded considerably when there were only few images per individual available for enrollment in the gallery sets, or when the probe set included faces under large variations in illumination and expression. In this paper, we proposed the Probabilistic Facial Trait Code (PFTC) with a novel encoding scheme and a probabilistic codeword distance measure. The impact made by illumination and expression variations were also handled in the construction of PFTC. The proposed PFTC was evaluated and compared with state-of-the-art algorithms, including the FTC, the algorithm using sparse representation, and the one using Local Binary Pattern. PFTC out-performed the algorithms compared in this study in most scenarios. Ping-Han Lee, Szu-Wei Wu, Gee-Sern Hsu, Yi-Ping Hung |
ICIP | 3 |
| 2009 | Polymorphous Facial Trait Code
Ping-Han Lee, Gee-Sern Hsu, Yi-Ping Hung |
ACCV (3) | 2 |
| 2009 | Face verification and identification using Facial Trait CodeabstractWe propose the Facial Trait Code (FTC) to encode human facial images. The proposed FTC is motivated by the discovery of some basic patterns existing in certain local facial features. We call these basic patterns Distinctive Trait Patterns (DTP), which can be extracted from a large number of faces. We have also found that the fusion of these DTP's can accurately capture the appearance of a face. The extraction of DTP involves clustering and boosting for maximizing the discrimination between human faces. The extracted DTP's can be symbolized and used to make up the n-ary facial trait codes. A given face can be encoded at some prescribed facial traits to render an n-ary facial trait code with each symbol in its codeword corresponding to the closest DTP. We applied FTC to a face identification and verification problems with 3575 facial images from 840 people under different illumination conditions, and it yielded satisfactory results. Ping-Han Lee, Gee-Sern Hsu, Yi-Ping Hung |
CVPR | 2 |
| 2006 | Distinctive Personal Traits for Face Recognition Under OcclusionabstractExisting local feature methods for face recognition utilize visually salient regions around eye, nose, and mouth to model the characteristics of a person. The premise of such an approach is that there exists a set of features that are common in all human faces and yet distinct to tell one from the rest apart. In this paper we present an algorithm that selects the best set of features or templates for each individual, and uses these distinct personal traits to boost face recognition performance even when they are partially occluded. Borne out by numerous experiments and comparisons, we demonstrate that the proposed method is effective in recognizing faces with partial occlusion and variation in expression. Ping-Han Lee, Yun-Wen Wang, Ming-Hsuan Yang 0001, Gee-Sern Hsu, Yi-Ping Hung |
SMC | 4 |