Terence Sim

dblp:78/392 · DBLP profile ↗
← Back
90ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0002-0198-094XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 76 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 40 · 5 first-author · 9 since 2021Security and privacy · 8 · 3 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 NullFace: Training-Free Localized Face Anonymization
abstract
Privacy concerns around ever increasing number of cameras are increasing in today's digital age. Although existing anonymization methods are able to obscure identity information, they often struggle to preserve the utility of the images. In this work, we introduce a training-free method for face anonymization that preserves key non-identity-related attributes. Our approach utilizes a pre-trained text-to-image diffusion model without requiring optimization or training. It begins by inverting the input image to recover its initial noise. The noise is then denoised through an identity-conditioned diffusion process, where modified identity embeddings ensure the anonymized face is distinct from the original identity. Our approach also supports localized anonymization, giving users control over which facial regions are anonymized or kept intact. Comprehensive evaluations against state-of-the-art methods show our approach excels in anonymization, attribute preservation, and image quality. Its flexibility, robustness, and practicality make it well-suited for real-world applications. Code and data can be found at https://github.com/hanweikung/nullface .
Han-Wei Kung, Tuomas Varanka, Terence Sim, Nicu Sebe
FG3
2026 WRATH: Turning Watermark Robustness Against Itself via a Watermark-Agnostic Black-Box Invalidation Attack
Bangjie Sun, Terence Sim, Jun Han 0001
SP4
2025 Seeing Through Deepfakes: A Human-Inspired Framework for Multi-Face Detection
abstract
Multi-face deepfake videos are becoming increasingly prevalent, often appearing in natural social settings that challenge existing detection methods. Most current approaches excel at single-face detection but struggle in multi-face scenarios, due to a lack of awareness of crucial contextual cues. In this work, we develop a novel approach that leverages human cognition to analyze and defend against multi-face deepfake videos. Through a series of human studies, we systematically examine how people detect deepfake faces in social settings. Our quantitative analysis reveals four key cues humans rely on: scene-motion coherence, inter-face appearance compatibility, interpersonal gaze alignment, and face-body consistency. Guided by these insights, we introduce \textsf{HICOM}, a novel framework designed to detect every fake face in multi-face scenarios. Extensive experiments on benchmark datasets show that \textsf{HICOM} improves average accuracy by 3.3\% in in-dataset detection and 2.8\% under real-world perturbations. Moreover, it outperforms existing methods by 5.8\% on unseen datasets, demonstrating the generalization of human-inspired cues. \textsf{HICOM} further enhances interpretability by incorporating an LLM to provide human-readable explanations, making detection results more transparent and convincing. Our work sheds light on involving human factors to enhance defense against deepfakes.
Shaojing Fan, Terence Sim
ICCV3
2025 Face Anonymization Made Simple
abstract
Current face anonymization techniques often depend on identity loss calculated by face recognition models, which can be inaccurate and unreliable. Additionally, many methods require supplementary data such as facial landmarks and masks to guide the synthesis process. In contrast, our approach uses diffusion models with only a reconstruction loss, eliminating the need for facial landmarks or masks while still producing images with intricate, fine-grained details. We validated our results on two public benchmarks through both quantitative and qualitative evaluations. Our model achieves state-of-the-art performance in three key areas: identity anonymization, facial attribute preservation, and image quality. Beyond its primary function of anonymization, our model can also perform face swapping tasks by incorporating an additional facial image as input, demonstrating its versatility and potential for diverse applications. Our code and models are available at https://github.com/hanweikung/face_anon_simple.
Han-Wei Kung, Tuomas Varanka, Sanjay Saha, Terence Sim, Nicu Sebe
WACV4
2025 KVC-onGoing: Keystroke Verification Challenge
abstract
This article presents the Keystroke Verification Challenge - onGoing (KVC-onGoing) 1 1 https://sites.google.com/view/bida-kvc/ . , on which researchers can easily benchmark their systems in a common platform using large-scale public databases, the Aalto University Keystroke databases, and a standard experimental protocol. The keystroke data consist of tweet-long sequences of variable transcript text from over 185,000 subjects, acquired through desktop and mobile keyboards simulating real-life conditions. The results on the evaluation set of KVC-onGoing have proved the high discriminative power of keystroke dynamics, reaching values as low as 3.33% of Equal Error Rate (EER) and 11.96% of False Non-Match Rate (FNMR) @1% False Match Rate (FMR) in the desktop scenario, and 3.61% of EER and 17.44% of FNMR @1% at FMR in the mobile scenario, significantly improving previous state-of-the-art results. Concerning demographic fairness, the analyzed scores reflect the subjects’ age and gender to various extents, not negligible in a few cases. The framework runs on CodaLab 2 2 https://codalab.lisn.upsaclay.fr/competitions/14063 . . • We set up a novel framework for developing and evaluating keystroke biometrics. • We designed a unified experimental protocol with desktop and mobile scenarios. • We employ the biggest databases of keystroke dynamics, with over 185,000 subjects. • We provide a competitive performance baseline based on a limited-time challenge. • We provide a first exploration of the biometric fairness of keystroke dynamics.
Giuseppe Stragapede, Rubén Vera-Rodríguez, Ruben Tolosana, Aythami Morales, Ivan DeAndres-Tame, Naser Damer, Julian Fierrez, Javier Ortega-Garcia, Alejandro Acien, Nahuel González, Andrei Shadrikov, Dmitrii Gordin, Leon Schmitt, Daniel Wimmer, Christoph Großmann, Joerdis Krieger, Florian Heinz, Ron Krestel, Christoffer Mayer, Simon Haberl, Helena Gschrey, Yosuke Yamagishi, Sanjay Saha, Sanka Rasnayaka, Sandareka Wickramanayake, Terence Sim, Weronika Gutfeter, Adam Baran, Mateusz Krzyszton, Przemyslaw Jaskola
Pattern Recognit.26
2024 Can I Hear Your Face? Pervasive Attack on Voice Authentication Systems with a Single Face Image
Bangjie Sun, Terence Sim, Jun Han 0001
USENIX Security Symposium3
2023 IEEE BigData 2023 Keystroke Verification Challenge (KVC)
abstract
Institute, Warsaw, Poland This paper describes the results of the IEEE BigData 2023 Keystroke Verification Challenge1(KVC), that considers the biometric verification performance of Keystroke Dynamics (KD), captured as tweet-long sequences of variable transcript text from over 185,000 subjects. The data are obtained from two of the largest public databases of KD up to date, the Aalto Desktop and Mobile Keystroke Databases, guaranteeing a minimum amount of data per subject, age and gender annotations, absence of corrupted data, and avoiding excessively unbalanced subject distributions with respect to the considered demographic attributes. Several neural architectures were proposed by the participants, leading to global Equal Error Rates (EERs) as low as 3.33% and 3.61% achieved by the best team respectively in the desktop and mobile scenario, outperforming the current state of the art biometric verification performance for KD. Hosted on CodaLab2, the KVC will be made ongoing to represent a useful tool for the research community to compare different approaches under the same experimental conditions and to deepen the knowledge of the field.
Giuseppe Stragapede, Rubén Vera-Rodríguez, Ruben Tolosana, Aythami Morales, Ivan DeAndres-Tame, Naser Damer, Julian Fierrez, Javier Ortega-Garcia, Nahuel González, Andrei Shadrikov, Dmitrii Gordin, Leon Schmitt, Daniel Wimmer, Christoph Großmann, Joerdis Krieger, Florian Heinz, Ron Krestel, Christoffer Mayer, Simon Haberl, Helena Gschrey, Yosuke Yamagishi, Sanjay Saha, Sanka Rasnayaka, Sandareka Wickramanayake, Terence Sim, Weronika Gutfeter, Adam Baran, Mateusz Krzyszton, Przemyslaw Jaskola
IEEE Big Data25
2022 Action Invariant IMU-Gait for Continuous Authentication
abstract
Continuous Authentication (CA) is proposed as an alter-native authentication scheme for modern personal devices. Gait is a suitable biometric for CA due to its availability and low resource requirements. However, the drastic change in the gait pattern with changes in actions (such as walking, running or climbing stairs) and changes in terrain (such as walking on a flat surface or down an incline) makes it chal-lenging to deploy in a real-world CA system. We show that standard gait features are influenced by different actions. The gait pattern of an action is also in-fluenced by the actions performed before and immediately after. Therefore, gait features are usually not robust to these action variations. We propose action invariant gait features to address this robustness issue. Our proposed method learns action invariant gait features utilizing a Siamese Net-work architecture with triplet loss and a unique triplet mining protocol. Our evaluations highlight that our action in-variant features are robust to pre and post action impacts and real world action variations. These features allows for a CA system to be enrolled using a single action (walk) and be used across multiple different actions encountered throughout the day.
Sanka Rasnayaka, Terence Sim
IJCB2
2022 Virtual Tasks but Real Gains: Improving Multi-Task Learning
abstract
In supervised multi-task learning, the choice of auxiliary tasks is usually decided manually. This immediately raises the issue of how to choose suitable tasks, which in turn requires the need to measure the similarity between tasks. In this paper, we propose a task-similarity metric that depends solely on the task labels, not on the machine learning models. We show that auxiliary tasks may be synthetically generated from the main tasks with any desired similarity. Labels for these virtual tasks are also generated. Finally, we show that learning these virtual tasks along with the main task leads to real performance gains.
Theivendiram Pranavan, Terence Sim, Jianshu Li
ICPR2
2022 Shakespeer: Verifying the Co-presence of Smart Devices and Users via Vibration
abstract
Securely and unobtrusively authenticating a user is an important problem given the pervasiveness of smartphones. Existing approaches, such as password, fingerprints, or facial recognition, are vulnerable to various attacks, and/or degrade usability. To overcome this problem, we propose Shakespeer, which differentiates users based on uniqueness in the propagation of haptic vibrations through hand, forearm muscles and bones. These vibrations are generated by the user’s smartphone and sensed by their smartphone and smartwatch. The unobtrusive haptic vibrational response makes this biometric feature hard to be replicated. Meanwhile, it provides the co-presence detection function, which allows the devices to confirm the co-presence on the user’s body. We implement Shakespeer using smartphones and smartwatches and tested it across 32 subjects under real-world settings. From our preliminary exploratory evaluation, Shakespeer achieves an equal error rate (EER) of 0.59 %, demonstrating its feasibility.
Gucheng Wang, Jay Prakash, Terence Sim, Jun Han 0001
ICPR3
2022 Localizing Fake Segments in Speech
abstract
Accelerated progress in voice cloning technology is making phone scams easier and exposing potential threats to politicians. Previous spoofing detection technology focused more on the fully faked speech. In this work, we create a Partial Synthetic Detection (Psynd) dataset and propose a fake segments localization system of the partially faked speech. Psynd dataset is a multi-speaker English corpus of approximately 13 hours in total at 24kHz sampling rate read English speech injected with synthetic speech. The fake segments are generated by state-of-art multi-speaker text-to-speech models with high similarity to the real speech to be injected. Our fake segments localization system consists of 3 parts: acoustic feature extraction, classification and post-processing. Frame level CQCC features are extracted and forwarded to a spoofing-discriminant ANN to predict real or fake label sequence. Assuming that the fake or real segments in the partially faked speech cannot be shorter than the duration of a phoneme, the labels of extreme short fake or real segments are flipped. We use 1-D IoU to evaluate the localization performance and get the result of 98.58% during the test, much higher than a random guess of $\frac{1}{3}$. We also explore extreme cases like fully faked, fully real and multi-fake-segments speech and degraded partially faked audio. Some benchmark results are presented on this dataset and show that a more robust detector is needed.
Terence Sim
ICPR2
2021 35th Anniversary of IJPRAI
Patrick Shen-Pei Wang, Xiaoyi Jiang 0001, Frank Y. Shih, Terence Sim
Int. J. Pattern Recognit. Artif. Intell.4
2021 Multi-human Parsing with a Graph-based Generative Adversarial Model
abstract
Human parsing is an important task in human-centric image understanding in computer vision and multimedia systems. However, most existing works on human parsing mainly tackle the single-person scenario, which deviates from real-world applications where multiple persons are present simultaneously with interaction and occlusion. To address such a challenging multi-human parsing problem, we introduce a novel multi-human parsing model named MH-Parser, which uses a graph-based generative adversarial model to address the challenges of close-person interaction and occlusion in multi-human parsing. To validate the effectiveness of the new model, we collect a new dataset named Multi-Human Parsing (MHP), which contains multiple persons with intensive person interaction and entanglement. Experiments on the new MHP dataset and existing datasets demonstrate that the proposed method is effective in addressing the multi-human parsing problem compared with existing solutions in the literature.
Jianshu Li, Jian Zhao 0006, Congyan Lang, Yidong Li, Yunchao Wei, Guodong Guo, Terence Sim, Shuicheng Yan, Jiashi Feng
ACM Trans. Multim. Comput. Commun. Appl.7
2020 Your Tattletale Gait Privacy Invasiveness of IMU Gait Data
abstract
Modern personal devices measure and store vast amounts of sensory data such as Inertial Measurement Unit (IMU) data. These on-body sensor data can be used as a biometric by observing human movement (gait). People are less cautious about privacy vulnerabilities of such sensory data. We highlight which personal characteristics can be derived from on-body sensor data and the effect of sensor location towards these privacy invasions. By analyzing sensor locations with respect to privacy and utility we discover sensor locations which preserve utility such as biometric authentication while reducing privacy vulnerability. We have collected (1) a multi-stream on-body IMU dataset using 3 IMU sensors, consisting of 6 sensor locations, 6 actions along with various physical, personality and socio-economic characteristics from 53 participants. (2) an opinion survey of the relative importance of each attribute from 566 participants. Using these datasets we show that gait data reveals a lot of personal information, which maybe a privacy concern. The opinion survey reveals a ranking of the physical characteristics based on the perceived importance. Using a privacy vulnerability index we show that sensors located in the front pocket/wrist are more privacy invasive compared to back-pocket/bag which are less privacy invasive without a significant loss of utility as a biometric.
Sanka Rasnayaka, Terence Sim
IJCB2
2020 Is Face Recognition Safe from Realizable Attacks?
abstract
Face recognition is a popular form of biometric authentication and due to its widespread use, attacks have become more common as well. Recent studies show that Face Recognition Systems are vulnerable to attacks and can lead to erroneous identification of faces. Interestingly, most of these attacks are white-box, or they are manipulating facial images in ways that are not physically realizable. In this paper, we propose an attack scheme where the attacker can generate realistic synthesized face images with subtle perturbations and physically realize that onto his face to attack black-box face recognition systems. Comprehensive experiments and analyses show that subtle perturbations realized on attackers face can create successful attacks on state-of-the-art face recognition systems in black-box settings. Our study exposes the underlying vulnerability posed by the Face Recognition Systems against realizable black-box attacks.
Sanjay Saha, Terence Sim
IJCB2
2020 Learning with Delayed Feedback
abstract
We propose a novel supervised machine learning strategy, inspired by human learning, that enables an Agent to learn continually over its lifetime. A natural consequence is that the Agent must be able to handle an input whose label is delayed until a later time, or may not arrive at all. Our Agent learns in two steps: a short Seeding phase, in which the Agent's model is initialized with labelled inputs, and an indefinitely long Growing phase, in which the Agent refines and assesses its model if the label is given for an input, but stores the input in a finite-length queue if the label is missing. Queued items are matched against future input-label pairs that arrive, and the model is then updated. Our strategy also allows for the delayed feedback to take a different form. For example, in an image captioning task, the feedback could be a semantic segmentation rather than a textual caption. We show with many experiments that our strategy enables an Agent to learn flexibly and efficiently.
Theivendiram Pranavan, Terence Sim
ICPR2
2019 Learning Controllable Face Generator from Disjoint Datasets
Jing Li 0050, Yongkang Wong, Terence Sim
CAIP (1)3
2019 Task Relation Networks
abstract
Multi-task learning is popular in machine learning and computer vision. In multitask learning, properly modeling task relations is important for boosting the performance of jointly learned tasks. Task covariance modeling has been successfully used to model the relations of tasks but is limited to homogeneous multi-task learning. In this paper, we propose a feature based task relation modeling approach, suitable for both homogeneous and heterogeneous multi-task learning. First, we propose a new metric to quantify the relations between tasks. Based on the quantitative metric, we then develop the task relation layer, which can be combined with any deep learning architecture to form task relation networks to fully exploit the relations of different tasks in an online fashion. Benefiting from the task relation layer, the task relation networks can better leverage the mutual information from the data. We demonstrate our proposed task relation networks are effective in improving the performance in both homogeneous and heterogeneous multi-task learning settings through extensive experiments on computer vision tasks.
Jianshu Li, Pan Zhou 0002, Yunpeng Chen, Jian Zhao 0006, Sujoy Roy, Shuicheng Yan, Jiashi Feng, Terence Sim
WACV8
2019 Which Body Is Mine?
abstract
In the light of the human studies that report a strong correlation between head circumference and body size, we propose a new research problem: head-body matching. Given an image of a person's head, we want to match it with his body (headless) image. We propose a dual-pathway framework which computes head and body discriminating features independently, and learns the correlation between such features. We introduce a comprehensive evaluation of our proposed framework for this problem using different features including anthropometric features and deep-CNN features, different experimental setting such as head-body scale variations, and different body parts. We demonstrate the usefulness of our framework with two novel applications: head/body recognition, and T-shirt sizing from a head image. Our evaluations for head/body recognition application on the challenging large scale PIPA dataset (contains high variations of pose, viewpoint, and occlusion) show up to 53% of performance improvement using deep-CNN features, over the global model features in which head and body features are not separated or correlated. For T-shirt sizing application, we use anthropometric features for head-body matching. We achieve promising experimental results on small and challenging datasets.
Mona Ragab, Terence Sim, Joo-Hwee Lim, Keng Teck Ma
WACV2
2019 Toward a Comprehensive Face Detector in the Wild
abstract
In this paper, we aim to build a comprehensive face detection system which provides a one-stop solution to various practical challenges for face detection in realistic scenarios, e.g., detecting faces from multiple-views, faces with occlusions, exaggerated expressions or blurred faces. Moreover, we introduce an automatic data harvest algorithm to effectively improve the generalization performance of the system even when collecting training faces containing various challenging patterns is difficult. In particular, we introduce three critical components to build the system, i.e., a recently widely used deep convolutional neural network (CNN), a novel blur-aware bi-channel network architecture, and a new self-learning mechanism capable of exploiting video contexts continuously. The aforementioned challenges except for detecting blurred faces can potentially be addressed by the CNN component owing its robustness to local deformation of target faces. The more challenging problem of detecting blurred faces is addressed by the bi-channel architecture component which processes blurred and clear faces adaptively. In addition, to address the difficulties in improving the generalization performance of the learning-based face detection system, we introduce a video-context-based self-learning mechanism into the system, which enables the system to continuously enhance its performance by harvesting faces with challenging training patterns automatically. To exploit video context, the detector is applied to massive unlabeled videos, and challenging faces are captured based on temporal inference. These recaptured faces, generally corresponding to one or multiple challenges mentioned above, are fed into the detection system to further improve its performance. Extensive experiments with the proposed detection system provide new state-of-the-art performance on FDDB data set, PASCAL face data set, AFW data set, and WIDER Face data set.
Jianshu Li, Luoqi Liu, Jianan Li 0001, Jiashi Feng, Shuicheng Yan, Terence Sim
IEEE Trans. Circuits Syst. Video Technol.6
2018 Multi-Human Parsing Machines
abstract
Human parsing is an important task in human-centric analysis. Despite the remarkable progress in single-human parsing, the more realistic case of multi-human parsing remains challenging in terms of the data and the model. Compared with the considerable number of available single-human parsing datasets, the datasets for multi-human parsing are very limited in number mainly due to the huge annotation effort required. Besides the data challenge to multi-human parsing, the persons in real-world scenarios are often entangled with each other due to close interaction and body occlusion, making it difficult to distinguish body parts from different person instances. In this paper we propose the Multi-Human Parsing Machines (MHPM) system, which contains an MHP Montage model and an MHP Solver, to address both challenges in multi-human parsing. Specifically, the MHP Montage model in MHPM generates realistic images with multiple persons together with the parsing labels. It intelligently composes single persons onto background scene images while maintaining the structural information between persons and the scene. The generated images can be used to train better multi-human parsing algorithms. On the other hand, the MHP Solver in MHPM solves the bottleneck of distinguishing multiple entangled persons with close interaction. It employs a Group-Individual Push and Pull (GIPP) loss function, which can effectively separate persons with close interaction. We experimentally show that the proposed MHPM can achieve state-of-the-art performance on the multi-human parsing benchmark and the person individualization benchmark, which distinguishes closely entangled person instances.
Jianshu Li, Jian Zhao 0006, Yunpeng Chen, Sujoy Roy, Shuicheng Yan, Jiashi Feng, Terence Sim
ACM Multimedia7
2018 Understanding Humans in Crowded Scenes: Deep Nested Adversarial Learning and A New Benchmark for Multi-Human Parsing
abstract
Despite the noticeable progress in perceptual tasks like detection, instance segmentation and human parsing, computers still perform unsatisfactorily on visually understanding humans in crowded scenes, such as group behavior analysis, person re-identification and autonomous driving, etc. To this end, models need to comprehensively perceive the semantic information and the differences between instances in a multi-human image, which is recently defined as the multi-human parsing task. In this paper, we present a new large-scale database "Multi-Human Parsing (MHP)" for algorithm development and evaluation, and advances the state-of-the-art in understanding humans in crowded scenes. MHP contains 25,403 elaborately annotated images with 58 fine-grained semantic category labels, involving 2-26 persons per image and captured in real-world scenes from various viewpoints, poses, occlusion, interactions and background. We further propose a novel deep Nested Adversarial Network (NAN) model for multi-human parsing. NAN consists of three Generative Adversarial Network (GAN)-like sub-nets, respectively performing semantic saliency prediction, instance-agnostic parsing and instance-aware clustering. These sub-nets form a nested structure and are carefully designed to learn jointly in an end-to-end way. NAN consistently outperforms existing state-of-the-art solutions on our MHP and several other datasets, and serves as a strong baseline to drive the future research for multi-human parsing.
Jian Zhao 0006, Jianshu Li, Yu Cheng 0009, Terence Sim, Shuicheng Yan, Jiashi Feng
ACM Multimedia4
2018 Landmark Free Face Attribute Prediction
abstract
Face attribute prediction in the wild is important for many facial analysis applications yet it is very challenging due to ubiquitous face variations. In this paper, we address face attribute prediction in the wild by proposing a novel method, lAndmark Free Face AttrIbute pRediction (AFFAIR). Unlike traditional face attribute prediction methods that require facial landmark detection and face alignment, AFFAIR uses an endto- end learning pipeline to jointly learn a hierarchy of spatial transformations that optimize facial attribute prediction with no reliance on landmark annotations or pre-trained landmark detectors. AFFAIR achieves this through simultaneously 1) learning a global transformation which effectively alleviates negative effect of global face variation for the following attribute prediction tailored for each face, 2) locating the most relevant facial part for attribute prediction and 3) aggregating the global and local features for robust attribute prediction. Within AFFAIR, a new competitive learning strategy is developed that effectively enhances global transformation learning for better attribute prediction. We show that with zero information about landmarks, AFFAIR achieves state-of-the-art performance on three face attribute prediction benchmarks, which simultaneously learns the face-level transformation and attribute-level localization within a unified framework.
Jianshu Li, Fang Zhao 0006, Jiashi Feng, Sujoy Roy, Shuicheng Yan, Terence Sim
IEEE Trans. Image Process.6
2017 Multi-view Separation of Background and Reflection by Coupled Low-Rank Decomposition
Jian Lai, Wee Kheng Leow, Terence Sim
CAIP (2)3
2017 BoxFlow: Unsupervised Face Detector Adaptation from Images to Videos
abstract
Face detectors are usually trained on static images but deployed in the wild such as surveillance videos. Due to the domain shift between images and videos, directly applying the image-based face detectors onto videos usually gives unsatisfactory performance. In this paper, we introduce the BoxFlow - a new unsupervised detector adaptation method that can effectively adapt a face detector pre-trained on static images to videos. BoxFlow unsupervisedly adapts face detectors through fully exploiting the motion contexts across video frames. In particular, BoxFlow introduces three novel components: (1) generalized heat map representation of face locations with augmented shape flexibility; (2) motion based temporal contextual regularization among adjacent frames for unsupervised face detection refinement; (3) a self-paced learning strategy that adapts face detectors from easy data samples to challenging ones progressively. With these key components, we develop a systematic unsupervised face detector adaptation framework to help face detectors adapt to various deployed environments. Extensive experiments on the IDA dataset clearly demonstrate the superiority of our proposed method. Without utilizing any annotation, the BoxFlow achieves about 10%-20% performance gain in terms of Average Precision than directly applying image-based face detectors.
Jianshu Li, Jiashi Feng, Luoqi Liu, Terence Sim
FG4
2017 Integrated Face Analytics Networks through Cross-Dataset Hybrid Training
abstract
Face analytics benefits many multimedia applications. It consists of a number of tasks, such as facial emotion recognition and face parsing, and most existing approaches generally treat these tasks independently, which limits their deployment in real scenarios. In this paper we propose an integrated Face Analytics Network (iFAN), which is able to perform multiple tasks jointly for face analytics with a novel carefully designed network architecture to fully facilitate the informative interaction among different tasks. The proposed integrated network explicitly models the interactions between tasks so that the correlations between tasks can be fully exploited for performance boost. In addition, to solve the bottleneck of the absence of datasets with comprehensive training data for various tasks, we propose a novel cross-dataset hybrid training strategy. It allows "plug-in and play'' of multiple datasets annotated for different tasks without the requirement of a fully labeled common dataset for all the tasks. We experimentally show that the proposed iFAN achieves state-of-the-art performance on multiple face analytics tasks using a single integrated model. Specifically, iFAN achieves an overall F-score of 91.15% on the Helen dataset for face parsing, a normalized mean error of 5.81% on the MTFL dataset for facial landmark localization and an accuracy of 45.73% on the BNU dataset for emotion recognition with a single model.
Jianshu Li, Shengtao Xiao, Fang Zhao 0006, Jian Zhao 0006, Jianan Li 0001, Jiashi Feng, Shuicheng Yan, Terence Sim
ACM Multimedia8
2017 Lighting transfer across multiple views through local color transforms
abstract
We present a method for transferring lighting between photographs of a static scene. Our method takes as input a photo collection depicting a scene with varying viewpoints and lighting conditions. We cast lighting transfer as an edit propagation problem, where the transfer of local illumination across images is guided by sparse correspondences obtained through multi-view stereo. Instead of directly propagating color, we learn local color transforms from corresponding patches in pairs of images and propagate these transforms in an edge-aware manner to regions with no correspondences. Our color transforms model the large variability of appearance changes in local regions of the scene, and are robust to missing or inaccurate correspondences. The method is fully automatic and can transfer strong shadows between images. We show applications of our image relighting method for enhancing photographs, browsing photo collections with harmonized lighting, and generating synthetic time-lapse sequences.
Qian Zhang 0065, Pierre-Yves Laffont, Terence Sim
Comput. Vis. Media3
2016 Happiness level prediction with sequential inputs via multiple regressions
abstract
This paper presents our solution submitted to the Emotion Recognition in the Wild (EmotiW 2016) group-level happiness intensity prediction sub-challenge. The objective of this sub-challenge is to predict the overall happiness level given an image of a group of people in a natural setting. We note that both the global setting and the faces of the individuals in the image influence the group-level happiness intensity of the image. Hence the challenge lies in building a solution that incorporates both these factors and also considers their right combination. Our proposed solution incorporates both these factors as a combination of global and local information. We use a convolutional neural network to extract discriminative face features, and a recurrent neural network to selectively memorize the important features to perform the group-level happiness prediction task. Experimental evaluations show promising performance improvements, resulting in Root Mean Square Error (RMSE) reduction of about 0.5 units on the test set compared to the baseline algorithm that uses only global information.
Jianshu Li, Sujoy Roy, Jiashi Feng, Terence Sim
ICMI4
2016 Towards protecting biometric templates without sacrificing performance
abstract
The ideal biometric template protection scheme possesses the properties of irreversibility, revocability, unlinkability, and good performance. These properties protect the security of the biometrics system as well as users' privacy. Practical systems, however, fall short of this ideal. In this paper, we present a novel protection scheme that achieves this ideal under the circumstance that a subject's token and his biometric template are not concurrently exposed. Moreover, our scheme can add template protection to any face verifier. We do this by rendering virtual faces, rather than by devising new biometric features, which is the more common approach. Experimental evaluations using two public face recognition systems show that accuracy is not adversely affected with our scheme.
Jing Li 0050, Yongkang Wong, Terence Sim
ICPR3
2016 Robust Face Recognition with Deep Multi-View Representation Learning
abstract
This paper describes our proposed method targeting at the MSR Image Recognition Challenge MS-Celeb-1M. The challenge is to recognize one million celebrities from their face images captured in the real world. The challenge provides a large scale dataset crawled from the Web, which contains a large number of celebrities with many images for each subject. Given a new testing image, the challenge requires an identify for the image and the corresponding confidence score. To complete the challenge, we propose a two-stage approach consisting of data cleaning and multi-view deep representation learning. The data cleaning can effectively reduce the noise level of training data and thus improves the performance of deep learning based face recognition models. The multi-view representation learning enables the learned face representations to be more specific and discriminative. Thus the difficulties of recognizing faces out of a huge number of subjects are substantially relieved. Our proposed method achieves a coverage of 46.1% at 95% precision on the random set and a coverage of 33.0% at 95% precision on the hard set of this challenge.
Jianshu Li, Jian Zhao 0006, Fang Zhao 0006, Hao Liu 0003, Jing Li 0050, Shengmei Shen, Jiashi Feng, Terence Sim
ACM Multimedia8
2016 Correlation filter cascade for facial landmark localization
abstract
The application of correlation filters for the task of facial landmark detection has been studied by many vision works. Their success, however, is limited by the presence of large pose variations, expression and occlusion in face images. Moreover, existing correlation filters may suffer from poor discrimination to distinguish visually similar landmarks such as the right and left eyes. In this work, we present a new framework, referred to as Correlation Filter Cascade, to address the above limitations. The proposed framework consists of a set of correlation filters with different spatial supports (sizes) which are connected together in a cascade form. More specifically, the size of filters decreases from the lower to upper levels. Filters at lower levels implicitly code face shape information since they are trained using large patches stemmed from face images. This avoids ambiguous detections caused by landmarks with similar appearance. Detections in these levels, however, may not be accurate and suffer from small localization errors, mainly caused by face pose, expression and occlusion. Therefore, locations detected by lower levels will be further used by the higher levels to narrow down their search regions. Since the filters at higher levels have smaller size, they are less affected by pose, expression and occlusion, and thus, can perform more accurately. The evaluation on BioID and LFPW shows the superiority of our method compared to prior correlation filters and leading facial landmark detectors.
Hamed Kiani Galoogahi, Terence Sim
WACV2
2016 Think big, solve small: Scaling up robust PCA with coupled dictionaries
abstract
Recent advances in robust principle component analysis offers a powerful method for solving a wide variety of low-level vision problems. However, if the input data is very large, especially when high-resolution images are involved, it makes RPCA computationally prohibitive for many real applications. To tackle this problem, we propose a fixed-rank RPCA method that uses coupled dictionaries (FRPCA-CD) to handle high-resolution images. FRPCA-CD downsamples high-resolution images into low-resolution images, performs FRPCA on the low-level images to obtain the low-rank matrix, which is reconstructed at original resolution by coupled dictionaries. Comprehensive tests performed on video background recovery, noise reduction in photometric stereo, and image reflection removal problems show that FRPCA-CD can reduce computation time and memory space drastically without sacrificing accuracy.
Jian Lai, Wee Kheng Leow, Terence Sim, Vaishali Sharma
WACV3
2016 Hide and seek: Uncovering facial occlusion with variable-threshold robust PCA
abstract
Face images are very important in human social activities, which can be severely hampered when they are corrupted by occluders such as eyeglasses, face marks, and scarfs. Existing methods for removing occlusions in face images can be grouped into three broad categories, namely PCA, robust PCA (RPCA), and sparse coding. The major weaknesses of these methods are inconsistent performance across test conditions and possible corruption of unoccluded part of the recovered target image. This paper presents variable-threshold RPCA (VRPCA) method based on RPCA with variable thresholding. Comprehensive tests show that VRPCA is able to preserve the unoccluded parts of the target image with practically zero error. Compared to existing methods, it is more accurate, reliable, and consistent across various test conditions.
Wee Kheng Leow, Jian Lai, Terence Sim, Vaishali Sharma
WACV4
2015 Incremental Fixed-Rank Robust PCA for Video Background Recovery
Jian Lai, Wee Kheng Leow, Terence Sim
CAIP (2)3
2015 Correlation filters with limited boundaries
abstract
Correlation filters take advantage of specific properties in the Fourier domain allowing them to be estimated efficiently: O(N D log D) in the frequency domain, versus O(D3+ N D2) spatially where D is signal length, and N is the number of signals. Recent extensions to correlation filters, such as MOSSE, have reignited interest of their use in the vision community due to their robustness and attractive computational properties. In this paper we demonstrate, however, that this computational efficiency comes at a cost. Specifically, we demonstrate that only 1/D proportion of shifted examples are unaffected by boundary effects which has a dramatic effect on detection/tracking performance. In this paper, we propose a novel approach to correlation filter estimation that: (i) takes advantage of inherent computational redundancies in the frequency domain, (ii) dramatically reduces boundary effects, and (iii) is able to implicitly exploit all possible patches densely extracted from training examples during learning process. Impressive object tracking and detection results are presented in terms of both accuracy and computational efficiency.
Hamed Kiani Galoogahi, Terence Sim, Simon Lucey
CVPR2
2014 Multi-channel correlation filters for human action recognition
abstract
In this work, we propose to employ multi-channel correlation filters for recognizing human actions (e.g. waking, riding) in videos. In our framework, each action sequence is represented as a multi-channel signal (frames) and the goal is to learn a multi-channel filter for each action class that produces a set of desired outputs when correlated with training examples. The experiments on the Weizmann and UCF sport datasets demonstrate superior computational cost (real-time), memory efficiency and very competitive performance of our approach compared to the state of the arts.
Hamed Kiani Galoogahi, Terence Sim, Simon Lucey
ICIP2
2014 What does computer vision say about face reading?
abstract
Face reading is an ancient Chinese practice that assesses a person's personality traits by applying a set of rules on facial features. In this paper, we take an image processing approach to validate these face reading rules. We collected facial features and personality traits from 54 human subjects, analyzed their correlation, and found only weak correlations between the facial features and personality traits. We also attempted to discover new, more complex, face reading rules based on the data we collected, but we found no new rules with significant correlation. In short, we found no evidence that supports this ancient Chinese practice.
Terence Sim, Wei Tsang Ooi
ICIP1
2014 Light montage for perceptual image enhancement
abstract
Abstract Recent photography techniques such as sculpting with light show great potential in compositing beautiful images from fixed‐viewpoint photos under multiple illuminations. The process relies heavily on the artists’ experience and skills using the available tools. An apparent trend in recent works is to facilitate the interaction making it less time‐consuming and addressable not only to experts, but also novices. We propose a method that automatically creates enhanced light montages that are comparable to those produced by artists. It detects and emphasizes cues that are important for perception by introducing a technique to extract depth and shape edges from an unconstrained light stack. Studies show that these cues are associated with silhouettes and suggestive contours which artists use to sketch and construct the layout of paintings. Textures, due to perspective distortion, offer essential cues that depict shape and surface slant. We balance the emphasis between depth edges and reflectance textures to enhance the sense of both shape and reflectance properties. Our light montage technique works perfectly with a few to hundreds of illuminations for each scene. Experiments show great results for static scenes making it practical for small objects, interiors and small‐scale outdoor scenes. Dynamic scenes may be captured using spatially distributed light setups such as light domes. The approach could also be applied to time‐lapse photos, with the sun as the main light source.
Vlad Hosu, Mai Lan Ha, Terence Sim
Comput. Graph. Forum3
2014 A talking profile to distinguish identical twins
Li Zhang 0005, Keng Teck Ma, Hossein Nejati, Lewis Foo, Terence Sim, Dong Guo 0001
Image Vis. Comput.5
2013 Background Recovery by Fixed-Rank Robust Principal Component Analysis
Wee Kheng Leow, Yuan Cheng 0005, Li Zhang 0005, Terence Sim, Lewis Foo
CAIP (1)4
2013 Eyewitness Face Sketch Recognition Based on Two-Step Bias Modeling
Hossein Nejati, Li Zhang 0005, Terence Sim
CAIP (2)3
2013 Hearing versus Seeing Identical Twins
Li Zhang 0005, Shenggao Zhu, Terence Sim, Wee Kheng Leow, Hossein Nejati, Dong Guo 0001
CAIP (1)3
2013 Multi-channel Correlation Filters
abstract
Modern descriptors like HOG and SIFT are now commonly used in vision for pattern detection within image and video. From a signal processing perspective, this detection process can be efficiently posed as a correlation/ convolution between a multi-channel image and a multi-channel detector/filter which results in a single channel response map indicating where the pattern (e.g. object) has occurred. In this paper, we propose a novel framework for learning a multi-channel detector/filter efficiently in the frequency domain, both in terms of training time and memory footprint, which we refer to as a multichannel correlation filter. To demonstrate the effectiveness of our strategy, we evaluate it across a number of visual detection/ localization tasks where we: (i) exhibit superior performance to current state of the art correlation filters, and (ii) superior computational and memory efficiencies compared to state of the art spatial detectors.
Hamed Kiani Galoogahi, Terence Sim, Simon Lucey
ICCV2
2012 Recognizing emotions of characters in movies
abstract
This work presents an investigation into recognizing emotions of people in near real life scenarios. Most existing studies on recognizing emotions of people have been conducted under controlled environments where the emotions are not spontaneous, rather highly exaggerated, and the number of modalities considered and their interactions is limited. The proposed bimodal approach fuses facial expression recognition (FER) with the “semantic orientation” of dialogs of actors to identify emotions under difficult illumination conditions, pose variations and occlusions in scenes. Experiments conducted on a dataset of 700 video clips from 17 movies demonstrate that the proposed fusion approach improves emotion recognition performance over unimodal approaches.
Ruchir Srivastava, Shuicheng Yan, Terence Sim, Sujoy Roy
ICASSP3
2012 Face sketch recognition by Local Radon Binary Pattern: LRBP
abstract
In this paper, we propose a new face descriptor to directly match face photos and sketches of different modalities, called Local Radon Binary Pattern (LRBP). LRBP is inspired by the fact that the shape of a face photo and its corresponding sketch is similar, even when the sketch is exaggerated by an artist. Therefore, the shape of face can be exploited to compute features which are robust against modality differences between face photo and sketch. In LRBP framework, the characteristics of face shape are captured by transforming face image into Radon space. Then, micro-information of face shape in new space is encoded by Local Binary Pattern (LBP). Finally, LRBP is computed by concatenating histograms of local LBPs. In order to capture both local and global characteristics of face shape, LRBP is extracted in a spatial pyramid fashion. Experiments on CUFS and CUFSF datasets indicate the efficiency of LRBP for face sketch recognition.
Hamed Kiani Galoogahi, Terence Sim
ICIP2
2012 Inter-modality Face Sketch Recognition
abstract
Automatic face sketch recognition plays an important role in law enforcement. Recently, various methods have been proposed to address the problem of face sketch recognition by matching face photos and sketches, which are of different modalities. However, their performance is strongly affected by the modality difference between sketches and photos. In this paper, we propose a new face descriptor based on gradient orientations to reduce the modality difference in feature extraction stage, called Histogram of Averaged Oriented Gradients (HAOG). Experiments on CUFS database show that the new descriptor outperforms the state-of-the-art approaches.
Hamed Kiani Galoogahi, Terence Sim
ICME2
2012 Wonder ears: Identification of identical twins from ear images
Hossein Nejati, Li Zhang 0005, Terence Sim, Elisa Martínez Marroquín, Dong Guo 0001
ICPR3
2012 Expressive deformation profiles for cross expression face recognition
Li Zhang 0005, Elisa Martínez Marroquín, Terence Sim
ICPR4
2012 Face photo retrieval by sketch example
abstract
Face photo-sketch matching has received great attention in recent years due to its vital role in law enforcement. The major challenge of matching face photo and sketch is difference of visual characteristics between face photo and sketch which is referred as modality gap. Earlier approaches have reduced the modality gap by synthesizing face photos and sketches in a same modality (photo or sketch). However, the effectiveness of these approaches is highly affected by synthesis results. That means a poor synthesis might degrade the performance of matching. Therefore, recent works have focused to directly match face photo and sketch of different modalities. However, the features used by these approaches are not robust against modality gap. In this paper, a modality-invariant face descriptor called Gabor Shape is proposed to retrieve face photos based on a probe sketch. Experiments on CUFS and CUFSF datasets show that the new descriptor outperforms the state-of-the-art approaches.
Hamed Kiani Galoogahi, Terence Sim
ACM Multimedia2
2012 Don't ask me what i'm like, just watch and listen
abstract
Traditional (based on psychology) approaches for personality assessment of an individual require him/her to fill up a questionnaire. This paper presents a novel way of utilizing multimodal cues to automatically fill up the questionnaire. The contributions of this work are three-fold. (1) Novel psychology-based audio/visual/lexical features are proposed and shown to be effective in predicting answers to a personality questionnaire, Big-Five Inventory-10 (BFI- 10). (2) Extracted features are used to learn linear and kernel versions of a novel regression model, 'SLoT', to automatically predict BFI-10 answers. The model is based on Sparse and Low-rank Transformation (SLoT). (3) Predicted answers are used to compute personality scores using standard BFI-10 scoring scheme. We evaluated our approach on a dataset of 3907 clips (for 50 characters from movies of diverse genres) manually labeled with BFI-10 answers and personality scores as ground-truth. Experiments indicate that the proposed 'SLoT' model effectively automates the answering process by emulating human understanding. We also conclude that predicting personality scores through predicting answers first is better than directly predicting scores based on audio/visual features (as studied in state-of-the art methods).
Ruchir Srivastava, Jiashi Feng, Sujoy Roy, Shuicheng Yan, Terence Sim
ACM Multimedia5
2012 New hope for recognizing twins by using facial motion
abstract
Distinguishing between identical twins is the Holy Grail in face recognition because of the great similarity between the faces of a pair of twins. Most existing face recognition systems choose to simply ignore it. However, as the population of twins increases quickly, such an “ostrich strategy” is no longer acceptable. The biometric systems that overlook the twins problem are presenting a serious security hole. Inspired by recent advances in motion-based face recognition techniques, we propose to use facial motion to address the twins problem. We collect a twins facial expression database and conduct a series of experiments in two assumed scenarios: the Social Party Scenario and the Access Control Scenario. The experimental results show that facial motion ourperforms facial appearance in distinguishing between twins. Based on this finding, we propose a two-stage cascaded General Access Control System, which combines facial appearance with facial motion. The experimental results show that, compared with an appearance-based face recognition system, this cascaded system is much more secure against an “evil-twin” imposter attack, while performing as good for normal population.
Li Zhang 0005, Elisa Martínez Marroquín, Dong Guo 0001, Terence Sim
WACV5
2011 Accumulated motion images for facial expression recognition in videos
abstract
This paper details the method and experiments conducted towards our submission to the FERA 2011 facial expression recognition benchmarking evaluations. The benchmarking evaluation task involves recognizing 5 emotion classes in videos. Our method for detecting facial expressions is a fusion of the decisions of two FER approaches based on two different feature representations, namely using motion information from facial regions and facial feature point displacement information. The main observation motivating the approach we took is that different feature representations are discriminative in detecting different facial expressions. Hence a fusion approach could complement each other to improve recognition performance. Experiments were conducted on the GEMEP-FERA data set provided by the organizers.
Ruchir Srivastava, Sujoy Roy, Shuicheng Yan, Terence Sim
FG4
2011 Do you see what i see? A more realistic eyewitness sketch recognition
abstract
Face sketches have been used in eyewitness testimonies for about a century. However, 30 years of research shows that current eyewitness testimony methods are highly unreliable. Nonetheless, current face sketch recognition algorithms assume that eyewitness sketches are reliable and highly similar to their respective target faces. As proven by psychological findings and a recent work on face sketch recognition, these assumptions are unrealistic and therefore, current algorithms cannot handle real world cases of eyewitness sketch recognition. In this paper, we address the eyewitness sketch recog nition problem with a two-pronged approach. We propose a more reliable eyewitness testimony method, and an accompanying face sketch recognition method that accounts for realistic assumptions on sketch-photo similarities and individual eyewitness differences. In our eyewitness testimony method we first ask the eyewitness to directly draw a sketch of the target face, and provide some ancillary information about the target face. Then we build a drawing profile of the eyewitness by asking him/her to draw a set of face photos. This drawing profile implicitly contains the eyewitness' mental bias. In our face sketch recognition method we first correct the sketch for the eyewitness' bias using the drawing profile. Then we recognize the resulting sketch based on an optimized combination of the detected features and ancillary information. Experimental results show that our method is 12 times better than the leading competing method at Rank-1 accuracy, and 6 times better at Rank-10. Our method also maintains its superiority as gallery size increases.
Hossein Nejati, Terence Sim, Elisa Martínez Marroquín
IJCB2
2011 Towards automated pose invariant 3D dental biometrics
abstract
A novel pose invariant 3D dental biometrics framework is proposed for human identification by matching dental plasters in this paper. Using 3D overcomes a number of key problems that plague 2D methods. As best as we can tell, our study is the first attempt at 3D dental biometrics. It includes a multi-scale feature extraction algorithm for extracting pose invariant feature points and a triplet-correspondence algorithm for pose estimation. Preliminary experimental result achieves 100% rank-1 accuracy by matching 7 postmortem (PM) samples against 100 ante-mortem (AM) samples. In addition, towards a fully automated 3D dental identification testing, the accuracy achieves 71.4% at rank-1 accuracy and 100% at rank-4 accuracy. Comparing with the existing algorithms, the feature point extraction algorithm and the triplet-correspondence algorithm are faster and more robust for pose estimation. In addition, the retrieval time for a single subject has been significantly reduced. Furthermore, we discover that the investigated dental features are discriminative and useful for identification. The high accuracy, fast retrieval speed and the facilitated identification process suggest that the developed 3D framework is more suitable for practical use in dental biometrics applications in the future. Finally, the limitations and future research directions are discussed.
Deping Yu, Kelvin Weng Chiong Foong, Terence Sim, Yoke San Wong, Ho-Lun Cheng
IJCB4
2011 Learning universal multi-view age estimator using video context
abstract
Many existing techniques for analyzing face images assume that the faces are at nearly frontal. Generalizing to non-frontal faces is often difficult, due to a dearth of ground truth for non-frontal faces and also to the inherent challenges in handling pose variations. In this work, we investigate how to learn a universal multi-view age estimator by harnessing 1) unlabeled web videos, 2) a publicly available labeled frontal face corpus, and 3) zero or more non-frontal faces with age labels. First, a large diverse human-involved video corpus is collected from online video sharing website. Then, multi-view face detection and tracking are performed to build a large set of frontal-vs-profile face bundles, each of which is from the same tracking sequence, and thus exhibiting the same age. These unlabeled face bundles constitute the so-called video context, and the parametric multi-view age estimator is trained by 1) enforcing the face-to-age relation for the partially labeled faces, 2) imposing the consistency of the predicted ages for the non-frontal and frontal faces within each face bundle, and 3) mutually constraining the multi-view age models with the spatial correspondence priors derived from the face bundles. Our multi-view age estimator performs well on a realistic evaluation dataset that contains faces under varying poses, and whose ground truth age was manually annotated.
Bingbing Ni, Dong Guo 0001, Terence Sim, Shuicheng Yan
ICCV4
2011 Multi-actor Emotion Recognition in Movies Using a Bimodal Approach
Ruchir Srivastava, Sujoy Roy, Shuicheng Yan, Terence Sim
MMM (2)4
2011 A study on recognizing non-artistic face sketches
abstract
Face sketches are being used in eyewitness testimonies for about a century. These sketches are crucial in finding suspects when no photo is available, but a mental image in the eyewitness's mind. However, research shows that current procedures used for eyewitness testimonies have two main problems. First, they can significantly disturb the memories of the eyewitness. Second, in many cases, these procedures result in face images far from their target faces. These two problems are related to the plasticity of the human visual system and the differences between face perception in humans (holistic) and current methods of sketch production (piecemeal). In this paper, we present some insights for more realistic sketch to photo matching. We describe how to retrieve identity specific information from crude sketches, directly drawn by the non-artistic eyewitnesses. The sketches we used merely contain facial component outlines and facial marks (e.g. wrinkles and moles). We compare results of automatically matching two types sketches (trace-over and user-provided, 25 each) to four types of faces (original, locally exaggerated, configurally exaggerated, and globally exaggerated, 249 each), using two methods (PDM distance comparison and PCA classification). Based on our results, we argue that for automatic non-artistic sketch to photo matching, the algorithms should compare the user-provided sketches with globally exaggerated faces, with a soft constraint on facial marks, to achieve the best matching rates. This is because the user-provided sketch from the user's mental image, seems to be caricatured both locally and configurally.
Hossein Nejati, Terence Sim
WACV2
2011 Defocus map estimation from a single image
Shaojie Zhuo, Terence Sim
Pattern Recognit.2
2010 Correcting over-exposure in photographs
abstract
This paper introduces a method to correct over-exposure in an existing photograph by recovering the color and lightness separately. First, the dynamic range of well exposed region is slightly compressed to make room for the recovered lightness of the over-exposed region. Then the lightness is recovered based on an over-exposure likelihood. The color of each pixel is corrected via neighborhood propagation and also based on the confidence of the original color. Previous methods make use of ratios between different color channels to recover the over-exposed ones, and thus can not handle regions where all three channels are over-exposed. In contrast, our method does not have this limitation. Our method is fully automatic and requires only one single input photo. We also provide users with the flexibility to control the amount of over-exposure correction. Experiment results demonstrate the effectiveness of the proposed method in correcting over-exposure.
Dong Guo 0001, Yuan Cheng 0005, Shaojie Zhuo, Terence Sim
CVPR4
2010 Towards general motion-based face recognition
abstract
Motion-based face recognition is a young research topic, inspired mainly by psychological studies on motion-based perception of human faces. Unlike its close relative, appearance-based face recognition, motion-based face recognition extracts personal characteristics from facial motion (e.g. smile) and uses the information to recognize human identity. However, existing studies in this field are limited to fixed motion, that is - a subject must perform a specific type of facial motion in order to be correctly recognized. In this paper, we try to overcome this limitation by investigating the patterns of local skin deformation exhibited in facial motion. We are pushing the state-of-the-art towards general motion-based face recognition. Our approach is able to extract identity evidence from various types of facial motion, as long as those facial motions are at least, in some part of the face, locally similar to the facial motions used in training. We call our approach Local Deformation Profile (or LDP). This approach is tested through several experiments conducted over a video database of facial expression. The experiment results demonstrate the potential of LDP to be used for biometrics. We also evaluate LDP under extremely heavy face makeup, showing its usefulness to recognize faces even in disguise.
Terence Sim
CVPR2
2010 Robust flash deblurring
abstract
10.1109/CVPR.2010.5539941
Shaojie Zhuo, Dong Guo 0001, Terence Sim
CVPR3
2010 Video stylization by single image example
abstract
In this paper, we propose a novel method for transferring artistic styles from a single image to a video clip. This new method can be considered as an extension of the image analogies technique in the domain of video processing. The method is built upon a previous fast texture transfer algorithm. The main contribution of our work lies in an extra constraint added to the original formulation, which preserves temporal coherence during frame-wise texture transfer. The experimental results show that this new constraint successfully helps suppress the artifacts of flickering, “stained-glass” and dragging, which are the main causes of temporal incoherence.
Terence Sim, Xiaoping Miao
ICIP2
2010 Enhancing low light images using near infrared flash images
abstract
In low light environment, photographs taken with a high ISO setting suffer from significant noise. In this paper, we propose to use a near infrared (NIR) flash image, instead of a normal visible flash image, to enhance its corresponding noisy visible image. We build a hybrid camera system to take an visible image and its NIR counterpart simultaneously. We introduce a new method to denoise an visible image and enhance its details using its corresponding NIR flash image. Experimental results show the superiority of our method compared with previous image denoising and detail enhancement methods.
Shaojie Zhuo, Xiaopeng Zhang 0001, Xiaoping Miao, Terence Sim
ICIP4
2010 Rotation invariant Facial Expression Recognition in image sequences
abstract
Facial Expression Recognition has mostly been done on frontal or near frontal faces. However, most of the faces in real life are non-frontal. This paper deals with in-plane rotation of faces in image sequences and considers the six universal facial expressions. The proposed approach does not need to rotate the image to frontal position. FER by rotating images to frontal is sensitive to determination of rotation angle and can involve errors in tracking facial points. Directions of motion of Facial Feature Points (FFPs) is used for feature extraction. In training for six expressions, Gaussian Mixture Models are fit to the distribution of angles representing these directions of motion. These models are used for further classification of test sequences using SVM. Gaussian Mixture Modeling is experimentally found to be robust to errors in position of FFPs. For dimensionality reduction, feature selection is performed using Fisher ratio test.
Ruchir Srivastava, Sujoy Roy, Terence Sim
ICME3
2009 Color Me Right-Seamless Image Compositing
Dong Guo 0001, Terence Sim
CAIP2
2009 Combining Facial Appearance and Dynamics for Face Recognition
Terence Sim
CAIP2
2009 On the Recovery of Depth from a Single Defocused Image
Shaojie Zhuo, Terence Sim
CAIP2
2009 Digital face makeup by example
abstract
This paper introduces an approach of creating face makeup upon a face image with another image as the style example. Our approach is analogous to physical makeup, as we modify the color and skin detail while preserving the face structure. More precisely, we first decompose the two images into three layers: face structure layer, skin detail layer, and color layer. Thereafter, we transfer information from each layer of one image to corresponding layer of the other image. One major advantage of the proposed method lies in that only one example image is required. This renders face makeup by example very convenient and practical. Equally, this enables some additional interesting applications, such as applying makeup by a portraiture. The experiment results demonstrate the effectiveness of the proposed approach in faithfully transferring makeup.
Dong Guo 0001, Terence Sim
CVPR2
2009 Simultaneous and orthogonal decomposition of data using Multimodal Discriminant Analysis
abstract
We present Multimodal Discriminant Analysis (MMDA), a novel method for decomposing variations in a dataset into independent factors (modes). For face images, MMDA effectively separates personal identity, illumination and pose into orthogonal subspaces. MMDA is based on maximizing the Fisher Criterion on all modes at the same time, and is therefore well-suited for multimodal and mode-invariant pattern recognition. We also show that MMDA may be used for dimension reduction, and for synthesizing images under novel illumination and even novel personal identity.
Terence Sim, Sheng Zhang 0007, Jianran Li
ICCV1
2008 Enhancing photographs with Near Infra-Red images
abstract
Near Infra-Red (NIR) images of natural scenes usually have better contrast and contain rich texture details that may not be perceived in visible light photographs (VIS). In this paper, we propose a novel method to enhance a photograph by using the contrast and texture information of its corresponding NIR image. More precisely, we first decompose the NIR/VIS pair into average and detail wavelet subbands. We then transfer the contrast in the average subband and transfer texture in the detail subbands. We built a special camera mount that optically aligns two consumer-grade digital cameras, one of which was modified to capture NIR. Our results exhibit higher visual quality than tone-mapped HDR images, showing that NIR imaging is useful for computational photography.
Xiaopeng Zhang 0001, Terence Sim, Xiaoping Miao
CVPR2
2008 Using targeted statistics for face regeneration
abstract
Face occlusion is a common problem that occurs in applications that analyze images for faces, e.g. detection, tracking and recognition. The presence of occlusion can adversely affect such face processing algorithms. This paper proposes a solution to the problem: we attempt to remove the occlusion by considering it as a damaged part that needs to be regenerated. More precisely, our technique learns the statistical correlation between different regions of the face without enforcing left-right symmetry. However, we learn only from face images that are similar to the target face (i.e. the face dataset is filtered to retain only similar faces). We show that such targeted statistics yield better results than statistics learned from faces in general. The occluded region is then regenerated by predicting its appearance from the most correlated unoccluded region of the same face. We also study how different factors influence our face regeneration technique: the effect of filtering the dataset; the presence/ absence of the target face during learning; the location of the occluded region; and the size of the occlusion. Our work can be used as a pre-processing step for face processing algorithms, or simply to enhance a face image for human viewing.
Terence Sim
FG2
2008 Smile, you're on identity camera
abstract
Inspired by recent advances in psychological studies on motion-based face perception, we examine in this paper, from the viewpoint of pattern recognition, the identity information behind a smile. A smile video database is collected, from which we compute dense optical flow fields and generate features by summing up the flow fields over time during the neutral-to-smile period. We investigate the relationship between smiles and identity by studying the class separability of the features. Our experiment results indicate a strong identity-specific characteristic of smile dynamics. Moreover, we compare the discriminating power of the features generated from different face regions as well as from different periods of motion.
Terence Sim
ICPR2
2008 Interactive Portrait Art
abstract
Traditionally, enjoying a portrait art, e.g. the Mona Lisa, is a passive activity. The spectator merely views the painting and admires the brush strokes, composition, etc. But now, with real-time computer vision and graphics algorithms, we can inject interactivity into portrait art, thereby bringing these art works back to life and giving a new dimension to art enjoyment. Specifically, in our art installation, a spectator is allowed to animate the face in a portrait art work to produce any expression she/he likes. The system consists of one personal computer and one camera. Given one frontal portrait picture as input, it generates an animation-ready avatar with minimum user intervention. The spectator then can perform before the camera any expression. And the system will capture the facial motion of the spectator and retarget the motion to the avatar. The animation of the avatar is rendered back to the original portrait picture. The motion retargetting is done in real-time.
Terence Sim
WACV2
2007 Are Digraphs Good for Free-Text Keystroke Dynamics?
abstract
Research in keystroke dynamics has largely focused on the typing patterns found in fixed text (e.g. userid and passwords). In this regard, digraphs and trigraphs have proven to be discriminative features. However, there is increasing interest in free-text keystroke dynamics, in which the user to be authenticated is free to type whatever he/she wants, rather than a pre-determined text. The natural question that arises is whether digraphs and trigraphs are just as discriminative for free text as they are for fixed text. We attempt to answer this question in this paper. We show that digraphs and trigraphs, if computed without regard to what word was typed, are no longer discriminative. Instead, word-specific digraphs/trigraphs are required. We also show that the typing dynamics for some words depend on whether they are part of a larger word. Our study is the first to investigate these issues, and we hope our work will help guide researchers looking for good features for free-text keystroke dynamics.
Terence Sim, Rajkumar Janakiraman
CVPR1
2007 VIM: Vision for Interactive Music
abstract
Traditionally, people were either producers of entertainment media, or else consumers of them. Today's digital entertainment, however, provides for a new dimension: that of interactivity. Instead of passive enjoyment, consumers can now control some elements of the media that were previously solely determined by the producer. This interactivity appears to enhance enjoyment. In this paper, we present a vision-based, interactive music playback system which allows anyone, even untrained musicians, to conduct music. The goal is to allow the user to dynamically influence how music is played back, much like what a real conductor would do. The tempo and volume of the music playback are controlled by the user's movements. In addition, our system projects colorful patterns that respond to the user, making the interaction truly multimedia
Terence Sim, Dennis Ng, Rajkumar Janakiraman
WACV1
2007 Continuous Verification Using Multimodal Biometrics
abstract
Conventional verification systems, such as those controlling access to a secure room, do not usually require the user to reauthenticate himself for continued access to the protected resource. This may not be sufficient for high-security environments in which the protected resource needs to be continuously monitored for unauthorized use. In such cases, continuous verification is needed. In this paper, we present the theory, architecture, implementation, and performance of a multimodal biometrics verification system that continuously verifies the presence of a logged-in user. Two modalities are currently used--face and fingerprint--but our theory can be readily extended to include more modalities. We show that continuous verification imposes additional requirements on multimodal fusion when compared to conventional verification systems. We also argue that the usual performance metrics of false accept and false reject rates are insufficient yardsticks for continuous verification and propose new metrics against which we benchmark our system.
Terence Sim, Sheng Zhang 0007, Rajkumar Janakiraman
IEEE Trans. Pattern Anal. Mach. Intell.1
2007 Discriminant Subspace Analysis: A Fukunaga-Koontz Approach
abstract
The Fisher Linear Discriminant (FLD) is commonly used in pattern recognition. It finds a linear subspace that maximally separates class patterns according to the Fisher Criterion. Several methods of computing the FLD have been proposed in the literature, most of which require the calculation of the so-called scatter matrices. In this paper, we bring a fresh perspective to FLD via the Fukunaga-Koontz Transform (FKT). We do this by decomposing the whole data space into four subspaces with different discriminability, as measured by eigenvalue ratios. By connecting the eigenvalue ratio with the generalized eigenvalue, we show where the Fisher Criterion is maximally satisfied. We prove the relationship between FLD and FKT analytically, and propose a unified framework to understanding some existing work. Furthermore, we extend our our theory to Multiple Discriminant Analysis (MDA). This is done by transforming the data into intra- and extra-class spaces, followed by maximizing the Bhattacharyya distance. Based on our FKT analysis, we identify the discriminant subspaces of MDA/FKT, and propose an efficient algorithm, which works even when the scatter matrices are singular, or too large to be formed. Our method is general and may be applied to different pattern recognition problems. We validate our method by experimenting on synthetic and real data.
Sheng Zhang 0007, Terence Sim
IEEE Trans. Pattern Anal. Mach. Intell.2
2006 When Fisher meets Fukunaga-Koontz: A New Look at Linear Discriminants
abstract
The Fisher Linear Discriminant (FLD) is commonly used in pattern recognition. It finds a linear subspace that maximally separates class patterns according to Fisher’s Criterion. Several methods of computing the FLD have been proposed in the literature, most of which require the calculation of the so-called scatter matrices. In this paper, we bring a fresh perspective to FLD via the Fukunaga-Koontz Transform (FKT). We do this by decomposing the whole data space into four subspaces, and show where Fisher’s Criterion is maximally satisfied. We prove the relationship between FLD and FKT analytically, and propose a method of computing the most discriminative subspace. This method is based on the QR decomposition, which works even when the scatter matrices are singular, or too large to be formed. Our method is general and may be applied to different pattern recognition problems. We validate our method by experimenting on synthetic and real data.
Sheng Zhang 0007, Terence Sim
CVPR (1)2
2005 Using Continuous Biometric Verification to Protect Interactive Login Sessions
abstract
In this paper we describe the theory, architecture, implementation, and performance of a multimodal passive biometric verification system that continually verifies the presence/participation of a logged-in user. We assume that the user logged in using strong authentication prior to the starting of the continuous verification process. While the implementation described in the paper combines a digital camera-based face verification with a mouse-based fingerprint reader, the architecture is generic enough to accommodate additional biometric devices with different accuracy of classifying a given user from an imposter. The main thrust of our work is to build a multimodal biometric feedback mechanism into the operating system so that verification failure can automatically lock up the computer within some estimate of the time it takes to subvert the computer. This must be done with low false positives in order to realize a usable system. We show through experimental results that combining multiple suitably chosen modalities in our theoretical framework can effectively do that with currently available off-the-shelf components
Terence Sim, Rajkumar Janakiraman, Sheng Zhang 0007
ACSAC2
2005 Realistic and efficient wrinkle simulation using an anatomy-based face model with adaptive refinement
abstract
This paper presents a geometric wrinkle model based on facial muscles for realistically and efficiently simulating wrinkles generated in facial expressions. Our method simulates wrinkles on an anatomy-based face model that mimics the layered structure of skin, muscles, and skull for facial animation. Corresponding to three types of facial muscles, a geometric model is developed to govern how the wrinkle amplitude evolves locally upon skin deformation. By taking into account the properties of real wrinkles, it provides intuitive parameters for easy control over wrinkle characteristics. During facial animation, wrinkles are generated in the local regions influenced by muscle contraction, simulating resistance to compression of tissues. In the simulation, an adaptive refinement automatically adapts the local resolution at which potential inaccuracies arc detected depending on local deformation. Thanks to its geometric nature, our method simulates wrinkles that can be dynamically rendered at interactive rates.
Yu Zhang 0075, Terence Sim
Computer Graphics International2
2005 Music transcription using an instrument model
abstract
We introduce a method to transcribe music with the help of an instrument model. One of the most important and difficult problems in music transcription is polyphonic pitch estimation. Common pitch estimation algorithms reported in the literature tend to have errors in the following three situations: missing fundamental; missing harmonics; shared frequencies. We believe an instrument model can make pitch estimation more robust in these three situations and thus can help to improve music transcription accuracy. We devise a spectrum subtraction algorithm to transcribe single and multiple instrument polyphonic music.
Terence Sim, Ye Wang 0007, Arun Shenoy
ICASSP (3)2
2005 Faces Alive: Reconstruction of Animated 3D Human Faces
Yu Zhang 0075, Terence Sim, Chew Lim Tan
ICCSA (3)2
2005 Ambient image recovery and rendering from flash photographs
abstract
It is challenging to capture good photographs in a low-light environment. Flash is often applied to increase the illumination, but flash ruins the natural ambient lighting and makes the scene look flat. Solutions without using flash include prolonging exposure time, enlarging the aperture or using high ISO film (or the equivalent setting for digital cameras). These brighten the image at the expense of image quality. In this paper, we present a method to avoids such problems. Our technique combines multiple photographs, taken under varying flash intensities, to recover the intrinsic ambient scene radiance. From this, we can re-render the scene under an arbitrary shutter speed to create a visually pleasant image. We can also simulate the effects of different white-balance settings, as well as different flash intensities. Experiments show that our method can produce high quality images compared to traditional non-flash solutions.
Xiaoping Miao, Terence Sim
ICIP (2)2
2004 Adaptation-Based Individualized Face Modeling for Animation Using Displacement Map
abstract
In this paper a new adaptation-based method is presented to reconstruct animatable facial models of human individuals from scan data. An anatomy-based generic control model serves as the starting point for our adaptation algorithm. Based on a series of measurements between the specified 3D landmarks, a global adaptation is carried out to align the generic control model to the scan surface. A local adaptation then deforms the geometry of the generic model to fit all of its vertices to the scan surface. The high-resolution geometry of the scan surface is represented as a displacement map which is generated using an offset-envelope mapping. Reconstruction of high-resolution geometry on the adapted generic mesh is achieved by hierarchical refinement using a surface subdivision scheme with resampling of the displacement map
Yu Zhang 0075, Terence Sim, Chew Lim Tan
Computer Graphics International2
2004 Stereo-based human read detection from crowd scenes
abstract
In this paper, a novel stereo-based head detection method is proposed for human detection in crowd scene. It contains three steps: (1) scale-adaptive filtering, (2) spurious clue suppression and (3) human head location. With the depth information, the sizes of human heads could be estimated. From this, 3D scale-adaptive filtering is proposed. It is applied for extracting the likelihood evidence of heads from the stereo image. In the second step, the extracted points whose positions in the real space are much higher or lower than the average human height above the ground surface are further suppressed. Finally, human heads are located by applying a mean-shift algorithm to the likelihood map. Good results of detecting human heads in crowds have been obtained from the experiments on real scene.
Liyuan Li, Terence Sim
ICIP3
2004 Reanimating real humans: automatic reconstruction of animated faces from range data
abstract
Advances in 3D scanning technology have enabled automatic capture of complex 3D models such as human faces with highly detailed surfaces. However, the range data cannot be used easily for animatable face modeling due to the absence of functional animation structure and dense surface data. The paper presents an automatic facial model adaptation algorithm for reconstruction of animatable individualized 3D facial models from range data. A generic model that represents both the face shape and anatomical structure serves as the starting point for the adaptation algorithm. The global adaptation transforms the generic model to align it with the scan data in the 3D space based on measurements between a set of 3D landmarks. The local adaptation then deforms the skin mesh of the generic model to fit all of its vertices to the scan surface. The underlying muscle structure is automatically adapted and facial texture is transferred. The reconstructed 3D face resembles the shape and color of a real individual and can be animated immediately with muscle parameters.
Yu Zhang 0075, Terence Sim, Chew Lim Tan
ICME2
2003 When Is the Shape of a Scene Unique Given Its Light-Field: A Fundamental Theorem of 3D Vision?
abstract
The complete set of measurements that could ever be used by a passive 3D vision algorithm is the plenoptic function or light-field. We give a concise characterization of when the light-field of a Lambertian scene uniquely determines its shape and, conversely, when the shape is inherently ambiguous. In particular, we show that stereo computed from the light-field is ambiguous if and only if the scene is radiating light of a constant intensity (and color, etc.) over an extended region.
Simon Baker, Terence Sim, Takeo Kanade
IEEE Trans. Pattern Anal. Mach. Intell.2
2003 The CMU Pose, Illumination, and Expression Database
abstract
In the Fall of 2000, we collected a database of more than 40,000 facial images of 68 people. Using the Carnegie Mellon University 3D Room, we imaged each person across 13 different poses, under 43 different illumination conditions, and with four different expressions. We call this the CMU pose, illumination, and expression (PIE) database. We describe the imaging hardware, the collection procedure, the organization of the images, several possible uses, and how to obtain the database.
Terence Sim, Simon Baker, Maan Bsat
IEEE Trans. Pattern Anal. Mach. Intell.1
2001 A Characterization of Inherent Stereo Ambiguities
Simon Baker, Terence Sim, Takeo Kanade
ICCV2
2000 Memory-Based Face Recognition for Visitor Identification
abstract
We show that a simple, memory-based technique for appearance-based face recognition, motivated by the real-world task of visitor identification, can outperform more sophisticated algorithms that use principal components analysis (PCA) and neural networks. This technique is closely related to correlation templates; however, we show that the use of novel similarity measures greatly improves performance. We also show that augmenting the memory base with additional, synthetic face images results in further improvements in performance. Results of extensive empirical testing on two standard face recognition datasets are presented, and direct comparisons with published work show that our algorithm achieves comparable (or superior) results. Our system is incorporated into an automated visitor identification system that has been operating successfully in an outdoor environment since January 1999.
Terence Sim, Rahul Sukthankar, Matthew D. Mullin, Shumeet Baluja
FG1