Vuong Le

dblp:09/7547 · DBLP profile ↗
← Back
32ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0003-1582-1269ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Confident and Trustworthy Model for Fidgety Movement Classification
abstract
General movements (GMs) are part of the spontaneous movement repertoire and are present from early fetal life onwards up to age five months. GMs are connected to infants' neurological development and can be qualitatively assessed via the General Movement Assessment (GMA). In particular, between the age of three to five months, typically developing infants produce Fidgety Movements (FM) and their absence provides strong evidence for the presence of cerebral palsy (CP). To improve accessibility to the GMA, automated GMA solutions have been a key research area with proposed models becoming increasingly more accurate and interpretable. However, current models cannot gauge their ability to make decisions, which may lead to overconfident mistakes. To address this issue, we propose a Deep learning-based approach that not only classifies movements as fidgety or non-fidgety but also selectively abstains from classification when uncertain. Through two novel regularization losses, our model maintains a balanced coverage across the two movement types, which prevents bias toward an easy-to-classify subset of movements. We show that our proposed model learns to gauge its own confidence on movement classification, and our proposed regularization losses effectively ensure that the model maintains a similar confidence across movement types. We also show that the local movement abstentions have little impact on the video-level coverage and that relying on the most confident predictions improves the video-level performance.
Romero F. A. B. de Morais, Thao Minh Le, Truyen Tran 0001, Caroline Alexander, Natasha Amery, Catherine Morgan, Alicia J. Spittle, Vuong Le, Nadia Badawi, Alison Salt, Jane Valentine, Catherine Elliott, Elizabeth M. Hurrion, Paul A. Dawson, Svetha Venkatesh
IEEE J. Biomed. Health Informatics8
2025 In-context Learning for Addressing User Cold-start in Sequential Movie Recommenders
Xurong Liang, Vu Nguyen 0001, Vuong Le, Paul Albert, Julien Monteil
RecSys3
2025 Fine-Grained Fidgety Movement Classification Using Active Learning
abstract
Typically developing infants, between the corrected age of 9-20 weeks, produce fidgety movements. These movements can be identified with the General Movement Assessment, but their identification requires trained professionals to conduct the assessment from video recordings. Since trained professionals are expensive and their demand may be higher than their availability, computer vision-based solutions have been developed to assist practitioners. However, most solutions to date treat the problem as a direct mapping from video to infant status, without modeling fidgety movements throughout the video. To address that, we propose to directly model infants' short movements and classify them as fidgety or non-fidgety. In this way, we model the explanatory factor behind the infant's status and improve model interpretability. The issue with our proposal is that labels for an infant's short movements are not available, which precludes us to train such a model. We overcome this issue with active learning. Active learning is a framework that minimizes the amount of labeled data required to train a model, by only labeling examples that are considered "informative" to the model. The assumption is that a model trained on informative examples reaches a higher performance level than a model trained with randomly selected examples. We validate our framework by modeling the movements of infants' hips on two representative cohorts: typically developing and at-risk infants. Our results show that active learning is suitable to our problem and that it works adequately even when the models are trained with labels provided by a novice annotator.
Romero F. A. B. de Morais, Truyen Tran 0001, Caroline Alexander, Natasha Amery, Catherine Morgan, Alicia J. Spittle, Vuong Le, Nadia Badawi, Alison Salt, Jane Valentine, Catherine Elliott, Elizabeth M. Hurrion, Paul A. Dawson, Svetha Venkatesh
IEEE J. Biomed. Health Informatics7
2024 Learning evolving relations for multivariate time series forecasting
abstract
Abstract Multivariate time series forecasting is essential in various fields, including healthcare and traffic management, but it is a challenging task due to the strong dynamics in both intra-channel relations (temporal patterns within individual variables) and inter-channel relations (the relationships between variables), which can evolve over time with abrupt changes. This paper proposes ERAN (Evolving Relational Attention Network), a framework for multivariate time series forecasting, that is capable to capture such dynamics of these relations. On the one hand, ERAN represents inter-channel relations with a graph which evolves over time, modeled using a recurrent neural network. On the other hand, ERAN represents the intra-channel relations using a temporal attentional convolution, which captures the local temporal dependencies adaptively with the input data. The elvoving graph structure and the temporal attentional convolution are intergrated in a unified model to capture both types of relations. The model is experimented on a large number of real-life datasets including traffic flows, energy consumption, and COVID-19 transmission data. The experimental results show a significant improvement over the state-of-the-art methods in multivariate time series forecasting particularly for non-stationary data.
Binh Nguyen-Thai, Vuong Le, Ngoc-Dung T. Tieu, Truyen Tran 0001, Svetha Venkatesh, Naeem Ramzan
Appl. Intell.2
2023 Persistent-Transient Duality: A Multi-mechanism Approach for Modeling Human-Object Interaction
abstract
Humans are highly adaptable, swiftly switching between different modes to progressively handle different tasks, situations and contexts. In Human-object interaction (HOI) activities, these modes can be attributed to two mechanisms: (1) the large-scale consistent plan for the whole activity and (2) the small-scale children interactive actions that start and end along the timeline. While neuroscience and cognitive science have confirmed this multi-mechanism nature of human behavior, machine modeling approaches for human motion are trailing behind. While attempting to use gradually morphing structures (e.g., graph attention networks) to model the dynamic HOI patterns, they miss the expeditious and discrete mode-switching nature of the human motion. To bridge that gap, this work proposes to model two concurrent mechanisms that jointly control human motion: the Persistent process that runs continually on the global scale, and the Transient sub-processes that operate intermittently on the local context of the human while interacting with objects. These two mechanisms form an interactive Persistent-Transient Duality that synergistically governs the activity sequences. We model this conceptual duality by a parent-child neural network of Persistent and Transient channels with a dedicated neural module for dynamic mechanism switching. The framework is trialed on HOI motion forecasting. On two rich datasets and a wide variety of settings, the model consistently delivers superior performances, proving its suitability for the challenge.
Vuong Le, Svetha Venkatesh, Truyen Tran 0001
ICCV2
2023 Guiding Visual Question Answering with Attention Priors
abstract
The current success of modern visual reasoning systems is arguably attributed to cross-modality attention mechanisms. However, in deliberative reasoning such as in VQA, attention is unconstrained at each step, and thus may serve as a statistical pooling mechanism rather than a semantic operation intended to select information relevant to inference. This is because at training time, attention is only guided by a very sparse signal (i.e. the answer label) at the end of the inference chain. This causes the cross-modality attention weights to deviate from the desired visual-language bindings. To rectify this deviation, we propose to guide the attention mechanism using explicit linguistic-visual grounding. This grounding is derived by connecting structured linguistic concepts in the query to their referents among the visual objects. Here we learn the grounding from the pairing of questions and images alone, without the need for answer annotation or external grounding supervision. This grounding guides the attention mechanism inside VQA models through a duality of mechanisms: pre-training attention weight calculation and directly guiding the weights at inference time on a case- by-case basis. The resultant algorithm is capable of probing attention-based reasoning models, injecting relevant associative knowledge, and regulating the core reasoning process. This scalable enhancement improves the performance of VQA models, fortifies their robustness to limited access to supervised data, and increases interpretability.
Thao Minh Le, Vuong Le, Sunil Gupta 0001, Svetha Venkatesh, Truyen Tran 0001
WACV2
2023 Robust and Interpretable General Movement Assessment Using Fidgety Movement Detection
abstract
Fidgety movements occur in infants between the age of 9 to 20 weeks post-term, and their absence are a strong indicator that an infant has cerebral palsy. Prechtl's General Movement Assessment method evaluates whether an infant has fidgety movements, but requires a trained expert to conduct it. Timely evaluation facilitates early interventions, and thus computer-based methods have been developed to aid domain experts. However, current solutions rely on complex models or high-dimensional representations of the data, which hinder their interpretability and generalization ability. To address that we propose [Formula: see text], a method that detects fidgety movements and uses them towards an assessment of the quality of an infant's general movements. [Formula: see text] is true to the domain expert process, more accurate, and highly interpretable due to its fine-grained scoring system. The main idea behind [Formula: see text] is to specify signal properties of fidgety movements that are measurable and quantifiable. In particular, we measure the movement direction variability of joints of interest, for movements of small amplitude in short video segments. [Formula: see text] also comprises a strategy to reduce those measurements to a single score that quantifies the quality of an infant's general movements; the strategy is a direct translation of the qualitative procedure domain experts use to assess infants. This brings [Formula: see text] closer to the process a domain expert applies to decide whether an infant produced enough fidgety movements. We evaluated [Formula: see text] on the largest clinical dataset reported, where it showed to be interpretable and more accurate than many methods published to date.
Romero F. A. B. de Morais, Vuong Le, Catherine Morgan, Alicia J. Spittle, Nadia Badawi, Jane Valentine, Elizabeth M. Hurrion, Paul A. Dawson, Truyen Tran 0001, Svetha Venkatesh
IEEE J. Biomed. Health Informatics2
2022 Video Dialog as Conversation About Objects Living in Space-Time
Thao Minh Le, Vuong Le, Tu Minh Phuong, Truyen Tran 0001
ECCV (39)3
2022 Real-Time Skill Discovery in Intelligent Virtual Assistants
Preeti Gopal, Sunil Gupta 0001, Santu Rana, Vuong Le, Trong Nguyen, Svetha Venkatesh
PAKDD (1)4
2022 The three ghosts of medical AI: Can the black-box present deliver?
Thomas P. Quinn, Stephan Jacobs, Manisha Senadeera, Vuong Le, Simon Coghlan
Artif. Intell. Medicine4
2021 Learning Asynchronous and Sparse Human-Object Interaction in Videos
abstract
Human activities can be learned from video. With effective modeling it is possible to discover not only the action labels but also the temporal structure of the activities, such as the progression of the sub-activities. Automatically recognizing such structure from raw video signal is a new capability that promises authentic modeling and successful recognition of human-object interactions. Toward this goal, we introduce Asynchronous-Sparse Interaction Graph Networks (ASSIGN), a recurrent graph network that is able to automatically detect the structure of interaction events associated with entities in a video scene. ASSIGN pioneers learning of autonomous behavior of video entities including their dynamic structure and their interaction with the coexisting neighbors. Entities’ lives in our model are asynchronous to those of others therefore more flexible in adapting to complex scenarios. Their interactions are sparse in time hence more faithful to the true underlying nature and more robust in inference and learning. ASSIGN is tested on human-object interaction recognition and shows superior performance in segmenting and labeling of human sub-activities and object affordances from raw videos. The native ability of ASSIGN in discovering temporal structure also eliminates the dependence on external segmentation that was previously mandatory for this task.
Romero F. A. B. de Morais, Vuong Le, Svetha Venkatesh, Truyen Tran 0001
CVPR2
2021 Hierarchical Object-oriented Spatio-Temporal Reasoning for Video Question Answering
abstract
Video Question Answering (Video QA) is a powerful testbed to develop new AI capabilities. This task necessitates learning to reason about objects, relations, and events across visual and linguistic domains in space-time. High-level reasoning demands lifting from associative visual pattern recognition to symbol like manipulation over objects, their behavior and interactions. Toward reaching this goal we propose an object-oriented reasoning approach in that video is abstracted as a dynamic stream of interacting objects. At each stage of the video event flow, these objects interact with each other, and their interactions are reasoned about with respect to the query and under the overall context of a video. This mechanism is materialized into a family of general-purpose neural units and their multi-level architecture called Hierarchical Object-oriented Spatio-Temporal Reasoning (HOSTR) networks. This neural model maintains the objects' consistent lifelines in the form of a hierarchically nested spatio-temporal graph. Within this graph, the dynamic interactive object-oriented representations are built up along the video sequence, hierarchically abstracted in a bottom-up manner, and converge toward the key information for the correct answer. The method is evaluated on multiple major Video QA datasets and establishes new state-of-the-arts in these tasks. Analysis into the model's behavior indicates that object-oriented reasoning is a reliable, interpretable and efficient approach to Video QA.
Long Hoang Dang, Thao Minh Le, Vuong Le, Truyen Tran 0001
IJCAI3
2021 Object-Centric Representation Learning for Video Question Answering
abstract
Video question answering (Video QA) presents a powerful testbed for human-like intelligent behaviors. The task demands new capabilities to integrate video processing, language understanding, binding abstract linguistic concepts to concrete visual artifacts, and deliberative reasoning over spacetime. Neural networks offer a promising approach to reach this potential through learning from examples rather than handcrafting features and rules. However, neural networks are predominantly feature-based - they map data to unstructured vectorial representation and thus can fall into the trap of exploiting shortcuts through surface statistics instead of true systematic reasoning seen in symbolic systems. To tackle this issue, we advocate for object-centric representation as a basis for constructing spatio-temporal structures from videos, essentially bridging the semantic gap between low-level pattern recognition and high-level symbolic algebra. To this end, we propose a new query-guided representation framework to turn a video into an evolving relational graph of objects, whose features and interactions are dynamically and conditionally inferred. The object lives are then summarized into résumés, lending naturally for deliberative relational reasoning that produces an answer to the query. The framework is evaluated on major Video QA datasets, demonstrating clear benefits of the object-centric approach to video reasoning.
Long Hoang Dang, Thao Minh Le, Vuong Le, Truyen Tran 0001
IJCNN3
2021 From Deep Learning to Deep Reasoning
abstract
The rise of big data and big compute has brought modern neural networks to many walks of digital life, thanks to the relative ease of constructing large models that scale to the real world. Current successes of Transformers and self-supervised pretraining on massive data have led some to believe that deep neural networks will be able to do almost everything once we have sufficient data and computational resources. However, neural networks are fast to exploit surface statistics but fail miserably to generalize to novel combinations. This is because they are not designed for deliberate reasoning -- the capacity to deliberately deduce new knowledge out of the contextualized data. This tutorial reviews recent developments to extend the capacity of neural networks to "learning-to-reason'' from data, where the task is to determine if the data entails a conclusion. This capacity opens up new ways to generate insights from data through arbitrary compositional querying without the need of predefining a narrow set of tasks. The tutorial consists of four parts. The first part covers the learning-to-reason framework, and explains how neural networks can serve as a strong backbone for reasoning through its natural operations such as binding, attention & dynamic computational graphs. The second part goes into more detail on how neural networks perform reasoning over unstructured and structured data, and across modalities. The third part reviews neural memories and their role in reasoning. The last part discusses generalization to novel combinations, under less supervision and with more knowledge.
Truyen Tran 0001, Vuong Le, Hung Le 0002, Thao Minh Le
KDD2
2021 Goal-driven Long-Term Trajectory Prediction
abstract
The prediction of humans' short-term trajectories has advanced significantly with the use of powerful sequential modeling and rich environment feature extraction. However, long-term prediction is still a major challenge for the current methods as the errors could accumulate along the way. Indeed, consistent and stable prediction far to the end of a trajectory inherently requires deeper analysis into the overall structure of that trajectory, which is related to the pedestrian's intention on the destination of the journey. In this work, we propose to model a hypothetical process that determines pedestrians' goals and the impact of such process on long-term future trajectories. We design Goal-driven Trajectory Prediction model - a dual-channel neural network that realizes such intuition. The two channels of the network take their dedicated roles and collaborate to generate future trajectories. Different than conventional goal-conditioned, planning-based methods, the model architecture is designed to generalize the patterns and work across different scenes with arbitrary geometrical and semantic structures. The model is shown to outperform the state-of-the-art in various settings, especially in large prediction horizons. This result is another evidence for the effectiveness of adaptive structured representation of visual and geometrical features in human behavior analysis.
Vuong Le, Truyen Tran 0001
WACV2
2021 Hierarchical Conditional Relation Networks for Multimodal Video Question Answering
Thao Minh Le, Vuong Le, Svetha Venkatesh, Truyen Tran 0001
Int. J. Comput. Vis.2
2021 Trust and medical AI: the challenges we face and the expertise needed to overcome them
abstract
Artificial intelligence (AI) is increasingly of tremendous interest in the medical field. How-ever, failures of medical AI could have serious consequences for both clinical outcomes and the patient experience. These consequences could erode public trust in AI, which could in turn undermine trust in our healthcare institutions. This article makes 2 contributions. First, it describes the major conceptual, technical, and humanistic challenges in medical AI. Second, it proposes a solution that hinges on the education and accreditation of new expert groups who specialize in the development, verification, and operation of medical AI technologies. These groups will be required to maintain trust in our healthcare institutions.
Thomas P. Quinn, Manisha Senadeera, Stephan Jacobs, Simon Coghlan, Vuong Le
J. Am. Medical Informatics Assoc.5
2021 A Spatio-Temporal Attention-Based Model for Infant Movement Assessment From Videos
abstract
The absence or abnormality of fidgety movements of joints or limbs is strongly indicative of cerebral palsy in infants. Developing computer-based methods for assessing infant movements in videos is pivotal for improved cerebral palsy screening. Most existing methods use appearance-based features and are thus sensitive to strong but irrelevant signals caused by background clutter or a moving camera. Moreover, these features are computed over the whole frame, thus they measure gross whole body movements rather than specific joint/limb motion. Addressing these challenges, we develop and validate a new method for fidgety movement assessment from consumer-grade videos using human poses extracted from short clips. Human poses capture only relevant motion profiles of joints and limbs and are thus free from irrelevant appearance artifacts. The dynamics and coordination between joints are modeled using spatio-temporal graph convolutional networks. Frames and body parts that contain discriminative information about fidgety movements are selected through a spatio-temporal attention mechanism. We validate the proposed model on the cerebral palsy screening task using a real-life consumer-grade video dataset collected at an Australian hospital through the Cerebral Palsy Alliance, Australia. Our experiments show that the proposed method achieves the ROC-AUC score of 81.87%, significantly outperforming existing competing methods with better interpretability.
Binh Nguyen-Thai, Vuong Le, Catherine Morgan, Nadia Badawi, Truyen Tran 0001, Svetha Venkatesh
IEEE J. Biomed. Health Informatics2
2020 Learning to Abstract and Predict Human Actions
Romero F. A. B. de Morais, Vuong Le, Truyen Tran 0001, Svetha Venkatesh
BMVC2
2020 Hierarchical Conditional Relation Networks for Video Question Answering
abstract
Video question answering (VideoQA) is challenging as it requires modeling capacity to distill dynamic visual artifacts and distant relations and to associate them with linguistic concepts. We introduce a general-purpose reusable neural unit called Conditional Relation Network (CRN) that serves as a building block to construct more sophisticated structures for representation and reasoning over video. CRN takes as input an array of tensorial objects and a conditioning feature, and computes an array of encoded output objects. Model building becomes a simple exercise of replication, rearrangement and stacking of these reusable units for diverse modalities and contextual information. This design thus supports high-order relational and multi-step reasoning. The resulting architecture for VideoQA is a CRN hierarchy whose branches represent sub-videos or clips, all sharing the same question as the contextual condition. Our evaluations on well-known datasets achieved new SoTA results, demonstrating the impact of building a general-purpose reasoning unit on complex domains such as VideoQA.
Thao Minh Le, Vuong Le, Svetha Venkatesh, Truyen Tran 0001
CVPR2
2020 Dynamic Language Binding in Relational Visual Reasoning
abstract
We present Language-binding Object Graph Network, the first neural reasoning method with dynamic relational structures across both visual and textual domains with applications in visual question answering. Relaxing the common assumption made by current models that the object predicates pre-exist and stay static, passive to the reasoning process, we propose that these dynamic predicates expand across the domain borders to include pair-wise visual-linguistic object binding. In our method, these contextualized object links are actively found within each recurrent reasoning step without relying on external predicative priors. These dynamic structures reflect the conditional dual-domain object dependency given the evolving context of the reasoning through co-attention. Such discovered dynamic graphs facilitate multi-step knowledge combination and refinements that iteratively deduce the compact representation of the final answer. The effectiveness of this model is demonstrated on image question answering demonstrating favorable performance on major VQA datasets. Our method outperforms other methods in sophisticated question-answering tasks wherein multiple object relations are involved. The graph structure effectively assists the progress of training, and therefore the network learns efficiently compared to other reasoning models.
Thao Minh Le, Vuong Le, Svetha Venkatesh, Truyen Tran 0001
IJCAI2
2020 Neural Reasoning, Fast and Slow, for Video Question Answering
abstract
What does it take to design a machine that learns to answer natural questions about a video? A Video QA system must simultaneously understand language, represent visual content over space-time, and iteratively transform these representations in response to lingual content in the query, and finally arriving at a sensible answer. While recent advances in lingual and visual question answering have enabled sophisticated representations and neural reasoning mechanisms, major challenges in Video QA remain on dynamic grounding of concepts, relations and actions to support the reasoning process. Inspired by the dual-process account of human reasoning, we design a dual process neural architecture, which is composed of a question-guided video processing module (System 1, fast and reactive) followed by a generic reasoning module (System 2, slow and deliberative). System 1 is a hierarchical model that encodes visual patterns about objects, actions and relations in space-time given the textual cues from the question. The encoded representation is a set of high-level visual features, which are then passed to System 2. Here multi-step inference follows to iteratively chain visual elements as instructed by the textual elements. The system is evaluated on the SVQA (synthetic) and TGIF-QA datasets (real), demonstrating competitive results, with a large margin in the case of multi-step reasoning.
Thao Minh Le, Vuong Le, Svetha Venkatesh, Truyen Tran 0001
IJCNN2
2020 Scalable Backdoor Detection in Neural Networks
Haripriya Harikumar, Vuong Le, Santu Rana, Sourangshu Bhattacharya, Sunil Gupta 0001, Svetha Venkatesh
ECML/PKDD (2)2
2019 Learning Regularity in Skeleton Trajectories for Anomaly Detection in Videos
abstract
Appearance features have been widely used in video anomaly detection even though they contain complex entangled factors. We propose a new method to model the normal patterns of human movements in surveillance video for anomaly detection using dynamic skeleton features. We decompose the skeletal movements into two sub-components: global body movement and local body posture. We model the dynamics and interaction of the coupled features in our novel Message-Passing Encoder-Decoder Recurrent Network. We observed that the decoupled features collaboratively interact in our spatio-temporal model to accurately identify human-related irregular events from surveillance video sequences. Compared to traditional appearance-based models, our method achieves superior outlier detection performance. Our model also offers “open-box” examination and decision explanation made possible by the semantically understandable features and a network architecture supporting interpretability.
Romero F. A. B. de Morais, Vuong Le, Truyen Tran 0001, Budhaditya Saha, Moussa Reda Mansour, Svetha Venkatesh
CVPR2
2019 Memorizing Normality to Detect Anomaly: Memory-Augmented Deep Autoencoder for Unsupervised Anomaly Detection
abstract
Deep autoencoder has been extensively used for anomaly detection. Training on the normal data, the autoencoder is expected to produce higher reconstruction error for the abnormal inputs than the normal ones, which is adopted as a criterion for identifying anomalies. However, this assumption does not always hold in practice. It has been observed that sometimes the autoencoder "generalizes" so well that it can also reconstruct anomalies well, leading to the miss detection of anomalies. To mitigate this drawback for autoencoder based anomaly detector, we propose to augment the autoencoder with a memory module and develop an improved autoencoder called memory-augmented autoencoder, i.e. MemAE. Given an input, MemAE firstly obtains the encoding from the encoder and then uses it as a query to retrieve the most relevant memory items for reconstruction. At the training stage, the memory contents are updated and are encouraged to represent the prototypical elements of the normal data. At the test stage, the learned memory will be fixed, and the reconstruction is obtained from a few selected memory records of the normal data. The reconstruction will thus tend to be close to a normal sample. Thus the reconstructed errors on anomalies will be strengthened for anomaly detection. MemAE is free of assumptions on the data type and thus general to be applied to different tasks. Experiments on various datasets prove the excellent generalization and high effectiveness of the proposed MemAE.
Dong Gong, Lingqiao Liu, Vuong Le, Budhaditya Saha, Moussa Reda Mansour, Svetha Venkatesh, Anton van den Hengel
ICCV3
2014 Accurate Facial Landmarks Detection for Frontal Faces with Extended Tree-Structured Models
abstract
In this paper, we aim to improve one of the current state-of-the-art models for facial components detection/localization. The objectives are to increase the amount of landmark points detected and improve the landmark extraction accuracy for frontal faces. The model is following Zhu and Ramanan's approach with a tree-structure. The popular AR dataset is chosen as an alternative training dataset as it provides more landmark points requested. Our extension models are compared with Zhu and Ramanan's frontal face models in terms of detection accuracy. We also compare our models with another robust facial components detector called CompASM. Our experiments show that our models can achieve lower error rate on some fiducial points by providing more landmarks, and these accurate fiducial points will provide more accurate features for some applications related to facial shapes. The impact of image colour spaces other than RGB on the proposed detector is also investigated.
Antoni Liang, Wanquan Liu, Ling Li 0006, Mir Rizwan Farid, Vuong Le
ICPR5
2012 Interactive Facial Feature Localization
Vuong Le, Jonathan Brandt, Zhe Lin 0001, Lubomir D. Bourdev, Thomas S. Huang
ECCV (3)1
2012 Recognizing Emotions From an Ensemble of Features
abstract
This paper details the authors' efforts to push the baseline of emotion recognition performance on the Geneva Multimodal Emotion Portrayals (GEMEP) Facial Expression Recognition and Analysis database. Both subject-dependent and subject-independent emotion recognition scenarios are addressed in this paper. The approach toward solving this problem involves face detection, followed by key-point identification, then feature generation, and then, finally, classification. An ensemble of features consisting of hierarchical Gaussianization, scale-invariant feature transform, and some coarse motion features have been used. In the classification stage, we used support vector machines. The classification task has been divided into person-specific and person-independent emotion recognitions using face recognition with either manual labels or automatic algorithms. We achieve 100% performance for the person-specific one, 66% performance for the person-independent one, and 80% performance for overall results, in terms of classification rate, for emotion recognition with manual identification of subjects.
Usman Tariq, Kai-Hsiang Lin, Zhen Li 0028, Vuong Le, Thomas S. Huang, Xutao Lv, Tony X. Han
IEEE Trans. Syst. Man Cybern. Part B6
2011 Expression recognition from 3D dynamic faces using robust spatio-temporal shape features
abstract
This paper proposes a new method for comparing 3D facial shapes using facial level curves. The pair- and segment-wise distances between the level curves comprise the spatio-temporal features for expression recognition from 3D dynamic faces. The paper further introduces universal background modeling and maximum a posteriori adaptation for hidden Markov models, leading to a decision boundary focus classification algorithm. Both techniques, when combined, yield a high overall recognition accuracy of 92.22% on the BU-4DFE database in our preliminary experiments. Noticeably, our feature extraction method is very efficient, requiring simple preprocessing, and robust to variations of the input data quality.
Vuong Le, Hao Tang 0001, Thomas S. Huang
FG1
2011 Emotion recognition from an ensemble of features
abstract
This work details the authors' efforts to push the baseline of expression recognition performance on a realistic database. Both subject-dependent and subject-independent emotion recognition scenarios are addressed in this work. These two happen frequently in real life settings. The approach towards solving this problem involves face detection, followed by key point identification, then feature generation and then finally classification. An ensemble of features comprising of Hierarchial Gaussianization (HG), Scale Invariant Feature Transform (SIFT) and Optic Flow have been incorporated. In the classification stage we used SVMs. The classification task has been divided into person specific and person independent emotion recognition. Both manual labels and automatic algorithms for person verification have been attempted. They both give similar performance.
Usman Tariq, Kai-Hsiang Lin, Zhen Li 0028, Vuong Le, Thomas S. Huang, Xutao Lv, Tony X. Han
FG6
2010 Accurate and efficient reconstruction of 3D faces from stereo images
abstract
In this paper, we propose a novel algorithm for reconstructing the 3D shape and texture of human faces from two stereo images, which are captured from calibrated cameras. Our approach works in a sparse to dense manner: we first build a coarse shape estimation based on 3D keypoints, and then use a linear morphable model to efficiently match the detail shape and texture. Compared with the previous works, our algorithm can reconstruct the 3D face shape in a speed comparable with that of the fastest algorithm available, but gives a higher accuracy. It can also recover the texture with more complete, realistic looking. Our results show that the new algorithm possesses significant characteristics of a 3D face model reconstruction system, and is especially useful for face recognition and animation applications in practice.
Vuong Le, Hao Tang 0001, Liangliang Cao, Thomas S. Huang
ICIP1
2009 A quantitative evaluation for 3D face reconstruction algorithms
abstract
In this work, we proposed to use quantitative method to evaluate the accuracy of 3D face reconstruction algorithms. The reconstructed 3D faces are first aligned to the ground truth by iterative closest point (ICP) algorithm and then the shape difference between the two 3D faces is described by signal to noise ratio (SNR). Finally, the error maps (EM) illustrated the reconstruction errors on corresponded vertices in different dimensions. Comparing with the subjective and indirect evaluation methods, the proposed method provides more precise and detailed evaluations for face shape reconstruction. Based on the SNR, different 3D face reconstruction algorithms can be compared directly and the EM also can suggest guidance for feature extraction.
Vuong Le, Yuxiao Hu 0001, Thomas S. Huang
ICASSP1