Liming Chen 0002

dblp:32/7029-2 · DBLP profile ↗
← Back
144ranked-venue papers
1as first author
25since 2021 · last 2026
0000-0002-3654-9498ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 92 · 1 first-author · 12 since 2021Artificial intelligence and machine learning · 67 · 19 since 2021Human-computer interaction and ubiquitous computing · 14Security and privacy · 7Databases, data management, data science and information retrieval · 7Systems, architecture and hardware · 5 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Computer networks · 2
YearPublicationVenuePosition
2026 DCFANet: Merging dynamic context clustering mamba and context-to-focus attention for medical image segmentation
Xiaoyan Kui, Zhipeng Hu, Zexin Ji, Shen Jiang, Qianmu Xiao, Ziwei Zou, Qinsong Li, Yang Li 0111, Beiji Zou 0001, Liming Chen 0002
Neurocomputing10
2026 Alzheimer's disease classification based on multimodal consistent distribution and trusted fusion
Xiaoyan Kui, Yulan Dai, Beiji Zou 0001, Chengzhang Zhu, Yang Li 0111, Zexin Ji, Liming Chen 0002, Miguel Bordallo López
Neural Networks7
2025 NoPain: No-box Point Cloud Attack via Optimal Transport Singular Boundary
abstract
Adversarial attacks exploit the vulnerability of deep models against adversarial samples. Existing point cloud attackers are tailored to specific models, iteratively optimizing perturbations based on gradients in either a white-box or black-box setting. Despite their promising attack performance, they often struggle to produce transferable adversarial samples due to overfitting to the specific parameters of surrogate models. To overcome this issue, we shift our focus to the data distribution itself and introduce a novel approach named NoPain, which employs optimal transport (OT) to identify the inherent singular boundaries of the data manifold for cross-network point cloud attacks. Specifically, we first calculate the OT mapping from noise to the target feature space, then identify singular boundaries by locating non-differentiable positions. Finally, we sample along singular boundaries to generate adversarial point clouds. Once the singular boundaries are determined, NoPain can efficiently produce adversarial samples without the need of iterative updates or guidance from the surrogate classifiers. Extensive experiments demonstrate that the proposed end-to-end method outperforms baseline approaches in terms of both transferability and efficiency, while also maintaining notable advantages even against defense strategies. Code and model are available at https://github.com/cognaclee/nopain.
Zezeng Li, Na Lei, Liming Chen 0002, Weimin Wang 0007
CVPR4
2025 Facilitate and Scale Up the Creation of 3D Meshes, 6D Category-Based Datasets and Grasping with Generative Models: GenVegeFruits3D
abstract
Despite significant advances in 2D image, enabled by foundation models, progress in 3D understanding, particularly in 6D pose estimation and shape reconstruction, remains limited by the scarcity of datasets. In particular, in category 6D pose estimation, the high costs of real-world data collection have resulted in datasets with restricted categories, low diversity, and minimal instance variability. Recent methods have attempted to address this gap by leveraging synthetic image generation tools. However, these approaches are constrained by the limitations of available mesh datasets, which hinder the diversity, scalability, and inclusion of novel objects in generated samples. In this work, we propose a first automatic pipeline for the large-scale generation of 3D category-based datasets. Our approach uses 3D generative models guided by textual input to produce diverse and scalable datasets. To demonstrate its efficacy, we generated a new dataset named GenVegeFruits3D comprising 100 categories of fruits and vegetables, each containing over 1000 unique meshes. This significantly enhances the scale and diversity of existing category-based 3D datasets while reducing reliance on pre-existing 3D meshes. Additionally, we trained a 3D generative model, a 3D understanding model, and a grasping model, including on a real robotic setup. The dataset and code are available at: GenVegeFruits3D.
Guillaume Duret, Younes Bourennane, Danylo Mazurak, Anna Samsonenko, Florence Zara, Jan Peters 0001, Liming Chen 0002
ICIP7
2025 Online Continual Learning of Diffusion Models: Multi-Mode Adaptive Generative Distillation
abstract
Continual learning typically relies on storing real data, which is impractical in privacy-sensitive settings. Generative replay with diffusion models offers a high-fidelity alternative. However, in online continual learning (OCL), these models struggle with catastrophic forgetting and incur high computational costs from frequent updates and sampling. Existing distillation methods reduce generation steps but rely on a fixed teacher model, limiting their effectiveness as data distributions evolve. To address these, we introduce Multi-Mode Adaptive Generative Distillation (MAGD), which incorporates two innovative techniques: Noisy Intermediate Generative Distillation (NIGD) and SNR-Guided Generative Distillation (SGGD). NIGD leverages intermediate noisy images, created during the reverse process rather than by adding noise post-generation, to enhance knowledge transfer. SGGD uses a signal-to-noise ratio (SNR) based threshold to optimize the sampling of time steps, reducing unnecessary generation. Guided by an Exponential Moving Average (EMA) teacher, MAGD effectively mitigates catastrophic forgetting as it adapts to new data streams. Experiments on Fashion-MNIST, CIFAR-10, and CIFAR-100 show that MAGD reduces generation overhead by up to 25% relative to standard generative distillation and 92% compared to DDGR-1000, while maintaining generating quality. Furthermore, in class-conditioned diffusion models, MAGD outperforms memory-based methods in terms of classification accuracy.
Matthieu Grard, Emmanuel Dellandréa, Liming Chen 0002
ICIP4
2025 Noise Optimized Conditional Diffusion for Domain Adaptation
abstract
Pseudo-labeling is a cornerstone of Unsupervised Domain Adaptation (UDA), yet the scarcity of High-Confidence Pseudo-Labeled Target Domain Samples (hcpl-tds) often leads to inaccurate cross-domain statistical alignment, causing DA failures. To address this challenge, we propose Noise Optimized Conditional Diffusion for Domain Adaptation (NOCDDA), which seamlessly integrates the generative capabilities of conditional diffusion models with the decision-making requirements of DA to achieve task-coupled optimization for efficient adaptation. For robust cross-domain consistency, we modify the DA classifier to align with the conditional diffusion classifier within a unified optimization framework, enabling forward training on noise-varying cross-domain samples. Furthermore, we argue that the conventional N(0,I) initialization in diffusion models often generates class-confused hcpl-tds, compromising discriminative DA. To resolve this, we introduce a class-aware noise optimization strategy that refines sampling regions for reverse class-specific hcpl-tds generation, effectively enhancing cross-domain alignment. Extensive experiments across 5 benchmark datasets and 29 DA tasks demonstrate significant performance gains of NOCDDA over 31 state-of-the-art methods, validating its robustness and effectiveness.
Lingkun Luo, Shiqiang Hu, Liming Chen 0002
IJCAI3
2025 PK-Net: A prior knowledge-driven dual-path network for enhanced glaucoma screening
Xiaoyan Kui, Zeru Hai, Beiji Zou 0001, Yang Li 0111, Wei Liang 0005, Zuheng Ming, Liming Chen 0002
Knowl. Based Syst.7
2025 Beyond Batch Learning: Global Awareness Enhanced Domain Adaptation
abstract
In domain adaptation (DA), the effectiveness of deep learning-based models is often constrained by batch learning strategies that fail to fully apprehend the global statistical and geometric characteristics of data distributions. Addressing this gap, we introduce "Global Awareness Enhanced Domain Adaptation" (GAN-DA), a novel approach that transcends traditional batch-based limitations. GAN-DA integrates a unique predefined feature representation (PFR) to facilitate the alignment of cross-domain distributions, thereby achieving a comprehensive global statistical awareness. This representation is innovatively expanded to encompass orthogonal and common feature aspects, which enhances the unification of global manifold structures and refines decision boundaries for more effective DA. Our extensive experiments, encompassing 27 diverse cross-domain image classification tasks, demonstrate GAN-DA's remarkable superiority, outperforming 24 established DA methods by a significant margin. Furthermore, our in-depth analyses shed light on the decision-making processes, revealing insights into the adaptability and efficiency of GAN-DA. This approach not only addresses the limitations of existing DA methodologies but also sets a new benchmark in the realm of domain adaptation, offering broad implications for future research and applications in this field.
Lingkun Luo, Shiqiang Hu, Liming Chen 0002
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 ImFace++: A Sophisticated Nonlinear 3D Morphable Face Model With Implicit Neural Representations
abstract
Accurate representations of 3D faces are of paramount importance in various computer vision and graphics applications. However, the challenges persist due to the limitations imposed by data discretization and model linearity, which hinder the precise capture of identity and expression clues in current studies. This paper presents a novel 3D morphable face model, named ImFace++, to learn a sophisticated and continuous space with implicit neural representations. ImFace++ first constructs two explicitly disentangled deformation fields to model complex shapes associated with identities and expressions, respectively, which simultaneously facilitate automatic learning of point-to-point correspondences across diverse facial shapes. To capture more sophisticated facial details, a refinement displacement field within the template space is further incorporated, enabling fine-grained learning of individual-specific facial details. Furthermore, a Neural Blend-Field is designed to reinforce the representation capabilities through adaptive blending of an array of local fields. In addition to ImFace++, we devise an improved learning strategy to extend expression embeddings, allowing for a broader range of expression variations. Comprehensive qualitative and quantitative evaluation demonstrates that ImFace++ significantly advances the state-of-the-art in terms of both face reconstruction fidelity and correspondence accuracy.
Mingwu Zheng, Hongyu Yang 0001, Liming Chen 0002, Di Huang 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Adaptive representation learning and sample weighting for low-quality 3D face recognition
Cuican Yu, Fengxun Sun, Huibin Li 0001, Liming Chen 0002, Jian Sun 0009, Zongben Xu
Pattern Recognit.5
2025 ChebMixer: Efficient Graph Representation Learning With MLP Mixer
abstract
Graph neural networks (GNNs) have achieved remarkable success in learning graph representations, especially graph Transformers, which have recently shown superior performance on various graph mining tasks. However, the graph Transformer generally treats nodes as tokens, which results in quadratic complexity regarding the number of nodes during self-attention computation. The graph multilayer perceptron (MLP) mixer addresses this challenge using the efficient MLP Mixer technique from computer vision. However, the time-consuming process of extracting graph tokens limits its performance. In this article, we present a novel architecture named ChebMixer, a newly proposed graph MLP Mixer that uses fast Chebyshev polynomials-based spectral filtering to extract a sequence of tokens. First, we produce multiscale representations of graph nodes via fast Chebyshev polynomial-based spectral filtering. Next, we consider each node's multiscale representations as a sequence of tokens and refine the node representation with an effective MLP Mixer. Finally, we aggregate the multiscale representations of nodes through Chebyshev interpolation. Owing to the powerful representation capabilities and fast computational properties of the MLP Mixer, we can quickly extract more informative node representations to improve the performance of downstream tasks. The experimental results prove our significant improvements in various scenarios, ranging from homogeneous and heterophilic graph node classification to medical image segmentation. Compared with NAGphormer, the average performance improved by 1.45% on homogeneous graphs and 4.15% on heterophilic graphs. And the average performance improved by 1.39% on medical image segmentation tasks compared with VM-UNet. We will release the source code after this article is accepted.
Xiaoyan Kui, Haonan Yan, Qinsong Li, Liming Chen 0002, Beiji Zou 0001
IEEE Trans. Neural Networks Learn. Syst.5
2024 Imbalanced Data Robust Online Continual Learning Based on Evolving Class Aware Memory Selection and Built-In Contrastive Representation Learning
abstract
We introduce Memory Selection with Contrastive Learning (MSCL), an advanced Continual Learning (CL) approach, addressing challenges in dynamic and imbalanced environments. MSCL combines Feature-Distance Based Sample Selection (FDBS) for memory management, focusing on inter-class similarities and intra-class diversity, with a contrastive learning loss (IWL) for adaptive data representation. Our evaluations on datasets like MNIST, Cifar-100, miniImageNet, PACS, and DomainNet show that MSCL not only competes with but often surpasses existing memory-based CL methods, particularly in imbalanced scenarios, enhancing both balanced and imbalanced learning performance.
Emmanuel Dellandréa, Matthieu Grard, Liming Chen 0002
ICIP4
2024 Discriminative Noise Robust Sparse Orthogonal Label Regression-Based Domain Adaptation
Lingkun Luo, Shiqiang Hu, Liming Chen 0002
Int. J. Comput. Vis.3
2023 Weakly-Supervised Photo-realistic Texture Generation for 3D Face Reconstruction
abstract
Although much progress has been made recently in 3D face reconstruction, most previous work has been devoted to predicting accurate and fine-grained 3D shapes. In contrast, relatively little work has focused on generating high-fidelity face textures. Compared with the prosperity of photo-realistic 2D face image generation, high-fidelity 3D face texture generation has yet to be studied. In this paper, we propose a novel UV map generation model that predicts the UV map from a single face image. The model consists of a UV sampler and a UV generator. By selectively sampling the input face image's pixels and adjusting their relative locations, the UV sampler generates an incomplete UV map that could faithfully reconstruct the original face. Missing textures in the incomplete UV map are further full-filled by the UV generator. The training is based on pseudo ground truth blended by the 3DMM texture and the input face texture, thus weakly supervised. To deal with the artifacts in the imperfect pseudo UV map, multiple UV map and face image discriminators are leveraged.
Xiangnan Yin, Di Huang 0001, Zehua Fu, Yunhong Wang 0001, Liming Chen 0002
FG5
2023 Segmentation-Reconstruction-Guided Facial Image De-occlusion
abstract
Occlusions are very common in face images in the wild, leading to the degraded performance of face-related tasks. Although much effort has been devoted to removing occlusions from face images, the varying shapes and textures of occlusions still challenge the robustness of current methods. As a result, current methods either rely on manual occlusion masks or only apply to specific occlusions. This paper proposes a novel face de-occlusion model based on face segmentation and 3D face reconstruction, which is robust to arbitrary kinds of face occlusions. The proposed model consists of a 3D face reconstruction module, a face segmentation module, and an image generation module. With the face prior and the occlusion mask predicted by the first two, respectively, the image generation module can faithfully recover the missing facial textures. To supervise the training, we further build a large occlusion dataset, with both manually labeled and synthetic occlusions. Qualitative and quantitative results demonstrate the effectiveness and robustness of the proposed method.
Xiangnan Yin, Di Huang 0001, Zehua Fu, Yunhong Wang 0001, Liming Chen 0002
FG5
2022 ImFace: A Nonlinear 3D Morphable Face Model with Implicit Neural Representations
abstract
Precise representations of 3D faces are beneficial to various computer vision and graphics applications. Due to the data discretization and model linearity, however, it remains challenging to capture accurate identity and expression clues in current studies. This paper presents a novel 3D morphable face model, namely ImFace, to learn a nonlinear and continuous space with implicit neural representations. It builds two explicitly disentangled deformation fields to model complex shapes associated with identities and expressions, respectively, and designs an improved learning strategy to extend embeddings of expressions to allow more diverse changes. We further introduce a Neural Blend-Field to learn sophisticated details by adaptively blending a series of local fields. In addition to ImFace, an effective pre-processing pipeline is proposed to address the issue of watertight input requirement in implicit representations, enabling them to work with common facial surfaces for the first time. Extensive experiments are performed to demonstrate the superiority of ImFace.
Mingwu Zheng, Hongyu Yang 0001, Di Huang 0001, Liming Chen 0002
CVPR4
2022 Non-Deterministic Face Mask Removal Based on 3d Priors
abstract
This paper presents a novel image inpainting framework for face mask removal. Although current methods have demonstrated their impressive ability in recovering damaged face images, they suffer from two main problems: the dependence on manually labeled missing regions and the deterministic result corresponding to each input. The proposed approach tackles these problems by integrating a multi-task 3D face reconstruction module with a face inpainting module. Given a masked face image, the former predicts a 3DMM-based reconstructed face together with a binary occlusion map, providing dense geometrical and textural priors that greatly facilitate the inpainting task of the latter. By gradually controlling the 3D shape parameters, our method generates high-quality dynamic in-painting results with different expressions and mouth movements. Qualitative and quantitative experiments verify the effectiveness of the proposed method. Our code: https://github.com/face3d0725/face_de_mask
Xiangnan Yin, Di Huang 0001, Liming Chen 0002
ICIP3
2022 Look Beyond Bias with Entropic Adversarial Data Augmentation
abstract
Deep neural networks do not discriminate between spurious and causal patterns, and will only learn the most predictive ones while ignoring the others. This shortcut learning behaviour is detrimental to a network’s ability to generalize to an unknown test-time distribution in which the spurious correlations do not hold anymore. Debiasing methods were developed to make networks robust to such spurious biases but require to know in advance if a dataset is biased and make heavy use of minority counter-examples that do not display the majority bias of their class. In this paper, we argue that such samples should not be necessarily needed because the “hidden” causal information is often also contained in biased images. To study this idea, we propose 3 publicly released synthetic classification benchmarks, exhibiting predictive classification shortcuts, each of a different and challenging nature, without any minority samples acting as counter-examples. First, we investigate the effectiveness of several state-of-the-art strategies on our benchmarks and show that they do not yield satisfying results on them. Then, we propose an architecture able to succeed on our benchmarks, despite their unusual properties, using an entropic adversarial data augmentation training scheme. An encoder-decoder architecture is tasked to produce images that are not recognized by a classifier, by maximizing the conditional entropy of its outputs, and keep as much as possible of the initial content. A precise control of the information destroyed, via a disentangling process, enables us to remove the shortcut and leave everything else intact. Furthermore, results competitive with the state-of-the-art on the BAR dataset ensure the applicability of our method in real-life situations.
Thomas Duboudin, Emmanuel Dellandréa, Corentin Abgrall, Gilles Hénaff, Liming Chen 0002
ICPR5
2022 MES-P: An Emotional Tonal Speech Dataset in Mandarin with Distal and Proximal Labels
abstract
Emotion shapes all aspects of our interpersonal and intellectual experiences. Its automatic analysis has therefore many applications. In this paper, we propose an emotional tonal speech dataset, Mandarin Chinese Emotional Speech Dataset-Portrayed (MES-P), with both distal and proximal labels. In contrast with state of the art datasets which only focused on perceived emotions, MES-P includes not only perceived emotions (proximal labels) but also intended emotions (distal labels), to make it possible to study human emotional intelligence, i.e., emotion expression/understanding ability, and emotional misunderstandings in real life. Furthermore, MES-P also captures a main feature of tonal languages, and provides emotional speech samples matching the tonal distribution in real life Mandarin. MES-P dataset also features emotion intensity variations, by introducing both moderate and intense versions for joy, anger, and sadness, in addition to neutral. Ratings of the collected speech samples are made in valence-arousal space through continuous coordinate locations, resulting in an emotional distribution pattern in 2D VA space. High consistency between the speakers emotional intentions and the listeners perceptions is also proved by Cohens Kappa coefficients. Finally, extensive experiments are carried out as a baseline on MES-P for automatic emotion recognition and with comparison to human emotion intelligence.
Zhongzhe Xiao, Weibei Dou, Liming Chen 0002
IEEE Trans. Affect. Comput.5
2022 Attention Regularized Laplace Graph for Domain Adaptation
abstract
In leveraging manifold learning in domain adaptation (DA), graph embedding-based DA methods have shown their effectiveness in preserving data manifold through the Laplace graph. However, current graph embedding DA methods suffer from two issues: 1). they are only concerned with preservation of the underlying data structures in the embedding and ignore sub-domain adaptation, which requires taking into account intra-class similarity and inter-class dissimilarity, thereby leading to negative transfer; 2). manifold learning is proposed across different feature/label spaces separately, thereby hindering unified comprehensive manifold learning. In this paper, starting from our previous DGA-DA, we propose a novel DA method, namely A ttention R egularized Laplace G raph-based D omain A daptation (ARG-DA), to remedy the aforementioned issues. Specifically, by weighting the importance across different sub-domain adaptation tasks, we propose the A ttention R egularized Laplace Graph for class aware DA, thereby generating the attention regularized DA. Furthermore, using a specifically designed FEEL strategy, our approach dynamically unifies alignment of the manifold structures across different feature/label spaces, thus leading to comprehensive manifold learning. Comprehensive experiments are carried out to verify the effectiveness of the proposed DA method, which consistently outperforms the state of the art DA methods on 7 standard DA benchmarks, i.e., 37 cross-domain image classification tasks including object, face, and digit images. An in-depth analysis of the proposed DA method is also discussed, including sensitivity, convergence, and robustness.
Lingkun Luo, Liming Chen 0002, Shiqiang Hu
IEEE Trans. Image Process.2
2022 The Devil is in the Details: An Efficient Convolutional Neural Network for Transport Mode Detection
abstract
Transport mode detection is a classification problem aiming to design an algorithm that can infer the transport mode of a user given multimodal signals (GPS and/or inertial sensors). It has many applications, such as carbon footprint tracking, mobility behaviour analysis, or real-time door-to-door smart planning. Most current approaches rely on a classification step using Machine Learning techniques, and, like in many other classification problems, deep learning approaches usually achieve better results than traditional machine learning ones using handcrafted features. Deep models, however, have a notable downside: they are usually heavy, both in terms of memory space and processing cost. We show that a small, optimized model can perform as well as a current deep model. During our experiments on the GeoLife and SHL 2018 datasets, we obtain models with tens of thousands of parameters, that is, 10 to 1,000 times less parameters and operations than networks from the state of the art, which still reach a comparable performance. We also show, using the aforementioned datasets, that the current preprocessing used to deal with signals of different lengths is suboptimal, and we provide better replacements. Finally, we introduce a way to use signals with different lengths with the lighter Convolutional neural networks, without using the heavier Recurrent Neural Networks.
Hugues Moreau, Andrea Vassilev, Liming Chen 0002
IEEE Trans. Intell. Transp. Syst.3
2021 A 3D GAN for Improved Large-Pose Facial Recognition
abstract
Facial recognition using deep convolutional neural networks relies on the availability of large datasets of face images. Many examples of identities are needed, and for each identity, a large variety of images are needed in order for the network to learn robustness to intra-class variation. In practice, such datasets are difficult to obtain, particularly those containing adequate variation of pose. Generative Adversarial Networks (GANs) provide a potential solution to this problem due to their ability to generate realistic, synthetic images. However, recent studies have shown that current methods of disentangling pose from identity are inadequate. In this work we incorporate a 3D morphable model into the generator of a GAN in order to learn a nonlinear texture model from in-the-wild images. This allows generation of new, synthetic identities, and manipulation of pose, illumination and expression without compromising the identity. Our synthesised data is used to augment training of facial recognition networks with performance evaluated on the challenging CFP and CPLFW datasets.
Richard T. Marriott, Sami Romdhani, Liming Chen 0002
CVPR3
2021 Data Fusion for Deep Learning on Transport Mode Detection: A Case Study
Hugues Moreau, Andrea Vassilev, Liming Chen 0002
EANN3
2021 Scoring Graspability based on Grasp Regression for Better Grasp Prediction
abstract
Grasping objects is one of the most important abilities that a robot needs to master in order to interact with its environment. Current state-of-the-art methods rely on deep neural networks trained to jointly predict a graspability score together with a regression of an offset with respect to grasp reference parameters. However, these two predictions are performed independently, which can lead to a decrease in the actual graspability score when applying the predicted offset. Therefore, in this paper, we extend a state-of-the-art neural network with a scorer that evaluates the graspability of a given position, and introduce a novel loss function which correlates regression of grasp parameters with graspability score. We show that this novel architecture improves performance from 82.13% for a state-of-the-art grasp detection network to 85.74% on Jacquard dataset. When the learned model is transferred onto a real robot, the proposed method correlating graspability and grasp regression achieves a 92.4% rate compared to 88.1% for the baseline trained without the correlation.
Amaury Depierre, Emmanuel Dellandréa, Liming Chen 0002
ICRA3
2021 Intensity enhancement via GAN for multimodal face expression recognition
Hongyu Yang 0001, Kangkang Zhu, Di Huang 0001, Hebeizi Li, Yunhong Wang 0001, Liming Chen 0002
Neurocomputing6
2020 Taking Control of Intra-class Variation in Conditional GANs Under Weak Supervision
abstract
Generative Adversarial Networks (GANs) are able to learn mappings between simple, relatively low-dimensional, random distributions and points on the manifold of realistic images in image-space. The semantics of this mapping, however, are typically entangled such that meaningful image properties cannot be controlled independently of one another. Conditional GANs (cGANs) provide a potential solution to this problem, allowing specific semantics to be enforced during training. This solution, however, depends on the availability of precise labels, which are sometimes difficult or near impossible to obtain, e.g. labels representing lighting conditions or describing the background. In this paper we introduce a new formulation of the cGAN that is able to learn disentangled, multivariate models of semantically meaningful variation and which has the advantage of requiring only the weak supervision of binary attribute labels. For example, given only labels of ambient / non-ambient lighting, our method is able to learn multivariate lighting models disentangled from other factors such as the identity and pose. We coin the method intra-class variation isolation (IVI) and the resulting network the IVI-GAN. We evaluate IVI-GAN on the CelebA dataset and on synthetic 3D morphable model data, learning to disentangle attributes such as lighting, pose, expression, and even the background.
Richard T. Marriott, Sami Romdhani, Liming Chen 0002
FG3
2020 An Assessment of GANs for Identity-related Applications
abstract
Generative Adversarial Networks (GANs) are now capable of producing synthetic face images of exceptionally high visual quality. In parallel to the development of GANs themselves, efforts have been made to develop metrics to objectively assess the characteristics of the synthetic images, mainly focusing on visual quality and the variety of images. Little work has been done, however, to assess overfitting of GANs and their ability to generate new identities. In this paper we apply a state of the art biometric network to various datasets of synthetic images and perform a thorough assessment of their identity-related characteristics. We conclude that GANs can indeed be used to generate new, imagined identities meaning that applications such as anonymisation of image sets and augmentation of training datasets with distractor images are viable applications. We also assess the ability of GANs to disentangle identity from other image characteristics and propose a novel GAN triplet loss that we show to improve this disentanglement.
Richard T. Marriott, Safa Madiouni, Sami Romdhani, Stéphane Gentric, Liming Chen 0002
IJCB5
2020 Pixel Sampling for Style Preserving Face Pose Editing
abstract
The existing auto-encoder based face pose editing methods primarily focus on modeling the identity preserving ability during pose synthesis, but are less able to preserve the image style properly, which refers to the color, brightness, saturation, etc. In this paper, we take advantage of the well-known frontal/profile optical illusion and present a novel two-stage approach to solve the aforementioned dilemma, where the task of face pose manipulation is cast into face inpainting. By selectively sampling pixels from the input face and slightly adjust their relative locations with the proposed “Pixel Attention Sampling” module, the face editing result faithfully keeps the identity information as well as the image style unchanged. By leveraging high-dimensional embedding at the inpainting stage, finer details are generated. Further, with the 3D facial landmarks as guidance, our method is able to manipulate face pose in three degrees of freedom, i.e., yaw, pitch, and roll, resulting in more flexible face pose editing than merely controlling the yaw angle as usually achieved by the current state-of-the-art. Both the qualitative and quantitative evaluations validate the superiority of the proposed approach.
Xiangnan Yin, Di Huang 0001, Hongyu Yang 0001, Zehua Fu, Yunhong Wang 0001, Liming Chen 0002
IJCB6
2020 Cross-Year Multi-Modal Image Retrieval Using Siamese Networks
abstract
This paper introduces a multi-modal network that learns to retrieve by content vertical aerial images of French urban and rural territories taken about 15 years apart. This means it should be invariant against a big range of changes as the (natural) landscape evolves over time. It leverages the original images and semantically segmented and labeled regions. The core of the method is a Siamese network that learns to extract features from corresponding image pairs across time. These descriptors are discriminative enough, such that a simple kNN classifier on top, suffices as final geo-matching criteria. The method outperformed SOTA “off-the-shelf' image descriptors GEM and ResNet50 on the new aerial images dataset.
Margarita Khokhlova, Valérie Gouet-Brunet, Nathalie Abadie, Liming Chen 0002
ICIP4
2020 Intensity Enhancement Via Gan for Multimodal Facial Expression Recognition
abstract
Face expression recognition (FER) on low intensity is not well studied in the literature. This paper investigates this new problem and presents a novel Generative Adversarial Network (GAN) based multimodal approach to it. The method models the tasks of intensity enhancement and expression recognition jointly, ensuring that the synthesize faces not only present expression of high intensity, but also truly contribute to promoting the performance of FER. Extensive experiments are conducted on the BU-3DFE and BU-4DFE datasets. State-of-the-art FER performance clearly validates the effectiveness of the proposed method.
Kangkang Zhu, Yunhong Wang 0001, Hongyu Yang 0001, Di Huang 0001, Liming Chen 0002
ICIP5
2020 Face and Gesture Analysis for Health Informatics
abstract
The goal of Face and Gesture Analysis for Health Informatics's workshop is to share and discuss the achievements as well as the challenges in using computer vision and machine learning for automatic human behavior analysis and modeling for clinical research and healthcare applications. The workshop aims to promote current research and support growth of multidisciplinary collaborations to advance this groundbreaking research. The meeting gathers scientists working in related areas of computer vision and machine learning, multi-modal signal processing and fusion, human centered computing, behavioral sensing, assistive technologies, and medical tutoring systems for healthcare applications and medicine.
Zakia Hammal, Di Huang 0001, Kevin Bailly, Liming Chen 0002, Mohamed Daoudi
ICMI4
2020 SUMAC 2020: The 2nd Workshop on Structuring and Understanding of Multimedia heritAge Contents
abstract
SUMAC 2020 is the second edition of the workshop on Structuring and Understanding of Multimedia heritAge Contents. It is held in Seattle, USA on October 12th, 2020 and is co-located with the 28th ACM International Conference on Multimedia; this year, due to the sanitary crisis, it is organized virtually. Its objective is to present and discuss the latest and most significant trends and challenges in the analysis, structuring and understanding of multimedia contents dedicated to the valorization of heritage, with the emphasis on the unlocking of and access to the big data of the past. A representative scope of Computer Science methodologies dedicated to the processing of multimedia heritage contents and their exploitation is covered by the works presented, with the ambition of advancing and raising awareness about this fully developing research field.
Valérie Gouet-Brunet, Margarita Khokhlova, Ronak Kosti, Liming Chen 0002, Xu-Cheng Yin
ACM Multimedia4
2020 Deep Multicameral Decoding for Localizing Unoccluded Object Instances from a Single RGB Image
Matthieu Grard, Emmanuel Dellandréa, Liming Chen 0002
Int. J. Comput. Vis.3
2020 Discriminative and Geometry-Aware Unsupervised Domain Adaptation
abstract
Domain adaptation (DA) aims to generalize a learning model across training and testing data despite the mismatch of their data distributions. In light of a theoretical estimation of the upper error bound, we argue, in this article, that an effective DA method for classification should: 1) search a shared feature subspace where the source and target data are not only aligned in terms of distributions as most state-of-the-art DA methods do but also discriminative in that instances of different classes are well separated and 2) account for the geometric structure of the underlying data manifold when inferring data labels on the target domain. In comparison with a baseline DA method which only cares about data distribution alignment between source and target, we derive three different DA models for classification, namely, close yet discriminative DA (CDDA), geometry-aware DA (GA-DA), and discriminative and GA-DA (DGA-DA), to highlight the contribution of CDDA based on 1), GA-DA based on 2), and, finally, DGA-DA implementing jointly 1) and 2). Using both the synthetic and real data, we show the effectiveness of the proposed approach which consistently outperforms the state-of-the-art DA methods over 49 image classification DA tasks through eight popular benchmarks. We further carry out an in-depth analysis of the proposed DA method in quantifying the contribution of each term of our DA model and provide insights into the proposed DA methods in visualizing both real and synthetic data.
Lingkun Luo, Liming Chen 0002, Shiqiang Hu, Ying Lu 0007
IEEE Trans. Cybern.2
2019 Discriminative Attention-based Convolutional Neural Network for 3D Facial Expression Recognition
abstract
3D Facial Expression Recognition (FER) is an active research area in computer vision. Although previous methods report promising results, two key issues still remain to be solved. On the one hand, different facial areas contribute unequally to performing various expressions, but most existing methods extract features from the entire 3D surface. On the other hand, the difference between expressions varies, while previous methods generally treat different emotions equally, making some of them extremely hard to be distinguished. To solve these problems, we propose a novel approach for 3D FER, namely Discriminative Attention-based Convolution Neural Network (DA-CNN), to generate more comprehensive expression related representations. DA-CNN introduces an attention module to the CNN models, which helps the deep model selectively focus on emotional salient regions in a learnable way. Furthermore, a novel loss named Dimensional Distribution (DD) loss is proposed to model the inter-expression relationship. Supervised by DD loss, DA-CNN can generate more discriminative expression representation. Extensive experiments are conducted on BU-3DFE dataset, and the results show that DA-CNN achieves significant improvement over the state-of-the-art.
Kangkang Zhu, Zhengyin Du, Weixin Li 0001, Di Huang 0001, Yunhong Wang 0001, Liming Chen 0002
FG6
2019 SUMAC 2019: The 1st workshop on Structuring and Understanding of Multimedia heritAge Contents
abstract
SUMAC 2019 is the first workshop on Structuring and Understanding of Multimedia heritAge Contents. It is held in Nice, France on October 21, 2019 and is co-located with the 27th ACM International Conference on Multimedia. Its objective is to present and discuss the latest and most significant trends and challenges in the analysis, structuring and understanding of multimedia contents dedicated to the valorization of heritage, with the emphasis on the unlocking of and access to the big data of the past. A representative scope of Computer Science methodologies dedicated to the processing of multimedia heritage contents and their exploitation is covered by the works presented, with the ambition of advancing and raising awareness about this fully developing research field.
Valérie Gouet-Brunet, Margarita Khokhlova, Liming Chen 0002, Sander Münster
ACM Multimedia3
2019 Improving Shadow Suppression for Illumination Robust Face Recognition
abstract
2D face analysis techniques, such as face landmarking, face recognition and face verification, are reasonably dependent on illumination conditions which are usually uncontrolled and unpredictable in the real world. The current massive data-driven approach, e.g., deep learning-based face recognition, requires a huge amount of labeled training face data that hardly cover the infinite lighting variations that can be encountered in real-life applications. An illumination robust preprocessing method thus remains a very interesting but also a significant challenge in reliable face analysis. In this paper we propose a novel model driven approach to improve lighting normalization of face images. Specifically, we propose to build the underlying reflectance model which characterizes interactions between skin surface, lighting source and camera sensor, and elaborate the formation of face color appearance. The proposed illumination processing pipeline enables generation of the Chromaticity Intrinsic Image (CII) in a log chromaticity space which is robust to illumination variations. Moreover, as an advantage over most prevailing methods, a photo-realistic color face image is subsequently reconstructed, which eliminates a wide variety of shadows whilst retaining the color information and identity details. Experimental results under different scenarios and using various face databases show the effectiveness of the proposed approach in dealing with lighting variations, including both soft and hard shadows, in face recognition.
Wuming Zhang, Xi Zhao 0001, Jean-Marie Morvan, Liming Chen 0002
IEEE Trans. Pattern Anal. Mach. Intell.4
2019 Expression Robust 3D Facial Landmarking via Progressive Coarse-to-Fine Tuning
abstract
Facial landmarking is a fundamental task in automatic machine-based face analysis. The majority of existing techniques for such a problem are based on 2D images; however, they suffer from illumination and pose variations that may largely degrade landmarking performance. The emergence of 3D data theoretically provides an alternative to overcome these weaknesses in the 2D domain. This article proposes a novel approach to 3D facial landmarking, which combines both the advantages of feature-based methods as well as model-based ones in a progressive three-stage coarse-to-fine manner (initial, intermediate, and fine stages). For the initial stage, a few fiducial landmarks (i.e., the nose tip and two inner eye corners) are robustly detected through curvature analysis, and these points are further exploited to initialize the subsequent stage. For the intermediate stage, a statistical model is learned in the feature space of three normal components of the facial point-cloud rather than the smooth original coordinates, namely Active Normal Model (ANM). For the fine stage, cascaded regression is employed to locally refine the landmarks according to their geometry attributes. The proposed approach can accurately localize dozens of fiducial points on each 3D face scan, greatly surpassing the feature-based ones, and it also improves the state of the art of the model-based ones in two aspects: sensitivity to initialization and deficiency in discrimination. The proposed method is evaluated on the BU-3DFE, Bosphorus, and BU-4DFE databases, and competitive results are achieved in comparison with counterparts in the literature, clearly demonstrating its effectiveness.
Di Huang 0001, Yunhong Wang 0001, Liming Chen 0002
ACM Trans. Multim. Comput. Commun. Appl.4
2018 JEMImE: A Serious Game to Teach Children with ASD How to Adequately Produce Facial Expressions
abstract
Being able to produce facial expressions (FEs) that are adequate given a social context is key to harmonious social development, particularly in the case of children plagued with autism spectrum disorder (ASD). In this paper, we introduce JEMImE, a serious game solution that aims at teaching children how to produce FEs. JEMImE is based on a FE recognition module that is learned on a large video corpus of children performing FEs. This module is validated and incorporated through multiple scenarios of gradual difficulty, ranging from a training phase where children have to perform the FEs on request, with or without an avatar model, to an in-context phase that involves many emotion-eliciting social situations with virtual characters.
Arnaud Dapogny, Charline Grossard, Stéphanie Hun, Sylvie Serret, Jeremy Bourgeois, Hedy Jean-Marie, Pierre Foulon, Huaxiong Ding, Liming Chen 0002, Séverine Dubuisson, Ouriel Grynszpan, Kevin Bailly
FG9
2018 Accurate Facial Parts Localization and Deep Learning for 3D Facial Expression Recognition
abstract
Meaningful facial parts can convey key cues for both facial action unit detection and expression prediction. Textured 3D face scan can provide both detailed 3D geometric shape and 2D texture appearance cues of the face which are beneficial for Facial Expression Recognition (FER). However, accurate facial parts extraction as well as their fusion are challenging tasks. In this paper, a novel system for 3D FER is designed based on accurate facial parts extraction and deep feature fusion of facial parts. Experiments are conducted on the BU-3DFE database, demonstrating the effectiveness of combing different facial parts, texture and depth cues and reporting the state-of-the-art results in comparison with all existing methods under the same setting.
Asim Jan, Huaxiong Ding, Hongying Meng, Liming Chen 0002, Huibin Li 0001
FG4
2018 Unsupervised Domain Adaptation with Regularized Optimal Transport for Multimodal 2D+3D Facial Expression Recognition
abstract
Since human expressions have strong flexibility and personality, subject-independent facial expression recognition is a typical data bias problem. To address this problem, we propose a novel approach, namely unsupervised domain adaptation with regularized optimal transport for multimodal 2D+3D Facial Expression Recognition (FER). In particular, Wasserstein distance is employed to measure the distribution inconsistency between the training samples (i.e. source domain) and test samples (i.e. target domain). Minimization of this Wasserstein distance is equivalent to finding an optimal transport mapping from training to test samples. Once we find this mapping, original training samples can be transformed into a new space in which the distributions of the mapped training samples and the test samples can be well-aligned. In this case, classifier learned from the transformed training samples can be well generalized to the test samples for expression prediction. In practice, approximate optimal transport can be effectively solved by adding entropy regularization. To fully explore the class label information of training samples, group sparsity regularizer is also used to enforce that the training samples from the same expression class can be mapped to the same group. Experimental results evaluated on the BU-3DFE and Bosphorus databases demonstrate that the proposed approach can achieve superior performance compared with the state-of-the-art methods.
Xiaofan Wei, Huibin Li 0001, Jian Sun 0009, Liming Chen 0002
FG4
2018 Jacquard: A Large Scale Dataset for Robotic Grasp Detection
abstract
Grasping skill is a major ability that a wide number of real-life applications require for robotisation. State-of-the-art robotic grasping methods perform prediction of object grasp locations based on deep neural networks. However, such networks require huge amount of labeled data for training making this approach often impracticable in robotics. In this paper, we propose a method to generate a large scale synthetic dataset with ground truth, which we refer to as the Jacquard grasping dataset. Jacquard is built on a subset of ShapeNet, a large CAD models dataset, and contains both RGB-D images and annotations of successful grasping positions based on grasp attempts performed in a simulated environment. We carried out experiments using an off-the-shelf CNN, with three different evaluation metrics, including real grasping robot trials. The results show that Jacquard enables much better generalization skills than a human labeled dataset thanks to its diversity of objects and grasping positions. For the purpose of reproducible research in robotics, we are releasing along with the Jacquard dataset a web interface for researchers to evaluate the successfulness of their grasping position detections using our dataset.
Amaury Depierre, Emmanuel Dellandréa, Liming Chen 0002
IROS3
2018 Fast and Light Manifold CNN based 3D Facial Expression Recognition across Pose Variations
abstract
This paper proposes a novel approach to 3D Facial Expression Recognition (FER), and it is based on a Fast and Light Manifold CNN model, namely FLM-CNN. Different from current manifold CNNs, FLM-CNN adopts a human vision inspired pooling structure and a multi-scale encoding strategy to enhance geometry representation, which highlights shape characteristics of expressions and runs efficiently. Furthermore, a sampling tree based preprocessing method is presented, and it sharply saves memory when applied to 3D facial surfaces, without much information loss of original data. More importantly, due to the property of manifold CNN features of being rotation-invariant, the proposed method shows a high robustness to pose variations. Extensive experiments are conducted on BU-3DFE, and state-of-the-art results are achieved, indicating its effectiveness.
Di Huang 0001, Yunhong Wang 0001, Liming Chen 0002
ACM Multimedia4
2018 Visual and Semantic Knowledge Transfer for Large Scale Semi-Supervised Object Detection
abstract
Deep CNN-based object detection systems have achieved remarkable success on several large-scale object detection benchmarks. However, training such detectors requires a large number of labeled bounding boxes, which are more difficult to obtain than image-level annotations. Previous work addresses this issue by transforming image-level classifiers into object detectors. This is done by modeling the differences between the two on categories with both image-level and bounding box annotations, and transferring this information to convert classifiers to detectors for categories without bounding box annotations. We improve this previous work by incorporating knowledge about object similarities from visual and semantic domains during the transfer process. The intuition behind our proposed method is that visually and semantically similar categories should exhibit more common transferable properties than dissimilar categories, e.g. a better detector would result by transforming the differences between a dog classifier and a dog detector onto the cat class, than would by transforming from the violin class. Experimental results on the challenging ILSVRC2013 detection dataset demonstrate that each of our proposed object similarity based knowledge transfer methods outperforms the baseline methods. We found strong evidence that visual similarity and semantic relatedness are complementary for the task, and when combined notably improve detection, achieving state-of-the-art detection performance in a semi-supervised setting.
Yuxing Tang, Josiah Wang, Boyang Gao, Emmanuel Dellandréa, Robert J. Gaizauskas, Liming Chen 0002
IEEE Trans. Pattern Anal. Mach. Intell.7
2018 Affective Video Content Analysis: A Multidisciplinary Insight
abstract
In our present society, the cinema has become one of the major forms of entertainment providing unlimited contexts of emotion elicitation for the emotional needs of human beings. Since emotions are universal and shape all aspects of our interpersonal and intellectual experience, they have proved to be a highly multidisciplinary research field, ranging from psychology, sociology, neuroscience, etc., to computer science. However, affective multimedia content analysis work from the computer science community benefits but little from the progress achieved in other research fields. In this paper, a multidisciplinary state-of-the-art for affective movie content analysis is given, in order to promote and encourage exchanges between researchers from a very wide range of fields. In contrast to other state-of-the-art papers on affective video content analysis, this work confronts the ideas and models of psychology, sociology, neuroscience, and computer science. The concepts of aesthetic emotions and emotion induction, as well as the different representations of emotions are introduced, based on psychological and sociological theories. Previous global and continuous affective video content analysis work, including video emotion recognition and violence detection, are also presented in order to point out the limitations of affective video content analysis work.
Yoann Baveye, Christel Chamaret, Emmanuel Dellandréa, Liming Chen 0002
IEEE Trans. Affect. Comput.4
2018 Discriminative Transfer Learning Using Similarities and Dissimilarities
abstract
Transfer learning (TL) aims at solving the problem of learning an effective classification model for a target category, which has few training samples, by leveraging knowledge from source categories with far more training data. We propose a new discriminative TL (DTL) method, combining a series of hypotheses made by both the model learned with target training samples and the additional models learned with source category samples. Specifically, we use the sparse reconstruction residual as a basic discriminant and enhance its discriminative power by comparing two residuals from a positive and a negative dictionary. On this basis, we make use of similarities and dissimilarities by choosing both positively correlated and negatively correlated source categories to form additional dictionaries. A new Wilcoxon-Mann-Whitney statistic-based cost function is proposed to choose the additional dictionaries with unbalanced training data. Also, two parallel boosting processes are applied to both the positive and negative data distributions to further improve classifier performance. On two different image classification databases, the proposed DTL consistently outperforms other state-of-the-art TL methods while at the same time maintaining very efficient runtime.
Ying Lu 0007, Liming Chen 0002, Alexandre Saidi, Emmanuel Dellandréa, Yunhong Wang 0001
IEEE Trans. Neural Networks Learn. Syst.2
2018 Texture and Geometry Scattering Representation-Based Facial Expression Recognition in 2D+3D Videos
abstract
Facial Expression Recognition (FER) is one of the most important topics in the domain of computer vision and pattern recognition, and it has attracted increasing attention for its scientific challenges and application potentials. In this article, we propose a novel and effective approach to FER using multi-model two-dimensional (2D) and 3D videos, which encodes both static and dynamic clues by scattering convolution network. First, a shape-based detection method is introduced to locate the start and the end of an expression in videos; segment its onset, apex, and offset states; and sample the important frames for emotion analysis. Second, the frames in Apex of 2D videos are represented by scattering, conveying static texture details. Those of 3D videos are processed in a similar way, but to highlight static shape details, several geometric maps in terms of multiple order differential quantities, i.e., Normal Maps and Shape Index Maps, are generated as the input of scattering, instead of original smooth facial surfaces. Third, the average of neighboring samples centred at each key texture frame or shape map in Onset is computed, and the scattering features extracted from all the average samples of 2D and 3D videos are then concatenated to capture dynamic texture and shape cues, respectively. Finally, Multiple Kernel Learning is adopted to combine the features in the 2D and 3D modalities and compute similarities to predict the expression label. Thanks to the scattering descriptor, the proposed approach not only encodes distinct local texture and shape variations of different expressions as by several milestone operators, such as SIFT, HOG, and so on, but also captures subtle information hidden in high frequencies in both channels, which is quite crucial to better distinguish expressions that are easily confused. The validation is conducted on the BU-4DFE and BP-4D databa ses, and the accuracies reached are very competitive, indicating its competency for this issue.
Yongqiang Yao, Di Huang 0001, Yunhong Wang 0001, Liming Chen 0002
ACM Trans. Multim. Comput. Commun. Appl.5
2017 3D Facial Geometric Attributes Based Anti-Spoofing Approach against Mask Attacks
abstract
3D scanning and 3D printing techniques, as the technical impetus of 3D face recognition, also boost unconsciously the security threat against it from the spoofing attacks via manufactured mask. In order to improve the robustness of 3D face recognition system, several countermeasures against mask attacks based on photometric features have been reported in recent years. However, the anti-spoofing approach involving 3D meshed face scan and the related 3D facial features have not been studied yet. For filling this gap, in this paper, we propose to exploit the anti-spoofing performance of geometric attributes based 3D facial description. It synthesises the advantages of the selected geometric attributes, named principal curvature measures, and the meshSIFT-based feature descriptor. Specifically, the estimation of geometric attributes is coherent to the property of discrete surface, and the feature related to them can accurately describe the shape of facial surface. These characteristics are beneficial to discovering the geometry-based dissimilarity between genuine face and fraud mask. In the experiment part, the baselines of verification and anti-spoofing performance are evaluated on the Morpho database. Furthermore, for simulating a real-world scenario and testing the corresponding anti-spoofing performance, the size of genuine face set is massively extended by uniting the Morpho database and the FRGC v2.0 database to increase the ratio of genuine faces to fraud masks. The evaluation results prove that the proposed 3D face verification system can guarantee competitive verification accuracy for genuine faces and promising anti-spoofing performance against mask attacks.
Yinhang Tang, Liming Chen 0002
FG2
2017 Location-sensitive sparse representation of deep normal patterns for expression-robust 3D face recognition
abstract
This paper presents a straight-forward yet efficient, and expression-robust 3D face recognition approach by exploring location sensitive sparse representation of deep normal patterns (DNP). In particular, given raw 3D facial surfaces, we first run 3D face pre-processing pipeline, including nose tip detection, face region cropping, and pose normalization. The 3D coordinates of each normalized 3D facial surface are then projected into 2D plane to generate geometry images, from which three images of facial surface normal components are estimated. Each normal image is then fed into a pre-trained deep face net to generate deep representations of facial surface normals, i.e., deep normal patterns. Considering the importance of different facial locations, we propose a location sensitive sparse representation classifier (LS-SRC) for similarity measure among deep normal patterns associated with different 3D faces. Finally, simple score-level fusion of different normal components are used for the final decision. The proposed approach achieves significantly high performance, and reporting rank-one scores of 98.01%, 97.60%, and 96.13% on the FRGC v2.0, Bosphorus, and BU-3DFE databases when only one sample per subject is used in the gallery. These experimental results reveals that the performance of 3D face recognition would be constantly improved with the aid of training deep models from massive 2D face images, which opens the door for future directions of 3D face recognition.
Huibin Li 0001, Jian Sun 0009, Liming Chen 0002
IJCB3
2017 Interactive image segmentation based on samples reconstruction and FLDA
Lingkun Luo, Shiqiang Hu, Liming Chen 0002
J. Vis. Commun. Image Represent.5
2017 Guest EditorialSpecial Issue on Situation, Activity, and Goal Awareness in Cyber-Physical Human-Machine Systems
abstract
The papers in this special section focus on cyber-physical man-machine systems with particular emphasis on situation, activity, and goal awareness deployed in these systems. Recent advances in sensing technologies, the Internet of Things, pervasive computing, smart environments have transformed traditional embedded ICT systems into an ecosystem of interconnected and collaborating smart objects, devices, embedded systems, and most importantly humans. Such systems, often referred to as cyber-physical systems (CPS), are usually human user-driven or user-centered, and are aimed at providing people and businesses with a wide range of innovative applications and services. For example, a “smart home” can monitor and analyze the daily activities of its inhabitants, usually the elderly or individuals with disabilities, so that personalized context-aware assistance can be provided. A “smart city” can monitor, manage, and potentially control all basic city functionalities such as transport, energy supply, and waste collection, at a higher level of automation by collecting and harnessing sensor data across a large geographic expanse. In order to respond in real-time to an individual user’s specific needs in dynamic and complex situations, and to support ergonomics and user-friendliness through consideration of human factors such as privacy, dignity, and behavior characteristics, cyber-physical human–machine systems need to be aware of the physical environment and human participant behavior. This awareness enables effective and fast feedback loops between sensing and actuation, possibly with cognitive and learning capabilities adapting to participant preferences, capabilities, and the modality of human–machine interaction as well as dynamic situations.
Liming Chen 0001, Diane J. Cook, Bin Guo 0001, Liming Chen 0002, Wolfgang Leister
IEEE Trans. Hum. Mach. Syst.4
2017 Multimodal 2D+3D Facial Expression Recognition With Deep Fusion Convolutional Neural Network
abstract
This paper presents a novel and efficient deep fusion convolutional neural network (DF-CNN) for multimodal 2D+3D facial expression recognition (FER). DF-CNN comprises a feature extraction subnet, a feature fusion subnet, and a softmax layer. In particular, each textured three-dimensional (3D) face scan is represented as six types of 2D facial attribute maps (i.e., geometry map, three normal maps, curvature map, and texture map), all of which are jointly fed into DF-CNN for feature learning and fusion learning, resulting in a highly concentrated facial representation (32-dimensional). Expression prediction is performed by two ways: 1) learning linear support vector machine classifiers using the 32-dimensional fused deep features, or 2) directly performing softmax prediction using the six-dimensional expression probability vectors. Different from existing 3D FER methods, DF-CNN combines feature learning and fusion learning into a single end-to-end training framework. To demonstrate the effectiveness of DF-CNN, we conducted comprehensive experiments to compare the performance of DFCNN with handcrafted features, pre-trained deep features, finetuned deep features, and state-of-the-art methods on three 3D face datasets (i.e., BU-3DFE Subset I, BU-3DFE Subset II, and Bosphorus Subset). In all cases, DF-CNN consistently achieved the best results. To the best of our knowledge, this is the first work of introducing deep CNN to 3D FER and deep learning-based featurelevel fusion for multimodal 2D+3D FER.
Huibin Li 0001, Jian Sun 0009, Zongben Xu, Liming Chen 0002
IEEE Trans. Multim.4
2017 Weakly Supervised Learning of Deformable Part-Based Models for Object Detection via Region Proposals
abstract
The success of deformable part-based models (DPMs) for visual object detection relies on a large number of labeled bounding boxes. With only image-level annotations, our goal is to propose a model enhancing the weakly supervised DPMs by emphasizing the importance of location and size of the initial class-specific root filter. To adaptively select a discriminative set of candidate bounding boxes as this root filter estimate, first, we explore the generic objectness measurement to combine the most salient regions and “good” region proposals. Second, we propose learning of the latent class label of each candidate window as a binary classification problem, by training category-specific classifiers used to coarsely classify a candidate window into either a target object or a nontarget class. Finally, we design a flexible enlarging-and-shrinking postprocessing procedure to modify the DPMs outputs, which can effectively match the approximative object aspect ratios and further improve final accuracy. Extensive experimental results on the challenging PASCAL Visual Object Class 2007 and the Microsoft Common Objects in Context 2014 dataset demonstrate that our proposed framework is effective for initialization of the DPM's root filter. It also shows competitive final localization performance with state-of-the-art weakly supervised object detection methods, particularly for the object categories that are relatively salient in the images and deformable in structures.
Yuxing Tang, Emmanuel Dellandréa, Liming Chen 0002
IEEE Trans. Multim.4
2016 Large Scale Semi-Supervised Object Detection Using Visual and Semantic Knowledge Transfer
abstract
Deep CNN-based object detection systems have achieved remarkable success on several large-scale object detection benchmarks. However, training such detectors requires a large number of labeled bounding boxes, which are more difficult to obtain than image-level annotations. Previous work addresses this issue by transforming image-level classifiers into object detectors. This is done by modeling the differences between the two on categories with both imagelevel and bounding box annotations, and transferring this information to convert classifiers to detectors for categories without bounding box annotations. We improve this previous work by incorporating knowledge about object similarities from visual and semantic domains during the transfer process. The intuition behind our proposed method is that visually and semantically similar categories should exhibit more common transferable properties than dissimilar categories, e.g. a better detector would result by transforming the differences between a dog classifier and a dog detector onto the cat class, than would by transforming from the violin class. Experimental results on the challenging ILSVRC2013 detection dataset demonstrate that each of our proposed object similarity based knowledge transfer methods outperforms the baseline methods. We found strong evidence that visual similarity and semantic relatedness are complementary for the task, and when combined notably improve detection, achieving state-of-the-art detection performance in a semi-supervised setting.
Yuxing Tang, Josiah Wang, Boyang Gao, Emmanuel Dellandréa, Robert J. Gaizauskas, Liming Chen 0002
CVPR6
2016 Facial expression recognition based on a mlp neural network using constructive training algorithm
Hayet Boughrara, Mohamed Chtourou, Chokri Ben Amar, Liming Chen 0002
Multim. Tools Appl.4
2016 Active colloids segmentation and tracking
Boyang Gao, Simon Masnou, Liming Chen 0002, Isaac Theurkauff, Cécile Cottin-Bizonne, Frank Y. Shih
Pattern Recognit.4
2016 Automatic 2.5-D Facial Landmarking and Emotion Annotation for Social Interaction Assistance
abstract
People with low vision, Alzheimer's disease, and autism spectrum disorder experience difficulties in perceiving or interpreting facial expression of emotion in their social lives. Though automatic facial expression recognition (FER) methods on 2-D videos have been extensively investigated, their performance was constrained by challenges in head pose and lighting conditions. The shape information in 3-D facial data can reduce or even overcome these challenges. However, high expenses of 3-D cameras prevent their widespread use. Fortunately, 2.5-D facial data from emerging portable RGB-D cameras provide a good balance for this dilemma. In this paper, we propose an automatic emotion annotation solution on 2.5-D facial data collected from RGB-D cameras. The solution consists of a facial landmarking method and a FER method. Specifically, we propose building a deformable partial face model and fit the model to a 2.5-D face for localizing facial landmarks automatically. In FER, a novel action unit (AU) space-based FER method has been proposed. Facial features are extracted using landmarks and further represented as coordinates in the AU space, which are classified into facial expressions. Evaluated on three publicly accessible facial databases, namely EURECOM, FRGC, and Bosphorus databases, the proposed facial landmarking and expression recognition methods have achieved satisfactory results. Possible real-world applications using our algorithms have also been discussed.
Xi Zhao 0001, Jianhua Zou, Huibin Li 0001, Emmanuel Dellandréa, Ioannis A. Kakadiaris, Liming Chen 0002
IEEE Trans. Cybern.6
2016 Muscular Movement Model-Based Automatic 3D/4D Facial Expression Recognition
abstract
Facial expression is an important channel for human nonverbal communication. This paper presents a novel and effective approach to automatic 3D/4D facial expression recognition based on the muscular movement model (MMM). In contrast to most of existing methods, the MMM deals with such an issue in the viewpoint of anatomy. It first automatically segments the input 3D face (frame) by localizing the corresponding points within each muscular region of the reference using iterative closest normal point. A set of features with multiple differential quantities, including coordinate, normal, values, are then extracted to describe the geometry deformation of each segmented region. Meanwhile, we analyze the importance of these muscular areas, and a score level fusion strategy is exploited to optimize their weights by the genetic algorithm in the learning step. The support vector machine and the hidden Markov model are finally used to predict the expression label in 3D and 4D, respectively. The experiments are conducted on the BU-3DFE and BU-4DFE databases, and the results achieved clearly demonstrate the effectiveness of the proposed method.
Qingkai Zhen, Di Huang 0001, Yunhong Wang 0001, Liming Chen 0002
IEEE Trans. Multim.4
2015 Deep learning vs. kernel methods: Performance for emotion prediction in videos
abstract
Recently, mainly due to the advances of deep learning, the performances in scene and object recognition have been progressing intensively. On the other hand, more subjective recognition tasks, such as emotion prediction, stagnate at moderate levels. In such context, is it possible to make affective computational models benefit from the breakthroughs in deep learning? This paper proposes to introduce the strength of deep learning in the context of emotion prediction in videos. The two main contributions are as follow: (i) a new dataset, composed of 30 movies under Creative Commons licenses, continuously annotated along the induced valence and arousal axes (publicly available) is introduced, for which (ii) the performance of the Convolutional Neural Networks (CNN) through supervised fine-tuning, the Support Vector Machines for Regression (SVR) and the combination of both (Transfer Learning) are computed and discussed. To the best of our knowledge, it is the first approach in the literature using CNNs to predict dimensional affective scores from videos. The experimental results show that the limited size of the dataset prevents the learning or finetuning of CNN-based frameworks but that transfer learning is a promising solution to improve the performance of affective movie content analysis frameworks as long as very large datasets annotated along affective dimensions are not available.
Yoann Baveye, Emmanuel Dellandréa, Christel Chamaret, Liming Chen 0002
ACII4
2015 Depth edge based trilateral filter method for stereo matching
abstract
Recently, trilateral filter based method is more accurate than other local stereo matching methods on Middlebury benchmark. The boundary strength term is the key part of this method, which is computed from a color edge map. In this paper, we prove that boundary strength term, computed from a depth edge map, produces more accurate results. Then, we propose a depth edge detection technique to obtain an accurate depth edge map and present a depth edge based trilateral filter method. The experimental evaluation on the Middlebury benchmark shows that the proposed method outperforms the original trilateral filter method and also other local methods.
Dongming Chen, Mohsen Ardabilian, Liming Chen 0002
ICIP3
2015 Muscular Movement Model Based Automatic 3D Facial Expression Recognition
Qingkai Zhen, Di Huang 0001, Yunhong Wang 0001, Liming Chen 0002
MMM (1)4
2015 An efficient multimodal 2D + 3D feature-based approach to automatic facial expression recognition
Huibin Li 0001, Huaxiong Ding, Di Huang 0001, Yunhong Wang 0001, Xi Zhao 0001, Jean-Marie Morvan, Liming Chen 0002
Comput. Vis. Image Underst.7
2015 Towards 3D Face Recognition in the Real: A Registration-Free Approach Using Fine-Grained Matching of 3D Keypoint Descriptors
Huibin Li 0001, Di Huang 0001, Jean-Marie Morvan, Yunhong Wang 0001, Liming Chen 0002
Int. J. Comput. Vis.5
2015 LIRIS-ACCEDE: A Video Database for Affective Content Analysis
abstract
Research in affective computing requires ground truth data for training and benchmarking computational models for machine-based emotion understanding. In this paper, we propose a large video database, namely LIRIS-ACCEDE, for affective content analysis and related applications, including video indexing, summarization or browsing. In contrast to existing datasets with very few video resources and limited accessibility due to copyright constraints, LIRIS-ACCEDE consists of 9,800 good quality video excerpts with a large content diversity. All excerpts are shared under creative commons licenses and can thus be freely distributed without copyright issues. Affective annotations were achieved using crowdsourcing through a pair-wise video comparison protocol, thereby ensuring that annotations are fully consistent, as testified by a high inter-annotator agreement, despite the large diversity of raters' cultural backgrounds. In addition, to enable fair comparison and landmark progresses of future affective computational models, we further provide four experimental protocols and a baseline for prediction of emotions using a large set of both visual and audio features. The dataset (the video clips, annotations, features and protocols) is publicly available at: http://liris-accede.ec-lyon.fr/.
Yoann Baveye, Emmanuel Dellandréa, Christel Chamaret, Liming Chen 0002
IEEE Trans. Affect. Comput.4
2015 A Fast Trilateral Filter-Based Adaptive Support Weight Method for Stereo Matching
abstract
Adaptive support weight (ASW) methods represent the state of the art in local stereo matching, while the bilateral filter-based ASW method achieves outstanding performance. However, this method fails to resolve the ambiguity induced by nearby pixels at different disparities but with similar colors. In this paper, we introduce a novel trilateral filter (TF)-based ASW method that remedies such ambiguities by considering the possible disparity discontinuities through color discontinuity boundaries, i.e., the boundary strength between two pixels, which is measured by a local energy model. We also present a recursive TF-based ASW method whose computational complexity is O(N) for the cost aggregation step, and O(NLog2(N)) for boundary detection, where N denotes the input image size. This complexity is thus independent of the support window size. The recursive TF-based method is a nonlocal cost aggregation strategy. The experimental evaluation on the Middlebury benchmark shows that the proposed method, whose average error rate is 4.95%, outperforms other local methods in terms of accuracy. Equally, the average runtime of the proposed TF-based cost aggregation is roughly 260 ms on a 3.4-GHz Inter Core i7 CPU, which is comparable with state-of-the-art efficiency.
Dongming Chen, Mohsen Ardabilian, Liming Chen 0002
IEEE Trans. Circuits Syst. Video Technol.3
2015 Hand-Dorsa Vein Recognition by Matching Local Features of Multisource Keypoints
abstract
As an emerging biometric for people identification, the dorsal hand vein has received increasing attention in recent years due to the properties of being universal, unique, permanent, and contactless, and especially its simplicity of liveness detection and difficulty of forging. However, the dorsal hand vein is usually captured by near-infrared (NIR) sensors and the resulting image is of low contrast and shows a very sparse subcutaneous vascular network. Therefore, it does not offer sufficient distinctiveness in recognition particularly in the presence of large population. This paper proposes a novel approach to hand-dorsa vein recognition through matching local features of multiple sources. In contrast to current studies only concentrating on the hand vein network, we also make use of person dependent optical characteristics of the skin and subcutaneous tissue revealed by NIR hand-dorsa images and encode geometrical attributes of their landscapes, e.g., ridges, valleys, etc., through different quantities, such as cornerness and blobness, closely related to differential geometry. Specifically, the proposed method adopts an effective keypoint detection strategy to localize features on dorsal hand images, where the speciality of absorption and scattering of the entire dorsal hand is modeled as a combination of multiple (first-, second-, and third-) order gradients. These features comprehensively describe the discriminative clues of each dorsal hand. This method further robustly associates the corresponding keypoints between gallery and probe samples, and finally predicts the identity. Evaluated by extensive experiments, the proposed method achieves the best performance so far known on the North China University of Technology (NCUT) Part A dataset, showing its effectiveness. Additional results on NCUT Part B illustrate its generalization ability and robustness to low quality data.
Di Huang 0001, Yinhang Tang, Liming Chen 0002, Yunhong Wang 0001
IEEE Trans. Cybern.4
2015 A Global/Local Affinity Graph for Image Segmentation
abstract
Construction of a reliable graph capturing perceptual grouping cues of an image is fundamental for graph-cut based image segmentation methods. In this paper, we propose a novel sparse global/local affinity graph over superpixels of an input image to capture both short- and long-range grouping cues, and thereby enabling perceptual grouping laws, including proximity, similarity, continuity, and to enter in action through a suitable graph-cut algorithm. Moreover, we also evaluate three major visual features, namely, color, texture, and shape, for their effectiveness in perceptual segmentation and propose a simple graph fusion scheme to implement some recent findings from psychophysics, which suggest combining these visual features with different emphases for perceptual grouping. In particular, an input image is first oversegmented into superpixels at different scales. We postulate a gravitation law based on empirical observations and divide superpixels adaptively into small-, medium-, and large-sized sets. Global grouping is achieved using medium-sized superpixels through a sparse representation of superpixels' features by solving a ℓ0-minimization problem, and thereby enabling continuity or propagation of local smoothness over long-range connections. Small- and large-sized superpixels are then used to achieve local smoothness through an adjacent graph in a given feature space, and thus implementing perceptual laws, for example, similarity and proximity. Finally, a bipartite graph is also introduced to enable propagation of grouping cues between superpixels of different scales. Extensive experiments are carried out on the Berkeley segmentation database in comparison with several state-of-the-art graph constructions. The results show the effectiveness of the proposed approach, which outperforms state-of-the-art graphs using four different objective criteria, namely, the probabilistic rand index, the variation of information, the global consistency error, and the boundary displacement error.
Yuxing Tang, Simon Masnou, Liming Chen 0002
IEEE Trans. Image Process.4
2015 Learning-Based Driving Events Recognition and Its Application to Digital Roads
abstract
Automatic recognition of driving events, e.g., approaching roundabouts, is important both for the truck design process based on simulated road data and for advanced driver assistance systems. However, the problem faced is extremely challenging as only in-vehicle driving data must be used, whereas the number of driving events is usually quite large. In this paper, we propose a learning-based driving events classification method, which is trained and tested with a real driving events database. The proposed method includes definition of driving events relevant to our final application, selection of discriminating features, and classification, using two machine-learning techniques, namely, decision trees and linear logistic regression. We then introduce the digital road concept. This consists of simulated road data used in the truck design process to quantify the behavior of a truck, particularly in terms of fuel consumption. While a digital road typically contains far less driving information, we show that we can still apply the proposed driving events recognition models learnt on real driving data and pave the way for a more realistic assessment of truck characteristics via simulation tools.
Claire D'Agostino, Alexandre Saidi, Gilles Scouarnec, Liming Chen 0002
IEEE Trans. Intell. Transp. Syst.4
2014 3D-Aided Face Recognition Robust to Expression and Pose Variations
abstract
Expression and pose variations are major challenges for reliable face recognition (FR) in 2D. In this paper, we aim to endow state of the art face recognition SDKs with robustness to facial expression variations and pose changes by using an extended 3D Morphable Model (3DMM) which isolates identity variations from those due to facial expressions. Specifically, given a probe with expression, a novel view of the face is generated where the pose is rectified and the expression neutralized. We present two methods of expression neutralization. The first one uses prior knowledge to infer the neutral expression image from an input image. The second method, specifically designed for verification, is based on the transfer of the gallery face expression to the probe. Experiments using rectified and neutralized view with a standard commercial FR SDK on two 2D face databases, namely Multi-PIE and AR, show significant performance improvement of the commercial SDK to deal with expression and pose variations and demonstrates the effectiveness of the proposed approach.
Baptiste Chu, Sami Romdhani, Liming Chen 0002
CVPR3
2014 A coarse-to-fine approach to robust 3D facial landmarking via curvature analysis and Active Normal Model
abstract
Facial landmarking is a fundamental step in machine-based face analysis. The majority of existing techniques handle such an issue based on 2D images; however, they suffer from illumination and pose variations that largely degrade landmarking performance. The emergence of 3D data provides us with an alternative to overcome these unsolved problems in the 2D domain. This paper proposes a novel approach to 3D facial landmarking, combining both the advantages of feature based methods as well as model based ones in a coarse-to-fine manner. For the coarse stage, three fiducial landmarks (the nose tip and two inner eye corners) are robustly detected through curvature analysis, and these points are further employed to initialize the subsequent model fitting. For the fine stage, a statistical model is constructed based on the normal information including the x, y, and z components of the facial point-cloud rather than the smooth coordinate information, thereby namely Active Normal Model (ANM), to highlight its shape characteristics for final landmark prediction. The proposed approach accurately localizes 83 fiducial points on each 3D face model, greatly surpassing those of feature based ones, while improving the state of the art model based ones in two aspects, i.e. sensitivity to initialization and deficiency in discrimination. Evaluated on the BU-3DFE database, very competitive results are achieved in comparison with those in the literature, clearly demonstrating its effectiveness.
Di Huang 0001, Yunhong Wang 0001, Liming Chen 0002
IJCB4
2014 Learning visual categories through a sparse representation classifier based cross-category knowledge transfer
abstract
To solve the challenging task of learning effective visual categories with limited training samples, we propose a new sparse representation classifier based transfer learning method, namely SparseTL, which propagates the cross-category knowledge from multiple source categories to the target category. Specifically, we enhance the target classification task in learning a both generative and discriminative sparse representation based classifier using pairs of source categories most positively and most negatively correlated to the target category. We further improve the discriminative ability of the classifier by choosing the most discriminative bins in the feature vector with a feature selection process. The experimental results show that the proposed method achieves competitive performance on the NUS-WIDE Scene database compared to several state of the art transfer learning algorithms while keeping a very efficient runtime.
Ying Lu 0007, Liming Chen 0002, Alexandre Saidi, Zhaoxiang Zhang 0001, Yunhong Wang 0001
ICIP2
2014 Fusing generic objectness and deformable part-based models for weakly supervised object detection
abstract
In the context of lack of object-level annotation, we propose a model that enhances the weakly supervised deformable part model (DPM) by emphasizing the importance of size and aspect ratio of the initial class-specific root filter. For each image, to extract a reliable bounding box as this root filter estimate, we explore the generic objectness measurement to obtain a reference window based on the most salient region, and select a small set of candidate windows by adaptive thresholding and greedy Non-Maximum Suppression (NMS). The initial root filter estimate is decided by optimizing the score of overlap between the reference box and candidate boxes, as well as their corresponding objectness score. Then the derived window is treated as a positive training window for DPM training. Finally, we design a flexible enlarging-and-shrinking post-processing procedure to modify the output of DPM, which can effectively fit to the aspect ratio of the object and further improve the final accuracy. Experimental results on the challenging PASCAL VOC 2007 database demonstrate that our proposed framework is effective and competitive with the state-of-the-arts.
Yuxing Tang, Emmanuel Dellandréa, Simon Masnou, Liming Chen 0002
ICIP5
2014 3D assisted face recognition via progressive pose estimation
abstract
Most existing pose-independent Face Recognition (FR) techniques take advantage of 3D model to guarantee the naturalness while normalizing or simulating pose variations. Two nontrivial problems to be tackled are accurate measurement of pose parameters and computational efficiency. In this paper, we introduce an effective and efficient approach to estimate human head pose, which fundamentally ameliorates the performance of 3D aided FR systems. The proposed method works in a progressive way: firstly, a random forest (RF) is constructed utilizing synthesized images derived from 3D models; secondly, the classification result obtained by applying well-trained RF on a probe image is considered as the preliminary pose estimation; finally, this initial pose is transferred to shape-based 3D morphable model (3DMM) aiming at definitive pose normalization. Using such a method, similarity scores between frontal view gallery set and pose-normalized probe set can be computed to predict the identity. Experimental results achieved on the UHDB dataset outperform the ones so far reported. Additionally, it is much less time-consuming than prevailing 3DMM based approaches.
Wuming Zhang, Di Huang 0001, Dimitris Samaras, Jean-Marie Morvan, Yunhong Wang 0001, Liming Chen 0002
ICIP6
2014 MLP neural network using modified constructive training algorithm: Application to face recognition
abstract
This paper focuses on the study of modified constructive training algorithm for Multi Layer Perceptron “MLP” which is applied to face recognition applications. In general, constructive learning begins with a minimal structure, and increases the network by adding hidden neurons until a satisfactory solution is found. The contribution of this paper is to increment the output neurons simultaneously with incrementing the input patterns. In fact, the proposed algorithm started with a small number of output neurons and a single hidden-layer using an initial number of neurons. During neural network training, the hidden neurons number is increased while the Mean Square Error “MSE” threshold of the Training Data “TD” is not reduced to a predefined parameter. The output neurons number is increased as the input patterns are incrementally trained until all patterns of Training Data “TD” are presented and learned. The proposed algorithm is applied in the classification stage in face recognition system. For the feature extraction stage, a biological vision-based facial description, namely Perceived Facial Images “PFI” is applied to extract features from human face images. The proposed approach is tested on the Cohn-Kanade Facial Expression Database. Compared to the fixed “MLP” architecture and the constructive training algorithm, experimental results clearly demonstrate the efficiency of the proposed algorithm.
Hayet Boughrara, Mohamed Chtourou, Chokri Ben Amar, Liming Chen 0002
IPAS4
2014 Rational truck driving and its correlated driving features in extra-urban areas
abstract
Truck drivers typically display different behaviors when facing various driving events, e.g., approaching a roundabout, and thereby have a major impact both on the fuel consumption and the vehicle speed. Within the context where fuel is increasingly a major cost center for merchandise transport companies, it is important to recognize different driver behaviors in order to be able to simulate them as closely to the real data as possible during the truck development process. In this paper, we introduce, instead of economic driving, the notion of rational driving which seeks to decrease the average fuel consumption while respecting the transport companies' constraint, i.e., the delivery delay. Moreover, we also propose an indicator, namely rational driving index (RDI), which enables to quantify how good a driver behavior is with respect to the rational driving. We then investigate various driving features contributing to characterize a rational driver behavior, using real driving data collected from 34 different truck drivers on an extra-urban road section particularly representative of travel paths of trucks ensuring regional merchandise distribution. Given the fact that real driving data collected on an open road can differ in terms of environment, e.g., weather, traffic, we further study, through simulations on a digital representation of a roundabout, the impact of two major driving features, i.e., the use of coasting and crossing speed at roundabouts, with respect to rational driving. The experimental results from both real driving data and simulations show high correlations of these two driving features with respect to RDI and demonstrate that a good rational driver tends to decelerate slowly during braking periods (use of coasting) and have high crossing speed in roundabouts.
Claire D'Agostino, Alexandre Saidi, Gilles Scouarnec, Liming Chen 0002
Intelligent Vehicles Symposium4
2014 Face recognition based on perceived facial images and multilayer perceptron neural network using constructive training algorithm
abstract
This study presents a modified constructive training algorithm for multilayer perceptron (MLP) which is applied to face recognition problem. An incremental training procedure has been employed where the training patterns are learned incrementally. This algorithm starts with a small number of training patterns and a single hidden‐layer using an initial number of neurons. During the training, the hidden neurons number is increased when the mean square error (MSE) threshold of the training data (TD) is not reduced to a predefined value. Input patterns are trained incrementally until all patterns of TD are learned. The aim of this algorithm is to determine the adequate initial number of hidden neurons, the suitable number of training patterns in the subsets of each class and the number of iterations during the training step as well as the MSE threshold value. The proposed algorithm is applied in the classification stage in face recognition system. For the feature extraction stage, this paper proposes to use a biological vision‐based facial description, namely perceived facial images, applied to extract features from human face images. Gabor features and Zernike moment have been used in order to determine the best feature extractor. The proposed approach is tested on the Cohn‐Kanade Facial Expression Database. Experimental results indicate that a good architecture of neural network classifier can be obtained. The effectiveness of the proposed method compared with the fixed MLP architecture has been proved.
Hayet Boughrara, Mohamed Chtourou, Chokri Ben Amar, Liming Chen 0002
IET Comput. Vis.4
2014 Expression-robust 3D face recognition via weighted sparse representation of multi-scale and multi-component local normal patterns
Huibin Li 0001, Di Huang 0001, Jean-Marie Morvan, Liming Chen 0002, Yunhong Wang 0001
Neurocomputing4
2014 Local circular patterns for multi-modal facial gender and ethnicity classification
Di Huang 0001, Huaxiong Ding, Yunhong Wang 0001, Guangpeng Zhang, Liming Chen 0002
Image Vis. Comput.6
2014 HSOG: A Novel Local Image Descriptor Based on Histograms of the Second-Order Gradients
abstract
Recent investigations on human vision discover that the retinal image is a landscape or a geometric surface, consisting of features such as ridges and summits. However, most of existing popular local image descriptors in the literature, e.g., scale invariant feature transform (SIFT), histogram of oriented gradient (HOG), DAISY, local binary Patterns (LBP), and gradient location and orientation histogram, only employ the first-order gradient information related to the slope and the elasticity, i.e., length, area, and so on of a surface, and thereby partially characterize the geometric properties of a landscape. In this paper, we introduce a novel and powerful local image descriptor that extracts the histograms of second-order gradients (HSOGs) to capture the curvature related geometric properties of the neural landscape, i.e., cliffs, ridges, summits, valleys, basins, and so on. We conduct comprehensive experiments on three different applications, including the problem of local image matching, visual object categorization, and scene classification. The experimental results clearly evidence the discriminative power of HSOG as compared with its first-order gradient-based counterparts, e.g., SIFT, HOG, DAISY, and center-symmetric LBP, and the complementarity in terms of image representation, demonstrating the effectiveness of the proposed local descriptor.
Di Huang 0001, Chao Zhu 0003, Yunhong Wang 0001, Liming Chen 0002
IEEE Trans. Image Process.4
2014 Surface Meshing with Curvature Convergence
abstract
Surface meshing plays a fundamental role in graphics and visualization. Many geometric processing tasks involve solving geometric PDEs on meshes. The numerical stability, convergence rates and approximation errors are largely determined by the mesh qualities. In practice, Delaunay refinement algorithms offer satisfactory solutions to high quality mesh generations. The theoretical proofs for volume based and surface based Delaunay refinement algorithms have been established, but those for conformal parameterization based ones remain wide open. This work focuses on the curvature measure convergence for the conformal parameterization based Delaunay refinement algorithms. Given a metric surface, the proposed approach triangulates its conformal uniformization domain by the planar Delaunay refinement algorithms, and produces a high quality mesh. We give explicit estimates for the Hausdorff distance, the normal deviation, and the differences in curvature measures between the surface and the mesh. In contrast to the conventional results based on volumetric Delaunay refinement, our stronger estimates are independent of the mesh structure and directly guarantee the convergence of curvature measures. Meanwhile, our result on Gaussian curvature measure is intrinsic to the Riemannian metric and independent of the embedding. In practice, our meshing algorithm is much easier to implement and much more efficient. The experimental results verified our theoretical results and demonstrated the efficiency of the meshing algorithm.
Huibin Li 0001, Wei Zeng 0002, Jean-Marie Morvan, Liming Chen 0002, Xianfeng Gu
IEEE Trans. Vis. Comput. Graph.4
2013 A Large Video Data Base for Computational Models of Induced Emotion
abstract
To contribute to the need for emotional databases and affective tagging, the LIRIS-ACCEDE is proposed in this paper. LIRIS-ACCEDE is an Annotated Creative Commons Emotional DatabasE composed of 9800 video clips extracted from 160 movies shared under Creative Commons licenses. It allows to make this database publicly available without copyright issues. The 9800 video clips (each 8-12 seconds long) are sorted along the induced valence axis, from the video perceived the most negatively to the video perceived the most positively. The annotation was carried out by 1518 annotators from 89 different countries using crowd sourcing. A baseline late fusion scheme using ground truth from annotations is computed to predict emotion categories in video clips.
Yoann Baveye, Jean-Noel Bettinelli, Emmanuel Dellandréa, Liming Chen 0002, Christel Chamaret
ACII4
2013 A Novel Trilateral Filter based Adaptive Support Weight Method for Stereo Matching
abstract
International audience
Dongming Chen, Mohsen Ardabilian, Liming Chen 0002
BMVC3
2013 Sparse Coding and Mid-Level Superpixel-Feature for ℓ0-Graph Based Unsupervised Image Segmentation
Huibin Li 0001, Simon Masnou, Liming Chen 0002
CAIP (2)4
2013 Encoding Local Binary Descriptors by Bag-of-Features with Hamming Distance for Visual Object Categorization
Yu Zhang 0052, Chao Zhu 0003, Stéphane Bres, Liming Chen 0002
ECIR4
2013 A graph-cut approach to image segmentation using an affinity graph based on ℓ0-sparse representation of features
abstract
We propose a graph-cut based image segmentation method by constructing an affinity graph using ℓ0sparse representation. Computing first oversegmented images, we associate with all segments, that we call superpixels, a collection of features. We find the sparse representation of each set of features over the dictionary of all features by solving a ℓ0-minimization problem. Then, the connection information between superpixels is encoded as the non-zero representation coefficients, and the affinity of connected superpixels is derived by the corresponding representation error. This provides a ℓ0affinity graph that has interesting properties of long range and sparsity, and a suitable graph cut yields a segmentation. Experimental results on the BSD database demonstrate that our method provides perfectly semantic regions even with a constant segmentation number, but also that very competitive quantitative results are achieved.
Huibin Li 0001, Charles-Edmond Bichot, Simon Masnou, Liming Chen 0002
ICIP5
2013 An improved Non-Local Cost Aggregation method for stereo matching based on color and boundary cue
abstract
Recently, a novel Non-Local Cost Aggregation (NLCA) algorithm based on a minimum spanning tree (MST) has been proposed [1] for stereo matching providing extremely low computational complexity and outstanding performance. However, since the MST is constructed only based on color cue, this approach suffers from some improper connections at the boundaries of two objects with similar color distribution. In this paper, we propose a new weight function which includes not only the color cue but also the boundary cue of the reference image. Because the boundary cue is more discriminative than the color cue to distinguish the object from the background with similar color, the previous improper connections can be avoided and a more faithful MST can be constructed. Experimental results performed on Middlebury benchmark demonstrate the effectiveness of the improvements. The improved algorithm achieves rank 18th out of 143 submissions, while the original algorithm ranks 31th.
Dongming Chen, Mohsen Ardabilian, Liming Chen 0002
ICME4
2013 HSOG: a novel local descriptor based on histograms of second order gradients for object categorization
abstract
This paper presents a novel local image descriptor for object categorization that extracts the Histograms of the Second Order Gradients and is thereby named as HSOG. The HSOG descriptor is in contrast to the widely used ones in the literature, e.g. SIFT, DAISY, HOG, LBP, etc., which are based on the first order gradient information. The contributions of this work can be summarized as: (1) the design of HSOG; (2) the prove of its discriminative power and its complementation to the first order gradient based descriptors; (3) the analysis of performance variation caused by different parameter settings; and (4) the multi-scale extension which further improves the categorization accuracy. The experimental results achieved on the Caltech 101 and Caltech 256 databases clearly highlight the effectiveness of the proposed approach.
Di Huang 0001, Chao Zhu 0003, Charles-Edmond Bichot, Yunhong Wang 0001, Liming Chen 0002
ICMR5
2013 Multimodal recognition of visual concepts using histograms of textual concepts and selective weighted late fusion scheme
Ningning Liu, Emmanuel Dellandréa, Liming Chen 0002, Chao Zhu 0003, Yu Zhang 0052, Charles-Edmond Bichot, Stéphane Bres, Bruno Tellez
Comput. Vis. Image Underst.3
2013 A unified probabilistic framework for automatic 3D facial expression analysis based on a Bayesian belief inference and statistical feature models
Xi Zhao 0001, Emmanuel Dellandréa, Jianhua Zou, Liming Chen 0002
Image Vis. Comput.4
2013 Image region description using orthogonal combination of local binary patterns enhanced with color information
Chao Zhu 0003, Charles-Edmond Bichot, Liming Chen 0002
Pattern Recognit.3
2013 Learning the Spherical Harmonic Features for 3-D Face Recognition
abstract
In this paper, a competitive method for 3-D face recognition (FR) using spherical harmonic features (SHF) is proposed. With this solution, 3-D face models are characterized by the energies contained in spherical harmonics with different frequencies, thereby enabling the capture of both gross shape and fine surface details of a 3-D facial surface. This is in clear contrast to most 3-D FR techniques which are either holistic or feature based, using local features extracted from distinctive points. First, 3-D face models are represented in a canonical representation, namely, spherical depth map, by which SHF can be calculated. Then, considering the predictive contribution of each SHF feature, especially in the presence of facial expression and occlusion, feature selection methods are used to improve the predictive performance and provide faster and more cost-effective predictors. Experiments have been carried out on three public 3-D face datasets, SHREC2007, FRGC v2.0, and Bosphorus, with increasing difficulties in terms of facial expression, pose, and occlusion, and which demonstrate the effectiveness of the proposed method.
Peijiang Liu, Yunhong Wang 0001, Di Huang 0001, Zhaoxiang Zhang 0001, Liming Chen 0002
IEEE Trans. Image Process.5
2012 Hand Vein Recognition Based on Oriented Gradient Maps and Local Feature Matching
Di Huang 0001, Yinhang Tang, Liming Chen 0002, Yunhong Wang 0001
ACCV (4)4
2012 3D facial expression recognition via multiple kernel learning of Multi-Scale Local Normal Patterns
Huibin Li 0001, Liming Chen 0002, Di Huang 0001, Yunhong Wang 0001, Jean-Marie Morvan
ICPR2
2012 3-D Face Recognition Using eLBP-Based Facial Description and Local Feature Hybrid Matching
abstract
This paper presents an effective method for 3-D face recognition using a novel geometric facial representation along with a local feature hybrid matching scheme. The proposed facial surface description is based on a set of facial depth maps extracted by multiscale extended Local Binary Patterns (eLBP) and enables an efficient and accurate description of local shape changes; it thus enhances the distinctiveness of smooth and similar facial range images generated by preprocessing steps. The following matching strategy is SIFT-based and performs in a hybrid way that combines local and holistic analysis, robustly associating the keypoints between two facial representations of the same subject. As a result, the proposed approach proves robust to facial expression variations, partial occlusions, and moderate pose changes, and the last property makes our system registration-free for nearly frontal face models. The proposed method was experimented on three public datasets, i.e. FRGC v2.0, Gavab, and Bosphorus. It displays a rank-one recognition rate of 97.6% and a verification rate of 98.4% at a 0.001 FAR on the FRGC v2.0 database without any face alignment. Additional experiments on the Bosphorus dataset further highlight the advantages of the proposed method with regard to expression changes and external partial occlusions. The last experiment carried out on the Gavab database demonstrates that the entire system can also deal with faces under large pose variations and even partially occluded ones, when only aided by a coarse alignment process.
Di Huang 0001, Mohsen Ardabilian, Yunhong Wang 0001, Liming Chen 0002
IEEE Trans. Inf. Forensics Secur.4
2011 Associating Textual Features with Visual Ones to Improve Affective Image Classification
Ningning Liu, Emmanuel Dellandréa, Bruno Tellez, Liming Chen 0002
ACII (1)4
2011 3D Facial Expression Recognition Based on Histograms of Surface Differential Quantities
Huibin Li 0001, Jean-Marie Morvan, Liming Chen 0002
ACIVS3
2011 A Space-Time Depth Super-Resolution Scheme for 3D Face Scanning
Karima Ouji, Mohsen Ardabilian, Liming Chen 0002, Faouzi Ghorbel
ACIVS3
2011 Reconstructive and Discriminative Sparse Representation for Visual Object Categorization
abstract
International audience
Huanzhang Fu, Emmanuel Dellandréa, Liming Chen 0002
BMVC3
2011 Multi-camera 3D Scanning with a Non-rigid and Space-Time Depth Super-Resolution Capability
Karima Ouji, Mohsen Ardabilian, Liming Chen 0002, Faouzi Ghorbel
CAIP (2)3
2011 A novel geometric facial representation based on multi-scale extended local binary patterns
abstract
In this study, we present a novel geometric representation for 3D faces in order to enhance distinctiveness of generally smooth range images. This novel face representation is based on Multi-Scale Extended Local Binary Patterns (ELBP) and enables accurate and fast description of local shape variations on range faces. When associated with the proposed SIFT-based local feature matching scheme, this novel geometric facial representation shows its discriminative power in 3D face recognition, displaying a rank-one recognition rate up to 97.2% and a verification rate of 98.4% at a 0.001 FAR respectively on the FRGC v2.0 database. Moreover, costly registration is not needed thanks to the relative tolerance of the proposed representation and the SIFT methodology to moderate pose changes as the ones existing in FRGC v2.0. Finally, additional experiments demonstrate that the entire system is also robust to facial expression variations.
Di Huang 0001, Mohsen Ardabilian, Yunhong Wang 0001, Liming Chen 0002
FG4
2011 Learning weighted sparse representation of encoded facial normal information for expression-robust 3D face recognition
abstract
This paper proposes a novel approach for 3D face recognition by learning weighted sparse representation of encoded facial normal information. To comprehensively describe 3D facial surface, three components, in X, Y, and Z-plane respectively, of normal vector are encoded locally to their corresponding normal pattern histograms. They are finally fed to a sparse representation classifier enhanced by learning based spatial weights. Experimental results achieved on the FRGC v2.0 database prove that the proposed encoded normal information is much more discriminative than original normal information. Moreover, the patch based weights learned using the FRGC v1.0 and Bosphorus datasets also demonstrate the importance of each facial physical component for 3D face recognition.
Huibin Li 0001, Di Huang 0001, Jean-Marie Morvan, Liming Chen 0002
IJCB4
2011 Twins 3D face recognition challenge
abstract
Existing 3D face recognition algorithms have achieved high enough performances against public datasets like FRGC v2, that it is difficult to achieve further significant increases in recognition performance. However, the 3D TEC dataset is a more challenging dataset which consists of 3D scans of 107 pairs of twins that were acquired in a single session, with each subject having a scan of a neutral expression and a smiling expression. The combination of factors related to the facial similarity of identical twins and the variation in facial expression makes this a challenging dataset. We conduct experiments using state of the art face recognition algorithms and present the results. Our results indicate that 3D face recognition of identical twins in the presence of varying facial expressions is far from a solved problem, but that good performance is possible.
Vipin Vijayan, Kevin W. Bowyer, Patrick J. Flynn, Di Huang 0001, Liming Chen 0002, Mark Hansen, Omar Ocegueda, Shishir Shah 0001, Ioannis A. Kakadiaris
IJCB5
2011 Expression robust 3D face recognition via mesh-based histograms of multiple order surface differential quantities
abstract
This paper presents a mesh-based approach for 3D face recognition using a novel local shape descriptor and a SIFT-like matching process. Both maximum and minimum curvatures estimated in the 3D Gaussian scale space are employed to detect salient points. To comprehensively characterize 3D facial surfaces and their variations, we calculate weighted statistical distributions of multiple order surface differential quantities, including histogram of mesh gradient (HoG), histogram of shape index (HoS) and histogram of gradient of shape index (HoGS) within a local neighborhood of each salient point. The subsequent matching step then robustly associates corresponding points of two facial surfaces, leading to much more matched points between different scans of a same person than the ones of different persons. Experimental results on the Bosphorus dataset highlight the effectiveness of the proposed method and its robustness to facial expression variations.
Huibin Li 0001, Di Huang 0001, Pierre Lemaire 0002, Jean-Marie Morvan, Liming Chen 0002
ICIP5
2011 A mixture of gated experts optimized using simulated annealing for 3D face recognition
abstract
A commonly accepted fact in the biometrics related domain is that fusing multiple classifiers for decision making generally leads to improved recognition performance. Meanwhile, the search for an optimal fusion strategy remains extraordinarily complex since the cardinality of the space of possible fusion schemes is exponentially proportional to the number of competing classifiers. In this paper, we propose a mixture of gated experts for 3D face recognition using an ensemble of 24 different classifiers. The mixture of gated experts is optimized using a Simulated Annealing-based algorithm. It automatically selects and fuses the most relevant similarity measurements. The experimental results of 3D face recognition achieved on the FRGC v2.0 dataset illustrate the effectiveness and stability of the proposed method. Additionally, as a learning-based method, it also has a good robustness to the variations of training database.
Wael Ben Soltana, Di Huang 0001, Mohsen Ardabilian, Liming Chen 0002, Chokri Ben Amar
ICIP4
2011 Visual object recognition using DAISY descriptor
abstract
Visual content description is a key issue for the task of machine-based visual object categorization (VOC). A good visual descriptor should be both discriminative enough and computationally efficient while possessing some properties of robustness to viewpoint changes and lighting condition variations. The recent literature has featured local image descriptors, e.g. SIFT, as the main trend in VOC. However, it is well known that SIFT is computationally expensive, especially when the number of objects/concepts and learning data increase significantly. In this paper, we investigate the DAISY, which is a new fast local descriptor introduced for wide baseline matching problem, in the context of VOC. We carefully evaluate and compare the DAISY descriptor with SIFT both in terms of recognition accuracy and computation complexity on two standard image benchmarks - Caltech 101 and PASCAL VOC 2007. The experimental results show that DAISY outperforms the state-of-the-art SIFT while using shorter descriptor length and operating 3 times faster. When displaying a similar recognition accuracy to SIFT, DAISY can operate 12 times faster.
Chao Zhu 0003, Charles-Edmond Bichot, Liming Chen 0002
ICME3
2011 3D Face Recognition Based on Local Shape Patterns and Sparse Representation Classifier
Di Huang 0001, Karima Ouji, Mohsen Ardabilian, Yunhong Wang 0001, Liming Chen 0002
MMM (1)5
2011 Local Binary Patterns and Its Application to Facial Image Analysis: A Survey
abstract
Local binary pattern (LBP) is a nonparametric descriptor, which efficiently summarizes the local structures of images. In recent years, it has aroused increasing interest in many areas of image processing and computer vision and has shown its effectiveness in a number of applications, in particular for facial image analysis, including tasks as diverse as face detection, face recognition, facial expression analysis, and demographic classification. This paper presents a comprehensive survey of LBP methodology, including several more recent variations. As a typical application of the LBP approach, LBP-based facial image analysis is extensively reviewed, while its successful extensions, which deal with various tasks of facial image analysis, are also highlighted.
Di Huang 0001, Caifeng Shan, Mohsen Ardabilian, Yunhong Wang 0001, Liming Chen 0002
IEEE Trans. Syst. Man Cybern. Part C5
2011 Accurate Landmarking of Three-Dimensional Facial Data in the Presence of Facial Expressions and Occlusions Using a Three-Dimensional Statistical Facial Feature Model
abstract
Three-dimensional face landmarking aims at automatically localizing facial landmarks and has a wide range of applications (e.g., face recognition, face tracking, and facial expression analysis). Existing methods assume neutral facial expressions and unoccluded faces. In this paper, we propose a general learning-based framework for reliable landmark localization on 3-D facial data under challenging conditions (i.e., facial expressions and occlusions). Our approach relies on a statistical model, called 3-D statistical facial feature model, which learns both the global variations in configurational relationships between landmarks and the local variations of texture and geometry around each landmark. Based on this model, we further propose an occlusion classifier and a fitting algorithm. Results from experiments on three publicly available 3-D face databases (FRGC, BU-3-DFE, and Bosphorus) demonstrate the effectiveness of our approach, in terms of landmarking accuracy and robustness, in the presence of expressions and occlusions.
Xi Zhao 0001, Emmanuel Dellandréa, Liming Chen 0002, Ioannis A. Kakadiaris
IEEE Trans. Syst. Man Cybern. Part B3
2010 Automatic Asymmetric 3D-2D Face Recognition
abstract
3D Face recognition has been considered as a major solution to deal with unsolved issues of reliable 2D face recognition in recent years, i.e. lighting and pose variations. However, 3D techniques are currently limited by their high registration and computation cost. In this paper, an asymmetric 3D-2D face recognition method is presented, enrolling in textured 3D whilst performing automatic identification using only 2D facial images. The goal is to limit the use of 3D data to where it really helps to improve face recognition accuracy. The proposed approach contains two separate matching steps: Sparse Representation Classifier (SRC) is applied to 2D-2D matching, while Canonical Correlation Analysis (CCA) is exploited to learn the mapping between range LBP faces (3D) and texture LBP faces (2D). Both matching scores are combined for the final decision. Moreover, we propose a new preprocessing pipeline to enhance robustness to lighting and pose effects. The proposed method achieves better experimental results in the FRGC v2.0 dataset than 2D methods do, but avoiding the cost and inconvenience of data acquisition and computation of 3D approaches.
Di Huang 0001, Mohsen Ardabilian, Yunhong Wang 0001, Liming Chen 0002
ICPR4
2010 Unifying Approach for Fast License Plate Localization and Super-Resolution
abstract
This paper addresses the localization and super resolution of license plate in a unifying approach. Higher quality license plate can be obtained using super resolution on successive lower resolution plate images. All existing methods assume that plate zones are correctly extracted from every frame. However, the accurate localization needs a sufficient quality of the image, which is not always true in real video. Super-resolution on all pixels is a possible but much time consuming alternative. We propose a framework which interlaces successfully these two modules. First, coarse candidates are found by an weak but fast license plate detection based on edge map sub-sampling. Then, an improved fast MAP-based super-resolution, using local phase accurate registration and edge preserving prior, applied on these regions of interest. Finally, our robust ICHT-based localizer rejects false-alarms and localizes the high resolution license plate more accurately. Experiments which were conducted on synthetic and real data, proved the robustness of our approach with real-time possibility.
Chu Duc Nguyen, Mohsen Ardabilian, Liming Chen 0002
ICPR3
2010 Adaptive Feature and Score Level Fusion Strategy Using Genetic Algorithms
abstract
Classifier fusion is considered as one of the best strategies for improving performance of general purpose classification systems. On the other hand, fusion strategy space strongly depends on classifiers, features and data spaces. As the cardinality of this space is exponential, one needs to resort to a heuristic to find a sub-optimal fusion strategy. In this work, we present a new adaptive feature and score level fusion strategy (AFSFS) based on adaptive genetic algorithm. AFSFS tunes itself between feature and matching score level, and improves the final performance over the original on both levels, and as a fusion method, it does not only contain fusion strategy to combine the most relevant features so as to achieve adequate and optimized results, but also has the extensive ability to select the most discriminative features. Experiments are provided on the FRGC database showing that the proposed method produces significantly better results than the baseline fusion methods.
Wael Ben Soltana, Mohsen Ardabilian, Liming Chen 0002, Chokri Ben Amar
ICPR3
2010 Partial Face Biometry Using Shape Decomposition on 2D Conformal Maps of Faces
abstract
In this paper, we introduce a new approach for partial 3D face recognition, which makes use of shape decomposition over the rigid part of a face. To explore the descriptiveness of shape dissimilarity over an isometric part of a face, which has lower probability to be influenced by expression, we transform a 3D shape to a 2D domain using conformal mapping and use shape decomposition as a similarity measurement. In our work we investigate several classifiers as well as several shape descriptors for recognition purposes. Recognition tests on a subset of the FRGC data set show approximately 80% rank-one recognition rate using only the eyes and nose part of the face.
Przemyslaw Szeptycki, Mohsen Ardabilian, Liming Chen 0002, Wei Zeng 0002, Xianfeng Gu, Dimitris Samaras
ICPR3
2010 Bayesian GOETHE Tracking
abstract
Occlusions pose serious challenges when tracking multiple targets. By severly changing the measurement, they imply strong inter-target dependencies. Exact computation of these dependencies is not feasible. The GOETHE approximations preserve much of the information while staying computationally affordable.
Sebastian J. Wirkert, Emmanuel Dellandréa, Liming Chen 0002
ICPR3
2010 Automatic 3D Facial Expression Recognition Based on a Bayesian Belief Net and a Statistical Facial Feature Model
abstract
Automatic facial expression recognition on 3D face data is still a challenging problem. In this paper we propose a novel approach to perform expression recognition automatically and flexibly by combining a Bayesian Belief Net (BBN) and Statistical facial feature models (SFAM). A novel BBN is designed for the specific problem with our proposed parameter computing method. By learning global variations in face landmark configuration (morphology) and local ones in terms of texture and shape around landmarks, morphable Statistic Facial feature Model (SFAM) allows not only to perform an automatic landmarking but also to compute the belief to feed the BBN. Tested on the public 3D face expression database BU-3DFE, our automatic approach allows to recognize expressions successfully, reaching an average recognition rate over 82%.
Xi Zhao 0001, Di Huang 0001, Emmanuel Dellandréa, Liming Chen 0002
ICPR4
2010 Multi-scale Color Local Binary Patterns for Visual Object Classes Recognition
abstract
The Local Binary Pattern (LBP) operator is a computationally efficient yet powerful feature for analyzing local texture structures. While the LBP operator has been successfully applied to tasks as diverse as texture classification, texture segmentation, face recognition and facial expression recognition, etc., it has been rarely used in the domain of Visual Object Classes (VOC) recognition mainly due to its deficiency of power for dealing with various changes in lighting and viewing conditions in real-world scenes. In this paper, we propose six novel multi-scale color LBP operators in order to increase photometric invariance property and discriminative power of the original LBP operator. The experimental results on the PASCAL VOC 2007 image benchmark show significant accuracy improvement by the proposed operators as compared with both the original LBP and other popular texture descriptors such as Gabor filter.
Chao Zhu 0003, Charles-Edmond Bichot, Liming Chen 0002
ICPR3
2010 Multi-stage classification of emotional speech motivated by a dimensional emotion model
Zhongzhe Xiao, Emmanuel Dellandréa, Weibei Dou, Liming Chen 0002
Multim. Tools Appl.4
2009 Image Categorization Using ESFS: A New Embedded Feature Selection Method Based on SFS
Huanzhang Fu, Zhongzhe Xiao, Emmanuel Dellandréa, Weibei Dou, Liming Chen 0002
ACIVS5
2009 Pattern Analysis for an Automatic and Low-Cost 3D Face Acquisition Technique
Karima Ouji, Mohsen Ardabilian, Liming Chen 0002, Faouzi Ghorbel
ACIVS3
2009 A 3D Statistical Facial Feature Model and Its Application on Locating Facial Landmarks
Xi Zhao 0001, Emmanuel Dellandréa, Liming Chen 0002
ACIVS3
2009 Robust Car License Plate Localization Using a Novel Texture Descriptor
abstract
This paper presents a novel texture descriptor based on line-segment features for text detection in images and video sequences, which is applied to build a robust car license plate localization system. Unlike most of existing approaches which use low level features (color, edge) for text/non-text discrimination, our aim is to exploit more accurate perceptual information. A - scale and rotation invariant - texture descriptor which describes the directionality, regularity, similarity, alignment and connectivity of group of segments is proposed. A improved algorithm for feature extraction based on local connective Hough transform has been also investigated. The robustness of our approach is proved throughout a real-time detection/verification scheme of car license plate. First, all possible candidates are detected using a rule based method, which is very robust to illumination change and in varying poses. Then, true license plates are identified by the mean of a SVM classifier trained with proposed descriptor. Comparison and evaluation are conducted with two complex datasets.
Chu Duc Nguyen, Mohsen Ardabilian, Liming Chen 0002
AVSS3
2009 A People Counting System Based on Face Detection and Tracking in a Video
abstract
Vision-based people counting systems have wide potential applications including video surveillance and public resources management. Most works in the literature rely on moving object detection and tracking, assuming that all moving objects are people. In this paper, we present our people counting approach based on face detection, tracking and trajectory classification. While we have used a standard face detector, we achieve face tracking combining a new scale invariant Kalman filter with kernel based tracking algorithm. From each potential face trajectory an angle histogram of neighboring points is then extracted. Finally, an Earth Mover's Distance-based K-NN classification discriminates true face trajectories from the false ones. Experimented on a video dataset of more than 160 potential people trajectories, our approach displays an accuracy rate up to 93%.
Xi Zhao 0001, Emmanuel Dellandréa, Liming Chen 0002
AVSS3
2009 Visual Object Categorization via Sparse Representation
abstract
In this paper, we consider the problem of classifying a real world image to the corresponding object class based on its visual content via sparse representation, which is originally used as a powerful tool for acquiring, representing and compressing high-dimensional signals. Assuming the intuitive hypothesis that an image could be represented by a linear combination of the training images from the same class, we propose a novel approach for visual object categorization in which a sparse representation of the image is first of all obtained by solving a L1 (or L0)-minimization problem and then fed into a traditional classifier such as Support Vector Machine (SVM) to finally perform the specified task. Experimental results obtained on the SIMPLIcity database have shown that this new approach can improve the classification performance compared to standard SVM using directly features extracted from the image.
Huanzhang Fu, Chao Zhu 0003, Emmanuel Dellandréa, Charles-Edmond Bichot, Liming Chen 0002
ICIG5
2009 Asymmetric 3D/2D face recognition based on LBP facial representation and canonical correlation analysis
abstract
In the recent years, 3D Face recognition has emerged as a major solution to deal with the unsolved issues for reliable 2D face recognition, i.e. lighting condition and viewpoint variations. However, 3D method is currently limited by its registration and computation cost. In this paper, we propose to investigate a solution named asymmetric face recognition scheme, enrolling people in 3D environment but performing identification in 2D. The goal is to limit the use of 3D data to where it really helps to improve recognition performances. In our approach, Local Binary Patterns (LBP) is used as an efficient facial representation for both 2D texture images and 3D range images. A weighted Chi square distance is used as matching score between the 2D LBP facial representations; Canonical Correlation Analysis (CCA) is applied to learn the mapping between LBP-based range face images (3D) and LBP facial texture images (2D). Both matching scores are further fused to obtain the final result. Compared with the traditional 2D/2D algorithms, the proposed asymmetric face recognition scheme achieves better accuracy; while avoiding the high cost of data acquisition and computation in 3D/3D approaches.
Di Huang 0001, Mohsen Ardabilian, Yunhong Wang 0001, Liming Chen 0002
ICIP4
2009 3D Face Recognition Using R-ICP and Geodesic Coupled Approach
Karima Ouji, Boulbaba Ben Amor, Mohsen Ardabilian, Liming Chen 0002, Faouzi Ghorbel
MMM4
2009 Precise 2.5D facial landmarking via an analysis by synthesis approach
abstract
3D face landmarking aims at automatic localization of 3D facial features and has a wide range of applications, including face recognition, face tracking, facial expression analysis. Methods so far developed for pure 2D texture images were shown sensitive to lighting condition changes. In this paper, we present a statistical model-based technique for accurate 3D face landmarking, thus using an ¿analysis by synthesis¿ approach. Our model learns from a training set both variations of global face shapes as well as the local ones in terms of scale-free texture and range patches around each landmark. Given a shape instance, local regions of a new face can be approximated by synthesizing texture and range instances using respectively the texture and range models. By optimizing an objective function describing the similarity of the new face and instances, we can optimize the best shape in order to locate the landmarks. Experimented on more than 1860 face models from FRGC datasets, our method achieves an average of locating errors less than 7 mm for 15 feature points. Compared with a curvature analysis-based method also developed within our team, this learning-based method enables localization of more facial landmarks with a general better accuracy at the cost of a learning step.
Xi Zhao 0001, Przemyslaw Szeptycki, Emmanuel Dellandréa, Liming Chen 0002
WACV4
2008 3D Face Recognition Evaluation on Expressive Faces Using the IV2 Database
Joseph Colineau, Johan D'Hose, Boulbaba Ben Amor, Mohsen Ardabilian, Liming Chen 0002, Bernadette Dorizzi
ACIVS5
2008 Toward a region-based 3D face recognition approach
abstract
3D face has recently emerged as a major trend in face recognition. While 3D face is reputed to be relatively invariant to lighting conditions and pose, one still needs to cope with facial expression variations for a reliable face recognition solution. Many works in the literature try to define some invariants with respect to facial expressions. In this paper, we consider facial expressions as results of moving parts of a facial surface, and thus a careful study of anatomical face enables us to set up a region-based face surface matching, giving more importance to stable facial regions in the comparison. The very first experiments show the effectiveness of our approach.
Boulbaba Ben Amor, Mohsen Ardabilian, Liming Chen 0002
ICME3
2007 A general audio classifier based on human perception motivated model
Hadi Harb, Liming Chen 0002
Multim. Tools Appl.2
2006 Combining Short and Long Term Audio Features for TV Sports Highlight Detection
Weibei Dou, Liming Chen 0002
ECIR3
2006 A Smart Identification Card System Using Facial Biometric: From Architecture to Application
Liming Chen 0002, Su Ruan
UIC2
2006 WebGuard: A Web Filtering Engine Combining Textual, Structural, and Visual Content-Based Analysis
abstract
Along with the ever-growing Web comes the proliferation of objectionable content, such as sex, violence, racism, etc. We need efficient tools for classifying and filtering undesirable Web content. In this paper, we investigate this problem and describe WebGuard, an automatic machine learning-based pornographic Web site classification and filtering system. Unlike most commercial filtering products, which are mainly based on textual content-based analysis such as indicative keywords detection or manually collected black list checking, WebGuard relies on several major data mining techniques associated with textual, structural content-based analysis, and skin color related visual content-based analysis as well. Experiments conducted on a testbed of 400 Web sites including 200 adult sites and 200 nonpornographic ones showed WebGuard's filtering effectiveness, reaching a 97.4 percent classification accuracy rate when textual and structural content-based analysis was combined with visual content-based analysis. Further experiments on a black list of 12,311 adult Web sites manually collected and classified by the French Ministry of Education showed that WebGuard scored a 95.62 percent classification accuracy rate. The basic framework of WebGuard can apply to other categorization problems of Web sites which combine, as most of them do today, textual and visual content.
Mohamed Hammami, Youssef Chahir, Liming Chen 0002
IEEE Trans. Knowl. Data Eng.3
2005 A novel scheme of face verification using active appearance models
abstract
Face verification and face identification are two main applications for face recognition. Based on the face representation which they use, the current image-based face recognition techniques can be roughly classified into two groups: appearance-based methods and model-based methods. The method presented in this paper focuses on applying model-based methods to face verification. Unfortunately, the existing current model-based schemes are used for the applications of face identification and are not suitable for face verification. A novel scheme, which is custom-built for face verification applications, is therefore proposed in this paper. Active appearance model, a 2D morphable face model, is chosen and applied to realize the proposed scheme.
Liming Chen 0002, Su Ruan
AVSS2
2005 Features extraction and selection for emotional speech classification
abstract
The classification of emotional speech is a topic in speech recognition with more and more interest, and it has giant prospect in applications in a wide variety of fields. It is an important preparation for automatic classification and recognition of emotions to select a proper feature set as a description to the emotional speech, and to find a proper definition to the emotions in speech. The speech samples used in this paper come from Berlin database which contains 7 kinds of emotions, with 207 speech samples of male voice and 287 speech samples of female voice. A feature set of 50 potentially features is extracted and analyzed, and the best features are selected. A definition of emotions as 3-states emotions is also proposed in this paper.
Zhongzhe Xiao, Emmanuel Dellandréa, Weibei Dou, Liming Chen 0002
AVSS4
2005 Voice-Based Gender Identification in Multimedia Applications
Hadi Harb, Liming Chen 0002
J. Intell. Inf. Syst.2
2004 Adult content Web filtering and face detection using data-mining based kin-color model
abstract
The paper presents a novel approach for robust skin-color detection using data-mining techniques. The goal of skin-color detection is to select the appropriate color model that allows pixels to be verified under different lighting conditions and other variations. When the appropriate color model is selected, it is implied that we have good skin-color classifier properties for skin detection. This model has been successfully applied to face detection and Web based adult content filtering issues.
Mohamed Hammami, Dzmitry V. Tsishkou, Liming Chen 0002
ICME3
2004 Mixture of experts for audio classification: an application to male female classification and musical genre recognition
abstract
We report the experimental results obtained when applying a mixture of experts to the problem of audio classification for multimedia applications. The mixture of experts is based on multilayer perceptron neural networks as individual experts and piecewise Gaussian modeling was used for audio signal representation. Experimental results on two audio classification problems, male/female classification and musical genre recognition, show a clear improvement in using a mixture of experts in comparison to one individual expert
Hadi Harb, Liming Chen 0002, Jean-Yves Auloge
ICME2
2004 Combining Text And Image Analysis in The Web Filtering System "WEBGUARD"
Mohamed Hammami, Youssef Chahir, Liming Chen 0002
iiWAS3
2003 Gender identification using a general audio classifier
abstract
In the context of content-based multimedia indexing gender identification using speech signal is an important task. Existing techniques are dependent on the quality of the speech signal making them unsuitable for the video indexing problems. In this paper we introduce a novel gender identification approach based on a general audio classifier. The audio classifier models the audio signal by the first order spectrum's statistics in 1s windows and uses a set of neural networks as classifiers. The presented technique shows robustness to adverse audio compression and it is language independent. We show how practical considerations about the speech in audio-visual data, such as the continuity of speech, can further improve the classification results which attain 92%.
Hadi Harb, Liming Chen 0002
ICME2
2003 WebGuard: Web Based Adult Content Detection and Filtering System
abstract
@inproceedings{CI-CHAHIR-2003, author = {Hammami, M. and Chahir, Y. and Chen, L.}, title = {WebGuard: Web-Based Adult Content Detection and Filtering System}, booktitle = {IEEE/WIC International Conference on Web Intelligence (WI'03)}, pages = {574-578}, year = {2003}, address = {Halifax, Canada}, month = {October} }
Mohamed Hammami, Youssef Chahir, Liming Chen 0002
Web Intelligence3
2002 Self-Stabilizing Deterministic Network Decomposition
Fatima Belkouch, Marc Bui, Liming Chen 0002, Ajoy K. Datta
J. Parallel Distributed Comput.3
2000 Searching Images on the Basis of Color Homogeneous Objects and their Spatial Relationship
Youssef Chahir, Liming Chen 0002
J. Vis. Commun. Image Represent.2
1999 Self-Stabilizing Network Decomposition
Fatima Belkouch, Marc Bui, Liming Chen 0002, Ajoy K. Datta
HiPC3
1997 Multi-criteria video segmentation for TV news
abstract
In the near future, personalized TV news programs should become a relevant method to access TV news. They will be produced from selecting and reordering "stories" contained in news programs. Automatic segmentation of news programs into story is an important step to reach this goal. We present a set rules for such segmentation. Methods for implementing these rules are based on image and sound analysis.
Liming Chen 0002, Pascal Faudemay
MMSP1
1997 Self-Stabilizing Planar Quorums
Fatima Belkouch, Liming Chen 0002
OPODIS2