EDBT 2026 Demo / reviewers in the wild / expert
Hichem Sahli
dblp:33/4493
· DBLP profile ↗
116ranked-venue papers
0as first author
21since 2021 · last 2026
0000-0002-1774-2970ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 56 · 9 since 2021Artificial intelligence and machine learning · 27 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 22 · 4 since 2021Human-computer interaction and ubiquitous computing · 15Computer networks · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ADGT: Enhancing 3D human pose estimation with attention-driven graph-transformersabstract2D-to-3D lifting is a fundamental approach in 3D human pose estimation (3DHPE). This task is crucial in applications, including motion analysis and virtual reality. While Graph Convolutional Networks (GCNs) have demonstrated effectiveness in capturing spatial relationships in human skeletons, they suffer from over-smoothing and limited receptive fields. Transformer-based models provide global context but struggle with local feature extraction and computational efficiency. To address these challenges, we propose ADGT, a novel parallel GCN-transformer architecture combining the strengths of both approaches. Our method introduces three key innovations: Hop-Wise Scalable Adaptive GCN to refine local feature extraction, Attention-Based Local Feature Extractor to enhance the integration of local and global representations, and Register-Based Transformer Enhancement to improve feature separation. Extensive experiments on Human3.6M and MPI-INF-3DHP datasets demonstrate ADGT achieves state-of-the-art performance among frame-based methods while maintaining computational efficiency. These results highlight the potential of ADGT for real-time applications requiring accurate and efficient 3DHPE. The code is available at https://github.com/sYANGunique1111/ADGT . Anh Tuan Luu, Xuan Son Nguyen, Aymeric Histace, Bart Jansen 0001, Hichem Sahli |
J. Vis. Commun. Image Represent. | 6 |
| 2024 | Co-Occurrence Graph-Enhanced Hierarchical Prediction of ICD CodesabstractRecent healthcare applications of natural language processing involve multi-label classification of health records using the International Classification of Diseases (ICD). While prior research highlights intricate text models and explores external knowledge like hierarchical ICD ontology, fewer studies integrate code relationships from whole datasets to enhance ICD coding accuracy. This study presents a modular approach, sequentially combining graph-based integration of ICD code co-occurrence with a hard-coded hierarchical-enriched text representation drawn from the ICD ontology. Findings reveal: 1) significant performance gains in the combined model, aside from the significant performance gain in each enhancement module in isolation, 2) graph-based module’s efficacy is more pronounced when applied to enhanced features using the hierarchical ICD ontology, and 3) experiments demonstrate hierarchy depth’s impact on performance, concluding the deepest level’s enrichment. Soha Sadat Mahdi, Eirini Papagiannopoulou, Nikos Deligiannis, Hichem Sahli |
ICASSP | 4 |
| 2024 | Semi-supervised medical image classification via distance correlation minimization and graph attention regularization
Abel Díaz Berenguer, Maryna Kvasnytsia, Matías N. Bossa, Tanmoy Mukherjee, Nikos Deligiannis, Hichem Sahli |
Medical Image Anal. | 6 |
| 2024 | The STOIC2021 COVID-19 AI challenge: Applying reusable training methodologies to private dataabstractChallenges drive the state-of-the-art of automated medical image analysis. The quantity of public training data that they provide can limit the performance of their solutions. Public access to the training methodology for these solutions remains absent. This study implements the Type Three (T3) challenge format, which allows for training solutions on private data and guarantees reusable training methodologies. With T3, challenge organizers train a codebase provided by the participants on sequestered training data. T3 was implemented in the STOIC2021 challenge, with the goal of predicting from a computed tomography (CT) scan whether subjects had a severe COVID-19 infection, defined as intubation or death within one month. STOIC2021 consisted of a Qualification phase, where participants developed challenge solutions using 2000 publicly available CT scans, and a Final phase, where participants submitted their training methodologies with which solutions were trained on CT scans of 9724 subjects. The organizers successfully trained six of the eight Final phase submissions. The submitted codebases for training and running inference were released publicly. The winning solution obtained an area under the receiver operating characteristic curve for discerning between severe and non-severe COVID-19 of 0.815. The Final phase solutions of all finalists improved upon their Qualification phase solutions. Luuk H. Boulogne, Julian Lorenz, Daniel Kienzle, Robin Schön, Katja Ludwig, Rainer Lienhart, Simon Jégou, Derik Shi, Mayug Maniparambil, Dominik Müller, Silvan Mertes, Niklas Schröter, Fabio Hellmann, Miriam Elia, Ine Dirks, Matías N. Bossa, Abel Díaz Berenguer, Tanmoy Mukherjee, Jef Vandemeulebroucke, Hichem Sahli, Nikos Deligiannis, Panagiotis Gonidakis, Ngoc Dung Huynh, Muhammad Imran Razzak, Mohamed Reda Bouadjenek, Mario Verdicchio, Pasquale Borrelli, Marco Aiello 0003, James A. Meakin, Alexander Lemm, Christoph Russ, Razvan Ionasec, Nikos Paragios, Bram van Ginneken, Marie-Pierre Revel |
Medical Image Anal. | 23 |
| 2024 | Long-Term Regional Influenza-Like-Illness Forecasting Using Exogenous DataabstractDisease forecasting is a longstanding problem for the research community, which aims at informing and improving decisions with the best available evidence. Specifically, the interest in respiratory disease forecasting has dramatically increased since the beginning of the coronavirus pandemic, rendering the accurate prediction of influenza-like-illness (ILI) a critical task. Although methods for short-term ILI forecasting and nowcasting have achieved good accuracy, their performance worsens at long-term ILI forecasts. Machine learning models have outperformed conventional forecasting approaches enabling to utilize diverse exogenous data sources, such as social media, internet users' search query logs, and climate data. However, the most recent deep learning ILI forecasting models use only historical occurrence data achieving state-of-the-art results. Inspired by recent deep neural network architectures in time series forecasting, this work proposes the Regional Influenza-Like-Illness Forecasting (ReILIF) method for regional long-term ILI prediction. The proposed architecture takes advantage of diverse exogenous data, that are, meteorological and population data, introducing an efficient intermediate fusion mechanism to combine the different types of information with the aim to capture the variations of ILI from various views. The efficacy of the proposed approach compared to state-of-the-art ILI forecasting methods is confirmed by an extensive experimental study following standard evaluation measures. Eirini Papagiannopoulou, Matías N. Bossa, Nikos Deligiannis, Hichem Sahli |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | Fusing Event-based Camera and Radar for SLAM Using Spiking Neural Networks with Continual STDP LearningabstractThis work proposes a first-of-its-kind SLAM architecture fusing an event-based camera and a Frequency Modulated Continuous Wave (FMCW) radar for drone navigation. Each sensor is processed by a bio-inspired Spiking Neural Network (SNN) with continual Spike-Timing-Dependent Plasticity (STDP) learning, as observed in the brain. In contrast to most learning-based SLAM systems, our method does not require any offline training phase, but rather the SNN continuously learns features from the input data on the fly via STDP. At the same time, the SNN outputs are used as feature descriptors for loop closure detection and map correction. We conduct numerous experiments to benchmark our system against state-of-the-art RGB methods and we demonstrate the robustness of our DVS-Radar SLAM approach under strong lighting variations. Ali Safa, Tim Verbelen, Ilja Ocket, André Bourdoux, Hichem Sahli, Francky Catthoor, Georges Gielen |
ICRA | 5 |
| 2023 | Region Attentive Action Unit Intensity Estimation With Uncertainty Weighted Multi-Task LearningabstractFacial action units (AUs) refer to a comprehensive set of atomic facial muscle movements. Recent works have focused on exploring complementary information by learning the relationships among AUs. Most existing approaches process AU co-occurrence and enhance AU recognition by learning the dependencies among AUs from labels, however, the complementary information among features of different AUs are ignored. Moreover, ground truth annotations suffer from a large intra-class variance and their associated intensity levels may vary depending on the annotators’ experience. In this paper, we propose the Region Attentive AU intensity estimation method with Uncertainty Weighted Multi-task Learning (RA-UWML). A RoI-Net is first used to extract features from the pre-defined facial patches where the AUs locate. Then, we use the co-occurrence of AUs using both within patch and between patches representation learning. Within a given patch, we propose sharing representation learning in a multi-task manner. To achieve complementarity and avoid redundancy between different image patches, we propose to use a multi-head self-attention mechanism to adaptively and attentively encode each patch specific representation. Moreover, the AU intensity is represented as a Gaussian distribution, instead of a single value, where the mean value indicates the most likely AU intensity and the variance indicates the uncertainty of the estimated AU intensity. The estimated variances are leveraged to automatically weight the loss of each AU in the multitask learning model. In extensive experiments on the Disfa, Fera2015 and Feafa benchmarks, it is shown that the proposed AU intensity estimation model achieves better results compared to the state-of-the-art models. Dongmei Jiang, Xiaoyong Wei, Ke Lu 0002, Hichem Sahli |
IEEE Trans. Affect. Comput. | 6 |
| 2023 | A Bayesian Filtering Framework for Continuous Affect Recognition From Facial ImagesabstractContinuous affective state estimation from facial information is a task which requires the prediction of time series of emotional state outputs from a facial image sequence. Modeling the spatial-temporal evolution of facial information plays an important role in affective state estimation. One of the most widely used methods is Recurrent Neural Networks (RNN). RNNs provide an attractive framework for propagating information over a sequence using a continuous-valued hidden layer representation. In this work, we propose to instead learn rich affective state dynamics. We model human affect as a dynamical system and define the affective state in terms of valence, arousal and their higher-order derivatives. We then pose the affective state estimation problem as a jointly trained state estimator for high-dimensional input images, combining an RNN and a Bayesian Filter, i.e. Kalman filters (KF) and Extended Kalman filters (EKF), so that all weights in the resulting network can be trained using backpropagation. We use a recently proposed general framework for designing and learning discriminative state estimators framed as computational graphs. Such approach can handle high dimensional observations and efficiently optimize, in an end-to-end fashion, the state estimator. In addition, to deal with the asynchrony between emotion labels and input images, caused by the inherent reaction lag of the annotators, we introduce a convolutional layer that aligns features with emotion labels. Experimental results, on the RECOLA and SEMAINE datasets for continuous emotion prediction, illustrate the potential of the proposed framework compared to recent state-of-the-art models. Ercheng Pei, Meshia Cédric Oveneke, Dongmei Jiang, Hichem Sahli |
IEEE Trans. Multim. | 5 |
| 2022 | Uplink Payload Power Control in Cell-Free Communication and Radar NetworksabstractThis paper considers the cell-free (CF) massive multiple-input multiple-output (MIMO) architecture with an added virtual uplink radar integration. We consider joint communications and radar (JCR) in the uplink payload of the time-division duplex (TDD) frame and formulate a set of linear interference constraints in order to limit the user equipment (UE) interference imposed on the radar system at each accesspoint (AP) through power control. The constraints are incorporated into the sum spectral efficiency (SSE) and the sum-log-SNR (SLS) policies. Furthermore, a largest large-scale fading (LLSF) heuristic is formulated as an approximate low-complexity solution. Numerical simulations indicate that all methods provide similar performance in terms of spectral efficiency (SE) and that power control allows the communication system to be controlled to satisfy a given worst case radar performance. Adham Sakhnini, André Bourdoux, Mamoun Guenach, Hichem Sahli, Sofie Pollin |
GLOBECOM | 4 |
| 2022 | A Target Detection Analysis in Cell-Free Massive MIMO Joint Communication and Radar SystemsabstractThis paper considers the cell-free (CF) massive MIMO architecture from a joint communication and radar point of view. We propose a protocol for communication and sensing, where the network allocates a set of access-points (APs) to participate in the uplink together with the users (UEs) to serve the network with a radar signal. This realizes the radar system as a distributed bistatic radar system, where the objective is to recover the radar echoes from the multiuser interference. The imposed cost on the communication system is the loss of one AP and the need to schedule one additional virtual UE per allocated AP. We present two modes of radar sensing, occurring in either the uplink training segment or the data payload segment of the communication frame. A subspace signal model and its corresponding generalized likelihood ratio test is developed in order to evaluate the detection performance. We present expressions for the probably of detection and false alarm, and demonstrate the system numerically. Our main message is that coordinating the network to transmit in the uplink together with the UEs serves as an interesting approach in enabling radar sensing in CF massive MIMO systems. Adham Sakhnini, Mamoun Guenach, André Bourdoux, Hichem Sahli, Sofie Pollin |
ICC | 4 |
| 2022 | Event Camera Data Classification Using Spiking Networks with Spike-Timing-Dependent PlasticityabstractWe present an optimization-based theory describing spiking cortical ensembles equipped with Spike-Timing-Dependent Plasticity (STDP) learning, as empirically observed in the visual cortex. Using this generic framework, we build a class of global and action-based feature descriptors for event-based cameras that we assess on the N-MNIST and the IBM DVS128 Gesture datasets. We report significant accuracy improvements compared to state-of-the-art STDP-based systems (+9.3% on N-MNIST, +7.74% on IBM DVS128 Gesture). In addition to ultra-low-power learning in neuromorphic edge devices, our work contributes towards a biologically-plausible, optimization-based theory of cortical vision. Ali Safa, Ilja Ocket, André Bourdoux, Hichem Sahli, Francky Catthoor, Georges Gielen |
IJCNN | 4 |
| 2022 | Uncertainty-Aware Semi-Supervised Learning of 3D Face Rigging from Single ImageabstractWe present a method to rig 3D faces via Action Units (AUs), viewpoint and light direction, from single input image. Existing 3D methods for face synthesis and animation rely heavily on 3D morphable model (3DMM), which was built on 3D data and cannot provide intuitive expression parameters, while AU-driven 2D methods cannot handle head pose and lighting effect. We bridge the gap by integrating a recent 3D reconstruction method with 2D AU-driven method in a semi-supervised fashion. Built upon the auto-encoding 3D face reconstruction model that decouples depth, albedo, viewpoint and light without any supervision, we further decouple expression from identity for depth and albedo with a novel conditional feature translation module and pretrained critics for AU intensity estimation and image classification. Novel objective functions are designed using unlabeled in-the-wild images and in-door images with AU labels. We also leverage uncertainty losses to model the probably changing AU region of images as input noise for synthesis, and model the noisy AU intensity labels for intensity estimation of the AU critic. Experiments with face editing and animation on four datasets show that, compared with six state-of-the-art methods, our proposed method is superior and effective on expression consistency, identity similarity and pose similarity. Hichem Sahli, Ke Lu 0002, Dongmei Jiang |
ACM Multimedia | 3 |
| 2022 | A multi-scale multi-attention network for dynamic facial expression recognition
Xiaohan Xia, Le Yang 0009, Xiaoyong Wei, Hichem Sahli, Dongmei Jiang |
Multim. Syst. | 4 |
| 2022 | Leveraging the Deep Learning Paradigm for Continuous Affect Estimation from Facial ExpressionsabstractContinuous affect estimation from facial expressions has attracted increased attention in the affective computing research community. This paper presents a principled framework for estimating continuous affect from video sequences. Based on recent developments, we address the problem of continuous affect estimation by leveraging the Bayesian filtering paradigm, i.e., considering affect as a latent dynamical system corresponding to a general feeling of pleasure with a degree of arousal, and recursively estimating its state using a sequence of visual observations. To this end, we advance the state-of-the-art as follows: (i) Canonical face representation (CFR): a novel algorithm for two-dimensional face frontalization, (ii) Convex unsupervised representation learning (CURL): a novel frequency-domain convex optimization algorithm for unsupervised training of deep convolutional neural networks (CNN)s, and (iii) Deep extended Kalman filtering (DEKF): an extended Kalman filtering-based algorithm for affect estimation from a sequence of CNN observations. The performance of the resulting CFR-CURL-DEKF algorithmic framework is empirically evaluated on publicly available benchmark datasets for facial expression recognition (CK+) and continuous affect estimation (AVEC 2012 and 2014). Meshia Cédric Oveneke, Ercheng Pei, Abel Díaz Berenguer, Dongmei Jiang, Hichem Sahli |
IEEE Trans. Affect. Comput. | 6 |
| 2022 | Multipath Ghost Recognition for Indoor MIMO RadarabstractMultipath is a challenging problem for radar-based localization systems, especially in indoor scenarios. Multipath is caused by the bounces from static objects like walls and furniture in the room creating false alarms (“ghosts”) in target detections. Although solutions for the multipath effect have been proposed for a range of radar sensing problems, the specific case of multipath recognition and mitigation for a colocated multiple-input–multiple-output (MIMO) radar remains unsolved. For MIMO radar, the different direction-of-arrival (DoA) and direction-of-departure (DoD) angles inhibit the use of beamforming with a virtual array for localizing the first-order ghosts. Additionally, the prior knowledge of the multipath geometry model (room layout and boundary) is not always accessible. Classical ray tracing methods to resolve multipath are hence, not practical. In this work, we exploit a linear relationship between the target and multipath ghosts in the range-Doppler map to propose a Hough-transform-based multipath recognition solution. The algorithm does not require prior multipath geometry information and applies to the various indoor environments for an MIMO radar. Simulation and measurement results demonstrate the effectiveness of the proposed algorithm. Eddy De Greef, Maxim Rykunov, Hichem Sahli, Sofie Pollin, André Bourdoux |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Near-Field Coherent Radar Sensing Using a Massive MIMO Communication TestbedabstractThis paper considers the problem of radar sensing by using a large number of antennas. We use the orthogonal frequency division multiplexing (OFDM) waveform, and show that the large arrays used in massive multiple-input multiple-output (MIMO) communications enable accurate localization in the array near-field, even at the narrow bandwidths typically encountered at low carrier frequencies. We validate our findings experimentally with a massive MIMO testbed operating at 3.5 GHz carrier frequency and 18 MHz OFDM bandwidth in an indoor environment. We consider a single moving cylinder, and demonstrate a median accuracy of (3.4, 5.6) cm in ($x$,$y$) in the near-field. We show that the accuracy is maintained with only a single subcarrier, and that the resolution increases with an order of magnitude when combining all antennas, effectively surpassing the 16.67 m bistatic range resolution set by the OFDM waveform. We use a radar symbol duration of$71.88~\mu $s at an effective transmission period of 2.5 ms, which indicates that the radar and communication systems can be implemented in time-division with a capacity loss of only 2.9%. Our results suggest that near-field radar sensing can be integrated into future massive MIMO systems operating at low carrier frequencies and narrow bandwidths. Adham Sakhnini, Sibren De Bast, Mamoun Guenach, André Bourdoux, Hichem Sahli, Sofie Pollin |
IEEE Trans. Wirel. Commun. | 5 |
| 2021 | Action Unit Driven Facial Expression Synthesis from a Single Image with Patch Attentive GANabstractAbstract Recent advances in generative adversarial networks (GANs) have shown tremendous success for facial expression generation tasks. However, generating vivid and expressive facial expressions at Action Units (AUs) level is still challenging, due to the fact that automatic facial expression analysis for AU intensity itself is an unsolved difficult task. In this paper, we propose a novel synthesis‐by‐analysis approach by leveraging the power of GAN framework and state‐of‐the‐art AU detection model to achieve better results for AU‐driven facial expression generation. Specifically, we design a novel discriminator architecture by modifying the patch‐attentive AU detection network for AU intensity estimation and combine it with a global image encoder for adversarial learning to force the generator to produce more expressive and realistic facial images. We also introduce a balanced sampling approach to alleviate the imbalanced learning problem for AU synthesis. Extensive experimental results on DISFA and DISFA+ show that our approach outperforms the state‐of‐the‐art in terms of photo‐realism and expressiveness of the facial expression quantitatively and qualitatively. Le Yang 0009, Ercheng Pei, Meshia Cédric Oveneke, Mitchel Alioscha-Pérez, Dongmei Jiang, Hichem Sahli |
Comput. Graph. Forum | 8 |
| 2021 | Integrating Deep and Shallow Models for Multi-Modal Depression Analysis - Hybrid ArchitecturesabstractAt present, although great progress has been made in automatic depression assessment, most of the recent works only concern the audio and video paralinguistic information, rather than the linguistic information from the spoken content. In this work, we argue that beside developing good audio and video features, to build reliable depression detection systems, text-based content features are also of importance to analyse depression-related textual indicators. Furthermore, to improve the performance of automatic depression assessment systems, powerful models, capable of modelling the characteristics of depression embedded in the audio, visual and text descriptors, are also required. This paper proposes new text and video features and hybridizes deep and shallow models for depression estimation and classification from audio, video and text descriptors. The proposed hybrid framework consists of three main parts: 1) A Deep Convolutional Neural Network (DCNN) and Deep Neural Network (DNN) based audio-visual multi-modal depression recognition model for estimating the Patient Health Questionnaire depression scale (PHQ-8); 2) A Paragraph Vector (PV) and Support Vector Machine (SVM) based model for inferring the physical and mental conditions of the individual from the transcripts of the interview; 3) A Random Forest (RF) model for depression classification from the estimated PHQ-8 score and the inferred conditions of the individual. In the PV-SVM model, PV embedding is used to obtain fixed-length feature vectors from transcripts of the answers to the questions associated with psychoanalytic aspects of depression, which are subsequently fed into the SVM classifiers for detecting the presence/absence of the considered psychoanalytic symptoms. To our best knowledge, this approach is the first attempt to apply PV for depression analysis. Besides, we propose a new visual descriptor - Histogram of Displacement Range (HDR) to characterize the displacement and velocity of the facial landmarks in the video segment. Experiments have been carried out on the Audio Visual Emotion Challenge (AVEC2016) depression dataset, they demonstrate that: 1) The proposed hybrid framework effectively improves the accuracies of both depression estimation and depression classification, with an average F1 measure up to 0.746, which is higher than the best result (0.724) of the depression sub-challenge of AVEC2016. 2) HDR obtains better depression recognition performance than Bag-of-Words (BoW) and Motion History Histogram (MHH) features. Le Yang 0009, Dongmei Jiang, Hichem Sahli |
IEEE Trans. Affect. Comput. | 3 |
| 2021 | Context-Aware Human Trajectories Prediction via Latent Variational ModelabstractUnderstanding human-contextual interaction to predict human trajectories is a challenging problem. Most of previous trajectory prediction approaches focused on modeling the human-human interaction located in a near neighborhood and neglected the influence of individuals which are farther in the scene as well as the scene layout. To alleviate these limitations, in this article we propose a model to address pedestrian trajectory prediction using a latent variable model aware of the human-contextual interaction. Our proposal relies on contextual information that influences the trajectory of pedestrians to encode human-contextual interaction. We model the uncertainty about future trajectories via latent variational model and captures relative interpersonal influences among all the subjects within the scene and their interaction with the scene layout to decode their trajectories. In extensive experiments, on publicly available datasets, it is shown that using contextual information and latent variational model, our trajectory prediction model achieves competitive results compared to state-of-the-art models. Abel Díaz Berenguer, Mitchel Alioscha-Pérez, Meshia Cédric Oveneke, Hichem Sahli |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Transformer Encoder With Multi-Modal Multi-Head Attention for Continuous Affect RecognitionabstractContinuous affect recognition is becoming an increasingly attractive research topic in affective computing. Previous works mainly focused on modelling the temporal dependency within a sensor modality, or adopting early or late fusion for multi-modal affective state recognition. However, early fusion suffers from the curse of dimensionality, and late fusion ignores the complementarity and redundancy between multiple modal streams. In this paper, we first introduce the transformer-encoder with a self-attention mechanism and propose a Convolutional Neural Network-Transformer Encoder (CNN-TE) framework to model the temporal dependency for single modal affect recognition. Further, to effectively consider the complementarity and redundancy between multiple streams we propose a Transformer Encoder with Multi-modal Multi-head Attention (TEMMA) for multi-modal affect recognition. TEMMA allows to progressively and simultaneously refine the inter-modality interactions and intra-modality temporal dependency. The learned multi-modal representations are fed to an Inference Sub-network with fully connected layers to estimate the affective state. The proposed framework is trained in a nutshell and demonstrates its effectiveness on the AVEC2016 and AVEC2019 datasets. Compared to state-of-the-art models, our approach obtains remarkable improvements on both arousal and valence in terms of concordance correlation coefficient (CCC) reaching 0.583 for arousal and 0.564 for valence on the AVEC2019 test set. Dongmei Jiang, Hichem Sahli |
IEEE Trans. Multim. | 3 |
| 2021 | Monocular 3D Facial Expression Features for Continuous Affect RecognitionabstractAutomated facial expression analysis from image sequences for continuous emotion recognition is a very challenging task due to the loss of the three-dimensional information during the image formation process. State-of-the-art relied on estimating dynamic textures features and convolutional neural network features to derive spatio-temporal features. Despite their great success, such features are insensitive to micro facial muscle deformations and are affected by identity, face pose, illumination variation, and self-occlusion. In this work, we argue that retrieving, from image sequences, 3D facial spatio-temporal information, which describes the natural facial muscle deformation, provides a semantical and efficient way of representation and is useful for emotion recognition. In this paper, we propose a framework for extracting three-dimensional facial spatio-temporal features from monocular image sequences using an extended 3D Morphable Model (3DMM) which disentangles the identity factor from the facial expressions of a specific person. An LSTM model is used to evaluate the effectiveness of the proposed spatio-temporal features on video-based facial expression recognition task and continuous affect recognition task. Experimental results, on the AFEW6.0 datasets for facial expression recognition, and the RECOLA and SEMAINE datasets for continuous emotion prediction, illustrate the potential of the proposed 3D spatio-temporal features for facial expressions analysis and continuous affect recognition, as well as their efficiency compared to recent state-of-the-art features. Ercheng Pei, Meshia Cédric Oveneke, Dongmei Jiang, Hichem Sahli |
IEEE Trans. Multim. | 5 |
| 2020 | An efficient model-level fusion approach for continuous affect recognition from audiovisual signals
Ercheng Pei, Dongmei Jiang, Hichem Sahli |
Neurocomputing | 3 |
| 2020 | SVRG-MKL: A Fast and Scalable Multiple Kernel Learning Solution for Features Combination in Multi-Class Classification ProblemsabstractIn this paper, we present a novel strategy to combine a set of compact descriptors to leverage an associated recognition task. We formulate the problem from a multiple kernel learning (MKL) perspective and solve it following a stochastic variance reduced gradient (SVRG) approach to address its scalability, currently an open issue. MKL models are ideal candidates to jointly learn the optimal combination of features along with its associated predictor. However, they are unable to scale beyond a dozen thousand of samples due to high computational and memory requirements, which severely limits their applicability. We propose SVRG-MKL, an MKL solution with inherent scalability properties that can optimally combine multiple descriptors involving millions of samples. Our solution takes place directly in the primal to avoid Gram matrices computation and memory allocation, whereas the optimization is performed with a proposed algorithm of linear complexity and hence computationally efficient. Our proposition builds upon recent progress in SVRG with the distinction that each kernel is treated differently during optimization, which results in a faster convergence than applying off-the-shelf SVRG into MKL. Extensive experimental validation conducted on several benchmarking data sets confirms a higher accuracy and a significant speedup of our solution. Our technique can be extended to other MKL problems, including visual search and transfer learning, as well as other formulations, such as group-sensitive (GMKL) and localized MKL (LMKL) in convex settings. Mitchel Alioscha-Pérez, Meshia Cédric Oveneke, Hichem Sahli |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | FACS3D-Net: 3D Convolution based Spatiotemporal Representation for Action Unit DetectionabstractMost approaches to automatic facial action unit (AU) detection consider only spatial information and ignore AU dynamics. For humans, dynamics improves AU perception. Is same true for algorithms? To make use of AU dynamics, recent work in automated AU detection has proposed a sequential spatiotemporal approach: Model spatial information using a 2D CNN and then model temporal information using LSTM (Long-Short-Term Memory). Inspired by the experience of human FACS coders, we hypothesized that combining spatial and temporal information simultaneously would yield more powerful AU detection. To achieve this, we propose FACS3D-Net that simultaneously integrates 3D and 2D CNN. Evaluation was on the Expanded BP4D+ database of 200 participants. FACS3D-Net outperformed both 2D CNN and 2D CNN-LSTM approaches. Visualizations of learnt representations suggest that FACS3D-Net is consistent with the spatiotemporal dynamics attended to by human FACS coders. To the best of our knowledge, this is the first work to apply 3D CNN to the problem of AU detection. Le Yang 0009, Itir Önal, Jeffrey F. Cohn, Zakia Hammal, Dongmei Jiang, Hichem Sahli |
ACII | 6 |
| 2019 | Continuous affect recognition with weakly supervised learning
Ercheng Pei, Dongmei Jiang, Mitchel Alioscha-Pérez, Hichem Sahli |
Multim. Tools Appl. | 4 |
| 2019 | A video prediction approach for animating single face image
Meshia Cédric Oveneke, Dongmei Jiang, Hichem Sahli |
Multim. Tools Appl. | 4 |
| 2019 | Automatic Depression Analysis Using Dynamic Facial Appearance Descriptor and Dirichlet Process Fisher EncodingabstractDepression causes mood disorders with noticeable problems in day-to-day activities. Current methods of assessing depression depend almost entirely on clinical interviews or questionnaires. They lack systematic and efficient ways of incorporating behavioral observations that are strong indicators of a psychological disorder. To help clinicians effectively and efficiently diagnose depression severity, automated systems, using objective and quantifiable data for depression assessment, are being developed. This paper presents a framework toward estimating a clinical depression-specific score, namely the Beck Depression Inventory-II (BDI-II) score, based on the analysis of facial expressions features. To extract facial dynamic features, we propose a novel dynamic feature descriptor denoted as median robust local binary patterns from three orthogonal planes (MRLBP-TOP), which can capture both the microstructure and macrostructure of facial appearance and dynamics. To aggregate the MRLBP-TOP over an image sequence, we propose a variant to the Fisher vector (FV) encoding scheme, denoted as the Dirichlet process FV (DPFV). DPFV adopts Dirichlet process Gaussian mixture models (DPGMM) to automatically learn the number of GMM mixtures and model parameters. Experimental results on the AVEC2013 and AVEC2014 depression databases have demonstrated the effectiveness of the proposed method. Dongmei Jiang, Hichem Sahli |
IEEE Trans. Multim. | 3 |
| 2018 | Improving Bag-of-Visual-Words model using visual n-grams for human action classification
Ruber Hernández-García, Julián Ramos Cózar, Nicolás Guil, Edel B. García Reyes, Hichem Sahli |
Expert Syst. Appl. | 5 |
| 2018 | Hierarchical sparse coding framework for speech emotion recognition
Diana Torres, Meshia Cédric Oveneke, Fengna Wang, Dongmei Jiang, Werner Verhelst, Hichem Sahli |
Speech Commun. | 6 |
| 2018 | Leveraging the Bayesian Filtering Paradigm for Vision-Based Facial Affective State EstimationabstractEstimating a person's affective state from facial information is an essential capability for social interaction. Automatizing such a capability has therefore increasingly driven multidisciplinary research for the past decades. At the heart of this issue are very challenging signal processing and artificial intelligence problems driven by the inherent complexity of human affect. We therefore propose a principled framework for designing automated systems capable of continuously estimating the human affective state from an incoming stream of images. First, we model human affect as a dynamical system and define the affective state in terms of valence, arousal and their higher-order derivatives. We then pose the affective state estimation problem as a Bayesian filtering problem and provide a solution based on Kalman filtering (KF) for probabilistic reasoning overtime, combined with multiple instance sparse Gaussian processes (MI-SGP) for inferring affect-related measurements from image sequences. We quantitatively and qualitatively evaluate our proposed framework on the AVEC 2012 and AVEC 2014 benchmark datasets and obtain state-of-the-art results using the baseline features as input to our MI-SGP-KF model. We therefore believe that leveraging the Bayesian filtering paradigm can pave the way for further enhancing the design of automated systems for affective state estimation. Meshia Cédric Oveneke, Isabel Gonzalez, Valentin Enescu, Dongmei Jiang, Hichem Sahli |
IEEE Trans. Affect. Comput. | 5 |
| 2017 | DCNN and DNN based multi-modal depression recognitionabstractIn this paper, we propose an audio visual multimodal depression recognition framework composed of deep convolutional neural network (DCNN) and deep neural network (DNN) models. For each modality, corresponding feature descriptors are input into a DCNN to learn high-level global features with compact dynamic information, which are then fed into a DNN to predict the PHQ-8 score. For multi-modal depression recognition, the predicted PHQ-8 scores from each modality are integrated in a DNN for the final prediction. In addition, we propose the Histogram of Displacement Range as a novel global visual descriptor to quantify the range and speed of the facial landmarks' displacements. Experiments have been carried out on the Distress Analysis Interview Corpus-Wizard of Oz (DAIC-WOZ) dataset for the Depression Sub-challenge of the Audio-Visual Emotion Challenge (AVEC 2016), results show that the proposed multi-modal depression recognition framework obtains very promising results on both the development set and test set, which outperforms the state-of-the-art results. Le Yang 0009, Dongmei Jiang, Wenjing Han, Hichem Sahli |
ACII | 4 |
| 2016 | A hybrid land-use mapping approach based on multi-scale spatial contextabstractMulti-scale spatial context which integrates spatial metrics and textural metrics is used to characterize land-use parcel and a hybrid land-use mapping approach is proposed in this paper. In terms of land-use characterization, the contributions of textural and spatial metrics are evaluated quantitatively. In terms of land-use categorization, a hybrid land-use classification scheme which combines Pairwise Decision Tree based Support Vector Machine (PDTSVM) and rule based decision tree is designed to classify parcels into construction, cultivated and uncultivated agricultural parcels. Experiment show that applying the presented technique can facilitate land-use mapping. Jingbo Chen, Hichem Sahli, Chengyi Wang 0001, Dong-xu He, Anzhi Yue |
IGARSS | 2 |
| 2016 | Improving unsupervised flood detection with spatio-temporal context on HJ-1B CCD dataabstractThe study of flood detection is significant to human life and social economy. In this paper, a completely unsupervised flood detection approach is presented, which combines spatio-temporal context and histogram thresholding. A global thresholding algorithm can be used in most of the cases to distinguish flood from non-flood pixels, but it may not distinguish local grey-level changes when the method is unsupervised. In this work, we introduce a kind of local context information to improve the results. A statistical model is used to establish the spatial relationships between each pixel and its surrounding regions, then a confidence map is computed. If the context structure changes significantly, the pixel is then considered potentially abnormal. Experimental investigations performed on HJ-1B CCD data from Northeast China during large-scale flooding in August 2013 showed higher precision of the proposed approach. Jiancheng Li, Hichem Sahli, Yu Meng 0002 |
IGARSS | 3 |
| 2016 | Vehicles detection using GF-2 imagery based on watershed image segmentationabstractRoad traffic volume monitoring plays an important role in transportation planning and spatial development, particularly in urban areas. The high-resolution satellite imagery provides a new data source to detect vehicles. Meanwhile, Satellite image covers large areas instantaneously, providing a possibility for snapshotting road traffic conditions. In this paper, we proposed an approach based on watershed image segmentation to detect the urban road vehicles from GF-2 imagery. The vehicles detection involves the two main steps: Firstly, a GIS road vector map and vegetation masks were applied to the image to guide vehicle detection by restricting the roads only. Secondly, watershed image segmentation was performed to separate bright and dark vehicles from the background in the road region. Then, a rule-based classifier was established to classify the image objects into the vehicle and the non-vehicle objects by using the spectral and shape feature information of image objects. Finally, the overall performance of the vehicle detection were compared with the manually counts, yielding overall accuracy of 81% with 93% classification accuracy. This detection accuracy may be considered acceptable for operational use in traffic monitoring. Yu Meng 0002, Hichem Sahli, Anzhi Yue, Jingbo Chen, Dong-xu He |
IGARSS | 3 |
| 2016 | Compressive Tracking based on Superpixel Segmentation
Ting Chen 0004, Hichem Sahli, Yanning Zhang 0001, Tao Yang 0006, Lingyan Ran |
MoMM | 2 |
| 2016 | Tracking with dynamic weighted compressive model
Ting Chen 0004, Yanning Zhang 0001, Tao Yang 0006, Hichem Sahli |
J. Vis. Commun. Image Represent. | 4 |
| 2016 | A Robust Actin Filaments Image Analysis FrameworkabstractThe cytoskeleton is a highly dynamical protein network that plays a central role in numerous cellular physiological processes, and is traditionally divided into three components according to its chemical composition, i.e. actin, tubulin and intermediate filament cytoskeletons. Understanding the cytoskeleton dynamics is of prime importance to unveil mechanisms involved in cell adaptation to any stress type. Fluorescence imaging of cytoskeleton structures allows analyzing the impact of mechanical stimulation in the cytoskeleton, but it also imposes additional challenges in the image processing stage, such as the presence of imaging-related artifacts and heavy blurring introduced by (high-throughput) automated scans. However, although there exists a considerable number of image-based analytical tools to address the image processing and analysis, most of them are unfit to cope with the aforementioned challenges. Filamentous structures in images can be considered as a piecewise composition of quasi-straight segments (at least in some finer or coarser scale). Based on this observation, we propose a three-steps actin filaments extraction methodology: (i) first the input image is decomposed into a 'cartoon' part corresponding to the filament structures in the image, and a noise/texture part, (ii) on the 'cartoon' image, we apply a multi-scale line detector coupled with a (iii) quasi-straight filaments merging algorithm for fiber extraction. The proposed robust actin filaments image analysis framework allows extracting individual filaments in the presence of noise, artifacts and heavy blurring. Moreover, it provides numerous parameters such as filaments orientation, position and length, useful for further analysis. Cell image decomposition is relatively under-exploited in biological images processing, and our study shows the benefits it provides when addressing such tasks. Experimental validation was conducted using publicly available datasets, and in osteoblasts grown in two different conditions: static (control) and fluid shear stress. The proposed methodology exhibited higher sensitivity values and similar accuracy compared to state-of-the-art methods. Mitchel Alioscha-Pérez, Carine Benadiba, Katty Goossens, Sandor Kasas, Giovanni Dietler, Ronnie Willaert, Hichem Sahli |
PLoS Comput. Biol. | 7 |
| 2016 | Towards long-term social child-robot interaction: using multi-activity switching to engage young usersabstractSocial robots have the potential to provide support in a number of practical domains, such as learning and behaviour change. This potential is particularly relevant for children, who have proven receptive to interactions with social robots. To reach learning and therapeutic goals, a number of issues need to be investigated, notably the design of an effective child-robot interaction (cHRI) to ensure the child remains engaged in the relationship and that educational goals are met. Typically, current cHRI research experiments focus on a single type of interaction activity (e.g. a game). However, these can suffer from a lack of adaptation to the child, or from an increasingly repetitive nature of the activity and interaction. In this paper, we motivate and propose a practicable solution to this issue: an adaptive robot able to switch between multiple activities within single interactions. We describe a system that embodies this idea, and present a case study in which diabetic children collaboratively learn with the robot about various aspects of managing their condition. We demonstrate the ability of our system to induce a varied interaction and show the potential of this approach both as an educational tool and as a research method for long-term cHRI. Miranda Coninx, Paul Baxter 0001, Elettra Oleari, Sara Bellini, Bert P. B. Bierman, Olivier A. Blanson Henkemans, Lola Cañamero, Piero Cosi, Valentin Enescu, Raquel Ros, Antoine Hiolle, Rémi Humbert, Bernd Kiefer, Ivana Kruijff-Korbayová, Rosemarijn Looije, Marco Mosconi, Mark A. Neerincx, Giulio Paci, Yorgos Patsis, Clara Pozzi, Francesca Sacchitelli, Hichem Sahli, Alberto Sanna, Giacomo Sommavilla, Fabio Tesser, Yiannis Demiris, Tony Belpaeme |
J. Hum. Robot Interact. | 22 |
| 2016 | Adaptive Real-Time Emotion Recognition from Body MovementsabstractWe propose a real-time system that continuously recognizes emotions from body movements. The combined low-level 3D postural features and high-level kinematic and geometrical features are fed to a Random Forests classifier through summarization (statistical values) or aggregation (bag of features). In order to improve the generalization capability and the robustness of the system, a novel semisupervised adaptive algorithm is built on top of the conventional Random Forests classifier. The MoCap UCLIC affective gesture database (labeled with four emotions) was used to train the Random Forests classifier, which led to an overall recognition rate of 78% using a 10-fold cross-validation. Subsequently, the trained classifier was used in a stream-based semisupervised Adaptive Random Forests method for continuous unlabeled Kinect data classification. The very low update cost of our adaptive classifier makes it highly suitable for data stream applications. Tests performed on the publicly available emotion datasets (body gestures and facial expressions) indicate that our new classifier outperforms existing algorithms for data streams in terms of accuracy and computational costs. Valentin Enescu, Hichem Sahli |
ACM Trans. Interact. Intell. Syst. | 3 |
| 2015 | Framework for combination aware AU intensity recognitionabstractWe present a framework for combination aware AU intensity recognition. It includes a feature extraction approach that can handle small head movements which does not require face alignment. A three layered structure is used for the AU classification. The first layer is dedicated to independent AU recognition, and the second layer incorporates AU combination knowledge. At a third layer, AU dynamics are handled based on variable duration semi-Markov model. The first two layers are modeled using extreme learning machines (ELMs). ELMs have equal performance to support vector machines but are computationally more efficient, and can handle multi-class classification directly. Moreover, they include feature selection via manifold regularization. We show that the proposed layered classification scheme can improve results by considering AU combinations as well as intensity recognition. Isabel Gonzalez, Werner Verhelst, Meshia Cédric Oveneke, Hichem Sahli, Dongmei Jiang |
ACII | 4 |
| 2015 | Multimodal depression recognition with dynamic visual and audio cuesabstractIn this paper, we present our system design for audio visual multi-modal depression recognition. To improve the estimation accuracy of the Beck Depression Inventory (BDI) score, besides the Low Level Descriptors (LLD) features and the Local Gabor Binary Pattern-Three Orthogonal Planes (LGBP-TOP) features provided by the 2014 Audio/Visual Emotion Challenge and Workshop (AVEC2014), we extract extra features to capture key behavioural changes associated with depression. From audio we extract the speaking rate, and from video, the head pose features, the Space-Temporal Interesting Point (STIP) features, and local kinematic features via the Divergence-Curl-Shear descriptors. These features describe body movements, and spatio-temporal changes within the image sequence. We also consider global dynamic features, obtained using motion history histogram (MHH), bag of words (BOW) features and vector of local aggregated descriptors (VLAD). To capture the complementary information within the used features, we evaluate two fusion systems - the feature fusion scheme, and the model fusion scheme via local linear regression (LLR). Experiments are carried out on the training set and development set of the Depression Recognition Sub-Challenge (DSC) of AVEC2014, we obtain root mean square error (RMSE) of 7.6697, and mean absolute error (MAE) of 6.1683 on the development set, which are better or comparable with the state of the art results of the AVEC2014 challenge. Dongmei Jiang, Hichem Sahli |
ACII | 3 |
| 2015 | Monocular 3D facial information retrieval for automated facial expression analysisabstractUnderstanding social signals is a very important aspect of human communication and interaction and has therefore attracted increased attention from various research areas. Among the different types of social signals, particular attention has been paid to facial expression of emotions and its automated analysis from image sequences. Automated facial expression analysis is a very challenging task due to the complex three-dimensional deformation and motion of the face associated to the facial expressions and the loss of 3D information during the image formation process. As a consequence, retrieving 3D spatio-temporal facial information from image sequences is essential for automated facial expression analysis. In this paper, we propose a framework for retrieving three-dimensional facial structure, motion and spatio-temporal features from monocular image sequences. First, we estimate monocular 3D scene flow by retrieving the facial structure using shape-from-shading (SFS) and combine it with 2D optical flow. Secondly, based on the retrieved structure and motion of the face, we extract spatio-temporal features for automated facial expression analysis. Experimental results illustrate the potential of the proposed 3D facial information retrieval framework for facial expression analysis, i.e. facial expression recognition and facial action-unit recognition on a benchmark dataset. This paves the way for future research on monocular 3D facial expression analysis. Meshia Cédric Oveneke, Isabel Gonzalez, Dongmei Jiang, Hichem Sahli |
ACII | 5 |
| 2015 | Multimodal dimensional affect recognition using deep bidirectional long short-term memory recurrent neural networksabstractIn this paper we propose the deep bidirectional long short-term memory recurrent neural network (DBLSTM-RNN) based single modal and multi-modal affect recognition frameworks. In the single modal framework DBLSTM with moving average (MA), audio or visual features are input into the DBLSTM-RNN model, whose output estimations of a dimension are smoothed by the moving average filter. After the smoothed estimations are expanded to the frame rate of the ground truth labels, another MA is adopted for smoothing the final results. In the multi-modal framework DBLSTM-DBLSTM-MA, the initial estimations from the audio and visual modalities via the first layer of DBLSTM-RNNs are input into a second layer of DBLSTM-RNN, whose outputs are smoothed by MA. The smoothed estimations are then expanded to the frame rate of the ground truth labels and smoothed again by another MA. Affect recognition experiments are carried out on the training set and development set of the AVEC2014 database, results show that the proposed DBLSTM-MA framework outperforms linear regression, support vector regression (SVR), and BLSTM for single modal dimension estimation. For audio visual multi-modal affect recognition, DBLSTM-DBLSTM-MA obtains better or comparable performance than the state of the art results in the competition of AVEC2014, with the average correlation coefficient (COR) reaches 0.599 on the Freeform database, 0.630 on the Northwind database, and 0.615 on the Freeform-Northwind database. Ercheng Pei, Le Yang 0009, Dongmei Jiang, Hichem Sahli |
ACII | 4 |
| 2015 | 3D emotional facial animation synthesis with factored conditional Restricted Boltzmann MachinesabstractThis paper presents a 3D emotional facial animation synthesis approach based on the Factored Conditional Restricted Boltzmann Machines (FCRBM). Facial Action Parameters (FAPs) extracted from 2D face image sequences, are adopted to train the FCRBM model parameters. Based on the trained model, given an emotion label sequence and several initial frames of FAPs, the corresponding FAP sequence is generated via the Gibbs sampling, and then used to construct the MPEG-4 compliant 3D facial animation. Emotion recognition and subjective evaluation on the synthesized animations show that the proposed method can obtain natural facial animations representing well the dynamic process of emotions. Besides, facial animation with smooth emotion transitions can be obtained by blending the emotion labels. Dongmei Jiang, Hichem Sahli |
ACII | 3 |
| 2015 | Multi-Object Tracking in Airborne Video Imagery based on Compressive Tracking Detection ResponsesabstractMulti-object tracking (MOT) in airborne video is a challenging problem due to the uncertain airborne vehicle motion as well as mounted camera vibrations. Most approaches addressing tracking in such type of scenario, use data association based on motion detection responses. Such approaches fail tracking objects with low speed or static ones. To alleviate the motion detection failures, in this paper we propose a multi-object tracking system based on combining motion-detection and Compressive Tracking detection responses. Ting Chen 0004, Hichem Sahli, Yanning Zhang 0001, Tao Yang 0006 |
MoMM | 2 |
| 2015 | Robust speaker localization for real-world robots
Georgios Athanasopoulos, Werner Verhelst, Hichem Sahli |
Comput. Speech Lang. | 3 |
| 2015 | Recognition of facial actions and their temporal segments based on duration models
Isabel Gonzalez, Francesco Cartella, Valentin Enescu, Hichem Sahli |
Multim. Tools Appl. | 4 |
| 2015 | Relevance units machine based dimensional and continuous speech emotion prediction
Fengna Wang, Hichem Sahli, Junbin Gao, Dongmei Jiang, Werner Verhelst |
Multim. Tools Appl. | 2 |
| 2015 | Gibberish speech as a tool for the study of affective expressiveness for robotic agents
Selma Yilmazyildiz, Werner Verhelst, Hichem Sahli |
Multim. Tools Appl. | 3 |
| 2014 | Behavioral accommodation towards a dance robot tutorabstractWe report first results on children adaptive behavior towards a dance tutoring robot. We can observe that children behavior rapidly evolves through few sessions in order to accommodate with the robotic tutor rhythm and instructions. Raquel Ros, Miranda Coninx, Yiannis Demiris, Yorgos Patsis, Valentin Enescu, Hichem Sahli |
HRI | 6 |
| 2014 | Object Segmentation Based on Contour-Skeleton DualityabstractThis paper presents a novel algorithm for performing integrated object segmentation from a single image. Unlike other state of the art methods which focus on either using contour-based or skeleton-based methods, our approach considers the duality of the two representations (contour/skeleton) and an iterative segmentation procedure that alternates between contour recovery and skeleton fitting. The contour recovery extracts the object contour by adopting the skeleton prior, while the skeleton fitting employs the contour to infer the optimal representation of the object shape. In our approach, the object contour can be directly recovered with no iteration if a detected skeleton is given. Although the proposed method is evaluated for human pose segmentation experiments, it can also be applied to other applications. Fengna Wang, Valentin Enescu, Hichem Sahli |
ICPR | 4 |
| 2014 | Object Tracking using Reformative Transductive Learning with Sample Variational CorrespondenceabstractTracking-by-learning strategies have effectively solved many challenging problems for visual tracking. When labeled samples are limited, the learning performance can be improved by exploiting unlabeled ones. Thus, a key issue for semi-supervised learning is the label assignment of the unlabeled samples, which is the principal focus of transductive learning. Unfortunately, the optimization scheme employed by the transductive learning is hard to be applied to online tracking because of its large amount of computation for sample labeling. In this paper, a reformative transductive learning was proposed with the variational correspondence between the learning samples, which are utilized to build an effective matching cost function for more efficient label assignment during the learning of representative separators. By using a weighted accumulative average to update the coefficients via a fixed budget of support vectors, the proposed tracking has been demonstrated to outperform most of the state-of-art trackers. Tao Zhuo, Peng Zhang 0005, Yanning Zhang 0001, Wei Huang 0013, Hichem Sahli |
ACM Multimedia | 5 |
| 2014 | Augmented Lagrangian-based approach for dense three-dimensional structure and motion estimation from binocular image sequencesabstractIn this study, the authors propose a framework for stereo–motion integration for dense depth estimation. They formulate the stereo–motion depth reconstruction problem into a constrained minimisation one. A sequential unconstrained minimisation technique, namely, the augmented Lagrange multiplier (ALM) method has been implemented to address the resulting constrained optimisation problem. ALM has been chosen because of its relative insensitivity to whether the initial design points for a pseudo‐objective function are feasible or not. The development of the method and results from solving the stereo–motion integration problem are presented. Although the authors work is not the only one adopting the ALMs framework in the computer vision context, to thier knowledge the presented algorithm is the first to use this mathematical framework in a context of stereo–motion integration. This study describes how the stereo–motion integration problem was cast in a mathematical context and solved using the presented ALM method. Results on benchmark and real visual input data show the validity of the approach. Geert De Cubber, Hichem Sahli |
IET Comput. Vis. | 2 |
| 2014 | A Segmentation Framework for phase Contrast and Fluorescence microscopy ImagesabstractThe noninvasive imaging of unstained living cells is a widely used technique in biotechnology for determining biological and biochemical role of proteins, since it allows studying living specimens without altering them. Usually, fluorescence and contrast (or transmission) images are both used complementarily, as their combination allows possible better outcomes. However, segmentation of contrast images is particularly difficult due to the presence of defocused scans, lighting/shade-off artifacts and cells overlapping. In this work, we investigate the optical properties intervening during the image formation process, and propose different segmentation strategies that can benefit from these properties. The proposed scheme (i) combines the estimated phase and the fluorescence information in order to obtain initial markers for a latter segmentation stage; and (ii) use the shear oriented polar snakes, an active contour model that implicitly involves phase information on its energy functional. The obtained contour can be used as region of interest estimation, as data for a latter shape-fitting process, or as smart markers for a more detailed segmentation process (i.e. watershed). Experimental results provide a comparison of the different segmentation schemes, and confirm the suitability of the proposed strategy and model for cell images segmentation. Mitchel Alioscha-Pérez, Ronnie Willaert, Hichem Sahli |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2014 | Speech driven photo realistic facial animation based on an articulatory DBN model and AAM features
Dongmei Jiang, Hichem Sahli, Yanning Zhang 0001 |
Multim. Tools Appl. | 3 |
| 2013 | Hybrid Deep Neural Network-Hidden Markov Model (DNN-HMM) Based Speech Emotion RecognitionabstractDeep Neural Network Hidden Markov Models, or DNN-HMMs, are recently very promising acoustic models achieving good speech recognition results over Gaussian mixture model based HMMs (GMM-HMMs). In this paper, for emotion recognition from speech, we investigate DNN-HMMs with restricted Boltzmann Machine (RBM) based unsupervised pre-training, and DNN-HMMs with discriminative pre-training. Emotion recognition experiments are carried out on these two models on the eNTERFACE'05 database and Berlin database, respectively, and results are compared with those from the GMM-HMMs, the shallow-NN-HMMs with two layers, as well as the Multi-layer Perceptrons HMMs (MLP-HMMs). Experimental results show that when the numbers of the hidden layers as well hidden units are properly set, the DNN could extend the labeling ability of GMM-HMM. Among all the models, the DNN-HMMs with discriminative pre-training obtain the best results. For example, for the eNTERFACE'05 database, the recognition accuracy improves 12.22% from the DNN-HMMs with unsupervised pre-training, 11.67% from the GMM-HMMs, 10.56% from the MLP-HMMs, and even 17.22% from the shallow-NN-HMMs, respectively. Dongmei Jiang, Yanning Zhang 0001, Fengna Wang, Isabel Gonzalez, Valentin Enescu, Hichem Sahli |
ACII | 8 |
| 2013 | Oriented Polar Snakes for Phase Contrast Cell Images Segmentation
Mitchel Alioscha-Pérez, Ronnie Willaert, Helene Tournu, Patrick Van Dijck, Hichem Sahli |
CIARP (2) | 5 |
| 2013 | Biologically Inspired Anomaly Detection in Pap-Smear Images
Maykel Orozco-Monteagudo, Alberto Taboada-Crispí, Hichem Sahli |
CIARP (2) | 3 |
| 2013 | Non-rigid target tracking based on 'flow-cut' in pair-wise frames with online hough forestsabstractIn conventional online learning based tracking studies, fixed-shape appearance modeling is often incorporated for training samples generation, as it is simple and convenient to be applied. However, for more general non-rigid and articulated object, this strategy may regard some background areas as foreground, which is likely to deteriorate the learning process. Recently published works utilize more than one patches to represent non-rigid object with foreground object segmentation, but most of these segmentation for target representation are performed only in single frame manner. Since the motion information between the consecutive frames was not considered by these approaches, when the backgrounds are similar to the target, accurate segmentation is hard to be achieved. In this work, we propose a novel model for non-rigid object segmentation by incorporating consecutive gradients flow between pair-wise frames into a Gibbs energy function. With help from motion information, the irregular target areas can be segmented more accurately during precise boundary convergence. The proposed segmentation model is incorporated into a semi-supervised online tracking framework for training samples generation. We test the proposed tracking on challenging videos involving heavy intrinsic variations and occlusions. As a result, the experiments demonstrate a significant improvement in tracking accuracy and robustness in comparison with other state-of-art tracking works. Yanning Zhang 0001, Peng Zhang 0005, Wei Huang 0013, Hichem Sahli |
ACM Multimedia | 5 |
| 2013 | Dynamic Compressive TrackingabstractReal-Time Compressive Tracking utilizes a very spare measurement matrix to extract the features for the appearance model. Such model performs well when the tracked objects are well defined. However, when the objects are low-grain, low-resolution, or small, a fixed size sparse measurement matrix is not sufficient enough to preserve the image structure of the object. In this work, we propose a Dynamic Compressive Tracking algorithm that employs adaptive random projections that preserve the image structure of the objects during tracking. The proposed tracker uses a dynamic importance ranking weight to evaluate the classification results obtained by each of the sparse measurement matrices and complete the tracking with the optimal sparse matrix. Extensive experimental results, on challenging publicly available data sets, shows that the proposed dynamic compressible tracking algorithm outperforms conventional compressive tracker. Ting Chen 0004, Yanning Zhang 0001, Tao Yang 0006, Hichem Sahli |
MoMM | 4 |
| 2013 | Evaluation of Attention Levels in a Tetris Game Using a Brain Computer Interface
Yorgos Patsis, Hichem Sahli, Werner Verhelst, Olga De Troyer |
UMAP | 2 |
| 2012 | Real-Time Dance Pattern Recognition Invariant to Anthropometric and Temporal Differences
Meshia Cédric Oveneke, Valentin Enescu, Hichem Sahli |
ACIVS | 3 |
| 2012 | Adaptive nonlinear probabilistic filter for Positron Emission TomographyabstractRadiologists face difficulties when reading and interpreting Positron Emission Tomography (PET) images because of the high noise level in the raw-projection data (i.e. the sinogram). The later may lead to erroneous diagnoses. Aiming at finding a suitable denoising technique for PET images, in our first work, we investigated filtering the sinogram with a constraint curvature motion filter where we computed the edge stopping function in terms of edge probability under a marginal prior on the noise free gradient. In this paper, we show that the Chi-square is the appropriate prior for finding the edge probability in the sinogram noise-free gradient. Since the sinogram noise is uncorrelated and follows a Poisson distribution, we then propose an adaptive probabilistic diffusivity function where the edge probability is computed at each pixel. We demonstrate quantitatively and qualitatively through simulations that the performance of the proposed method substantially surpasses that of state-of-art methods, both visually and in terms of statistical measures. Musa Alrefaya, Hichem Sahli |
BIBE | 2 |
| 2012 | Combined spatial point pattern analysis and remote sensing for assessing landmine affected areasabstractThis work extends our previous work on risk estimation of a mine contaminated area using mine action information such as minefield records and mine incidence, and feature indicators extracted from satellite images. Changes in temporal NDVI's and access to main transport routes and protective landscape such as forest edge are used to indicate potentially risk areas. All features are harmonized in an GIS environment. Kernel density function is performed on landmine and Earth Observation indicators. Safe areas are identified by land-in-use. A final risk map is produced by merging safe areas and risk estimation layer. Jonathan Cheung-Wai Chan, Aura Cecilia Alegria, Maria G. Veratelli, Marco Folegani, Hichem Sahli |
IGARSS | 5 |
| 2012 | Local linear spectral unmixing via cluster analysis and non-negative matrix factorization for hyperspectral (CHRIS/PROBA) imageryabstractWe present a novel approach for spectral unmixing in hyperspectral imagery called local linear spectral unmixing (LLSU). Our new proposal relies on an existing strategy for non-linear data modelling where a general non-linear model is approximated via several piecewise linear models. The algorithm is the result of hybridizing two well-known strategies for exploratory high-dimensional data analysis: cluster analysis and non-negative matrix factorization. It has been proposed to answer the limitations of global linear unmixing models which are widely used for spectral unmixing. Our strategy is to first group similar pixels from the hyperspctral image via cluster analysis and then to consider a local linear unmixing model for each cluster in order to obtain the constituent endmembers. Subsequently, the resulted local endmembers from each cluster of mixed pixels are at their turn clustered based on their spectral similarity; the final solution is given by the clusters' centroids. Cosmin Lazar, Luca Demarchi, David Steenhoff, Jonathan Cheung-Wai Chan, Ann Nowé, Hichem Sahli |
IGARSS | 6 |
| 2011 | Kalman Filter-Based Facial Emotional Expression Recognition
Isabel Gonzalez, Valentin Enescu, Hichem Sahli, Dongmei Jiang |
ACII (1) | 4 |
| 2011 | Context-Independent Facial Action Unit Recognition Using Shape and Gabor Phase Information
Isabel Gonzalez, Hichem Sahli, Valentin Enescu, Werner Verhelst |
ACII (1) | 2 |
| 2011 | Audio Visual Emotion Recognition Based on Triple-Stream Dynamic Bayesian Network Models
Dongmei Jiang, Yulu Cui, Isabel Gonzalez, Hichem Sahli |
ACII (1) | 6 |
| 2011 | Relevance Vector Machine Based Speech Emotion Recognition
Fengna Wang, Werner Verhelst, Hichem Sahli |
ACII (2) | 3 |
| 2011 | Automatic Occlusion Removal from Facades for 3D Urban Reconstruction
Chris Engels, David Tingdahl, Mathias Vercruysse, Tinne Tuytelaars, Hichem Sahli, Luc Van Gool |
ACIVS | 5 |
| 2011 | Soil thickness zonation approach using Landsat ETM+, geological maps, DEM data, and field investigationabstractSoil thickness is extensively concerned by engineering geology and slope hazards prevention. But it is not available from geological maps and traditional methods to obtain soil thickness information are costly and time consuming, such as drilling and geophysics methods. In this paper, we proposed a method for regional soil thickness zonation based on multi-sources data. Three essential parameters were chosen for soil thickness analysis, include slope, lithology and NDVI which extracted from DEM data, geological maps and Landsat image respectively. In combining the parameters with field survey data, Support Vector Machine was employed for soil thickness classification. Experiment results in the Three Gorges region show the usefulness of the approach. Runqing Ye, Ruiqing Niu, Qiying Jiang, Hichem Sahli |
IGARSS | 4 |
| 2011 | Smooth adaptive fitting of 3D face model for the estimation of rigid and nonrigid facial motion in video sequences
Yunshu Hou, Ilse Ravyse, Valentin Enescu, Hichem Sahli |
Signal Process. Image Commun. | 5 |
| 2010 | A Quality Analysis on JPEG 2000 Compressed Leukocyte Images by Means of Segmentation Algorithms
Alexander Falcón-Ruiz, Juan Paz-Viera, Alberto Taboada-Crispí, Hichem Sahli |
CIARP | 4 |
| 2010 | Hardware and software architecture for AUV based on low-cost sensorsabstractThe use of Autonomous Underwater Vehicles (AUV) as robots for exploration and oceanology science has been a field of interest of several universities and research center's around the world in the last decade. Cuba being a country surrounded by the Caribbean Sea, having most of it's resources in it. Researchers from the Central University of Las Villas (UCLV) and the Hydrographic Research Center (HRC) have joined forces in the development of the HRC-AUV project, whit the objective of the implementation of a AUV capable to operate in the Cuban sea for exploration and supervision. The platform is based on low cost sensors and hardware. In the present document the hardware and software architectures of the AUV are presented. Also a set of experimental results are presented to validate the system. Alain Sebastian Martinez, Yidier Rodriguez, Luis Hernández Santana, Hichem Sahli |
ICARCV | 5 |
| 2010 | Realistic mouth animation based on an articulatory DBN model with constrained asynchronyabstractIn this paper, we propose an approach to convert acoustic speech to video realistic mouth animation based on an articulatory dynamic Bayesian network model with constrained asynchrony (AF_AVDBN). Conditional probability distributions are defined to control the asynchronies between the articulators such as lips, tongue and glottis/velum. An EM-based conversion algorithm is also presented to learn the optimal visual features given an auditory input and the trained AF_AVDBN parameters. In the training of the AF_AVDBN models, downsampled YUV spatial frequency features of the interpolated mouth image sequences are extracted as visual features. For reproducing the mouth animation sequence, from the learned visual features, a spatial upsampling and a temporal downsampling are applied. Both qualitative and quantitative results show that the proposed method is capable of producing more natural and realistic mouth animations, and the accuracy is further improved compared to the state of the art multi-stream Hidden Markov Model (MSHMM) and articulatory DBN model without asynchrony constraint (AF_DBN). Dongmei Jiang, Ilse Ravyse, Peizhen Liu, Hichem Sahli, Werner Verhelst |
ICASSP | 4 |
| 2010 | Multi scale representation for remotely sensed images using fast anisotropic diffusion filteringabstractObject based image analysis has gained on the traditional per-pixel multi-spectral based approaches. The main pitfall of using anisotropic diffusion for creating a multi scale representation of a remotely sensed image remains the computational burden. Producing the coarser scales in a multi scale representation or, diffusing spatially large images involves significant time and resources. This paper proposes a fast approach for anisotropic diffusion that overcomes spatial size limitations by distributing the diffusion as individual sub-processes over several overlapping sub-images. The overlap areas are synchronized at specific diffusion time ensuring that the fast approximation does not deviate too much from its single process equivalent. This demonstrated for an image, which can be diffused using a traditional sequential approach. In addition, experimental data for very large images that can not efficiently be processed using a sequential approach is illustrated. Iris Vanhamel, Musa Alrefaya, Hichem Sahli |
IGARSS | 3 |
| 2009 | Perception-Based Lighting Adjustment of Image Sequences
Xiaoyue Jiang, Ilse Ravyse, Hichem Sahli, Jianguo Huang, Rongchun Zhao, Yanning Zhang 0001 |
ACCV (3) | 4 |
| 2009 | 3D Face Alignment via Cascade 2D Shape Alignment and Constrained Structure from Motion
Yunshu Hou, Ilse Ravyse, Hichem Sahli |
ACIVS | 4 |
| 2009 | Audio-Visual Emotion Recognition Based on a DBN Model with Constrained AsynchronyabstractThis paper presents an audio visual multi-stream DBN model (Asy_DBN) for emotion recognition with constraint asynchrony, in which audio state and visual state transit individually in their corresponding stream but the transition is constrained by the allowed maximum audio visual asynchrony. Emotion recognition experiments of Asy_DBN with different asynchrony constraints are carried out on an audio visual speech database of four emotions, and compared with the single stream HMM, state synchronous HMM (Syn_HMM) and state synchronous DBN model, as well the state asynchronous DBN model without asynchrony constraint. Results show that by setting the appropriate maximum asynchrony constraint between audio and visual streams, the proposed audio visual asynchronous DBN model gets the highest emotion recognition performance, with an improvement of 15% over Syn_HMM. Danqi Chen 0001, Dongmei Jiang, Ilse Ravyse, Hichem Sahli |
ICIG | 4 |
| 2009 | Rule-Based Video Interpretation Framework: Application to Automated SurveillanceabstractVideo applications usually involve a large number of moving objects. Moving objects refer to semantic realworld entities denoting a coherent spatial region and being automatically computed by the continuity of spatial and temporal low-level features, such as color and motion. In surveillance application, spatial and temporal relationships among these objects should be efficiently supported and retrieved for abnormal event recognition. In this paper we emphasize on analyzing and interpreting video object motions for advanced surveillance application. We propose a rule-based system allowing moving objects description at the low level (spatio-temporal features and relationships) as well as at the semantic level (actions, events and interaction). The result of the proposed system is illustrated to detect context dependent events in surveillance. Thomas Geerinck, Valentin Enescu, Ilse Ravyse, Hichem Sahli |
ICIG | 4 |
| 2009 | A Visual Silence Detector Constraining Speech Source SeparationabstractWe propose an audiovisual source separation algorithm for speech signals. In our proposed algorithm we first extract the time segments with low activity of the mouth region from synchronous video recordings. An automatically selected optimal classifier is used to detect silent intervals in these instants of low visual mouth activity. Then, the source separation problem is formulated and solved for the entire signal duration. Our approach was tested on two challenging speech corpora with two speakers and two microphones, namely in the first corpus separate source signals were mixed in a simulated room, and the second corpus contains recorded conversations. The results are promising on both corpora: with the visual silence detector the performance of the source separation algorithm, measured by the signal to noise inference ratio increases. Isabel Gonzalez, Ilse Ravyse, Henk Brouckxon, Werner Verhelst, Dongmei Jiang, Hichem Sahli |
ICIG | 6 |
| 2009 | Smooth Adaptive Fitting of 3D Face Model for the Estimation of Rigid and Non-rigid Facial Motion in Video SequencesabstractIn this paper, we propose a two-level integrated model for accurate 3D tracking of rigid head motion and non-rigid facial animation. At the lowest level, the 2D shape of facial features is robustly extracted using a regularized shape model and a cascade multi-stage algorithm. At the highest level, we estimate both the facial animation and 3D pose parameters via minimizing an energy function comprising three terms. The first quantifies the error in matching a 3D wireframe face model (Candide) to the image sequence, while the remaining terms impose temporal and spatial motion-smoothness constraints over the 3D model points. Through extensive experiments, we demonstrate the feasibility and effectiveness of the proposed method and also show it outperforms the state-of-the-art algorithms such as the Bayes Tangent Shape Model. Yunshu Hou, Ilse Ravyse, Valentin Enescu, Rongchun Zhao, Hichem Sahli |
ICIG | 6 |
| 2009 | Video Realistic Mouth Animation Based on an Audio Visual DBN Model with Articulatory Features and Constrained AsynchronyabstractThis paper presents a mouth animation construction method based on the DBN models with articulatory features (AF_AVDBN), in which the articulatory features of lips, tongue, glottis/velum can be asynchronous within a maximum asynchrony constraint to describe the speech production process more reasonably. Given an audio input and the trained AF_AVDBN models, the optimal visual feature learning algorithm is deduced based on the Maximum Likelihood Estimation criterion. The learned visual features are then used to construct the mouth images for the input speech. Objective and subjective evaluations on the mouth animations of 110 speech sentences show that the learned visual features from the AF_AVDBN models track the real visual features very closely, and the constructed mouth images from the AF_AVDBN models are very much like the real ones. Dongmei Jiang, Peizhen Liu, Ilse Ravyse, Hichem Sahli, Werner Verhelst |
ICIG | 4 |
| 2009 | Manifold Analysis for Subject Independent Dynamic Emotion Recognition in Video SequencesabstractThis paper proposes subject independent manifold features for dynamic emotion recognition. Facial action features, based on FACs, are firstly embedded into a low-dimensional manifold space using the ISOMAP algorithm, then the manifold features from different subjects are aligned into a global coordinate space by the supervised ISOMAP algorithm for recognition. To validate and evaluate the proposed manifold representation for emotion recognition, experiments with GMMs are presented. Given a new expression sequence, and tracked facial features, we are able to pin-point the actual occurrence of specific expressions, while characterizing its intensity by considering different expression temporal transition characteristics. Finally, experimental results show that our approach is able to separate different expressions successfully. Dongmei Jiang, Fengna Wang, Ilse Ravyse, Hichem Sahli |
ICIG | 5 |
| 2009 | Multispectral Data Classification based on Spectral Indices and Cascaded Fuzzy C-mean ClassifiersabstractLand Use and Land Cover (LULC) are characterized by a large variety of spectrally distinct LULC classes. The diagnostic and evaluation of the spectral separability measure yields the potential for automated identification and mapping of these classes. This study proposes a new cascaded fuzzy C-mean classification method for a rough classification of (LULC) classes. The method is based on the use of spectral indices as innovative features to provide a coarse classification of remotely sensed data. The robustness and accuracy of the defined classification schema is performed based on the computation of confusion matrices and Kappa coefficient. The Kappa statistic ranges from 0.90 to 0.98 for the set of the evaluated images which infer the good accuracy of the new rough classification scheme. Mohamed Jabloun, Cosmin Mihai, Iris Vanhamel, Thomas Geerinck, Hichem Sahli |
IGARSS (3) | 5 |
| 2009 | Scale Selection for Compact Scale-Space Representation of Vector-Valued Images
Iris Vanhamel, Cosmin Mihai, Hichem Sahli, Antonis Katartzis, Ioannis Pratikakis |
Int. J. Comput. Vis. | 3 |
| 2008 | Simple solution for visual servoing of camera-in-hand robots in the 3d Cartesian spaceabstractIn this paper, the control problem, in a 3D cartesian space, of camera-in-hand robotic systems is considered. In this approach, a camera is mounted on the robot, at the hand, which provides an image of objects located in the robot environment. The aim of this approach is to move the robot arm in such a way that the image of the considered object, a sphere in our case, coincides with the center of the image, and its apparent radius is constant. We propose a simple image-based direct visual servo controller which requires the knowledge of the objects' radius, but it does not need to use the inverse Jacobian matrix. The robotic and vision system are modelled for small variations around the operating point for position control. In these conditions, the stability of the whole system is balanced and the characteristics for steady-state response for object trajectory are obtained. Both simulation and experimental results, using an ASEA-IRB6 robot manipulator, are presented to illustrate the proposed controller's performance. Luis Hernández Santana, Rene Gonzalez, Hichem Sahli, Yoani Guerra |
ICARCV | 3 |
| 2008 | A biomechanical model for image-based estimation of 3D face deformationsabstractTo reproduce a face motion from an image sequence, natural motion parameters provide a semantical and efficient way of representation. State-of-the art techniques for 3D face motion estimation employ a limited set of predefined key-shapes of face structure, and thereby restrict the possible face motion which can cause distortions. We propose a new approach in which such distortions are avoided by augmenting a 3D structural surface face model with a physical motion model originating from continuum mechanics. Implementation with a displacement-based FEM does not only describe, but also explain the motion of the face's skin tissue caused by the muscle force parameterized actuations. The correct usage of the motion model and the mapping of the 3D scene flow to the 2D optical flow, allow posing the 3D deformation estimation as an inverse problem, for which a solution has been obtained using numerical solvers. Ilse Ravyse, Hichem Sahli |
ICASSP | 2 |
| 2008 | Accurate visual speech synthesis based on diviseme unit selection and concatenationabstractThis paper presents a novel speech driven accurate realistic visual speech synthesis approach. Firstly, an audio visual instance database is built for different viseme context combinations, i.e. diviseme units, using 100 audio visual speech sentences of a female speaker. Then a diviseme instance selection algorithm is introduced to choose the optimal diviseme instances for the viseme contexts in the input speech, considering both the concatenation smoothness of the image sequences, and matching of the mouth movements to the acoustic pronunciation process, as well the intensity of the input speech. Finally mouth image sequences of corresponding viseme segments in the selected diviseme instances are time warped and blended to construct the mouth images of the final animation. Visual speech synthesis experiments and subjective evaluation results show that mouth animations can be obtained which are not only realistic with clear and smooth mouth images, but also in good accordance with the acoustic pronunciation and intensity of the input speech. Dongmei Jiang, Ilse Ravyse, Hichem Sahli, Yanning Zhang 0001 |
MMSP | 3 |
| 2008 | A Stochastic Framework for the Identification of Building Rooftops Using a Single Remote Sensing ImageabstractThe identification of building rooftops from a single image, without the use of auxiliary 3-D information like stereo pairs or digital elevation models, is a very challenging and difficult task in the area of remote sensing. The existing methodologies rarely tackle the problem of 3-D object identification, like buildings, from a purely stochastic viewpoint. Our approach is based on a stochastic image interpretation model, which combines both 2-D and 3-D contextual information of the imaged scene. Building rooftop hypotheses are extracted using a contour-based grouping hierarchy that emanates from the principles of perceptual organization. We use a Markov random field model to describe the dependencies between all available hypotheses with regard to a globally consistent interpretation. The hypothesis verification step is treated as a stochastic optimization process that operates on the whole grouping hierarchy to find the globally optimal configuration for the locally interacting grouping hypotheses, providing also an estimate of the height of each extracted rooftop. This paper describes the main principles of our method and presents building detection results on a set of synthetic and airborne images. Antonis Katartzis, Hichem Sahli |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2008 | A Nonlinear Iterative Reconstruction and Analysis Approach to Shape-Based Approximate Electromagnetic TomographyabstractA nonlinear Helmholtz-equation-modeled electromagnetic tomographic reconstruction problem is solved for the object boundary and inhomogeneity parameters in a damped Tikhonov-regularized Gauss-Newton (DTRGN) solution framework. In this paper, the object is represented in a suitable global basis, whereas the boundary is expressed as the zero level set of a signed-distance function. For an explicit parameterized boundary-representation-based reconstruction scheme, analytical Jacobian and Hessian calculations are made to express the changes in scattered field values w.r.t. changes in the inhomogeneity parameters and the control points in a spline representation of the object boundary, via the use of a level-set representation of the object. Even though, in this paper, a homogeneous dielectric is considered and a spline representation has been used to represent the boundary, the formulation can be used for a general global basis representation of the inhomogeneity as well as arbitrary parameterizations of the boundary, and is generalizable to three dimensions. Reconstruction results are presented for test cases of landminelike dielectric objects embedded in the ground under noisy data conditions. To confirm convergence and, at times, to know which of the obtained iterates are closer to the actual unknown solution, using a perturbation theory framework, a local (Hessian-based) convergence analysis is applied to the DTRGN scheme for the reconstruction, yielding estimates of convergence rates in the residual and parameter spaces. Naren Naik, Jerry Eriksson, Pieter de Groen, Hichem Sahli |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2008 | Infrared Thermography for Buried Landmine Detection: Inverse Problem SettingabstractThis paper deals with an inverse problem arising in infrared (IR) thermography for buried landmine detection. It is aimed at using a thermal model and measured IR images to detect the presence of buried objects and characterize them in terms of thermal and geometrical properties. The inverse problem is mathematically stated as an optimization one using the well-known least-square approach. The main difficulty in solving this problem comes from the fact that it is severely ill posed due to lack of information in measured data. A two-step algorithm is proposed for solving it. The performance of the algorithm is illustrated using some simulated and real experimental data. The sensitivity of the proposed algorithm to various factors is analyzed. A data processing chain including anomaly detection and characterization is also introduced and discussed. Nguyen Trung Thành, Hichem Sahli, Dinh Nho Hào |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2007 | Visual Tracking by Hypothesis Testing
Valentin Enescu, Ilse Ravyse, Hichem Sahli |
ACIVS | 3 |
| 2007 | Robust Shape-Based Head Tracking
Yunshu Hou, Hichem Sahli, Ilse Ravyse, Yanning Zhang 0001, Rongchun Zhao |
ACIVS | 2 |
| 2007 | Graph Cuts Approach to MRF Based Linear Feature Extraction in Satellite Images
Anesto del-Toro-Almenares, Cosmin Mihai, Iris Vanhamel, Hichem Sahli |
CIARP | 4 |
| 2007 | Investigation of Time-Frequency Features for GPR Landmine DiscriminationabstractGround-penetrating radar (GPR) is capable to detect plastic antipersonnel landmines as well as other subsurface targets. In order to reduce false alarms, an option of automatic landmine discrimination from neutral minelike targets would be very useful. This paper presents a possibility for such discrimination and analyzes it experimentally. The authors investigate time-frequency features of an ultrawideband (UWB) target response for the discrimination between buried landmines and other objects. The discrimination method includes the extraction of an early-time target impulse response, its time-frequency transformation, and the extraction of time-frequency features based on a singular value decomposition of the transformed image. In order to take into account the changes in the UWB target signals, the experimental conditions are completely controlled to focus on the behavior of the target's response with respect to its depth and the horizontal position of the GPR above it. The dependence of the features on the GPR bandwidth is analyzed as well. The Mahalanobis distance is used as a criterion for optimal discrimination. The obtained results define the best features and conditions when the landmine discrimination is successful. For comparison, the discriminant power of the proposed features has been tested on a dataset, acquired during a field campaign in Angola Timofey Grigorievich Savelyev, Luc Van Kempen, Hichem Sahli, Jürgen Sachs, Motoyuki Sato |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2007 | Finite-Difference Methods and Validity of a Thermal Model for Landmine Detection With Soil Property EstimationabstractIn this paper, we introduce and validate a 3-D linear thermal model for landmine detection. A finite-difference approximation of generalized solutions to the model is proposed, and its convergence properties are proved. An efficient numerical algorithm based on splitting methods is suggested for solving the discretized problem. Moreover, we introduce methods to estimate the (bare) soil and air-soil interface thermal properties. These parameters depend strongly on weather, environmental conditions, and soil type; their accuracy affects strongly the thermal modeling. The validity of the thermal model with the estimated soil properties is verified by comparing the simulations with data sets acquired in outdoor minefields Nguyen Trung Thành, Hichem Sahli, Dinh Nho Hào |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2006 | Facial Analysis and Synthesis Scheme
Ilse Ravyse, Hichem Sahli |
ACIVS | 2 |
| 2006 | MRF-Based Foreground Detection in Image Sequences from a Moving CameraabstractThis paper presents a Bayesian approach for simultaneously detecting the moving objects (foregrounds) and estimating their motion in image sequences taken with a moving camera mounted on the top of a mobile robot. To model the background, the algorithm uses the GMM approach [1] for its simplicity and capability to adapt to illumination changes and small motions in the scene. To over-come come the limitations of the GMM approach with its pixel-wise processing, the background model is combined with the motion cue in a maximum a posteriori probability (MAP)-MRF framework. This enables us to exploit the advantages of spatio-temporal dependencies that moving objects impose on pixels and the interdependence of motion and segmentation fields. As a result, the detected moving objects have visually attractive silhouettes and they are more accurate and less affected by noise than those obtained with simple pixel-wise methods. To enhance the segmentation accuracy, the background model is re-updated using the MAP-MRF results. Experimental results and a qualitative study of the proposed approach are presented on image sequences with a static camera as well as with a moving camera. Sid Ahmed Berrabah, Geert De Cubber, Valentin Enescu, Hichem Sahli |
ICIP | 4 |
| 2006 | Multiscale Graph Theory Based Color SegmentationabstractIn this paper, image segmentation is addressed within the framework of nonlinear multiscale watersheds in combination with graph theory. First, a graph is created which decomposes the image in scale and space using the concept of multiscale watersheds. In the subsequent step the obtained graph is partitioned using recursive graph cuts in a coarse to fine manner. In this way, we combine scale and feature measures in a flexible way. The dissimilarity between graph-nodes is estimated by using the Earth Mover's Distance on a featureset that combines color, scale and contrast. Experimental results demonstrate the efficiency of the proposed method for natural scene images. Iris Vanhamel, Ioannis Pratikakis, Hichem Sahli |
ICIP | 3 |
| 2006 | A framework for integrating MPEG-7 knowledge templates into video surveillance applicationsabstractIn this paper we propose a method of representing knowledge structures for visual content categorization and annotation in video surveillance applications. In order to achieve this purpose, we designed and implemented an ontology database for storing templates of domain-dependent moving objects and scene regions descriptions. The novelty of the system consists in the employment of a double layer, visual and semantic knowledge representation at the beginning of the video processing chain. As in similar approaches, bottom-up low-level feature extraction mechanisms are used, followed by top-down object classification using hypothesis proposal extracted from the ontology database. Instead of using top-down inference after the low-level processing steps, the two processes are entangled, such that the feature extraction algorithms are driven by logical and scenario-bound object and scene regions relations. The resulting content description is converted into MPEG-7 compliant representation and archived for later retrieval purposes Mike Barais, Tom Caljon, Valentin Enescu, Hichem Sahli |
MMSP | 4 |
| 2006 | A Bayesian formulation of edge-stopping functions in nonlinear diffusionabstractWe propose a novel, Bayesian formulation of the edge-stopping (diffusivity) function in a nonlinear diffusion scheme in terms of edge probability under a marginal prior on noise-free gradient. This formulation differs from the existing probabilistic diffusion approaches that give stochastic formulations for the conductivity but not for the diffusivity function of the gradient. In particular, we impose a Laplacian prior for the ideal gradient, but the proposed formulation is general and can be used with other marginal distributions. We also make links to related works that treat correspondences between nonlinear diffusion and wavelet shrinkage. Aleksandra Pizurica, Iris Vanhamel, Hichem Sahli, Wilfried Philips, Antonis Katartzis |
IEEE Signal Process. Lett. | 3 |
| 2005 | An Offline Bidirectional Tracking Scheme
Tom Caljon, Valentin Enescu, Peter Schelkens, Hichem Sahli |
ACIVS | 4 |
| 2005 | Active stereo vision-based mobile robot navigation for person tracking
Valentin Enescu, Geert De Cubber, Kenny Cauwerts, Sid Ahmed Berrabah, Hichem Sahli, Marnix Nuttin |
ICINCO | 5 |
| 2005 | Tele-robots with shared autonomy: tele-presence for high level operability
Thomas Geerinck, Valentin Enescu, Ioan Alexandru Salomie, Sid Ahmed Berrabah, Kenny Cauwerts, Hichem Sahli |
ICINCO | 6 |
| 2005 | Kernel-based head tracker for videophonyabstractAn approach for automatically segmenting and tracking a face in a sequence of color images is presented. The face detection in the initial image frame consists of a two-step process: the face candidates selection, using skin color clustering, and the face verification, yielding the best face candidate based on shape and color cues. The tracking of the head in the subsequent frames is performed via a kernel-based method wherein a joint spatial-color probability density characterizes the head region. In this context, the novelty of our tracking approach lies in the introduction of two parametric models: a geometric transformation enabling the rotation, scaling, and translation of the target, and an affine illumination change model. The parameters of these models are estimated by minimizing the similarity between the predicted and the current head appearance. The proposed algorithms achieve reliable detection and tracking results. Ilse Ravyse, Valentin Enescu, Hichem Sahli |
ICIP (3) | 3 |
| 2005 | A nonlinear multigrid diffusion model for efficient dense optical flow estimationabstractThis paper presents a nonlinear multigrid isotropic diffusion model to estimate 2D dense motion. Multigrid provides an efficient multi-level nonlinear relaxation method that accelerates the evolution process of the nonlinear isotropic diffusion model. The resulting system of the two coupled nonlinear isotropic diffusion equations, for the two components of the optical flow, is successively transferred to coarser levels, and coarse-grid correction scheme is used to improve the performance. At each level, the Gauss-Seidel Newton method is used for nonlinear relaxation. Experimental results on both synthetic and real image sequences are reported to demonstrate the efficiency and accuracy of the proposed model. Hichem Sahli |
ICIP (1) | 2 |
| 2005 | A hierarchical Markovian model for multiscale region-based classification of vector-valued imagesabstractWe propose a new classification method for vector-valued images, based on: 1) a causal Markovian model, defined on the hierarchy of a multiscale region adjacency tree (MRAT), and 2) a set of nonparametric dissimilarity measures that express the data likelihoods. The image classification is treated as a hierarchical labeling of the MRAT, using a finite set of interpretation labels (e.g., land cover classes). This is accomplished via a noniterative estimation of the modes of posterior marginals (MPM), inspired from existing approaches for Bayesian inference on the quadtree. The paper describes the main principles of our method and illustrates classification results on a set of artificial and remote sensing images, together with qualitative and quantitative comparisons with a variety of pixel-based techniques that follow the Bayesian-Markovian framework either on hierarchical structures or the original image lattice. Antonis Katartzis, Iris Vanhamel, Hichem Sahli |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2004 | Improved thermal analysis of buried landminesabstractIn this paper, we address the problem of the detection and identification of surface-laid and shallowly buried landmines from measured infrared images. A three-dimensional thermal model has been developed to study the effect of the presence of landmines in the thermal signature of the bare soil. Based on this model, a target identification procedure is proposed aiming at detecting and classifying the anomalies found on the soil thermal signature. In our approach, landmines are thought of as a thermal barrier in the natural flow of the heat inside the soil, which produces a perturbation of the expected thermal pattern on the surface. The detection of these perturbations will put into evidence the presence of potential mine targets. We propose an iterative procedure to classify the detected perturbations as mines or nonmines and to estimate their depth of burial. This paper describes the main principles of our method and illustrates classification results on a set of acquired images. Qualitative and quantitative comparisons with independent component analysis are also given. Paula López Martinez 0001, Luc Van Kempen, Hichem Sahli, Diego Cabello |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2003 | Hierarchical segmentation via a diffusion scheme in color/texture feature spaceabstractThis paper presents a segmentation scheme for images containing both smooth regions and textures. It is based on a vector-valued anisotropic diffusion on a combined color/Gabor feature space, followed by a hierarchical segmentation using dynamics of multiscale 'generalized gradient' watersheds. The proposed method gives good segmentation results and is shown to be more effective than its counterpart, which uses only multiscale color information. Iris Vanhamel, Antonis Katartzis, Hichem Sahli |
ICIP (1) | 3 |
| 2003 | Multiscale gradient watersheds of color imagesabstractWe present a new framework for the hierarchical segmentation of color images. The proposed scheme comprises a nonlinear scale-space with vector-valued gradient watersheds. Our aim is to produce a meaningful hierarchy among the objects in the image using three image components of distinct perceptual significance for a human observer, namely strong edges, smooth segments and detailed segments. The scale-space is based on a vector-valued diffusion that uses the Additive Operator Splitting numerical scheme. Furthermore, we introduce the principle of the dynamics of contours in scale-space that combines scale and contrast information. The performance of the proposed segmentation scheme is presented via experimental results obtained with a wide range of images including natural and artificial scenes. Iris Vanhamel, Ioannis Pratikakis, Hichem Sahli |
IEEE Trans. Image Process. | 3 |
| 2002 | A Multi-stage Online Signature Verification System
Edgard Nyssen, Hichem Sahli |
Pattern Anal. Appl. | 2 |
| 2001 | Scale space segmentation of color images using watersheds and fuzzy region mergingabstractA multi-resolution segmentation approach for color images is proposed. The scale space is generated using the Perona-Malik diffusion approach and the watershed algorithm is employed to produce the regions in each scale. The dynamics of contours and the relative entropy of color region distribution are estimated as region dissimilarity features across the scale-space stack, and combined using a fuzzy rule based system. A minima-linking process by downward projection is carried out and subsequently the region dissimilarity, combining color, scale and homogeneity is estimated for the finer scale (localization scale). The final segmentation is derived using a previously presented merging process. To validate its performance qualitative and quantitative results are provided. Sokratis Makrogiannis, Iris Vanhamel, Hichem Sahli, Spiros Fotopoulos |
ICIP (1) | 3 |
| 2001 | A model-based approach to the automatic extraction of linear features from airborne imagesabstractThe authors describe a model-based method for the automatic extraction of linear features, like roads and paths, from aerial images. The paper combines and extends two earlier approaches for road detection in SAR satellite images and presents the modifications needed for the application domain of airborne image analysis together with representative results. Antonis Katartzis, Hichem Sahli, Veselin Pizurica, Jan Cornelis 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2000 | Eye Activity Detection and Recognition Using Morphological Scale-Space DecompositionabstractAutomatic recovery of eye gestures from image sequences is one of the important topics for face recognition and model-based coding of videophone sequences. Usually, complicated models of the eye and its motion are used. In this paper an eye gesture parameter estimation is described. A previously published automatic eye detection/tracking algorithm, based on template matching, is used for the eye pose detection. The eye gesture analysis is realised with a mathematical morphology scale-space approach, forming spatio-temporal curves out of scale measurement statistics. The resulting curves provide a direct measure of the eye gesture, which can then be used as an eye animation parameter. Experimental results demonstrate the efficiency and robustness of the method. Ilse Ravyse, Hichem Sahli, Jan Cornelis 0001, Marcel J. T. Reinders |
ICPR | 2 |
| 1999 | Low level image partitioning guided by the gradient watershed hierarchy
Ioannis Pratikakis, Hichem Sahli, Jan Cornelis 0001 |
Signal Process. | 2 |