VLDB 2026 Research / reviewers in the wild / expert
Hichem Sahbi
dblp:10/252
· DBLP profile ↗
95ranked-venue papers
42as first author
30since 2021 · last 2026
0000-0001-6813-9146ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 63 · 28 first-author · 21 since 2021Artificial intelligence and machine learning · 33 · 18 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 8 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Less Is Better: Sparse Instance Learning for Cross-Domain Few-Shot Object DetectionabstractCross-Domain Few-Shot Object Detection (CD-FSOD) is an extremely challenging task due to the inherent data scarcity and substantial domain shift between the source and target domains. Existing methods often suffer from overfitting and noisy feature representations, which hinder the construction of discriminative class prototypes in the target domain. In this paper, we propose a novel framework with sparse instance learning (SI-ViTO) for CD-FSOD, which leverages instance sparsity to achieve a better detection with less representation. SI-ViTO adopts a dual-stage sparsity module, consisting of instance feature sparsity not only on the few-shot support images but also on the query images. This dual sparsity enables the model to effectively preserve salient foreground semantics and simultaneously to filter out redundant or noisy information. Furthermore, a new prototype calibration strategy is also used to dynamically refine the class prototypes with query instances to accelerate prototype adaptation. Extensive experimental results on CD-FSOD benchmarks show that SI-ViTO outperforms the state-of-the-art methods, demonstrating that less discriminative representations yield better cross-domain few-shot object detection performance than more abundant ones. Yali Huang, Hongru Zhao, Mingyuan Jiu, Hichem Sahbi |
AAAI | 7 |
| 2026 | Calibrate and Aggregate: Cross-Modal Retrieval with Distribution Alignment and Token ReductionabstractCross-modal image-text retrieval remains a fundamental challenge at the intersection of vision and language. While existing methods leveraging large-scale pre-trained models such as CLIP and BERT have achieved notable progress, they often struggle with the inherent modality gap, background clutter in images, and computational inefficiency. In this paper, we propose an end-to-end framework for efficient cross-modal retrieval using Distribution Alignment and Token Reduction (DATR). The model employs a feature extraction network to extract multi-level representations, which are preliminarily aligned via a distribution calibration module. This module maps features into a Gaussian latent space and maximizes inter-modal mutual information through InfoNCE loss, effectively bridging the semantic gap. Furthermore, we incorporate a differentiable token reduction module inspired by GroupViT to dynamically cluster redundant visual tokens into semantic groups, significantly reducing computational overhead. Similarity computation integrates global and local alignments for robust matching. Extensive experiments on Flickr30K and MSCOCO datasets demonstrate state-of-the-art performance, with image-to-text retrieval R@1 reaching 89.6% and text-to-image retrieval R@1 achieving 77.2% on Flickr30K. The source codes are available at https://github.com/Wuziyi123/DATR. Mingyuan Jiu, Hongru Zhao, Hichem Sahbi, Mingliang Xu 0001 |
ICMR | 5 |
| 2025 | Label-Efficient Skeleton-Based Recognition with Stable Graph ConvnetsabstractDespite the notable success of graph convolutional networks (GCNs) in skeleton-based action recognition, their performance often depends on large volumes of labeled data, which are frequently scarce in practical settings. To address this limitation, we propose a novel label-efficient GCN model. Our work makes two primary contributions. First, we develop a novel acquisition function that employs an adversarial strategy to identify a compact set of informative exemplars for labeling. This selection process balances representativeness, diversity, and uncertainty. Second, we introduce bidirectional and stable GCN architectures. These enhanced networks facilitate a more effective mapping between the ambient and latent data spaces, enabling a better understanding of the learned exemplar distribution. Extensive evaluations on two challenging skeleton-based action recognition benchmarks reveal significant improvements achieved by our label-efficient GCNs compared to prior work. Hichem Sahbi |
CBMI | 1 |
| 2025 | Stable-Invertible Graph Convolutional Networks for Label-Efficient Skeleton-Based RecognitionabstractSkeleton-based action recognition is a hotspot in image processing. A key challenge of this task lies in its dependence on large, manually labeled datasets whose acquisition is costly and time-consuming. This paper devises a novel, label-efficient method for skeleton-based action recognition using graph convolutional networks (GCNs). The contribution of the proposed method resides in learning a novel acquisition function — scoring the most informative subsets for labeling — as the optimum of an objective function mixing data rep-resentativity, diversity and uncertainty. We also extend this approach by learning the most informative subsets using an invertible GCN which allows mapping data from ambient to latent spaces where the inherent distribution of the data is more easily captured. Extensive experiments, conducted on two challenging skeleton-based recognition datasets, show the effectiveness and the outperformance of our label-frugal GCNs against the related work. Hichem Sahbi |
ICIP | 1 |
| 2024 | Learning Classwise Untangled Continuums for Conditional Normalizing Flows
Victor Enescu, Hichem Sahbi |
ACCV (5) | 2 |
| 2024 | Learning conditionally untangled latent spaces using Fixed Point Iteration
Victor Enescu, Hichem Sahbi |
BMVC | 2 |
| 2024 | Coarse-To-Fine Pruning of Graph Convolutional Networks for Skeleton-Based RecognitionabstractMagnitude Pruning is a staple lightweight network design method which seeks to remove connections with the smallest magnitude. This process is either achieved in a structured or unstructured manner. While structured pruning allows reaching high efficiency, unstructured one is more flexible and leads to better accuracy, but this is achieved at the expense of low computational performance. In this paper, we devise a novel coarse-to-fine (CTF) method that gathers the advantages of structured and unstructured pruning while discarding their inconveniences to some extent. Our method relies on a novel CTF parametrization that models the mask of each connection as the Hadamard product involving four parametrizations which capture channel-wise, column-wise, rowwise and entry-wise pruning respectively. Hence, fine-grained pruning is enabled only when the coarse-grained one is disabled, and this leads to highly efficient networks while being effective. Extensive experiments conducted on the challenging task of skeleton-based recognition, using the standard SBU and FPHA datasets, show the clear advantage of our CTF approach against different baselines as well as the related work. Index Terms-Coarse and fine-grained pruning, graph convolutional networks, skeleton-based recognition. Hichem Sahbi |
CBMI | 1 |
| 2024 | TCMP: End-to-End Topologically Consistent Magnitude Pruning for Miniaturized Graph Convolutional NetworksabstractMagnitude pruning is one of the mainstream methods in lightweight architecture design whose goal is to extract sub-networks with the largest weight connections. This method is known to be successful, but under very high pruning regimes, it suffers from topological inconsistency which renders the extracted subnetworks disconnected, and this hinders their generalization ability. In this paper, we devise TCMP a novel end-to-end Topologically Consistent Magnitude Pruning method that allows extracting subnetworks while guaranteeing their topological consistency. The latter ensures that only accessible and co-accessible — impactful — connections are kept in the resulting lightweight networks. Our solution is based on a novel reparametrization and two supervisory bi-directional networks which implement accessibility/co-accessibility and guarantee that only connected subnetworks will be selected during training. This solution allows enhancing generalization significantly, under very high pruning regimes, as corroborated through extensive experiments, involving graph convolutional networks, on the challenging task of skeleton-based action recognition. Hichem Sahbi |
ICASSP | 1 |
| 2024 | DAMP: Distribution-Aware Magnitude Pruning for Budget-Sensitive Graph Convolutional NetworksabstractGraph convolutional networks (GCNs) are nowadays becoming mainstream in solving many image processing tasks including skeleton-based recognition. Their general recipe consists in learning convolutional and attention layers that maximize classification performances. With multi-head attention, GCNs are highly accurate but oversized, and their deployment on edge devices requires their pruning. Among existing methods, magnitude pruning (MP) is relatively effective but its design is clearly suboptimal as network topology selection and weight retraining are achieved independently. In this paper, we devise a novel lightweight GCN design dubbed as Distribution-Aware Magnitude Pruning (DAMP). The latter is variational and proceeds by aligning the weight distribution of the learned networks with an a priori distribution. This allows implementing any targeted pruning rate while maintaining high generalization of the designed lightweight GCNs particularly at the highest (most interesting) pruning regimes. Extensive experiments conducted on the challenging task of skeleton-based recognition show a substantial gain of our DAMP compared to MP as well as related methods. Hichem Sahbi |
ICASSP | 1 |
| 2024 | One-Shot Multi-Rate Pruning Of Graph Convolutional Networks For Skeleton-Based RecognitionabstractIn this paper, we devise a novel lightweight Graph Convolutional Network (GCN) design dubbed as Multi-Rate Magnitude Pruning (MRMP) that jointly trains network topology and weights. Our method is variational and proceeds by aligning the weight distribution of the learned networks with an a priori distribution. In the one hand, this allows implementing any fixed pruning rate, and also enhancing the generalization performances of the designed lightweight GCNs. In the other hand, MRMP achieves a joint training of multiple GCNs, on top of shared weights, in order to extrapolate accurate networks at any targeted pruning rate without retraining their weights. Extensive experiments conducted on the challenging task of skeleton-based recognition show a substantial gain of our lightweight GCNs particularly at very high pruning regimes. Hichem Sahbi |
ICIP | 1 |
| 2024 | Deep Multi-order Context-Aware Kernel Network for Multi-label Classification
Mingyuan Jiu, Hailong Zhu, Hichem Sahbi |
ICPR (3) | 3 |
| 2024 | Semi-structured Pruning of Graph Convolutional Networks for Skeleton-Based Recognition
Hichem Sahbi |
ICPR (2) | 1 |
| 2024 | Conditional Normalizing Flows for Nonlinear Remote Sensing Image Augmentation and ClassificationabstractDeep neural networks have recently shown outstanding performances in remote sensing image classification. The success of these models is highly reliant on the availability of large collections of hand-labeled training images which are usually scarce. Data augmentation mitigates this scarcity by enriching labeled training sets using different geometric and photometric transformations, or by relying on deep generative models. In this paper, we investigate the potential of generative models, and particularly normalizing flows (NFs), in remote sensing image augmentation and classification. The main contribution relies on a novel conditional NF model that achieves a bidirectional mapping of images between ambient and latent spaces with the particularity of learning disentangled multi-modal distributions through image classes. The proposed NF also achieves nonlinear augmentations in highly intricate ambient spaces by mapping images to latent spaces where augmentations become linear and more tractable. Extensive experiments conducted on the EuroSAT benchmark show the benefit of our NF-based augmentation when learning vision transformers. Victor Enescu, Hichem Sahbi |
IGARSS | 2 |
| 2024 | Label-Efficient Display-Size Selection for Satellite Image Change DetectionabstractWe introduce a novel label-efficient satellite image change detection algorithm based on active and reinforcement learning. The proposed method is interactive and consists in frugally probing the user (oracle) about the labels of the most critical images, and according to the oracle’s annotations, updates change detection results. First, we consider a probabilistic model which assigns to each unlabeled sample a relevance measure modeling how critical is that sample when training change detection functions. We obtain these relevance measures by minimizing an objective function mixing uncertainty, diversity and representativity; when combined, these three criteria allow exploring different data modes and also refining change detections. Then, we further explore the potential of this objective function, by considering a reinforcement learning approach that finds the best combination of these three criteria together with display-sizes through active learning iterations, leading to better generalization as shown through experiments in interactive satellite image change detection. Hichem Sahbi |
IGARSS | 1 |
| 2024 | Sparse Context Transformer for Few-Shot Object Detection
Mingyuan Jiu, Hichem Sahbi, Xiaoheng Jiang, Mingliang Xu 0001 |
PRICAI (4) | 3 |
| 2024 | MonoProb: Self-Supervised Monocular Depth Estimation with Interpretable UncertaintyabstractSelf-supervised monocular depth estimation methods aim to be used in critical applications such as autonomous vehicles for environment analysis. To circumvent the potential imperfections of these approaches, a quantification of the prediction confidence is crucial to guide decision-making systems that rely on depth estimation. In this paper, we propose MonoProb, a new unsupervised monocular depth estimation method that returns an interpretable uncertainty, which means that the uncertainty reflects the expected error of the network in its depth predictions. We rethink the stereo or the structure-from-motion paradigms used to train unsupervised monocular depth models as a probabilistic problem. Within a single forward pass inference, this model provides a depth prediction and a measure of its confidence, without increasing the inference time. We then improve the performance on depth and uncertainty with a novel self-distillation loss for which a student is supervised by a pseudo ground truth that is a probability distribution on depth output by a teacher. To quantify the performance of our models we design new metrics that, unlike traditional ones, measure the absolute performance of uncertainty predictions. Our experiments highlight enhancements achieved by our method on standard depth and uncertainty metrics as well as on our tailored metrics. https://github.com/CEA-LIST/MonoProb Rémi Marsal, Florian Chabot, Angélique Loesch, William Grolleau, Hichem Sahbi |
WACV | 5 |
| 2023 | Adversarial Label-Efficient Satellite Image Change DetectionabstractSatellite image change detection aims at finding occurrences of targeted changes in a given scene taken at different instants. This task is highly challenging due to the acquisition conditions and also to the subjectivity of changes. In this paper, we investigate satellite image change detection using active learning. Our method is interactive and relies on a question & answer model which asks the oracle (user) questions about the most informative display (dubbed as virtual exemplars), and according to the user’s responses, updates change detections. The main contribution of our method consists in a novel adversarial model that allows frugally probing the oracle only with the most representative, diverse and uncertain virtual exemplars. The latter are learned to challenge (the most) the trained change decision criteria which ultimately leads to a better re-estimate of these criteria in the following iterations of active learning. Conducted experiments show the out-performance of our proposed adversarial display model against related display strategies as well as the related work. Hichem Sahbi, Sebastien Deschamps |
IGARSS | 1 |
| 2023 | Deep-Net Inversion for Frugal Satellite Image Change DetectionabstractChange detection in satellite imagery seeks to find occurrences of targeted changes in a given scene taken at different instants. This task has several applications ranging from land-cover mapping, to anthropogenic activity monitory as well as climate change and natural hazard damage assessment. However, change detection is highly challenging due to the acquisition conditions and also to the subjectivity of changes.In this paper, we devise a novel algorithm for change detection based on active learning. The proposed method is based on a question & answer model that probes an oracle (user) about the relevance of changes only on a small set of critical images (referred to as virtual exemplars), and according to oracle’s responses updates deep neural network (DNN) classifiers. The main contribution resides in a novel adversarial model that allows learning the most representative, diverse and uncertain virtual exemplars (as inverted preimages of the trained DNNs) that challenge (the most) the trained DNNs, and this leads to a better re-estimate of these networks in the subsequent iterations of active learning. Experiments show the out-performance of our proposed deep-net inversion against the related work. Hichem Sahbi, Sebastien Deschamps |
IGARSS | 1 |
| 2023 | BrightFlow: Brightness-Change-Aware Unsupervised Learning of Optical FlowabstractUnsupervised optical flow estimation relies on the assumption that pixels characterizing the same observed object should exhibit a stable appearance across video frames. With this assumption, the long-standing principle behind flow estimation consists in optimizing a photometric loss that maximizes the similarity between paired pixels in successive frames. However, these frames could be subject to strong brightness changes due to the radiometric properties of scenes as well as their viewing conditions.In this paper, we present BrightFlow, a new method to train any optical flow estimation network in an unsupervised manner. It consists in training two networks that jointly estimate optical flow and brightness changes. These changes are then compensated in the photometric loss so that reconstruction errors due to shadows or reflections will not affect negatively the training. As this compensation mechanism is only used at training stage, our method does not impact the number of parameters or the complexity at inference. Extensive experiments conducted on standard datasets and optical flow architectures show a consistent gain of our method. Source code is available at https://github.com/CEA-LIST/BrightFlow. Rémi Marsal, Florian Chabot, Angélique Loesch, Hichem Sahbi |
WACV | 4 |
| 2022 | Extracting Effective Subnetworks with Gumbel-SoftmaxabstractLarge and performant neural networks are often overparameterized and can be drastically reduced in size and complexity thanks to pruning. Pruning is a group of methods, which seeks to remove redundant or unnecessary weights or groups of weights in a network. These techniques allow the creation of lightweight networks, which are particularly critical in embedded or mobile applications.In this paper, we devise an alternative pruning method that allows extracting effective subnetworks from larger untrained ones. Our method is stochastic and extracts subnetworks by exploring different topologies which are sampled using Gumbel Softmax. The latter is also used to train probability distributions which measure the relevance of weights in the sampled topologies. The resulting subnetworks are further enhanced using a highly efficient rescaling mechanism that reduces training time and improves performance. Extensive experiments conducted on CIFAR show the outperformance of our subnetwork extraction method against the related work. Robin Dupont, Mohammed Amine Alaoui, Hichem Sahbi, Alice Lebois |
ICIP | 3 |
| 2022 | Topologically-Consistent Magnitude Pruning for Very Lightweight Graph Convolutional NetworksabstractGraph convolution networks (GCNs) are currently main-stream in learning with irregular data. These models rely on message passing and attention mechanisms that capture context and node-to-node relationships. With multi-head attention, GCNs become highly accurate but oversized, and their deployment on cheap devices requires their pruning. However, pruning at high regimes usually leads to topologically inconsistent networks with weak generalization.In this paper, we devise a novel method for lightweight GCN design. Our proposed approach parses and selects subnet-works with the highest magnitudes while guaranteeing their topological consistency. The latter is obtained by selecting only accessible and co-accessible connections which actually contribute in the evaluation of the selected subnetworks. Experiments conducted on the challenging FPHA dataset show the substantial gain of our topologically consistent pruning method especially at very high pruning regimes. Hichem Sahbi |
ICIP | 1 |
| 2022 | Reinforcement-based Display Selection for Frugal LearningabstractMost of the existing learning models, particularly deep neural networks, are reliant on large datasets whose hand-labeling is expensive and time demanding. A current trend is to make the learning of these models frugal and less dependent on large collections of labeled data. Among the existing solutions, deep active learning is currently witnessing a major interest and its purpose is to train deep networks using as few labeled samples as possible. However, the success of active learning is highly dependent on how critical are these samples when training models. In this paper, we devise a novel active learning approach for label-efficient training. The proposed method is iterative and aims at minimizing a constrained objective function that mixes diversity, representativity and uncertainty criteria. The proposed approach is probabilistic and unifies all these criteria in a single objective function whose solution models the probability of relevance of samples (i.e., how critical) when learning a decision function. We also introduce a novel weighting mechanism based on reinforcement learning, which adaptively balances these criteria at each training iteration, using a particular stateless Q-learning model. Extensive experiments conducted on staple image classification data, including Object-DOTA, show the effectiveness of our proposed model w.r.t. several baselines including random, uncertainty and flat as well as other work. Sebastien Deschamps, Hichem Sahbi |
ICPR | 2 |
| 2022 | Reinforcement-Based Frugal Learning for Interactive Satellite Image Change DetectionabstractIn this paper, we introduce a novel interactive satellite im-age change detection algorithm based on active learning. The proposed approach is iterative and asks the user (oracle) questions about the targeted changes and according to the oracle's responses updates change detections. We consider a proba-bilistic framework which assigns to each unlabeled sample a relevance measure modeling how critical is that sample when training change detection functions. These relevance mea-sures are obtained by minimizing an objective function mixing diversity, representativity and uncertainty. These criteria when combined allow exploring different data modes and also refining change detections. To further explore the potential of this objective function, we consider a reinforcement learning approach that finds the best combination of diversity, repre-sentativity and uncertainty, through active learning iterations, leading to better generalization as corroborated through ex-periments in interactive satellite image change detection. Sebastien Deschamps, Hichem Sahbi |
IGARSS | 2 |
| 2022 | Learning Virtual Exemplars for Label-Efficient Satellite Image Change DetectionabstractIn this paper, we devise a novel interactive satellite image change detection algorithm based on active learning. The proposed framework is iterative and relies on a question & answer model which asks the oracle (user) questions about the most informative display (subset of critical images), and according to the user's responses, updates change detections. The contribution of our framework resides in a novel display model which selects the most representative and diverse vir-tual exemplars that adversely challenge the learned change detection functions, thereby leading to highly discriminating functions in the subsequent iterations of active learning. Ex-periments, conducted on the challenging task of interactive satellite image change detection, show the superiority of the proposed virtual display model against the related work. Hichem Sahbi, Sebastien Deschamps |
IGARSS | 1 |
| 2022 | Context-aware deep kernel networks for image annotation
Mingyuan Jiu, Hichem Sahbi |
Neurocomputing | 2 |
| 2021 | FFNB: Forgetting-Free Neural Blocks for Deep Continual Learning
Hichem Sahbi, Haoming Zhan |
BMVC | 1 |
| 2021 | DHCN: Deep Hierarchical Context Networks For Image AnnotationabstractContext modeling is one of the most fertile sub-fields of visual recognition which aims at designing discriminant image representations while incorporating their intrinsic and extrinsic relationships. However, the potential of context modeling is currently under-explored and most of the existing solutions are either context-free or restricted to simple handcrafted geometric relationships.We introduce in this paper DHCN: a novel Deep Hierarchical Context Network that leverages different sources of contexts including geometric and semantic relationships. The proposed method is based on the minimization of an objective function mixing a fidelity term, a context criterion and a regularizer. The solution of this objective function defines the architecture of a bi-level hierarchical context network; the first level of this network captures scene geometry while the second one corresponds to semantic relationships. We solve this representation learning problem by training its underlying deep network whose parameters correspond to the most influencing bi-level contextual relationships and we evaluate its performances on image annotation using the challenging ImageCLEF benchmark. Mingyuan Jiu, Hichem Sahbi |
ICASSP | 2 |
| 2021 | Weight Reparametrization for Budget-Aware Network PruningabstractPruning seeks to design lightweight architectures by removing redundant weights in overparameterized networks. Most of the existing techniques first remove structured sub-networks (filters, channels,…) and then fine-tune the resulting networks to maintain a high accuracy. However, removing a whole structure is a strong topological prior and recovering the accuracy, with fine-tuning, is highly cumbersome.In this paper, we introduce an “end-to-end” lightweight network design that achieves training and pruning simultaneously without fine-tuning. The design principle of our method relies on reparametrization that learns not only the weights but also the topological structure of the lightweight sub-network. This reparametrization acts as a prior (or regularizer) that defines pruning masks implicitly from the weights of the underlying network, without increasing the number of training parameters. Sparsity is induced with a budget loss that provides an accurate pruning. Extensive experiments conducted on the CIFAR10 and the TinyImageNet datasets, using standard architectures (namely Conv4, VGG19 and ResNet18), show compelling results without fine-tuning. Robin Dupont, Hichem Sahbi, Guillaume Michel |
ICIP | 2 |
| 2021 | Lightweight Connectivity In Graph Convolutional Networks For Skeleton-Based RecognitionabstractGraph convolutional networks (GCNs) aim at extending deep learning to arbitrary irregular domains, namely graphs. Their success is highly dependent on how the topology of input graphs is defined and most of the existing GCN architectures rely on predefined or handcrafted graph structures. In this paper, we introduce a novel method that learns the topology (or connectivity) of input graphs as a part of GCN design. The main contribution of our method resides in building an orthogonal connectivity basis that optimally aggregates nodes, through their neighborhood, prior to achieve convolution. Our method also considers a stochasticity criterion which acts as a regularizer that makes the learned basis and the underlying GCNs lightweight while still being highly effective. Experiments conducted on the challenging task of skeleton-based hand-gesture recognition show the high effectiveness of the learned GCNs w.r.t. the related work. Hichem Sahbi |
ICIP | 1 |
| 2021 | Frugal Learning for Interactive Satellite Image Change DetectionabstractWe introduce in this paper a novel active learning algorithm for satellite image change detection. The proposed solution is interactive and based on a question & answer model, which asks an oracle (annotator) the most informative questions about the relevance of sampled satellite image pairs, and according to the oracle's responses, updates a decision function iteratively. We investigate a novel framework which models the probability that samples are relevant; this probability is obtained by minimizing an objective function capturing representativity, diversity and ambiguity. Only data with a high probability according to these criteria are selected and displayed to the oracle for further annotation. Extensive experiments on the task of satellite image change detection after natural hazards (namely tornadoes) show the relevance of the proposed method against the related work. Hichem Sahbi, Sebastien Deschamps, Andrei Stoian |
IGARSS | 1 |
| 2020 | End-to-End Deep Kernel Map Design for Image AnnotationabstractDeep kernel map networks have shown excellent performances in various classification problems including image annotation. Their general recipe consists in aggregating several layers of singular value decompositions (SVDs) - that map data from input spaces into high dimensional spaces - while preserving the similarity of the underlying kernels. However, the potential of these deep map networks has not been fully explored as the original setting of these networks focuses mainly on the approximation quality of their kernels and ignores their discrimination power.In this paper, we introduce a novel “end-to-end” design for deep kernel map learning that balances the approximation quality of kernels and their discrimination power. Our method proceeds in two steps; first, layerwise SVD is applied in order to build initial deep kernel map approximations and then an “end-to-end” supervised learning is employed to further enhance their discrimination power while maintaining their efficiency. Extensive experiments, conducted on the challenging ImageCLEF annotation benchmark, show the high efficiency and the out-performance of this two-step process with respect to different related methods. Mingyuan Jiu, Hichem Sahbi |
ICIP | 2 |
| 2020 | Coarse-to-Fine Aggregation for Cross-Granularity Action RecognitionabstractIn this paper, we introduce a novel hierarchical aggregation design that captures different levels of temporal granularity in action recognition. Our design principle is coarse-to-fine and achieved using a tree-structured network; as we traverse this network top-down, pooling operations are getting less invariant but timely more resolute and well localized. Learning the combination of operations in this network - which best fits a given ground-truth - is obtained by solving a constrained minimization problem whose solution corresponds to the distribution of weights that capture the contribution of each level (and thereby temporal granularity) in the global hierarchical pooling process. Besides being principled and well grounded, the proposed hierarchical pooling is also video-length agnostic and resilient to misalignments in actions. Extensive experiments conducted on the challenging UCF-101 database corroborate these statements. Ahmed Mazari, Hichem Sahbi |
ICIP | 2 |
| 2020 | Kernel-based Graph Convolutional NetworksabstractLearning graph convolutional networks (GCNs) is an emerging field which aims at generalizing deep learning to arbitrary non-regular domains. Most of the existing GCNs follow a neighborhood aggregation scheme, where the representation of a node is recursively obtained by aggregating its neighboring node representations using averaging or sorting operations. However, these operations are either ill-posed or weak to be discriminant or increase the number of training parameters and thereby the computational complexity and the risk of overfitting. In this paper, we introduce a novel GCN framework that achieves spatial graph convolution in a reproducing kernel Hilbert space. The latter makes it possible to design, via implicit kernel representations, convolutional graph filters in a high dimensional and more discriminating space without increasing the number of training parameters. The particularity of our GCN model also resides in its ability to achieve convolutions without explicitly realigning nodes in the receptive fields of the learned graph filters with those of the input graphs, thereby making convolutions permutation agnostic and well defined. Experiments conducted on the challenging task of skeleton-based action recognition show the superiority of the proposed method against different baselines as well as the related work. Hichem Sahbi |
ICPR | 1 |
| 2020 | Learning Connectivity with Graph Convolutional NetworksabstractLearning graph convolutional networks (GCNs) is an emerging field which aims at generalizing convolutional operations to arbitrary non-regular domains. In particular, GCNs operating on spatial domains show superior performances compared to spectral ones, however their success is highly dependent on how the topology of input graphs is defined. In this paper, we introduce a novel framework for graph convolutional networks that learns the topological properties of graphs. The design principle of our method is based on the optimization of a constrained objective function which learns not only the usual convolutional parameters in GCNs but also a transformation basis that conveys the most relevant topological relationships in these graphs. Experiments conducted on the challenging task of skeleton-based action recognition shows the superiority of the proposed method compared to handcrafted graph design as well as the related work. Hichem Sahbi |
ICPR | 1 |
| 2019 | MLGCN: Multi-Laplacian Graph Convolutional Networks for Human Action Recognition
Ahmed Mazari, Hichem Sahbi |
BMVC | 2 |
| 2019 | Deep Temporal Pyramid Design for Action RecognitionabstractDeep convolutional neural networks (CNNs) are nowadays achieving significant leaps in different pattern recognition tasks including action recognition. Current CNNs are increasingly deeper, data-hungrier and this makes their success tributary of the abundance of labeled training data. CNNs also rely on max/average pooling which reduces dimensionality of output layers and hence attenuates their sensitivity to the availability of labeled data. However, this process may dilute the information of upstream convolutional layers and thereby affect the discrimination power of the trained representations, especially when the learned categories are fine-grained. In this paper, we introduce a novel hierarchical aggregation design, for final pooling, that controls granularity of the learned representations w.r. t the actual granularity of action categories. Our solution is based on a tree-structured temporal pyramid that aggregates outputs of CNNs at different levels. Top levels of this hierarchy are dedicated to coarse categories while deep levels are more suitable to fine-grained ones. The design of our temporal pyramid is based on solving a constrained minimization problem whose solution corresponds to the distribution of weights of different representations in the temporal pyramid. Experiments conducted using the challenging UCF101 database show the relevance of our hierarchical design w.r. t other related methods. Ahmed Mazari, Hichem Sahbi |
ICASSP | 2 |
| 2019 | Deep representation design from deep kernel networks
Mingyuan Jiu, Hichem Sahbi |
Pattern Recognit. | 2 |
| 2019 | Stochastic Graphlet EmbeddingabstractGraph-based methods are known to be successful in many machine learning and pattern classification tasks. These methods consider semistructured data as graphs where nodes correspond to primitives (parts, interest points, and segments) and edges characterize the relationships between these primitives. However, these nonvectorial graph data cannot be straightforwardly plugged into off-the-shelf machine learning algorithms without a preliminary step of-explicit/implicit-graph vectorization and embedding. This embedding process should be resilient to intraclass graph variations while being highly discriminant. In this paper, we propose a novel high-order stochastic graphlet embedding that maps graphs into vector spaces. Our main contribution includes a new stochastic search procedure that efficiently parses a given graph and extracts/samples unlimitedly high-order graphlets. We consider these graphlets, with increasing orders, to model local primitives as well as their increasingly complex interactions. In order to build our graph representation, we measure the distribution of these graphlets into a given graph, using particular hash functions that efficiently assign sampled graphlets into isomorphic sets with a very low probability of collision. When combined with maximum margin classifiers, these graphlet-based representations have a positive impact on the performance of pattern comparison and recognition as corroborated through extensive experiments using standard benchmark databases. Anjan Dutta 0001, Hichem Sahbi |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Structured Scene Decoding with Finite State MachinesabstractWe introduce in this work a novel stochastic inference process, for scene annotation and object class segmentation, based on finite state machines (FSMs). The design principle of our framework is generative and based on building, for a given scene, finite state machines that encode annotation lattices, and inference consists in finding and scoring the best configurations in these lattices. Different novel operations are defined using our FSM framework including reordering, segmentation, visual transduction, and label dependency modeling. All these operations are combined together in order to achieve annotation as well as object class segmentation. Hichem Sahbi |
ICIP | 1 |
| 2018 | Deep Context Networks for Image AnnotationabstractContext plays an important role in visual pattern recognition as it provides complementary clues for different learning tasks including image classification and annotation. In the particular scenario of kernel learning, the general recipe of context-based kernel design consists in learning positive semi-definite similarity functions that return high values not only when data share similar content but also similar context. However, in spite of having a positive impact on performance, the use of context in these kernel design methods has not been fully explored; indeed, context has been handcrafted instead of being learned. In this paper, we introduce a novel context-aware kernel design framework based on deep learning. Our method discriminatively learns spatial geometric context as the weights of a deep network (DN). The architecture of this network is fully determined by the solution of an objective function that mixes content, context and regularization, while the parameters of this network determine the most relevant (discriminant) parts of the learned context. We apply this context and kernel learning framework to image classification using the challenging ImageCLEF Photo Annotation benchmark; the latter shows that our deep context learning provides highly effective kernels for image classification as corroborated through extensive experiments. Mingyuan Jiu, Hichem Sahbi |
ICPR | 2 |
| 2018 | From Transductive to Inductive Semi-Supervised Attributes for Ship Category RecognitionabstractFine-grained ship category recognition is a data-hungry learning task that requires a lot of labeled data which are usually scarce. Alternative models, as transductive attributes, bypass this limitation by considering not only labeled data but also abundant unlabeled ones. However, these transductive methods are basically designed for observed data, and their extension to unobserved sets requires retraining the whole models. In this paper, we introduce a novel ship category recognition method based on semi-supervised learning; the strength of our method resides in its ability to leverage labeled and unlabeled observed data while being highly effective and efficient in order to handle unobserved ones. We consider two variants of our method, the first one is non-parametric and based on support vector regression while the second one is parametric and based on deep neural networks. Experiments conducted on the challenging fine-grained ship category recognition show that our semi-supervised method is highly effective and generalizes well across unobserved sets. Quentin Oliveau, Hichem Sahbi |
IGARSS | 2 |
| 2018 | Semi-Supervised Deep Attribute Networks for Fine-Grained Ship Category RecognitionabstractClassifying ships in satellite or aerial images is a challenging problem in remote sensing imagery. This task requires data-hungry learning algorithms (in particular deep models) that build discriminative representations which capture highly variable ship categories. As labeled training data are scarce and expensive, ship category recognition should also rely on abundant unlabeled data in order to enhance its effectiveness. In this paper, we introduce a novel representation learning algorithm based on semi-supervised attributes. Our method allows us to learn deep and discriminative image characteristics shared among different categories while taking into account both labeled and unlabeled data. We demonstrate the effectiveness of this method on the challenging ship category recognition problem and we show its out-performance with respect to related baselines, especially under the regime of scarce labeled data and abundant unlabeled ones. Quentin Oliveau, Hichem Sahbi |
IGARSS | 2 |
| 2018 | Superpixel Partitioning of Very High Resolution Satellite Images for Large-Scale Classification Perspectives with Deep Convolutional Neural NetworksabstractSupervised classification is the fundamental task for land-cover map generation. Deep neural networks recently outperformed other state-of-the-art classifiers in many machine learning challenges, from semantic segmentation to speech recognition. Such strategies are now commonly employed in the literature for the purpose of land-cover mapping. This paper develops the strategy for the use of deep networks to label very high resolution satellite images, with the perspective of mapping regions at country scale. Therefore, a superpixel based method is introduced in order to (i) ensure correct delineation of objects and (ii) perform the classification in a dense way but with decent computing times. Tristan Postadjian, Arnaud Le Bris, Clément Mallet, Hichem Sahbi |
IGARSS | 4 |
| 2018 | Domain Adaptation for Large Scale Classification of Very High Resolution Satellite Images with Deep Convolutional Neural NetworksabstractSemantic segmentation of remote sensing images enables in particular land-cover map generation for a given set of classes. Very recent literature has shown the superior performance of deep convolutional neural networks (DCNN) for many tasks, from object recognition to semantic labelling, including the classification of Very High Resolution (VHR) satellite images. However, while plethora of works aim at improving object delineation on geographically restricted areas, few tend to solve this classification task at very large scales. New issues occur such as intra-class class variability, diachrony between surveys, and the appearance of new classes in a specific area, that do not exist in the predefined set of labels. Therefore, this work intends to (i) perform large scale classification and to (ii) expand a set of land-cover classes, using the off-the-shelf model learnt in a specific area of interest and adapting it to unseen areas. Tristan Postadjian, Arnaud Le Bris, Hichem Sahbi, Clément Mallet |
IGARSS | 3 |
| 2017 | Transductive attributes for ship category recognitionabstractShip category recognition is one of the remote sensing applications that requires designing accurate image representation and classification models. Training these models is usually a data hungry process, that requires a lot of labeled data which are usually scarce and expensive. As unlabeled data are more abundant and relatively cheaper, transductive methods exploiting these data are highly preferred. We introduce in this paper a novel representation learning approach based on transductive attributes. The strength of our method resides in its ability to upgrade labeled training sets using attributes shared across categories, and also its ability to further push the benefit of attributes by taking into account not only the labeled data but also the abundant unlabeled ones. When plugging these learned attribute representations into transductive classifiers, we obtain a substantial gain of accuracy compared to several baselines as well as the related work, on the challenging task of ship category recognition, especially when labeled training data are very scarce. Quentin Oliveau, Hichem Sahbi |
IGARSS | 2 |
| 2017 | Interactive Satellite Image Change Detection With Context-Aware Canonical Correlation AnalysisabstractAutomatic change detection is one of the remote sensing applications that has received an increasing attention during the last years. However, fully automatic solutions reach their limitation; on the one hand, it is difficult to design general decision criteria able to select area of changes for images under various acquisition conditions, and on the other hand, the relevance of changes may differ from one user to another. In this letter, we introduce an alternative change detection method based on relevance feedback. The proposed algorithm is iterative and based on a query and answer model that: 1) asks the user questions about the relevance of his targeted changes and 2) according to these answers, updates change detection results. Our method is also based on a new formulation of canonical correlation analysis (CCA), referred to as context-aware CCA, that learns transformations, which map data from different input spaces (related to multitemporal satellite images) into a common latent space, which is sensitive to relevant changes while being resilient to irrelevant ones. These CCA transformations correspond to the optimum of a particular constrained maximization problem that mixes an alignment term with a context-based regularization criterion. The particularity of this novel CCA approach resides in its ability to exploit spatial geometric context resulting into better performances compared with other CCA approaches, as shown in experiments. Hichem Sahbi |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2017 | Nonlinear Deep Kernel Learning for Image AnnotationabstractMultiple kernel learning (MKL) is a widely used technique for kernel design. Its principle consists in learning, for a given support vector classifier, the most suitable convex (or sparse) linear combination of standard elementary kernels. However, these combinations are shallow and often powerless to capture the actual similarity between highly semantic data, especially for challenging classification tasks such as image annotation. In this paper, we redefine multiple kernels using deep multi-layer networks. In this new contribution, a deep multiple kernel is recursively defined as a multi-layered combination of nonlinear activation functions, each one involves a combination of several elementary or intermediate kernels, and results into a positive semi-definite deep kernel. We propose four different frameworks in order to learn the weights of these networks: supervised, unsupervised, kernel-based semisupervised and Laplacian-based semi-supervised. When plugged into support vector machines (SVMs), the resulting deep kernel networks show clear gain, compared to several shallow kernels for the task of image annotation. Extensive experiments and analysis on the challenging ImageCLEF photo annotation benchmark, the COREL5k database and the Banana dataset validate the effectiveness of the proposed method. Mingyuan Jiu, Hichem Sahbi |
IEEE Trans. Image Process. | 2 |
| 2016 | Laplacian deep kernel learning for image annotationabstractSemi-supervised learning seeks to build accurate classification machines by taking advantage of both labeled and unlabeled data. This learning scheme is useful especially when labeled data are scarce while unlabeled ones are abundant. Among the existing semi-supervised learning algorithms, Laplacian support vector machines (SVMs) are known to be particularly powerful but their success is highly dependent on the choice of kernels., In this paper, we propose an algorithm that designs kernels as a part of Laplacian SVM learning. The proposed kernels correspond to deep multi-layered combinations of elementary kernels which capture simple - linear - as well as intricate - nonlinear - relationships between data. Our optimization process finds both the parameters of the deep kernels and the Laplacian SVMs in a unified framework resulting into highly discriminative and accurate classifiers. When applied to the challenging ImageCLEF2013 Photo Annotation benchmark, the proposed deep kernels show significant and consistent gain compared to existing elementary kernels as well as standard multiple kernels. Mingyuan Jiu, Hichem Sahbi |
ICASSP | 2 |
| 2016 | Deep kernel map networks for image annotationabstractDeep multiple kernel learning is a powerful technique that selects and deeply combines multiple elementary kernels in order to provide the best performance on a given classification task. This technique, particularly effective, becomes intractable when handling large scale datasets; indeed, multiple nonlinear kernel combinations are time and memory demanding., In this paper, we propose a new framework that significantly reduces the complexity of deep multiple kernels. Given a deep kernel network (DKN), our method designs its equivalent deep map network (DMN), using multi-layer explicit maps that approximate the initial DKN with a high precision. When combined with support vector machines, the design of DMN preserves high classification accuracy compared to its underlying DKN while being (at least) an order of magnitude faster. Experiments conducted on the challenging Im-ageCLEF2013 annotation benchmark, show that the proposed DMN is indeed effective and highly efficient. Mingyuan Jiu, Hichem Sahbi |
ICASSP | 2 |
| 2016 | Semantic-free attributes for image classificationabstractAttributes are defined as mid-level image characteristics shared among different categories. These characteristics are suitable in order to handle classification problems especially when training data are scarce. In this paper, we design discriminative real-valued attributes by learning nonlinear inductive maps. Our method is based on solving a constrained optimization problem that mixes three criteria; the first one aims to predict real-valued attributes with a high precision while the second criterion maximizes their discrimination power. We also consider a particular smoothness term that provides gradual representations for data belonging to the same categories, resulting into highly discriminant and predictable attributes. Experiments conducted for the particular task of fine-grained image classification - with relatively small training sets - show that our attribute design approach is very competitive and outperforms related attribute learning methods. Quentin Oliveau, Hichem Sahbi |
ICPR | 2 |
| 2016 | Misalignment resilient CCA for interactive satellite image change detectionabstractChange detection, in multi-temporal satellite imagery, seeks to discover relevant changes and to discard irrelevant ones. This task is usually achieved by modeling accurate decision criteria that capture the user's intention while being resilient to many irrelevant changes including acquisition conditions. Among existing change detection solutions, correlation-based models - such as canonical correlation analysis (CCA) - are particularly successful, but their success is very dependent on the quality of alignments used to train these models. In this paper, we introduce a novel interactive change detection algorithm based on a new variant of CCA, referred to as misalignment resilient CCA. Given a small sample of “changes” and “no-changes” labeled by an oracle (user), our method learns transformation matrices that map these data from different input spaces, related to multi-temporal images, into a common latent space which is sensitive to relevant changes while being resilient to irrelevant ones including misalignments. These CCA transformations correspond to the optimum of a particular constrained maximization problem that mixes a new soft-alignment term and a context-based regularization criterion. Extensive experiments conducted in interactive satellite image change detection, show that our misalignment resilient CCA approach is highly effective. Hichem Sahbi |
ICPR | 1 |
| 2016 | Attribute learning for ship category recognition in remote sensing imageryabstractObject category recognition, in remote sensing imagery, usually relies on exemplar-based training. The latter is achieved by modeling intricate relationships between object categories and visual features. However, for real-world and fine grained object categories - exhibiting complex visual appearance and strong variability - these models may fail especially when training data are scarce. In this paper, we introduce an effective object category recognition approach that alleviates the limitation caused by small training sets. The method learns discriminant mid-level representations (a.k.a. attributes) through nonlinear mappings that make these attributes highly discriminant while being easily trainable and predictable. We demonstrate the effectiveness of our method, through extensive experiments on the challenging task of ship recognition in maritime environments, and we show how our attribute learning model generalizes well in spite of the scarcity of training data. Quentin Oliveau, Hichem Sahbi |
IGARSS | 2 |
| 2015 | Continuous visual speech recognition for audio speech enhancementabstractWe introduce in this paper a novel non-blind speech enhancement procedure based on visual speech recognition (VSR). The latter is based on a generative process that analyzes sequences of talking faces and classifies them into visual speech units known as visemes. We use an effective graphical model able to segment and label a given sequence of talking faces into a sequence of visemes. Our model captures unary potential as well as pairwise interaction; the former models visual appearance of speech units while the latter models their interactions using boundary and visual language model activations. Experiments conducted on a standard challenging dataset, show that when feeding the results of VSR to the speech enhancement procedure, it clearly outperforms baseline blind methods as well as related work. Éric Benhaim, Hichem Sahbi, Guillaume Vitte |
ICASSP | 2 |
| 2015 | Semi supervised deep kernel design for image annotationabstractIt is commonly agreed that the success of support vector machines (SVMs), is highly dependent on the choice of particular similarity functions referred to as kernels. The latter are usually handcrafted or designed using appropriate optimization schemes. Multiple kernel learning (MKL) is one possible scheme that designs kernels as sparse or convex linear combinations of existing elementary functions. However, this results into shallow kernels, which are powerless to capture the right similarity between data, especially when content of these data is highly semantic. In this paper, we redefine multiple kernels using a deep architecture. In this new formulation, a global kernel is learned as a multi-layered linear combination of activation functions, each one involves a combination of several elementary or intermediate functions on multiple features. We propose three different settings to learn the weights of these kernel combinations; supervised, unsupervised and semi-supervised. When plugged into SVMs, the resulting deep multiple kernels show a gain, compared to shallow kernels, for the challenging task of image annotation using the ImageCLEF benchmark. Mingyuan Jiu, Hichem Sahbi |
ICASSP | 2 |
| 2015 | Contextual kernel map learning for scene transductionabstractScene understanding, also known as object category segmentation, is one of the major trends in computer vision. It consists in modeling and inferring object categories through constellations of pixels belonging to a given test image. Many existing solutions suffer from (at least) two major limitations; on the one hand, they only use few (scarce) labeled training data, and on the other hand they rely on context-free learning models, thereby, their potential is not fully explored. In this paper, we adopt a transductive data-driven approach for scene understanding based on kernel machines. The main contribution of this work includes i) a novel transductive approach that exploits both labeled/unlabeled data and jointly learns classifiers and kernel maps for better discrimination, and ii) a context modeling approach that captures semantic as well as geometric relationships between object categories as a part of kernel learning. Experiments conducted in scene understanding, using the SiftFlow dataset, show that the proposed method is competitive against state of the art. Phong D. Vo, Hichem Sahbi |
ICIP | 2 |
| 2015 | Discriminant canonical correlation analysis for interactive satellite image change detectionabstractSatellite image change detection consists in defining criteria that localize relevant changes while discard irrelevant ones. This problem is known to be challenging due to several factors including illumination changes, occlusion and also subjectivity of users. In this paper, we introduce a novel change detection algorithm based on relevance feedback and a new variant of canonical correlation analysis (CCA), referred to as discriminant CCA. The contribution of this work includes: i) a relevance feedback method that adapts change detection criteria to content and acquisition conditions of satellite images as well as the user's intention, and ii) a novel discriminant CCA approach that further enhances the performance of relevance feedback by designing accurate and resilient change detection criteria. Conducted change detection experiments show a clear gain of our discriminant CCA approach compared to related methods. Hichem Sahbi |
IGARSS | 1 |
| 2014 | Continuous visual speech recognition for multimodal fusionabstractIt is admitted that human speech perception is a multimodal process that combines both visual and acoustic informations. In automatic speech perception, visual analysis is also crucial as it provides a complementary information in order to enhance the performances of audio systems especially in highly noisy environments. In this paper, we propose a unified probabilistic framework for speech unit recognition that combines both visual and audio informations. The method is based on the optimization of a criterion that achieves continuous speech unit segmentation and decoding using a learned (joint) phonetic-visemic model. Experiments conducted on the standard LIPS2008 dataset, show a clear and a consistent gain of our multimodal approach compared to others. Éric Benhaim, Hichem Sahbi, Guillaume Vitte |
ICASSP | 2 |
| 2014 | Semi-supervised learning using a graph-based phase field model for imbalanced data set classificationabstractIn this paper, we address the problem of semi-supervised learning for binary classification. This task is known to be challenging due to several issues including: the scarceness of labeled data, the large intra-class variability, and also the imbalanced class distributions. Our learning approach is transductive and built upon a graph-based phase field model that handles imbalanced class distributions. This method is able to encourage or penalize the memberships of data to different classes according to an explicit a priori model that avoids biased classifications. Experiments, conducted on real-world benchmarks, show the good performance of our model compared to several state of the art semi-supervised learning algorithms. Aymen El Ghoul, Hichem Sahbi |
ICASSP | 2 |
| 2014 | Modeling label dependencies in kernel learning for image annotationabstractWe introduce in this paper a novel image annotation approach based on maximum margin classification and a new class of kernels. The method goes beyond the naive use of existing kernels and their restricted combinations in order to design “model-free” transductive kernels applicable to interconnected image databases. In a first contribution of the method, we learn both a decision criterion and a kernel map that guarantee linear separability in a high dimensional space and good generalization performance. In the second contribution of this work, we extend this class of kernels in order to include label dependency statistics that model contextual relationships between concepts into images. Experiments conducted on MSRC and Corel5k databases show that our method achieves at least comparable results with related state of the art. Phong D. Vo, Hichem Sahbi |
ICIP | 2 |
| 2014 | Bags-of-daglets for action recognitionabstractRecent advances in human action recognition are focusing on fine-grained action categories in large video collections. With this current trend, one of the major issues is how to handle these large collections effectively and also efficiently. In this paper, we introduce a novel action recognition method based on mid-level components and directed acyclic graphs (DAGs). DAGs, taken from different videos, are efficiently processed in order to extract a large collection of spatio-temporal sub-patterns, of increasing complexities, referred to as daglets. The latter capture local appearances as well as causal structural relationships of interacting object-parts in video sequences. The main contribution of this work includes a daglet matching procedure and a DAG kernel that captures first and high order statistics of daglets into videos. When combined with support vector machines, this DAG kernel proved to be very effective in order to capture similarity between actions in videos and to successfully achieve action recognition on a standard challenging database. Hichem Sahbi |
ICIP | 2 |
| 2014 | Network-Dependent Image Annotation Based on Explicit Context-Dependent Kernel MapsabstractIt is commonly known that the success of support vector machines in image classification and annotation it highly dependent on the relevance of the chosen kernels. The latter, defined as symmetric positive semi-definite functions, take high values when images share similar visual content and vice-versa. However, usual kernels relying only on the visual content are not appropriate in order to capture the true semantics of images which are nowadays expressed through the rich contextual cues available in image collections. Relevant kernels should instead reserve high values not only when images share similar content but also similar context. In this paper, we introduce a novel method that upgrades usual kernels and makes them context-dependent. Our kernel solution corresponds to an optimum of an energy function that trades off content and context. We will show that the proposed kernel can be expressed with an explicit mapping which is computationally efficient and also effective for image annotation. We corroborate all these statements through our participation in the recent and challenging Image CLEF 2013 annotation benchmark which ranks our method first among 58 participants' runs. Hichem Sahbi |
ICPR | 1 |
| 2013 | Designing relevant features for visual speech recognitionabstractAutomatic speech analysis is currently evolving towards hybrid systems that combine both visual and acoustic information. This is due to limitations of existing acoustic-based approaches and the need for robust speech recognition systems working under extremely challenging conditions including noisy environments. We introduce in this paper a novel visual speech recognition approach, based on string kernels and support vector machines. The main contributions of this work include (i) the design of a similarity function, based on string kernels, that models the dynamics as well as the appearance of visual features in talking faces and (ii) a kernel combination procedure based on multiple kernel learning, that makes visual feature selection effective and also more tractable. Experiments conducted, on a standard digit database, show that the proposed algorithm outperforms current state-of-the-art methods. Éric Benhaim, Hichem Sahbi, Guillaume Vitte |
ICASSP | 2 |
| 2013 | Relevance feedback for satellite image change detectionabstractDue to the exponential growth of satellite image collections, there is an increasing need for automatic solutions that assist operators in different applications. Automatic change detection is one of these applications that received an increasing attention during the last years. Nevertheless, fully automatic solutions reach their limitation; on the one hand, it is difficult to build general decision criteria able to select area of changes for different images, and on the other hand, the relevance of changes may differ from one user to another. In this paper, we introduce an alternative change detection method based on relevance feedback. The proposed algorithm is iterative and based on a query and answer (Q&A) model that (i) asks the user the most informative questions about the relevance of his targeted changes, and (ii) according to these answers updates change detection results. Experiments conducted on large satellite images, show that indeed the approach is effective and allows the user to retrieve almost all his targeted changes while discarding the untargeted ones, with a negligible interaction effort. Hichem Sahbi |
ICASSP | 1 |
| 2013 | Directed Acyclic Graph Kernels for Action RecognitionabstractOne of the trends of action recognition consists in extracting and comparing mid-level features which encode visual and motion aspects of objects into scenes. However, when scenes contain high-level semantic actions with many interacting parts, these mid-level features are not sufficient to capture high level structures as well as high order causal relationships between moving objects resulting into a clear drop in performances. In this paper, we address this issue and we propose an alternative action recognition method based on a novel graph kernel. In the main contributions of this work, we first describe actions in videos using directed a cyclic graphs (DAGs), that naturally encode pair wise interactions between moving object parts, and then we compare these DAGs by analyzing the spectrum of their sub-patterns that capture complex higher order interactions. This extraction and comparison process is computationally tractable, resulting from the a cyclic property of DAGs, and it also defines a positive semi-definite kernel. When plugging the latter into support vector machines, we obtain an action recognition algorithm that overtakes related work, including graph-based methods, on a standard evaluation dataset. Hichem Sahbi |
ICCV | 2 |
| 2013 | Explicit Context-Aware Kernel Map Learning for Image Annotation
Hichem Sahbi |
ICVS | 1 |
| 2013 | Interactive satellite image change detectionabstractWith the rapid growth of satellite image collections there is an urgent need for automatic solutions that handle different applications. Change detection is one of these applications that recently received a particular attention but fully automatic solutions are still sensitive to the presence of many irrelevant changes (due to illumination, noise, etc.) as well as untargeted changes. In this paper, we introduce an alternative and semi-automatic change detection method based on relevance feedback. The proposed algorithm is interactive and based on a query and answer (Q&A) model that asks the user the most informative questions about the relevance of his targeted changes, and according to these answers updates change detection results. Experiments conducted on large satellite images, show that indeed the approach is accurate and allows the user to find almost all his targeted changes while discarding the untargeted ones, with a negligible interaction effort. Hichem Sahbi |
IGARSS | 1 |
| 2013 | Semantic subspace learning for mental search in satellite imagesabstractIn this paper we introduced a new approach, for mental satellite image search and visualization. We addressed the difficulties when visualizing high dimensional data, and we proposed a semantic subspace learning approach for effective exploration of large scale satellite image data. Our formulation is based on a quadratic programming algorithm that easily extend to large scale data. By plugging the embedding solution of this algorithm into Spacious, users can specify objects of interest, as mixtures of predefined semantics in the learned subspace, and retrieve their targets. As a future work, we are currently investigating the use of other low level features as well as the combination of our semantic subspace learning method with relevance feedback. Phong D. Vo, Hichem Sahbi |
IGARSS | 2 |
| 2013 | Spacious: an interactive mental search interfaceabstractWe introduce in this work a novel approach for semantic indexing and mental image search. Given semantic concepts defined by few training examples, our formulation is transductive and learns a mapping from an initial ambient space, related to low level visual features, to an output space spanned by a well defined semantic basis where data can be easily explored. With this method, searching for a mental visual target reduces to scanning data according to their coordinates in the learned semantic space. We illustrate the proposed method through our graphical user interface "Spacious", for the purpose of visualization and interactive navigation in generic image databases and satellite images. Phong D. Vo, Hichem Sahbi |
SIGIR | 2 |
| 2013 | Context-Dependent Logo Matching and RecognitionabstractWe contribute, through this paper, to the design of a novel variational framework able to match and recognize multiple instances of multiple reference logos in image archives. Reference logos and test images are seen as constellations of local features (interest points, regions, etc.) and matched by minimizing an energy function mixing: 1) a fidelity term that measures the quality of feature matching, 2) a neighborhood criterion that captures feature co-occurrence/geometry, and 3) a regularization term that controls the smoothness of the matching solution. We also introduce a detection/recognition procedure and study its theoretical consistency. Finally, we show the validity of our method through extensive experiments on the challenging MICC-Logos dataset. Our method overtakes, by 20%, baseline as well as state-of-the-art matching/recognition procedures. Hichem Sahbi, Lamberto Ballan, Giuseppe Serra 0001, Alberto Del Bimbo |
IEEE Trans. Image Process. | 1 |
| 2012 | Transductive Kernel Map Learning and Its Application Image AnnotationabstractInternational audience Dinh-Phong Vo, Hichem Sahbi |
BMVC | 2 |
| 2012 | Transductive inference & kernel design for object class segmentationabstractTransductive inference techniques are nowadays becoming standard in machine learning due to their relative success in solving many real-world applications. Among them, kernel-based methods are particularly interesting but their success remains highly dependent on the choice of kernels. The latter are usually handcrafted or designed in order to capture better similarity in training data. In this paper, we introduce a novel transductive learning algorithm for kernel design and classification. Our approach is based on the minimization of an energy function mixing i) a reconstruction term that factorizes a matrix of input data as a product of a learned dictionary and a learned kernel map ii) a fidelity term that ensures consistent label predictions with those provided in a ground-truth and iii) a smoothness term which guarantees similar labels for neighboring data and allows us to iteratively diffuse kernel maps and labels from labeled to unlabeled data. Solving this minimization problem makes it possible to learn both a decision criterion and a kernel map that guarantee linear separability in a high dimensional space and good generalization performance. Experiments conducted on object class segmentation, show improvements with respect to baseline as well as related work on the challenging VOC database. Dinh-Phong Vo, Hichem Sahbi |
ICIP | 2 |
| 2012 | Spatio-temporal interaction for aerial video change detectionabstractWith the growing capacity of video devices, human operators are nowadays overwhelmed by the huge volumes of data generated in different applications including surveillance. Therefore, automatic video processing techniques are required in order to filter out uninteresting data and to focus the attention of operators. However, reliability is still a challenging problem. In this paper, we show how spatio-temporal redundancy may be exploited to enhance the accuracy of automatic change detection in aerial videos. More precisely, we present an algorithm based on Belief Propagation in order to improve spatio-temporal consistency between successive change detection results. Experiments demonstrate that our method leads to increased accuracy in change detection. Nicolas Bourdis, Denis Marraud, Hichem Sahbi |
IGARSS | 3 |
| 2012 | Camera pose estimation using Visual Servoing for aerial video change detectionabstractAerial image change detection is highly dependent on the accuracy of camera pose and may be subject to false alarms caused by misregistrations. In this paper, we present a novel pose estimation approach based on Visual Servoing that combines aerial videos with 3D models. Firstly, we introduce a formulation that relates image registration with the poses of a moving camera observing a 3D plane. Then, we combine this formulation with Newton's algorithm in order to estimate camera poses in a given aerial video. Finally, we present and discuss experimental results which demonstrate the robustness and the accuracy of our method. Nicolas Bourdis, Denis Marraud, Hichem Sahbi |
IGARSS | 3 |
| 2012 | Mid-level features and spatio-temporal context for activity recognition
Fei Yuan 0003, Gui-Song Xia, Hichem Sahbi, Véronique Prinet |
Pattern Recognit. | 3 |
| 2011 | Superpixel-based object class segmentation using conditional random fieldsabstractObject class segmentation (OCS) is a key issue in semantic scene labeling and understanding. Its general principle consists of naming object entities into scenes according to their intrinsic visual features as well as their dependencies. In this paper, we propose a novel superpixel-based framework for object class segmentation using conditional random fields (CRFs). The framework proceeds in two steps: (i) superpixel label estimate; and (ii) CRF label propagation. Step (i) is achieved using multi-scale boosted classifiers over superpixels and makes it possible to find coarse estimates of initial labels. Fine labeling is afterward achieved in Step (ii), using an anisotropic contrast sensitive pairwise function designed in order to characterize the intrinsic interaction potentials between objects according to 4-neighborhoods. Finally, a higher-order criterion is applied to enforce region label consistency of OCS. Experimental results demonstrate the effectiveness of the proposed framework. Xi Li 0001, Hichem Sahbi |
ICASSP | 2 |
| 2011 | Constrained optical flow for aerial image change detectionabstractNowadays, the amount of video data acquired for observation or surveillance applications is overwhelming. Due to these huge volumes of video data, focusing the attention of operators on "areas of interest" requires change detection algorithms. In the particular task of aerial observation, camera motion and viewpoint differences introduce parallax effects, which may substantially affect the reliability and the efficiency of automatic change detection. In this paper, we introduce a novel approach for change detection that considers the geometric aspects of camera sensors as well as the statistical properties of changes. Indeed, our method is based on optical flow matching, constrained by the epipolar geometry, and combined with a statistical change decision criterion. The good performance of our method is demonstrated through our new public Aerial Imagery Change Detection (AICD) dataset of labeled aerial images. Nicolas Bourdis, Denis Marraud, Hichem Sahbi |
IGARSS | 3 |
| 2011 | Context-Dependent Kernels for Object ClassificationabstractKernels are functions designed in order to capture resemblance between data and they are used in a wide range of machine learning techniques, including support vector machines (SVMs). In their standard version, commonly used kernels such as the Gaussian one show reasonably good performance in many classification and recognition tasks in computer vision, bioinformatics, and text processing. In the particular task of object recognition, the main deficiency of standard kernels such as the convolution one resides in the lack in capturing the right geometric structure of objects while also being invariant. We focus in this paper on object recognition using a new type of kernel referred to as "context dependent.” Objects, seen as constellations of interest points, are matched by minimizing an energy function mixing 1) a fidelity term which measures the quality of feature matching, 2) a neighborhood criterion which captures the object geometry, and 3) a regularization term. We will show that the fixed point of this energy is a context-dependent kernel which is also positive definite. Experiments conducted on object recognition show that when plugging our kernel into SVMs, we clearly outperform SVMs with context-free kernels. Hichem Sahbi, Jean-Yves Audibert, Renaud Keriven |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2010 | Context-Based Support Vector Machines for Interconnected Image Annotation
Hichem Sahbi, Xi Li 0001 |
ACCV (1) | 1 |
| 2010 | Network-dependent kernels for image rankingabstractThe exponential growth of social networks (SN) currently makes them the standard way to share and explore data where users put informations (images, text, audio,...) and refer to other contents. This creates connected networks whose links provide valuable informations in order to enhance the performance of many tasks in information retrieval including ranking and annotation. We introduce in this paper a novel image retrieval framework based on a new class of kernels referred to as “network-dependent”. The main contribution of our method includes (i) a variational framework which helps designing a kernel using both the intrinsic image features and the underlying contextual informations resulting from different (e.g. social) links and (ii) the proof of convergence of the kernel to a fixed-point, that is positive definite and thus associated with a reproducing kernel Hilbert space (RKHS). Experiments conducted on different ground truths, including the ImageClef/Flickr set, show the outperformance and the substantial gain of our ranking kernel with respect to the use of classic “network-free” kernels. Hichem Sahbi, Jean-Yves Audibert |
ICIP | 1 |
| 2010 | Context dependent SVMs for interconnected image network annotationabstractThe exponential growth of interconnected networks, such as Flickr, currently makes them the standard way to share and explore data where users put contents and refer to others. These interconnections create valuable information in order to enhance the performance of many tasks in information retrieval including ranking and annotation. We introduce in this paper a novel image annotation framework based on support vector machines (SVMs) and a new class of kernels referred to as context-dependent. The method goes beyond the naive use of the intrinsic low level features (such as color, texture, shape, etc.) and context-free kernels, in order to design a kernel function applicable to interconnected databases such as social networks. The main contribution of our method includes a variational framework which helps designing this function using both intrinsic features and the underlying contextual information. This function also converges to a positive definite fixed-point, usable for SVM training and other kernel methods. When plugged in SVMs, our context-dependent kernel consistently improves the performance of image annotation, compared to context-free kernels, on hundreds of thousands of Flickr images. Hichem Sahbi, Xi Li 0001 |
ACM Multimedia | 1 |
| 2009 | Multi-view object matching and tracking using canonical correlation analysisabstractMulti-view tracking of objects in video surveillance consists in segmenting and automatically following them through different camera views. This may be achieved using geometric methods, e.g. by calibrating camera sensors and using their transformation matrices. However, in practice the precision of calibration is a major issue when trying to achieve this task robustly. In this paper, we present an alternative framework for multi-view object matching and tracking based on canonical correlation analysis. Our method is purely statistical and encodes intrinsic object appearances while being view-point invariant. We will show that our technique is (i) easy-to-set (ii) theoretically well grounded and (iii) provides robust matching and tracking results for traffic surveillance. Marin Ferecatu, Hichem Sahbi |
ICIP | 2 |
| 2009 | Content-based 3D object retrieval using 2D viewsabstract2D techniques have recently emerged as an important boost for 3D objects content-based retrieval in many real world applications such as photography, art, archeology and geolocalization thanks to its several complementary aspects. We introduce in this paper a new framework for 3D objects content-based retrieval based on a 2D photography approach. A new alignment process that is able to find canonical views consistently through scenes/objects and a new coarse-to-fine description and matching method used for ranking are our contributions. The results are presented through an international benchmarking and showing clearly the good performance of our framework with respect to the other participants. Thibault Napoléon, Hichem Sahbi |
ICIP | 2 |
| 2008 | Context-dependent kernel design for object matching and recognitionabstractThe success of kernel methods including support vector networks (SVMs) strongly depends on the design of appropriate kernels. While initially kernels were designed in order to handle fixed-length data, their extension to unordered, variable-length data became more than necessary for real pattern recognition problems such as object recognition and bioinformatics. We focus in this paper on object recognition using a new type of kernel referred to as ldquocontext-dependentrdquo. Objects, seen as constellations of local features (interest points, regions, etc.), are matched by minimizing an energy function mixing (1) a fidelity term which measures the quality of feature matching, (2) a neighborhood criteria which captures the object geometry and (3) a regularization term. We will show that the fixed-point of this energy is a ldquocontext-dependentrdquo kernel (ldquoCDKrdquo) which also satisfies the Mercer condition. Experiments conducted on object recognition show that when plugging our kernel in SVMs, we clearly outperform SVMs with ldquocontext-freerdquo kernels. Hichem Sahbi, Jean-Yves Audibert, Jaonary Rabarisoa, Renaud Keriven |
CVPR | 1 |
| 2008 | Manifold learning using robust Graph Laplacian for interactive image searchabstractInteractive image search or relevance feedback is the process which helps a user refining his query and finding difficult target categories. This consists in partially labeling a very small fraction of an image database and iteratively refining a decision rule using both the labeled and unlabeled data. Training of this decision rule is referred to as transductive learning. Our work is an original approach for relevance feedback based on Graph Laplacian. We introduce a new Graph Laplacian which makes it possible to robustly learn the embedding, of the manifold enclosing the dataset, via a diffusion map. Our approach is three-folds: it allows us (i) to integrate all the unlabeled images in the decision process (ii) to robustly capture the topology of the image set and (iii) to perform the search process inside the manifold. Relevance feedback experiments were conducted on simple databases including Olivetti and Swedish as well as challenging and large scale databases including Corel. Comparisons show clear and consistent gain, of our graph Laplacian method, with respect to state-of-the art relevance feedback approaches. Hichem Sahbi, Patrick Etyngier, Jean-Yves Audibert, Renaud Keriven |
CVPR | 1 |
| 2008 | Graph laplacian for interactive image retrievalabstractInteractive image search or relevance feedback is the process which helps a user refining his query and finding difficult target categories. This consists in a step-by-step labeling of a very small fraction of an image database and iteratively refining a decision rule using both the labeled and unlabeled data. Training of this decision rule is referred to as transductive learning. Our work is an original approach for relevance feedback based on Graph Laplacian. We introduce a new Graph Laplacian which makes it possible to robustly learn the embedding of the manifold enclosing the dataset via a diffusion map. Our approach is two-folds: it allows us (i) to integrate all the unlabeled images in the decision process and (ii) to robustly capture the topology of the image set. Relevance feedback experiments were conducted on simple databases including Olivetti and Swedish as well as challenging and large scale databases including Corel. Comparisons show clear and consistent gain of our graph Laplacian method with respect to state-of-the art relevance feedback approaches. Hichem Sahbi, Patrick Etyngier, Jean-Yves Audibert, Renaud Keriven |
ICASSP | 1 |
| 2008 | Robust matching and recognition using context-dependent kernelsabstractThe success of kernel methods including support vector machines (SVMs) strongly depends on the design of appropriate kernels. While initially kernels were designed in order to handle fixed-length data, their extension to unordered, variable-length data became more than necessary for real pattern recognition problems such as object recognition and bioinformatics. Hichem Sahbi, Jean-Yves Audibert, Jaonary Rabarisoa, Renaud Keriven |
ICML | 1 |
| 2008 | A particular Gaussian mixture model for clustering and its application to image retrieval
Hichem Sahbi |
Soft Comput. | 1 |
| 2007 | Consensus Network Decoding for Statistical Machine Translation System CombinationabstractThis paper presents a simple and robust consensus decoding approach for combining multiple machine translation (MT) system outputs. A consensus network is constructed from an N-best list by aligning the hypotheses against an alignment reference, where the alignment is based on minimising the translation edit rate (TER). The minimum Bayes risk (MBR) decoding technique is investigated for the selection of an appropriate alignment reference. Several alternative decoding strategies proposed to retain coherent phrases in the original translations. Experimental results are presented primarily based on three-way combination of Chinese-English translation outputs, and also presents results for six-way system combination. It is shown that worthwhile improvements in translation performance can be obtained using the methods discussed. Khe Chai Sim, William J. Byrne, Mark J. F. Gales, Hichem Sahbi, Philip C. Woodland |
ICASSP (4) | 4 |
| 2007 | Graph-Cut Transducers for Relevance Feedback in Content Based Image RetrievalabstractClosing the semantic gap in content based image retrieval (CBIR) basically requires the knowledge of the user's intention which is usually translated into a sequence of questions and answers (Q&A). The user's feedback to these questions provides a CBIR system with a partial labeling of the data and makes it possible to iteratively refine a decision rule on the unlabeled data. Training of this decision rule is referred to as transductive learning. This work is an original approach to relevance feedback (RF) based on graph-cuts. Training consists in implicitly modeling the manifold enclosing both the labeled and unlabeled dataset and finding a partition of this manifold using a min-cut. The contribution of this work is two-fold (i) this is the first comprehensive study of relevance feedback using graph cuts and (ii) our RF model exploits the structure of the data manifold by considering also the structure of the unlabeled data. Experiments conducted on generic as well as specific databases show that our graph-cut based approach is very effective, outperforms other existing methods and makes it possible to converge to almost all the images of the user's "class of interest" with a very small labeling effort. A demo is available through our image retrieval tool kit (IRTK). Hichem Sahbi, Jean-Yves Audibert, Renaud Keriven |
ICCV | 1 |
| 2007 | Kernel PCA for similarity invariant shape recognition
Hichem Sahbi |
Neurocomputing | 1 |
| 2006 | A Hierarchy of Support Vector Machines for Pattern DetectionabstractWe introduce a computational design for pattern detection based on a tree-structured network of support vector machines (SVMs). An SVM is associated with each cell in a recursive partitioning of the space of patterns (hypotheses) into increasingly finer subsets. The hierarchy is traversed coarse-to-fine and each chain of positive responses from the root to a leaf constitutes a detection. Our objective is to design and build a network which balances overall error and computation. Initially, SVMs are constructed for each cell with no constraints. This "free network" is then perturbed, cell by cell, into another network, which is "graded" in two ways: first, the number of support vectors of each SVM is reduced (by clustering) in order to adjust to a pre-determined, increasing function of cell depth; second, the decision boundaries are shifted to preserve all positive responses from the original set of training data. The limits on the numbers of clusters (virtual support vectors) result from minimizing the mean computational cost of collecting all detections subject to a bound on the expected number of false positives. When applied to detecting faces in cluttered scenes, the patterns correspond to poses and the free network is already faster and more accurate than applying a single pose-specific SVM many times. The graded network promotes very rapid processing of background regions while maintaining the discriminatory power of the free network. Hichem Sahbi, Donald Geman |
J. Mach. Learn. Res. | 1 |
| 2005 | Validity of Fuzzy Clustering Using Entropy RegularizationabstractWe introduce in this paper a new formulation of the regularized fuzzy c-means (FCM) algorithm which allows us to find automatically the actual number of clusters. The approach is based on the minimization of an objective function which mixes, via a particular parameter, a classical FCM term and a new entropy regularizer. The main contribution of the method is the introduction of a new exponential form of the fuzzy memberships which ensures the consistency of their bounds and makes it possible to interpret the mixing parameter as the variance (or scale) of the clusters. This variance closely related to the number of clusters, provides us with an intuitive and an easy to set parameter. We will discuss the proposed approach from the regularization point-of-view and we will demonstrate its validity both analytically and experimentally. We will show an extension of the method to nonlinearly separable data. Finally, we will illustrate preliminary results both on simple toy examples as well as database categorization problems Hichem Sahbi, Nozha Boujemaa |
FUZZ-IEEE | 1 |
| 2002 | Face detection using coarse-to-fine support vector classifiersabstractWe describe a new face detection algorithm based on a hierarchy of support vector classifiers (SVM) designed for efficient computation. The hierarchy serves as a platform for a coarse-to-fine search for faces: most of the image is quickly rejected as "background" and the processing naturally concentrates on regions containing faces and face-like structures. The hierarchy is tree-structured: In proceeding from the root to the leaves, the SVM gradually increase in complexity (measured by the number of support vectors) and discrimination (measured by the false alarm rate), but decrease in the level of invariance. Reduced complexity is achieved by clustering support vectors and shifting the decision boundary in order to satisfy a "conservation hypothesis" that preserves positive responses from the original set of support vectors. The computation is organized as a depth-first search and cancel strategy. The gain in efficiency is enormous. Hichem Sahbi, Donald Geman, Nozha Boujemaa |
ICIP (3) | 1 |
| 2001 | Robust matching by dynamic space warping for accurate face recognitionabstractThe utility of face recognition for multimedia indexing is enhanced by using accurate detection and alignment of salient invariant face features. The face recognition can be performed using template matching or a feature-based approach, but both these methods suffer from occlusion and require an a priori model for extracting information. To avoid these drawbacks, we present a complete scheme for face recognition based on salient feature extraction in challenging conditions, which is performed without an a priori or learned model. These features are used in a matching process that overcomes occlusion effects using the dynamic space warping which aligns each feature in the query image, if possible, with its corresponding feature in the gallery set. Thus, we make face recognition robust to low frequency variations (like the presence of occlusion, etc) as well as to high frequency variations (like expression, gender, etc). A maximum likelihood scheme is used to make the recognition process more precise, as is shown in the experiments. Hichem Sahbi, Nozha Boujemaa |
ICIP (1) | 1 |
| 2000 | From coarse to fine skin and face detection
Hichem Sahbi, Nozha Boujemaa |
ACM Multimedia | 1 |