Frédéric Precioso

dblp:83/1407 · DBLP profile ↗
← Back
71ranked-venue papers
5as first author
20since 2021 · last 2026
0000-0001-8712-1443ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 43 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 35 · 17 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021Software engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 2Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Lightweight Model Augmented by Expert Knowledge in Realistic Clinical Decision-Making on Colorectal Cancer Treatment
Célia D'Cruz, Frédéric Precioso, Jean-Marc Bereder, Michel Riveill
ICPR (6)2
2025 A Layer Selection Approach to Test Time Adaptation
abstract
Test Time Adaptation (TTA) addresses the problem of distribution shift by adapting a pretrained model to a new domain during inference. When faced with challenging shifts, most methods collapse and perform worse than the original pretrained model. In this paper, we find that not all layers are equally receptive to the adaptation, and the layers with the most misaligned gradients often cause performance degradation. To address this, we propose GALA, a novel layer selection criterion to identify the most beneficial updates to perform during test time adaptation. This criterion can also filter out unreliable samples with noisy gradients. Its simplicity allows seamless integration with existing TTA loss functions, thereby preventing degradation and focusing adaptation on the most trainable layers. This approach also helps to regularize adaptation to preserve the pretrained features, which are crucial for handling unseen domains. Through extensive experiments, we demonstrate that the proposed layer selection framework improves the performance of existing TTA approaches across multiple datasets, domain shifts, model architectures, and TTA losses.
Sabyasachi Sahoo, Mostafa ElAraby, Jonas Ngnawé, Yann Pequignot, Frédéric Precioso, Christian Gagné 0001
AAAI5
2025 Re-examining Concept-based Explainable Models for Multimodal Interpretative Tasks
abstract
Concept-based models have been proposed as a new line of research for explainable by-design deep learning models. However, those models show their whole power when applied to benchmarks where the concepts are well defined and the concepts' attributes easily extractable from the raw data. In this paper, we challenge the most recent concept-based model initially developed for image classification, on more complex interpretative tasks from a recently proposed video benchmark where they perform poorly. We conduct a root cause analysis of the poor performances of state-of-the-art explainable concept-based models for these multimodal interpretative tasks, and propose adaptations to design robust explainable models for detecting character objectification in this novel challenging video benchmark. We show that the optimal architectural choice may vary depending on the modality setting, thereby showing that designing multimodal concept-based approaches remains an open challenge and calls for further investigation.
Julie Tores, Elisa Ancarani, Rémy Sun, Lucile Sassatelli, Hui-Yin Wu, Frédéric Precioso
ACM Multimedia6
2025 WolBanking77: Wolof Banking Speech Intent Classification Dataset
abstract
Intent classification models have made a significant progress in recent years. However, previous studies primarily focus on high-resource language datasets, which results in a gap for low-resource languages and for regions with high rates of illiteracy, where languages are more spoken than read or written. This is the case in Senegal, for example, where Wolof is spoken by around 90\% of the population, while the national illiteracy rate remains at of 42\%. Wolof is actually spoken by more than 10 million people in West African region. To address these limitations, we introduce the Wolof Banking Speech Intent Classification Dataset (WolBanking77), for academic research in intent classification. WolBanking77 currently contains 9,791 text sentences in the banking domain and more than 4 hours of spoken sentences. Experiments on various baselines are conducted in this work, including text and voice state-of-the-art models. The results are very promising on this current dataset. In addition, this paper presents an in-depth examination of the dataset’s contents. We report baseline F1-scores and word error rates metrics respectively on NLP and ASR models trained on WolBanking77 dataset and also comparisons between models. Dataset and code available at: wolbanking77.
Abdou Karim Kandji, Frédéric Precioso, Cheikh Ba, Samba Ndiaye, Augustin Ndione
NeurIPS2
2025 Mind the Map! Accounting for Existing Maps When Estimating Online HDMaps from Sensors
abstract
While HDMaps are a crucial component of autonomous driving, they are expensive to acquire and maintain. Estimating these maps from sensors therefore promises to significantly lighten costs. These estimations however overlook existing HDMaps, with current methods at most geolocalizing low quality maps or considering a general database of known maps. In this paper, we propose to account for existing maps of the precise situation studied when estimating HDMaps. To prove this, we identify 3 reasonable types of useful existing maps (minimalist, noisy, and outdated). We then introduce MapEX, a novel online HDMap estimation framework that accounts for existing maps. MapEX achieves this by encoding map elements into query tokens and by refining the matching algorithm used to train classic query based map estimation models. We demonstrate that MapEX brings significant improvements on the nuScenes dataset. For instance, MapEX - given noisy maps - improves by 38% over the MapTRv2 detector it is based on and by 8% over the current SOTA.
Rémy Sun, Diane Lingrand, Frédéric Precioso
WACV4
2024 Visual Objectification in Films: Towards a New AI Task for Video Interpretation
abstract
In film gender studies, the concept of “male gaze” refers to the way the characters are portrayed on-screen as objects of desire rather than subjects. In this article, we introduce a novel video-interpretation task, to detect character objectification in films. The purpose is to reveal and quantify the usage of complex temporal patterns operated in cinema to produce the cognitive perception of objectification. We introduce the ObyGaze12 dataset, made of 1914 movie clips densely annotated by experts for objectification concepts identified in film studies and psychology. We evaluate recent vision models, show the feasibility of the task and where the challenges remain with concept bottleneck models. Our new dataset and code are made available to the community.
Julie Tores, Lucile Sassatelli, Hui-Yin Wu, Clement Bergman, Lea Andolfi, Victor Ecrement, Frédéric Precioso, Thierry Devars, Magali Guaresi, Virginie Julliard, Sarah Lecossais
CVPR7
2024 Domain-Specific Long Text Classification from Sparse Relevant Information
abstract
Large Language Models have undoubtedly revolutionized the Natural Language Processing field, the current trend being to promote one-model-for-all tasks (sentiment analysis, translation, etc.). However, the statistical mechanisms at work in the larger language models struggle to exploit the relevant information when it is very sparse, when it is a weak signal. This is the case, for example, for the classification of long domain-specific documents, when the relevance relies on a single relevant word or on very few relevant words from technical jargon. In the medical domain, it is essential to determine whether a given report contains critical information about a patient’s condition. This critical information is often based on one or few specific isolated terms. In this paper, we propose a hierarchical model which exploits a short list of potential target terms to retrieve candidate sentences and represent them into the contextualized embedding of the target term(s) they contain. A pooling of the term(s) embedding(s) entails the document representation to be classified. We evaluate our model on one public medical document benchmark in English and on one private French medical dataset. We show that our narrower hierarchical model is better than larger language models for retrieving relevant long documents in a domain-specific context.
Célia D'Cruz, Jean-Marc Bereder, Frédéric Precioso, Michel Riveill
ECAI3
2024 Attention Meets Post-hoc Interpretability: A Mathematical Perspective
abstract
Attention-based architectures, in particular transformers, are at the heart of a technological revolution. Interestingly, in addition to helping obtain state-of-the-art results on a wide range of applications, the attention mechanism intrinsically provides meaningful insights on the internal behavior of the model. Can these insights be used as explanations? Debate rages on. In this paper, we mathematically study a simple attention-based architecture and pinpoint the differences between post-hoc and attention-based explanations. We show that they provide quite different results, and that, despite their limitations, post-hoc methods are capable of capturing more useful insights than merely examining the attention weights.
Gianluigi Lopardo, Frédéric Precioso, Damien Garreau
ICML2
2024 Detecting Brittle Decisions for Free: Leveraging Margin Consistency in Deep Robust Classifiers
abstract
Despite extensive research on adversarial training strategies to improve robustness, the decisions of even the most robust deep learning models can still be quite sensitive to imperceptible perturbations, creating serious risks when deploying them for high-stakes real-world applications. While detecting such cases may be critical, evaluating a model's vulnerability at a per-instance level using adversarial attacks is computationally too intensive and unsuitable for real-time deployment scenarios. The input space margin is the exact score to detect non-robust samples and is intractable for deep neural networks. This paper introduces the concept of margin consistency -- a property that links the input space margins and the logit margins in robust models -- for efficient detection of vulnerable samples. First, we establish that margin consistency is a necessary and sufficient condition to use a model's logit margin as a score for identifying non-robust samples. Next, through comprehensive empirical analysis of various robustly trained models on CIFAR10 and CIFAR100 datasets, we show that they indicate high margin consistency with a strong correlation between their input space margins and the logit margins. Then, we show that we can effectively use the logit margin to confidently detect brittle decisions with such models. Finally, we address cases where the model is not sufficiently margin-consistent by learning a pseudo-margin from the feature representation. Our findings highlight the potential of leveraging deep representations to efficiently assess adversarial vulnerability in deployment scenarios.
Jonas Ngnawé, Sabyasachi Sahoo, Yann Pequignot, Frédéric Precioso, Christian Gagné 0001
NeurIPS4
2023 A Sea of Words: An In-Depth Analysis of Anchors for Text Data
abstract
Anchors (Ribeiro et al., 2018) is a post-hoc, rule-based interpretability method. For text data, it proposes to explain a decision by highlighting a small set of words (an anchor) such that the model to explain has similar outputs when they are present in a document. In this paper, we present the first theoretical analysis of Anchors, considering that the search for the best anchor is exhaustive. After formalizing the algorithm for text classification, we present explicit results on different classes of models when the vectorization step is TF-IDF, and words are replaced by a fixed out-of-dictionary token when removed. Our inquiry covers models such as elementary if-then rules and linear classifiers. We then leverage this analysis to gain insights on the behavior of Anchors for any differentiable classifiers. For neural networks, we empirically show that the words corresponding to the highest partial derivatives of the model with respect to the input, reweighted by the inverse document frequencies, are selected by Anchors.
Gianluigi Lopardo, Frédéric Precioso, Damien Garreau
AISTATS2
2023 Interpretable Neural-Symbolic Concept Reasoning
abstract
Deep learning methods are highly accurate, yet their opaque decision process prevents them from earning full human trust. Concept-based models aim to address this issue by learning tasks based on a set of human-understandable concepts. However, state-of-the-art concept-based models rely on high-dimensional concept embedding representations which lack a clear semantic meaning, thus questioning the interpretability of their decision process. To overcome this limitation, we propose the Deep Concept Reasoner (DCR), the first interpretable concept-based model that builds upon concept embeddings. In DCR, neural networks do not make task predictions directly, but they build syntactic rule structures using concept embeddings. DCR then executes these rules on meaningful concept truth degrees to provide a final interpretable and semantically-consistent prediction in a differentiable manner. Our experiments show that DCR: (i) improves up to +25% w.r.t. state-of-the-art interpretable concept-based models on challenging benchmarks (ii) discovers meaningful logic rules matching known ground truths even in the absence of concept supervision during training, and (iii), facilitates the generation of counterfactual examples providing the learnt rules as guidance.
Pietro Barbiero, Gabriele Ciravegna, Francesco Giannini, Mateo Espinosa Zarlenga, Lucie Charlotte Magister, Alberto Paolo Tonda, Pietro Liò, Frédéric Precioso, Mateja Jamnik, Giuseppe Marra
ICML8
2023 Another Point of View on Visual Speech Recognition
abstract
International audience
Baptiste Pouthier, Laurent Pilati, Giacomo Valenti, Charles Bouveyron, Frédéric Precioso
INTERSPEECH5
2023 Knowledge-Driven Active Learning
abstract
In the last few years, Deep Learning models have become increasingly popular. However, their deployment is still precluded in those contexts where the amount of supervised data is limited and manual labelling expensive. Active learning strategies aim at solving this problem by requiring supervision only on few unlabelled samples, which improve the most model performances after adding them to the training set. Most strategies are based on uncertain sample selection, and even often restricted to samples lying close to the decision boundary. Here we propose a very different approach, taking into consideration domain knowledge. Indeed, in the case of multi-label classification, the relationships among classes offer a way to spot incoherent predictions, i.e., predictions where the model may most likely need supervision. We have developed a framework where first-order-logic knowledge is converted into constraints and their violation is checked as a natural guide for sample selection. We empirically demonstrate that knowledge-driven strategy outperforms standard strategies, particularly on those datasets where domain knowledge is complete. Furthermore, we show how the proposed approach enables discovering data distributions lying far from training data. Finally, the proposed knowledge-driven strategy can be also easily used in object-detection problems where standard uncertainty-based techniques are difficult to apply.
Gabriele Ciravegna, Frédéric Precioso, Alessandro Betti, Kevin Mottin, Marco Gori
ECML/PKDD (1)2
2022 Revisiting Artistic Style Transfer for Data Augmentation in A Real-Case Scenario
abstract
A tremendous number of techniques have been proposed to transfer artistic style from one image to another. In particular, techniques exploiting neural representation of data; from Convolutional Neural Networks to Generative Adversarial Networks. However, most of these techniques do not accurately account for the semantic information related to the objects present in both images or require a considerable training set. In this paper, we provide a data augmentation technique that is as faithful as possible to the style of the reference artist, while requiring as few training samples as possible, as artworks containing the same semantics of an artist are usually rare. Hence, this paper aims to improve the state-of-the-art by first applying semantic segmentation on both images to then transfer the style from the painting to a photo while preserving common semantic regions. The method is exemplified on Van Gogh’s paintings, shown to be challenging to segment.
Stefano D'Angelo, Frédéric Precioso, Fabien Gandon
ICIP2
2022 Generalised Mutual Information for Discriminative Clustering
abstract
In the last decade, recent successes in deep clustering majorly involved the mutual information (MI) as an unsupervised objective for training neural networks with increasing regularisations. While the quality of the regularisations have been largely discussed for improvements, little attention has been dedicated to the relevance of MI as a clustering objective. In this paper, we first highlight how the maximisation of MI does not lead to satisfying clusters. We identified the Kullback-Leibler divergence as the main reason of this behaviour. Hence, we generalise the mutual information by changing its core distance, introducing the generalised mutual information (GEMINI): a set of metrics for unsupervised neural network training. Unlike MI, some GEMINIs do not require regularisations when training. Some of these metrics are geometry-aware thanks to distances or kernels in the data space. Finally, we highlight that GEMINIs can automatically select a relevant number of clusters, a property that has been little studied in deep clustering context where the number of clusters is a priori unknown.
Louis Ohl, Pierre-Alexandre Mattei, Charles Bouveyron, Warith Harchaoui, Mickaël Leclercq, Arnaud Droit, Frédéric Precioso
NeurIPS7
2022 Concept Embedding Models: Beyond the Accuracy-Explainability Trade-Off
abstract
Deploying AI-powered systems requires trustworthy models supporting effective human interactions, going beyond raw prediction accuracy. Concept bottleneck models promote trustworthiness by conditioning classification tasks on an intermediate level of human-like concepts. This enables human interventions which can correct mispredicted concepts to improve the model's performance. However, existing concept bottleneck models are unable to find optimal compromises between high task accuracy, robust concept-based explanations, and effective interventions on concepts---particularly in real-world conditions where complete and accurate concept supervisions are scarce. To address this, we propose Concept Embedding Models, a novel family of concept bottleneck models which goes beyond the current accuracy-vs-interpretability trade-off by learning interpretable high-dimensional concept representations. Our experiments demonstrate that Concept Embedding Models (1) attain better or competitive task accuracy w.r.t. standard neural models without concepts, (2) provide concept representations capturing meaningful semantics including and beyond their ground truth labels, (3) support test-time concept interventions whose effect in test accuracy surpasses that in standard concept bottleneck models, and (4) scale to real-world conditions where complete concept supervisions are scarce.
Mateo Espinosa Zarlenga, Pietro Barbiero, Gabriele Ciravegna, Giuseppe Marra, Francesco Giannini, Michelangelo Diligenti, Zohreh Shams, Frédéric Precioso, Stefano Melacci, Adrian Weller, Pietro Liò, Mateja Jamnik
NeurIPS8
2022 SMACE: A New Method for the Interpretability of Composite Decision Systems
Gianluigi Lopardo, Damien Garreau, Frédéric Precioso, Greger Ottosson
ECML/PKDD (1)3
2022 Semi-supervised consensus clustering based on closed patterns
Nicolas Pasquier, Frédéric Precioso
Knowl. Based Syst.3
2022 TRACK: A New Method From a Re-Examination of Deep Architectures for Head Motion Prediction in 360${}^{\circ }$∘ Videos
abstract
videos, with 2 modalities only: the past user's positions and the video content (not knowing other users' traces). We make two main contributions. First, we re-examine existing deep-learning approaches for this problem and identify hidden flaws from a thorough root-cause analysis. Second, from the results of this analysis, we design a new proposal establishing state-of-the-art performance. First, re-assessing the existing methods that use both modalities, we obtain the surprising result that they all perform worse than baselines using the user's trajectory only. A root-cause analysis of the metrics, datasets and neural architectures shows in particular that (i) the content can inform the prediction for horizons longer than 2 to 3 sec. (existing methods consider shorter horizons), and that (ii) to compete with the baselines, it is necessary to have a recurrent unit dedicated to process the positions, but this is not sufficient. Second, from a re-examination of the problem supported with the concept of Structural-RNN, we design a new deep neural architecture, named TRACK. TRACK achieves state-of-the-art performance on all considered datasets and prediction horizons, outperforming competitors by up to 20 percent on focus-type videos and horizons 2-5 seconds. The entire framework (codes and datasets) is online and received an ACM reproducibility badge https://gitlab.com/miguelfromeror/head-motion-prediction.
Miguel Fabián Romero Rondón, Lucile Sassatelli, Ramon Aparicio-Pardo, Frédéric Precioso
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 Active Speaker Detection as a Multi-Objective Optimization with Uncertainty-Based Multimodal Fusion
abstract
International audience
Baptiste Pouthier, Laurent Pilati, Leela K. Gudupudi, Charles Bouveyron, Frédéric Precioso
Interspeech5
2020 Semi-supervised Consensus Clustering Based on Frequent Closed Itemsets
abstract
Semi-supervised consensus clustering integrates supervised information into consensus clustering in order to improve the quality of clustering. In this paper, we study the novel Semi-MultiCons semi-supervised consensus clustering method extending the previous MultiCons approach. Semi-MultiCons aims to improve the clustering result by integrating pairwise constraints in the consensus creation process and infer the number of clusters K using frequent closed itemsets extracted from the ensemble members. Experimental results show that the proposed method outperforms other state-of-art semi-supervised consensus algorithms.
Nicolas Pasquier, Antoine Hom, Laurent Dollé, Frédéric Precioso
CIKM5
2020 Ensemble Clustering based Semi-supervised Learning for Revenue Accounting Workflow Management
abstract
Registered as Defensive Paper of the Intellectual Property Invention "Clustering Techniques for Revenue Accounting Error-Handling Automation", Patent number ID2326WW00, Intellectual Property Board, Amadeus S.A.S., Sophia Antipolis, France. July, 2020.
Nicolas Pasquier, Frédéric Precioso
DATA3
2020 Track: a Multi-Modal Deep Architecture for Head Motion Prediction in 360° Videos
abstract
Head motion prediction is an important problem with 360° videos, in particular to inform the streaming decisions. Various methods tackling this problem with deep neural networks have been proposed recently. In this article, we introduce a new deep architecture, named TRACK, that benefits both from the history of past positions and knowledge of the video content. We show that TRACK achieves state-of-the-art performance when compared against all recent approaches considering the same datasets and wider prediction horizons: from 0 to 5 seconds.
Miguel Fabián Romero Rondón, Lucile Sassatelli, Ramon Aparicio-Pardo, Frédéric Precioso
ICIP4
2020 Hierarchical Multimodal Attention for Deep Video Summarization
abstract
The way people consume sports on TV has drastically evolved in the last years, particularly under the combined effects of the legalization of sport betting and the huge increase of sport analytics. Several companies are nowadays sending observers in the stadiums to collect live data of all the events happening on the field during the match. Those data contain meaningful information providing a very detailed description of all the actions occurring during the match to feed the coaches and staff, the fans, the viewers, and the gamblers. Exploiting all these data, sport broadcasters want to generate extra content such as match highlights, match summaries, players and teams analytics, etc., to appeal subscribers. This paper explores the problem of summarizing professional soccer matches as automatically as possible using both the aforementioned event-stream data collected from the field and the content broadcasted on TV. We have designed an architecture, introducing first (1) a Multiple Instance Learning method that takes into account the sequential dependency among events and then (2) a hierarchical multimodal attention layer that grasps the importance of each event in an action. We evaluate our approach on matches from two professional European soccer leagues, showing its capability to identify the best actions for automatic summarization by comparing with real summaries made by human operators.
Melissa Sanabria, Frédéric Precioso, Thomas Menguy
ICPR2
2020 Profiling Actions for Sport Video Summarization: An attention signal analysis
abstract
Currently, in broadcast companies many human operators select which actions should belong to the summary based on multiple rules they have built upon their own experience using different sources of information. These rules define the different profiles of actions of interest that help the operator to generate better customized summaries. Most of these profiles do not directly rely on broadcast video content but rather exploit metadata describing the course of the match. In this paper, we show how the signals produced by the attention layer of a recurrent neural network can be seen as a learned representation of these action profiles and provide a new tool to support operators' work. The results in soccer matches show the capacity of our approach to transfer knowledge between datasets from different broadcasting companies, from different leagues, and the ability of the attention layer to learn meaningful action profiles.
Melissa Sanabria, Frédéric Precioso, Thomas Menguy
MMSP2
2020 A unified evaluation framework for head motion prediction methods in 360° videos
abstract
The streaming transmissions of 360° videos is a major challenge for the development of Virtual Reality, and require a reliable head motion predictor to identify which region of the sphere to send in high quality and save data rate. Different head motion predictors have been proposed recently. Some of these works have similar evaluation metrics or even share the same dataset, however, none of them compare with each other. In this article we introduce an open software that enables to evaluate heterogeneous head motion prediction methods on various common grounds. The goal is to ease the development of new head/eye motion prediction methods. We first propose an algorithm to create a uniform data structure from each of the datasets. We also provide the description of the algorithms used to compute the saliency maps either estimated from the raw video content or from the users' statistics. We exemplify how to run existing approaches on customizable settings, and finally present the targeted usage of our open framework: how to train and evaluate a new prediction method, and compare it with existing approaches and baselines in common settings. The entire material (code, datasets, neural network weights and documentation) is publicly available.
Miguel Fabián Romero Rondón, Lucile Sassatelli, Ramon Aparicio-Pardo, Frédéric Precioso
MMSys4
2019 Analysis of temporal alignment for Video Classification
abstract
Thanks to their sucess on image recognition, deep neural networks achieve best classification accuracy on videos. However, traditional methods or shallow architectures remain competitive and combinations of different network types are the usual chosen approach. A reason for this less important impact of deep methods for video recognition is the motion representation. The time has a stronger redundancy, and an important elasticity compared to the spatial dimensions. The temporal redundancy is evident, but the elasticity within an action class is well less considered. Several instances of the action still widely differ by their style and speed of execution. In this article, we analyze the temporal dimension by focusing on its singular dynamism, and we focus on the normalization of temporal elasticity on sequences to reduce speed variation within a class. We propose a framework to temporally align video instance in a classification task using the latest temporal warping method, Generalized Canonical Time Warping (GCTW). We evaluate our strategy on video datasets where the intra-class variations lie in temporal dimension rather than in spatial dimensions. Finally, we show the interest of accounting for temporal elasticity for a better video classification and we draw perspectives on more efficient ways to normalize simultaneously temporal and spatial intra-class variations.
Katy Blanc, Diane Lingrand, Antonio Paladini, Luca Coviello, Dane Mitrev, Emily Söhler, Leonardo Guzman, Frédéric Precioso
FG8
2019 Analysing the Impact of Rationality on the Italian Electricity Market
Sara Bevilacqua, Célia da Costa Pereira, Eric Guerci, Frédéric Precioso, Claudio Sartori 0001
MDAI4
2019 Impact of Saliency and Gaze Features on Visual Control: Gaze-Saliency Interest Estimator
abstract
Predicting user intent from gaze presents a challenging question for developing real-time interactive systems like interactive search engine, implicit annotations of large datasets or intelligent robot behavior. Indeed, solutions to annotate easily large sets of images while reducing the burden of annotators is a key aspect for current machine learning techniques. We propose in this paper to design an estimator of the user interest for a given visual content based on eye-tracker feature analysis. We revise existing gaze-based interest estimator, and analyze the impact of the intrinsic saliency of the content displayed for interest estimation. We first explore low-level saliency prediction and propose a new gaze and saliency interest estimator. Experimental results show the advantage of our method for the annotation task in a weakly supervised context. In partic- ular, we extend previous evaluation criteria on new experimental protocol displaying four images by frame as a first step towards "Google Image search-like" interfaces. Our Gaze and Saliency Inter-est Estimator (GSIE) reaches an overall accuracy of 83% in average of user interest prediction. If we consider the accuracy reached in a limited time, the GSIE is 70% in average within about 500ms and 80% in average within 1000ms. This result confirms our GSIE as an efficient real-time visual control solution.
Souad Chaabouni, Frédéric Precioso
ACM Multimedia2
2019 A Co-evolutionary Approach to Analyzing the Impact of Rationality on the Italian Electricity Market
Célia da Costa Pereira, Sara Bevilacqua, Eric Guerci, Frédéric Precioso, Claudio Sartori 0001
PRIMA4
2018 Textual Deconvolution Saliency (TDS) : a deep tool box for linguistic analysis
abstract
Laurent Vanni, Melanie Ducoffe, Carlos Aguilar, Frederic Precioso, Damon Mayaffre. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.
Laurent Vanni, Mélanie Ducoffe, Carlos Aguilar, Frédéric Precioso, Damon Mayaffre
ACL (1)4
2018 An Interactive Content-Based 3D Shape Retrieval System for on-Site Cultural Heritage Analysis
abstract
In this paper, we analyse the process of designing a Content-Based 3D shape Retrieval (CB3DR) adapted for non-experts. Our CB3DR solution aims at scanning an object on the fly with a low-cost 3D sensor and retrieve similar shapes from a database using the 3D point cloud acquired. Our system should meet the requirements of archaeologists who would like to be able to acquire artefacts without prior expertise in scanning, then query easily from the field knowledge bases for Cultural Heritage, and thus retrieve artefacts (i.e. objects or parts of objects) with similar shape without carrying or even moving the artefact found on the site. This context is definitely a retrieval context rather than a classification one since the found artefact may be unknown. At the era of Deep Learning techniques, we investigate in this paper how to design a system which could benefit from the latest advances in 3D object recognition. We thus design a classic CBR workflow starting with a similarity search step followed by an interactive learning SVM which requires very simple positive-negative annotations from the users to refine the ranking provided by the similarity search. The 3D shapes are first described using data representations. We compare first the “old fashion” hand-crafted 3D descriptors (which do not require to be learnt or trained) with PointNet, one of the most recent 3D object deep representation. We evaluate our retrieval baseline against recent Deep Learning approaches on ModelNet10 and explore the potential of deep representations in our interactive learning framework.
Lirone Samoun, Thomas Fisichella, Diane Lingrand, Lucas Malleus, Frédéric Precioso
ICIP5
2018 Foveated streaming of virtual reality videos
abstract
While Virtual Reality (VR) represents a revolution in the user experience, current VR systems are flawed on different aspects. The difficulty to focus naturally in current headsets incurs visual discomfort and cognitive overload, while high-end headsets require tethered powerful hardware for scene synthesis. One of the major solutions envisioned to address these problems is foveated rendering. We consider the problem of streaming stored 360° videos to a VR headset equipped with eye-tracking and foveated rendering capabilities. Our end research goal is to make high-performing foveated streaming systems allowing the playback buffer to build up to absorb the network variations, which is permitted in none of the current proposals. We present our foveated streaming prototype based on the FOVE, one of the first commercially available headsets with an integrated eye-tracker. We build on the FOVE's Unity API to design a gaze-adaptive streaming system using one low- and one high-resolution segment from which the foveal region is cropped with per-frame filters. The low- and high-resolution frames are then merged at the client to approach the natural focusing process.
Miguel Fabián Romero Rondón, Lucile Sassatelli, Frédéric Precioso, Ramon Aparicio-Pardo
MMSys3
2017 Handling Noisy Labels in Gaze-Based CBIR System
Stéphanie Lopez, Arnaud Revel, Diane Lingrand, Frédéric Precioso
ACIVS4
2017 Active learning strategy for CNN combining batchwise Dropout and Query-By-Committee
Mélanie Ducoffe, Frédéric Precioso
ESANN2
2016 Catching Relevance in One Glimpse: Food or Not Food?
abstract
Retrieving specific categories of images among billions of images usually requires an annotation step. Unfortunately, keywords-based techniques suffer from the semantic gap existing between a semantic concept and its digital representation. Content Based Image Retrieval (CBIR) systems tackle this issue simply considering semantic proximities can be mapped to similarities in the image space. Introducing relevance feedbacks involves the user in the task, but extends the annotation step.
Stéphanie Lopez, Arnaud Revel, Diane Lingrand, Frédéric Precioso, V. Dusaucy, Alain Giboin
AVI4
2016 Introducing Semantics in Short Text Classification
Ameni Bouaziz, Célia da Costa Pereira, Christel Dartigues-Pallez, Frédéric Precioso
CICLing (2)4
2016 Human action recognition based on 3D skeleton part-based pose estimation and temporal multi-resolution analysis
abstract
Human action recognition is a challenging field that have been addressed with many different classification techniques such as SVM or Random Decision Forests and by considering many different kinds of information joints, key poses, joints rotation matrix, angles for example. This paper presents our approach for action recognition that considers only information given by the 3D joints from the skeleton and trains a two stage random forest to classify them. We extract skeletal features by computing all angles between any triplet of joints and all distances between any pair of joints then organizing them into a feature vector for each static pose. Complex dynamic actions are then described by sequences of such feature vectors. We evaluate our approach on the most recent and the largest benchmark, MSRC-12 Kinect Gesture Dataset, and compare our results with the state-of-the-art methods on this dataset.
A. Aly Halim, Christel Dartigues-Pallez, Frédéric Precioso, Michel Riveill, Abderrahim Benslimane, Salma A. Ghoneim
ICIP3
2016 Using Frequent Closed Pattern Mining to Solve a Consensus Clustering Problem
abstract
Clustering is the process of partitioning a dataset into groups based on the similarity between the instances.Many clustering algorithms were proposed, but none of them proved to provide good quality partition in all situations.Consensus clustering aims to enhance the clustering process by combining different partitions obtained from different algorithms to yield a better quality consensus solution.In this work, we propose a new consensus method that uses a pattern mining technique in order to reduce the search space from instance-based into pattern-based space.Instead of finding one solution, our method generates multiple consensus candidates based on varying the number of base clusterings considered.The different solutions are then linked and presented as a tree that gives more insight about the similarities between the instances and the different partitions in the ensemble.
Atheer Al-Najdi, Nicolas Pasquier, Frédéric Precioso
SEKE3
2016 Using Closed Patterns to Solve the Consensus Clustering Problem
abstract
Clustering is the process of partitioning a dataset into groups based on the similarity between the instances. Many clustering algorithms were proposed, but none of them proved to provide good quality partition in all situations. Consensus clustering aims to enhance the clustering process by combining different partitions obtained from different algorithms to yield a better quality consensus solution. In this work, we propose a new consensus clustering method that uses a pattern mining technique in order to reduce the search space from instance-based into pattern-based space. Instead of finding one solution, our method generates multiple consensus candidates based on varying the number of base clusterings considered. The different solutions are then linked and presented as a tree that gives more insight about the similarities between the instances and the different partitions in the ensemble.
Atheer Al-Najdi, Nicolas Pasquier, Frédéric Precioso
Int. J. Softw. Eng. Knowl. Eng.3
2016 Fine-grained object recognition in underwater visual data
Concetto Spampinato, Simone Palazzo, Pierre-Hugues Joalland, Sébastien Paris, Hervé Glotin, Katy Blanc, Diane Lingrand, Frédéric Precioso
Multim. Tools Appl.8
2015 One gaze is worth ten thousand (key-)words
abstract
With the maturity of machine learning methods to provide satisfying Content-Based Image Retrieval systems (CBIR), research focus has recently turned back towards visual saliency analysis. The goal in these works is to extract even more efficient visual features than the existing ones. However, analyzing visual saliency is critically dependent on the task to be accomplished from the extracted visual features. A significant number of CBIR systems consider image retrieval as a binary classification problem: what is relevant for the user against what is irrelevant. In this paper, we focus on extracting relevant gaze features within the paradigm of visual preference in order to support the annotation by gaze for a CBIR system. We thus define a gaze acquisition protocol, design a benchmark from a subset of Pascal VOC database and present an in depth analysis of eye-tracking data for visual preference paradigm. Our paper provides new informations on relevant gaze features for image binary classification.
Stéphanie Lopez, Arnaud Revel, Diane Lingrand, Frédéric Precioso
ICIP4
2015 Assessment of algorithms for mitosis detection in breast cancer histopathology images
Mitko Veta, Paul J. van Diest, Stefan M. Willems, Anant Madabhushi, Angel Cruz-Roa, Fabio A. González 0001, Anders Boesen Lindbo Larsen, Jacob S. Vestergaard, Anders Bjorholm Dahl, Dan C. Ciresan, Jürgen Schmidhuber, Alessandro Giusti, Luca Maria Gambardella, Faik Boray Tek, Thomas Walter 0003, Ching-Wei Wang, Satoshi Kondo, Bogdan J. Matuszewski, Frédéric Precioso, Violet Snell, Josef Kittler, Teófilo Emídio de Campos, Adnan Mujahid Khan, Nasir M. Rajpoot, Evdokia Arkoumani, Miangela M. Lacle, Max A. Viergever, Josien P. W. Pluim
Medical Image Anal.20
2014 Short Text Classification Using Semantic Random Forest
Ameni Bouaziz, Christel Dartigues-Pallez, Célia da Costa Pereira, Frédéric Precioso, Patrick Lloret
DaWaK4
2014 Statistical region-based active contour using optimization of alpha-divergence family for image segmentation
abstract
This article deals with statistical region-based active contour segmentation using the alpha-divergence family as similarity measure between the density probability functions of the background and the object regions of interest. Following previous publications on that topic, main originality of this contribution is in the proposed joint optimization of the energy steering the evolution of the active curve and the parameter alpha related to the metric of the divergence and closely related to the statistical luminance distribution of the data. Experiments are shown on both synthetic noisy and textured data as well as on real images (natural and medical ones). We show that the joint optimization process leads to satisfying results for every targeted tasks: above all, it is shown that the proposed approach overcome classic statistical-based region active contour approach using Kullback-Leibler divergence as similarity measure, that can stuck in local extrema during the usual optimization process.
Leila Meziou, Aymeric Histace, Frédéric Precioso
ICIP3
2014 Boosted kernel for image categorization
Alexis Lechervy, Philippe Henri Gosselin, Frédéric Precioso
Multim. Tools Appl.3
2013 Standardized evaluation framework for evaluating coronary artery stenosis detection, stenosis quantification and lumen segmentation algorithms in computed tomography angiography
Hortense A. Kirisli, Michiel Schaap, Coert Metz, A. S. Dharampal, W. B. Meijboom, S. L. Papadopoulou, A. Dedic, K. Nieman, Michiel A. de Graaf, M. F. L. Meijs, M. J. Cramer, Alexander Broersen, Suheyla Cetin, Abouzar Eslami, Leonardo Floréz-Valencia, Kuo-Lung Lor, Bogdan J. Matuszewski, Imen Melki, Brian Mohr, Ilkay Öksüz, Rahil Khurram Shahzad, Chunliang Wang, Pieter H. Kitslaar, Gozde Unal, Amin Katouzian, Maciej Orkisz, Chung-Ming Chen, Frédéric Precioso, Laurent Najman, S. Masood, Devrim Ünay, Lucas J. van Vliet, Rodrigo Moreno, Roman Goldenberg, Erald Vuçini, Gabriel P. Krestin, Wiro J. Niessen, Theo van Walsum
Medical Image Anal.28
2012 Linear kernel combination using boosting
Alexis Lechervy, Philippe Henri Gosselin, Frédéric Precioso
ESANN3
2012 Alpha-divergence maximization for statistical region-based active contour segmentation with non-parametric PDF estimations
abstract
In this article, a complete original framework for unsupervised statistical region-based active contour segmentation is proposed. More precisely, the method is based on the maximization of alpha-divergences between non-paramterically estimated probability density functions (PDFs) of the inner and outer regions defined by the evolving curve. We define the variational context associated to distance maximization in the particular case of alpha-divergences and provide the complete derivation of the partial differential equation leading the segmentation. Results on synthetic data, corrupted with a high level of Gaussian and Poisson noises, but also on clinical X-ray images show that the proposed unsupervised method improves standard approaches of that kind.
Leila Meziou, Aymeric Histace, Frédéric Precioso
ICASSP3
2012 Boosting kernel combination for multi-class image categorization
abstract
In this paper, we propose a novel algorithm to design multi-class kernel functions based on an iterative combination of weak kernels in a scheme inspired from boosting framework. The method proposed in this article aims at building a new feature where the centroid for each class are optimally located. We evaluate our method for image categorization by considering a state-of-the-art image database and by comparing our results with reference methods. We show that on the Oxford Flower databases our approach achieves better results than previous state-of-the-art methods.
Alexis Lechervy, Philippe Henri Gosselin, Frédéric Precioso
ICIP3
2012 Hybdrid Content Based Image Retrieval combining multi-objective interactive genetic algorithm and SVM
Romaric Pighetti, Denis Pallez, Frédéric Precioso
ICPR3
2012 Locality-Sensitive Hashing for Chi2 Distance
abstract
In the past 10 years, new powerful algorithms based on efficient data structures have been proposed to solve the problem of Nearest Neighbors search (or Approximate Nearest Neighbors search). If the Euclidean Locality Sensitive Hashing algorithm, which provides approximate nearest neighbors in a euclidean space with sublinear complexity, is probably the most popular, the euclidean metric does not always provide as accurate and as relevant results when considering similarity measure as the Earth-Mover Distance and 2 distances. In this paper, we present a new LSH scheme adapted to 2 distance for approximate nearest neighbors search in high-dimensional spaces. We define the specific hashing functions, we prove their local-sensitivity, and compare, through experiments, our method with the Euclidean Locality Sensitive Hashing algorithm in the context of image retrieval on real image databases. The results prove the relevance of such a new LSH scheme either providing far better accuracy in the context of image retrieval than euclidean scheme for an equivalent speed, or providing an equivalent accuracy but with a high gain in terms of processing speed.
David Gorisse, Matthieu Cord, Frédéric Precioso
IEEE Trans. Pattern Anal. Mach. Intell.3
2011 Statistical Shape Model of Legendre Moments with Active Contour Evolution for Shape Detection and Segmentation
Yan Zhang 0020, Bogdan J. Matuszewski, Aymeric Histace, Frédéric Precioso
CAIP (1)4
2011 Segmentation of cellular structures in actin tagged fluorescence confocal microscopy images
abstract
The paper reports on a novel method for reconstruction of cellular features including cell nuclei and cellular boundaries from actin tagged fluorescence confocal microscopy images. Such reconstruction can provide spatial context for subsequent quantitative analysis of changes to actin organisation and cell morphology in both controlled and stressed cell cultures. The proposed method is fully automatic and is formulated within active contour multiphase level set framework. The derived level set evolution PDEs combine previously proposed curvature and advection flows with propagation flow defined by specially designed set of geodesic distance maps. Additionally the proposed PDEs include additional components to impose known inclusion/exclusion topological constraints between cellular structures. The paper gives an overview of the proposed methodology as well as reports on initial results obtained for monolayer of human prostate cells (PNT2) culture visualised using acting tagged fluorescence confocal microscopy.
Bogdan J. Matuszewski, Mark F. Murphy, David R. Burton, Tom Marchant, Christopher J. Moore, Aymeric Histace, Frédéric Precioso
ICIP7
2011 Confocal microscopy segmentation using active contour based on (α)-divergence
abstract
This paper describes a novel method for active contour segmentation based on foreground/background alpha-divergence histogram distance measure. In recent years a number of variational segmentation techniques have been proposed for a region based active contour segmentation utilising different distance measures between probability density functions (PDFs) describing foreground and background regions. The most common techniques use χ2, Hellinger/Bhattacharya distances or Kullback-Leibler divergence. In this paper, it is proposed to generalize these methods by using the alpha-divergences distance function. This distance function depending on the selected value of its parameter encompasses mentioned above classical distances. The paper defines a partial differential equation, associated with alpha-divergence variational criterion, that governs the iterative deformations of the active contour. The experimental results on a synthetic data demonstrate that the proposed method outperforms previously proposed histogram based methods in terms of segmentation accuracy and robustness with respect to type and level of noise. The potential of the proposed technique for segmentation of cellular structures in fluorescence confocal microscopy data is also illustrated.
Leila Meziou, Aymeric Histace, Frédéric Precioso, Bogdan J. Matuszewski, Mark F. Murphy
ICIP3
2011 Efficient Bag-of-Feature kernel representation for image similarity search
abstract
Although “Bag-of-Features” image models have shown very good potential for object matching and image retrieval, such a complex data representation requires computationally expensive similarity measure evaluation. In this paper, we propose a framework unifying dictionary-based and kernel-based similarity functions that highlights the tradeoff between powerful data representation and eff cient similarity computation. On the basis of this formalism, we propose a new kernel-based similarity approach for Bag-of-Feature descriptions. We introduce a method for fast similarity search in large image databases. The conducted experiments prove that our approach is very competitive among State-of-the-art methods for similarity retrieval tasks.
Frédéric Precioso, Matthieu Cord, David Gorisse, Nicolas Thome
ICIP1
2011 Spatio-Temporal Tube data representation and Kernel design for SVM-based video object retrieval system
Shuji Zhao, Frédéric Precioso, Matthieu Cord
Multim. Tools Appl.2
2011 SALSAS: Sub-linear active learning strategy with approximate k-NN search
David Gorisse, Matthieu Cord, Frédéric Precioso
Pattern Recognit.3
2011 Incremental kernel learning for active image retrieval without global dictionaries
Philippe Henri Gosselin, Frédéric Precioso, Sylvie Philipp-Foliguet
Pattern Recognit.2
2010 Scalable active learning strategy for object category retrieval
abstract
Since the digital revolution, the volume of images to be processed has grown exponentially. Interactive search systems have to deal with these huge databases to remain effective. As the complexity of on-line learning methods is at least linear in the size of the database, scalability is the major problem for these methods. Fast retrieval systems, with index structures for fast navigation, have hence become like a Holy Grail. In this article, we propose a strategy to overcome this scalability limitation. Our technique exploits ultra fast retrieval methods as Locally Sensitive Hashing to speed up active learning system. Experiments on database of 180 K images are reported. The results show that our method is 45 times faster than state of the art approaches for similar accuracy.
David Gorisse, Matthieu Cord, Frédéric Precioso
ICIP3
2010 STTK-based video object recognition
abstract
In this paper, we extend our video object recognition system to multiclass object recognition context, dealing with unbalanced data sets and comparing our resuls to state-of-the-art methods. Our approach is based on a Spatio-Temporal data representation, a dedicated kernel design and statistical learning techniques for object recognition. From video tracks made of segmented object regions in the successive frames, we extract sets of spatio-temporally coherent SIFT-based features, called Spatio-Temporal Tubes. To compare these complex tube objects, we integrate a Spatio-Temporal Tube Kernel (STTK) function into a multi-class classification framework with balancing process for unequal classes. Our approach is successfully evaluated on episodes from “Buffy, the Vampire Slayer” TV series which have been used in other works targeting same objectives. Our method proved to be more robust than dictionary based, facial feature based and key-frame based approaches. Our method is also tested on a small car database and preliminary results for car identification task illustrate its generalization potential.
Shuji Zhao, Frédéric Precioso, Matthieu Cord
ICIP2
2010 Active Boosting for Interactive Object Retrieval
abstract
This paper presents a new algorithm based on boosting for interactive object retrieval in images. Recent works propose ”online boosting” algorithms where weak classifier sets are iteratively trained from data. These algorithms are proposed for visual tracking in videos, and are not well adapted to ”online boosting” for interactive retrieval. We propose in this paper to iteratively build weak classifiers from images, labeled as positive by the user during a retrieval session. A novel active learning strategy for the selection of images for user annotation is also proposed. This strategy is used to enhance the strong classifier resulting from ”boosting” process, but also to build new weak classifiers. Experiments have been carried out on a generalist database in order to compare the proposed method to a SVM based reference approach.
Alexis Lechervy, Philippe Henri Gosselin, Frédéric Precioso
ICPR3
2009 Optimization on active learning strategy for object category retrieval
abstract
Active learning is a machine learning technique which has attracted a lot of research interest in the content-based image retrieval (CBIR) in recent years. To be effective, an active learning system must be fast and efficient using as few (relevance) feedback iterations as possible. Scalability is the major problem for such an on-line learning method, since the complexity of such methods on a database of size n is in the best case O(n * log(n)). In this article we propose a strategy to overcome this limitation. Our technique exploits ultra fast retrieval methods like Locality Sensitive Hashing (LSH), recently applied for unsupervised image retrieval. Combined with active selection, our method is able to achieve very fast active learning task in very large database. Experiments on VOC2006 database are reported, results are obtained four times faster while preserving the accuracy.
David Gorisse, Matthieu Cord, Frédéric Precioso
ICIP3
2009 Windows and facades retrieval using similarity on graph of contours
abstract
The development of street-level geoviewers become recently a very active and challenging research topic. In this context, the detection, representation and classification of windows can be beneficial for the identification of the respective facade. In this paper, a novel method for windows and facade retrieval is presented. This method, based on a similarity of graph of contours, introduces a new kernel on graph for inexact graph matching. We design a kernel similarity function for structured sets of contours which will take into account the variations of contour orientation inside a structure set, as well as spatial proximity. Then we are able to extract a window as a sub-graph of the graph of all contours of the facade image and to retrieve similar windows from a database of images of facades.
Jean-Emmanuel Haugeard, Sylvie Philipp-Foliguet, Frédéric Precioso
ICIP3
2009 Spatio-Temporal Tube Kernel for actor retrieval
abstract
This paper presents an actor video retrieval system based on face video-tubes extraction and representation with sets of temporally coherent features. Visual features, SIFT points, are tracked along a video shot, resulting in sets of feature point chains (spatio-temporal tubes). These tubes are then classified and retrieved using a kernel-based SVM learning framework for actor retrieval in a movie. In this paper, we present optimized feature tubes, we extend our feature representation with spatial location of SIFT points and we describe the new Spatio-Temporal Tube Kernel (STTK) of our content-based retrieval system. Our approach has been tested on a real movie and proved to be faster and more robust for actor retrieval task.
Shuji Zhao, Frédéric Precioso, Matthieu Cord
ICIP2
2008 Fast approximate kernel-based similarity search for image retrieval task
abstract
In content based image retrieval, the success of any distance-based indexing scheme depends critically on the quality of the chosen distance metric. We propose in this paper a kernel-based similarity approach working on sets of vectors to represent images. We introduce a method for fast approximate similarity search in large image databases with our kernel-based similarity metric. We evaluate our algorithm on image retrieval task and show it to be accurate and faster than linear scanning.
David Gorisse, Matthieu Cord, Frédéric Precioso, Sylvie Philipp-Foliguet
ICPR3
2005 A genetic algorithm-based approach to knowledge-assisted video analysis
abstract
Efficient video content management and exploitation requires extraction of the underlying semantics, a non-trivial task associating low-level features of the image domain and high-level semantic descriptions. In this paper, a knowledge-assisted approach for extracting semantics of domain-specific video content is presented. Domain knowledge considers both low-level features (color, motion, shape) and spatial behavior (topological and directional information). During the preprocessing step, a set of over-segmented homogenous atom-regions is generated and their low-level and spatial descriptions are extracted. A genetic algorithm is then applied in order to find the optimal interpretation according to a specific domain conceptualization. The proposed approach was tested on the formula one, tennis and beach vacations domains showing promising results.
Nikola Voisine, Stamatia Dasiopoulou, Frédéric Precioso, Vasileios Mezaris, Ioannis Kompatsiaris, Michael G. Strintzis
ICIP (3)3
2005 Robust real-time segmentation of images and videos using a smooth-spline snake-based algorithm
abstract
This paper deals with fast image and video segmentation using active contours. Region-based active contours using level sets are powerful techniques for video segmentation, but they suffer from large computational cost. A parametric active contour method based on B-Spline interpolation has been proposed in to highly reduce the computational cost, but this method is sensitive to noise. Here, we choose to relax the rigid interpolation constraint in order to robustify our method in the presence of noise: by using smoothing splines, we trade a tunable amount of interpolation error for a smoother spline curve. We show by experiments on natural sequences that this new flexibility yields segmentation results of higher quality at no additional computational cost. Hence, real-time processing for moving objects segmentation is preserved.
Frédéric Precioso, Michel Barlaud, Thierry Blu, Michael Unser
IEEE Trans. Image Process.1
2003 Smoothing B-spline active contour for fast and robust image and video segmentation
abstract
This paper deals with fast image and video segmentation using active contours. Region based active contours using level-sets are powerful techniques for video segmentation hut they suffer from large computational cost. A parametric active contour method based on B-Spline interpolation has been proposed in F. Precioso (2002) to highly reduce the computational cost but this method is sensitive to noise. Here, we choose to relax the rigid interpolation constraint in order to robustify our method in the presence of noise: by using smoothing splines, we trade a tunable amount of interpolation error for a smoother spline curve. We show by experiments on natural sequences that this new flexibility yields segmentation results of higher quality at no additional computational cost. Hence real time processing for moving objects segmentation is preserved.
Frédéric Precioso, Michel Barlaud, Thierry Blu, Michael Unser
ICIP (1)1
2002 Regular spatial B-spline active contour for fast video segmentation
abstract
This paper deals with fast video segmentation using active contours. Region-based active contours is a powerful technique for video segmentation. However most of these methods are implemented using level-sets. Although level-set methods provide accurate segmentation, they suffer from large computational cost. The proposed method uses a B-spline parametric method to highly improve the computation cost. Our method removes irregular sampling assumptions. It combines multi-resolution regular sampling and length penalty.
Frédéric Precioso, Michel Barlaud
ICIP (2)1
2001 B-spline active contours for fast video segmentation
abstract
Video segmentation is among the most important challenges of video processing and compression (MPEG4 and MPEG-7). A drawback of classical methods is the computational cost due to the model complexity. In this paper we propose to use a B-spline parametric contour to implement a region-based active contour segmentation Hence we get a fast variational method based on active contours with an intrinsic regularizing constraint. More precisely the evolution force follows from the minimization of a region-based criterion. The theory of B-splines allows analytical computation of the contour curvature at any point. The model complexity is fixed and depends on the desired level of detail. This complexity is highly compared to non-parametric methods. We compare our new approach to classical parametric polygon-based methods. We show experiments on real video sequences.
Frédéric Precioso, Michel Barlaud
ICIP (2)1