Georges Quénot

dblp:q/GeorgesQuenot · DBLP profile ↗
← Back
52ranked-venue papers
7as first author
8since 2021 · last 2024
0000-0003-2117-247XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 37 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 16 · 6 since 2021Databases, data management, data science and information retrieval · 10Systems, architecture and hardware · 3 · 2 first-authorComputer networks · 1Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2024 Combining Image and Region Uncertainty-Based Active Learning for Melanoma Segmentation
abstract
Segmentation of medical images using learning-based systems remains a challenge in medical computer vision, as training a segmentation model requires exhaustively annotated medical images by experts, which are difficult and expensive to obtain. In the context of melanoma segmentation, we explored active learning methods to combine human annotation for the most uncertain pixels with model predictions for the others. With only around$30 \%$of the images annotated, we achieved performance similar to that obtained with a fully annotated dataset. We also demonstrated that, after a few iterations, experts can focus on annotating only the most uncertain areas of the images, relying on the model for the rest. These approaches pave the way for accelerating the annotation of unlabeled medical datasets and optimizing the use of medical expertise in deep learning projects.
Jean-Pierre Chevallet, Philippe Mulhem, Georges Quénot
CBMI4
2023 Improving Causality in Interpretable Video Retrieval
abstract
This paper focuses on the causal relation between the detection scores of concept (or tag) classifiers and the ranking decisions based on these scores, paving the way for these tags to be used in the visual explanations. We first define a measure for quantifying a causality on a set of tags, typically those involved in visual explanations. We use this measure for evaluating the actual causality in the explanations generated using a recent interpretable video retrieval system (Dong et al. [4]), which we find to be quite low. We then propose and evaluate improvements for significantly increasing this causality without sacrificing the retrieval accuracy of the system.
Varsha Devi, Philippe Mulhem, Georges Quénot
CBMI3
2023 On the stability, correctness and plausibility of visual explanation methods based on feature importance
abstract
In the field of Explainable AI, multiples evaluation metrics have been proposed in order to assess the quality of explanation methods w.r.t. a set of desired properties. In this work, we study the articulation between the stability, correctness and plausibility of explanations based on feature importance for image classifiers. We show that the existing metrics for evaluating these properties do not always agree, raising the issue of what constitutes a good evaluation metric for explanations. Finally, in the particular case of stability and correctness, we show the possible limitations of some evaluation metrics and propose new ones that take into account the local behaviour of the model under test.
Romain Xu-Darme, Jenny Benois-Pineau, Romain Giot, Georges Quénot, Zakaria Chihani, Marie-Christine Rousset, Alexey Zhukov
CBMI4
2022 Analysis of the Complementarity of Latent and Concept Spaces for Cross-Modal Video Search
abstract
This paper focuses on studying the complementarity between the spaces from hybrid cross-modal state-of-the-art systems for video retrieval like [5]. We aim at investigating if these spaces really convey different features, or if they are representing the same things. We use PCA (Principal Component Analysis) to study the optimal dimensions, CCA (Canonical Correlation Analysis) to assess the similarity of the spaces, and check if such approach is in fact similar to ensemble learning. We achieve experiments on the MST-VTT corpus, and show that in fact these two spaces are indeed very similar, paving the way for new models that could enforce more dissimilar spaces.
Varsha Devi, Philippe Mulhem, Georges Quénot
CBMI3
2022 Segmenting partially annotated medical images
abstract
Segmentation of medical images using learning based systems remains a challenge in medical computer vision: training a segmentation model requires medical images exhaustively annotated by experts that are difficult and expensive to obtain. We propose to explore the usage of partially annotated images, i.e., all images are annotated but not all regions of a given class are annotated. In this paper, we propose several approaches and we experiment them on the segmentation of intra-oral images. First, we propose to modify the loss function to consider only the annotated areas, and second to integrate annotation from non-expert, as well as the combination of these methods. The experiments we conducted showed an improvement up to 33% on the segmentation performance. This approach allows to obtain better quality annotation masks than the initial human annotation using only partially annotated areas or non-expert annotations. In the future, these approaches can be extended by combination with active learning methods.
Jean-Pierre Chevallet, Georges Quénot
CBMI3
2022 A novel pattern-based edit distance for automatic log parsing
abstract
This work aims at inferring a set of regular expressions to parse a text file, like a system log. To this end, we propose a novel edit distance taking advantage of the pattern matching background. Edit distances are commonly used for fuzzy search and in bioinformatics, and compare two strings at the character level. By doing so, edit distances do not consider the nature of the data conveyed by the strings. To address this problem, we propose the following contributions. First, we propose to model strings at the pattern level using a dedicated data structure, called pattern automaton. Second, we design a novel edit distance, operating at the pattern level. Third, we derive a clustering algorithm optimized for this distance. Finally, we evaluate our proposal through experimental validation.
Maxime Raynal, Marc-Olivier Buob, Georges Quénot
ICPR3
2022 Overview of the Multimedia Grand Challenges 2022
abstract
The Multimedia Grand Challenge track was first presented as part of ACM Multimedia 2009 and has established itself as a prestigious competition in the multimedia community. The purpose of the Multimedia Grand Challenges is to engage the multimedia research community by establishing well-defined and objectively judged challenge problems intended to exercise the state-of-the-art methods and inspire future research directions. The key criteria for Grand Challenges are that they should be useful, interesting, and their solution should involve a series of research tasks over a long period of time, with pointers towards longer-term research. The 2022 edition of ACM Multimedia hosted 10 Grand Challenges covering all aspects of multimedia computing, from delivery systems to video retrieval, from video generation to audio recognition.
Miriam Redi, Georges Quénot
ACM Multimedia2
2021 Content-based multimedia indexing (CBMI 2018) [SI 1104]
Renaud Péteri, Georges Quénot, Philippe Joly
Multim. Tools Appl.2
2020 Coupled ensembles of neural networks
Anuvabh Dutt, Denis Pellerin, Georges Quénot
Neurocomputing3
2019 Classifier Training from a Generative Model
abstract
We investigate the samples derived from generative adversarial networks (GAN) from a classification perspective. We train a classifier on generated samples and on real data and see how they compared on a held out validation set. We see that recent GAN models which produce visually convincing samples are not yet able to match the training on real data. To analyse this we compare training a classifier on generated samples and various sizes of the real training set. We propose architectural and algorithmic changes to reduce this gap. First, we show that a modification to the GAN architecture is needed, which leads to improve generation of samples. Second, we use multiple GAN models as a way to cover the real data distribution, again leading to improvement in classifier training. We also show that in the case of training on small number of samples, a GAN model provides better compression in terms of storage requirements as compared to the real data.
Pham Thanh Dat, Anuvabh Dutt, Denis Pellerin, Georges Quénot
CBMI4
2019 Explaining Visual Classification using Attributes
abstract
The performance of deep Convolutional Neural Networks (CNN) has been reaching or even exceeding the human level on large number of tasks. Some examples are image classification, Mastering Go game, speech understanding etc. However, their lack of decomposability into intuitive and understandable components make them hard to interpret, i.e. no information is provided about what makes them arrive at their prediction. We propose a technique to interpret CNN classification task and justify the classification result with visual explanation and visual search. The model consists of two sub networks: a deep recurrent neural network for generating textual justification and a deep convolutional network for image analysis. This multimodal approach generates the textual justification about the classification decision. To verify the textual justification, we use the visual search to extract the similar content from the training set. We evaluate our strategy on a novel CUB dataset with the ground-truth attributes. We make use of these attributes to further strengthen the justification by providing the attributes of images.
Muneeb Ul Hassan 0003, Philippe Mulhem, Denis Pellerin, Georges Quénot
CBMI4
2019 Object instance identification with fully convolutional networks
Maxime Portaz, Matthias Kohl, Jean-Pierre Chevallet, Georges Quénot, Philippe Mulhem
Multim. Tools Appl.4
2018 Coupled Ensembles of Neural Networks
abstract
We investigate in this paper architectures of deep convolutional networks. Building on existing state of the art models, we propose a reconfiguration of the model parameters into several parallel branches at the global network level, each branch being a standalone CNN. We show that this arrangement is an efficient way to significantly reduce the number of parameters while at the same time improving the performance. The use of branches brings an additional form of regularization. In addition to splitting the parameters into parallel branches, we propose a tighter coupling of these branches by averaging their log-probabilities. The tighter coupling favours the learning of better representations, even at the level of the individual branches, as compared to when each branch is trained independently. We refer to this branched architecture as “coupled ensembles”. The approach is generic and can be applied to almost any neural network architecture. With coupled ensembles of DenseNet-BC networks and a 25M-parameter budget, we obtain error rates of 2.92%, 15.68% and 1.50% on CIFAR-10, CIFAR-100 and SVHN respectively. For the same parameter budget, DenseNet-Bchas an error rate of 3.46%, 17.18%, and 1.8% respectively. With ensembles of coupled ensembles of DenseNet-BC networks with a 50M-parameter budget, we obtain error rates of 2.72%, 15.13% and 1.42 % respectively on these tasks.
Anuvabh Dutt, Denis Pellerin, Georges Quénot
CBMI3
2017 Video Indexing, Search, Detection, and Description with Focus on TRECVID
abstract
There has been a tremendous growth in video data the last decade. People are using mobile phones and tablets to take, share or watch videos more than ever before. Video cameras are around us almost everywhere in the public domain (e.g. stores, streets, public facilities, ...etc). Efficient and effective retrieval methods are critically needed in different applications. The goal of TRECVID is to encourage research in content-based video retrieval by providing large test collections, uniform scoring procedures, and a forum for organizations interested in comparing their results. In this tutorial, we present and discuss some of the most important and fundamental content-based video retrieval problems such as recognizing predefined visual concepts, searching in videos for complex ad-hoc user queries, searching by image/video examples in a video dataset to retrieve specific objects, persons, or locations, detecting events, and finally bridging the gap between vision and language by looking into how can systems automatically describe videos in a natural language. A review of the state of the art, current challenges, and future directions along with pointers to useful resources will be presented by different regular TRECVID participating teams. Each team will present one of the following tasks:
George Awad, Duy-Dinh Le, Chong-Wah Ngo, Vinh-Tiep Nguyen, Georges Quénot, Cees Snoek, Shin'ichi Satoh 0001
ICMR5
2017 Improving Image Classification using Coarse and Fine Labels
abstract
The performance of classifiers is in general improved by designing models with a large number of parameters or by ensembles. We tackle the problem of classification of coarse and fine grained categories, which share a semantic relationship. On being given the predictions that a classifier has for a given test sample, we adjust the probabilities according to the semantics of the categories, on which the classifier was trained. We present an algorithm for doing such an adjustment and we demonstrate improvement for both coarse and fine grained classification. We evaluate our method using convolutional neural networks. However, the algorithm can be applied to any classifier which outputs category wise probabilities.
Anuvabh Dutt, Denis Pellerin, Georges Quénot
ICMR3
2017 Learned features versus engineered features for multimedia indexing
Mateusz Budnik, Efrain-Leonardo Gutierrez-Gomez, Bahjat Safadi, Denis Pellerin, Georges Quénot
Multim. Tools Appl.5
2016 The CAMOMILE Collaborative Annotation Platform for Multi-modal, Multi-lingual and Multi-media Documents
Johann Poignant, Mateusz Budnik, Hervé Bredin, Claude Barras, Mickaël Stefas, Pierrick Bruneau, Gilles Adda, Laurent Besacier, Hazim Kemal Ekenel, Gil Francopoulo, Javier Hernando, Joseph Mariani, Ramon Morros, Georges Quénot, Sophie Rosset, Thomas Tamisier
LREC14
2016 A comparative study for multiple visual concepts detection in images and videos
Abdelkader Hamadi, Philippe Mulhem, Georges Quénot
Multim. Tools Appl.3
2016 Guest Editorial: Content-based Multimedia Indexing
Harald Kosch, Georges Quénot
Multim. Tools Appl.2
2016 Naming multi-modal clusters to identify persons in TV broadcast
Johann Poignant, Guillaume Fortier, Laurent Besacier, Georges Quénot
Multim. Tools Appl.4
2015 Extended conceptual feedback for semantic multimedia indexing
Abdelkader Hamadi, Philippe Mulhem, Georges Quénot
Multim. Tools Appl.3
2015 Descriptor optimization for multimedia indexing and retrieval
Bahjat Safadi, Nadia Derbas, Georges Quénot
Multim. Tools Appl.3
2015 Unsupervised Speaker Identification in TV Broadcast Based on Written Names
abstract
Identifying speakers in TV broadcast in an unsupervised way (i.e., without biometric models) is a solution for avoiding costly annotations. Existing methods usually use pronounced names, as a source of names, for identifying speech clusters provided by a diarization step but this source is too imprecise for having sufficient confidence. To overcome this issue, another source of names can be used: the names written in a title block in the image track. We first compared these two sources of names on their abilities to provide the name of the speakers in TV broadcast. This study shows that it is more interesting to use written names for their high precision for identifying the current speaker. We also propose two approaches for finding speaker identity based only on names written in the image track. With the “late naming” approach, we propose different propagations of written names onto clusters. Our second proposition, “Early naming,” modifies the speaker diarization module (agglomerative clustering) by adding constraints preventing two clusters with different associated written names to be merged together. These methods were tested on the REPERE corpus phase 1, containing 3 hours of annotated videos. Our best “late naming” system reaches an F-measure of 73.1%. “early naming” improves over this result both in terms of identification error rate and of stability of the clustering stopping criterion. By comparison, a mono-modal, supervised speaker identification system with 535 speaker models trained on matching development data and additional TV and radio data only provided a 57.2% F-measure.
Johann Poignant, Laurent Besacier, Georges Quénot
IEEE ACM Trans. Audio Speech Lang. Process.3
2014 Joint Audio-Visual Words for Violent Scenes Detection in Movies
abstract
This paper presents an audio-visual data representation for violent scenes detection in movies. Existing works in this field consider either the audio or the visual information; or their shallow fusion. None has yet explored their joint dependence for violent scenes detection. We propose a feature which provides strong multi-modal audio and visual cues by first joining the audio and the visual features and then revealing statistically the joint multi-modal patterns. Experimental validation was conducted in the context of the Violent Scenes Detection task of the MediaEval 2013 Multimedia benchmark. The obtained results show the potential of the proposed approach in comparison to methods using audio and visual features separately and other fusion methods.
Nadia Derbas, Georges Quénot
ICMR2
2014 Infrequent concept pairs detection in multimedia documents
abstract
Single visual concept detection in videos is a hard task, especially for infrequent concepts or for those difficult to model. This question becomes even more difficult in the case of concept pairs. Two main directions may tackle this problem: 1) combine the predictions of their corresponding detectors in a way which is similar to usual information retrieval, or 2) build supervised learners for these pairs of concepts by generating annotations based on the occurrences of the two individual concepts. Each of these approaches have advantages and drawbacks. We evaluated them in the context of the concept pair detection subtask of the TRECVid 2013 semantic indexing (SIN) task and found that information retrieval-like fusions of concept detection scores outperforms the learning approaches. The described methods outperform the best official result of the evaluation campaign cited previously, by 9% in terms of relative improvement on MAP.
Abdelkader Hamadi, Philippe Mulhem, Georges Quénot
ICMR3
2013 Content-Based Re-ranking of Text-Based Image Search Results
Franck Thollard, Georges Quénot
ECIR2
2013 Unsupervised naming of speakers in broadcast TV: using written names, pronounced names or both?
abstract
International audience
Johann Poignant, Laurent Besacier, Viet Bac Le, Sophie Rosset, Georges Quénot
INTERSPEECH5
2012 A Local Temporal Context-Based Approach for TV News Story Segmentation
abstract
Users are often interested in retrieving only a particular passage on a topic of interest to them. It is therefore necessary to split videos into shorter segments corresponding to appropriate retrieval units. We propose here a method based on a local temporal context for the segmentation of TV news videos into stories. First, we extract multiple descriptors which are complementary and give good insights about story boundaries. Once extracted, these descriptors are expanded with a local temporal context and combined by an early fusion process. The story boundaries are then predicted using machine learning techniques. We investigate the system by experiments conducted using TRECVID 2003 data and protocol of the story boundary detection task and we show that the extension of multimodal descriptors by a local temporal context approach improves results and our method outperforms the state of the art.
Emilie Dumont, Georges Quénot
ICME2
2012 From Text Detection in Videos to Person Identification
abstract
We present in this article a video OCR system that detects and recognizes overlaid texts in video as well as its application to person identification in video documents. We proceed in several steps. First, text detection and temporal tracking are performed. After adaptation of images to a standard OCR system, a final post-processing combines multiple transcriptions of the same text box. The semi-supervised adaptation of this system to a particular video type (video broadcast from a French TV) is proposed and evaluated. The system is efficient as it runs 3 times faster than real time (including the OCR step) on a desktop Linux box. Both text detection and recognition are evaluated individually and through a person recognition task where it is shown that the combination of OCR and audio (speaker) information can greatly improve the performances of a state of the art audio based person identification system.
Johann Poignant, Laurent Besacier, Georges Quénot, Franck Thollard
ICME3
2012 Unsupervised Speaker Identification using Overlaid Texts in TV Broadcast
abstract
Poster Session: Speaker Recognition III
Johann Poignant, Hervé Bredin, Viet Bac Le, Laurent Besacier, Claude Barras, Georges Quénot
INTERSPEECH6
2012 Active Cleaning for Video Corpus Annotation
Bahjat Safadi, Stéphane Ayache, Georges Quénot
MMM3
2012 Preface for the special issue of MTAP following CBMI 2010
Georges Quénot, Jenny Benois-Pineau, Régine André-Obrecht
Multim. Tools Appl.1
2012 Active learning with multiple classifiers for multimedia indexing
Bahjat Safadi, Georges Quénot
Multim. Tools Appl.2
2011 Re-ranking by local re-scoring for video indexing and retrieval
abstract
Video retrieval can be done by ranking the samples according to their probability scores that were predicted by classifiers. It is often possible to improve the retrieval performance by re-ranking the samples. In this paper, we proposed a re-ranking method that improves the performance of semantic video indexing and retrieval, by re-evaluating the scores of the shots by the homogeneity and the nature of the video they belong to. Compared to previous works, the proposed method provides a framework for the re-ranking via the homogeneous distribution of video shots content in a temporal sequence. The experimental results showed that the proposed re-ranking method was able to improve the system performance by about 18% in average on the TRECVID 2010 semantic indexing task, videos collection with homogeneous contents. For TRECVID 2008, in the case of collections of videos with non-homogeneous contents, the system performance was improved by about 11-13%.
Bahjat Safadi, Georges Quénot
CIKM2
2011 Re-ranking for Multimedia Indexing and Retrieval
Bahjat Safadi, Georges Quénot
ECIR2
2011 Incremental Multiple Classifier Active Learning for Concept Indexing in Images and Videos
Bahjat Safadi, Yubing Tong, Georges Quénot
MMM (1)3
2010 Content-based search in multilingual audiovisual documents using the International Phonetic Alphabet
Georges Quénot, Tien Ping Tan, Viet Bac Le, Stéphane Ayache, Laurent Besacier, Philippe Mulhem
Multim. Tools Appl.1
2008 Video Corpus Annotation Using Active Learning
Stéphane Ayache, Georges Quénot
ECIR2
2007 Classifier Fusion for SVM-Based Multimedia Semantic Indexing
Stéphane Ayache, Georges Quénot, Jérôme Gensel
ECIR2
2007 Active learning for multimedia
abstract
Active learning improves the performance of classification or search systems by adding humans to the loop. It aims at optimizing the production of the class labels that are necessary for supervised learning. The proposed tutorial responds to a strong need for the integration of this technique in multimedia indexing and retrieval systems. It presents the basics of active learning and gives the necessary information for quickly and efficiently integrating it within a project. Several applications are considered, from relevance feedback to corpus annotation. Most illustrations are given in the context of the NIST benchmarks on video indexing and retrieval (TRECVID).
Georges Quénot
ACM Multimedia1
2007 Evaluation of active learning strategies for video indexing
Stéphane Ayache, Georges Quénot
Signal Process. Image Commun.2
2007 The ARGOS campaign: Evaluation of video analysis and indexing tools
Philippe Joly, Jenny Benois-Pineau, Ewa Kijak, Georges Quénot
Signal Process. Image Commun.4
2006 Context-Based Conceptual Image Indexing
abstract
Automatic semantic classification of image databases is very useful for users searching and browsing but it is at the same time a very challenging research problem. Local features based image classification is one of the promising way to bridge the semantic gap in detecting concepts. This paper proposes a framework for incorporating contextual information into the concept detection process. The proposed method combines local and global classifiers (SVMs) with stacking. We studied the impact of topologic and semantic contexts in concept detection performance and proposed solutions to handle the large amount of dimensions involved in classified data. We conducted experiments on TRECVID'04 data set with 48104 images and 5 concepts. We found that the use of context yields a significant improvement both for the topologic and semantic contexts
Stéphane Ayache, Georges Quénot, Shin'ichi Satoh 0001
ICASSP (2)2
2005 Video Shot Classification Using Lexical Context
Stéphane Ayache, Georges Quénot, Mbarek Charhad
ECIR2
2004 Recognizing emotions for the audio-visual document indexing
abstract
In this paper, we proposed using MFCC coefficients (mel-scaled cepstral coefficients) and a simple but efficient classifying method: vector quantification (VQ) to perform speaker-dependent emotion recognition. Many other features: energy, pitch, zero crossing, phonetic rate, LPC... and their derivatives are also tested and combined with MFCC coefficients in order to find the best combination. Other models, GMM and HMM (discrete and continuous hidden Markov model), are studied as well in the hope that the use of continuous distribution and the temporal evolution of this set of features will improve the quality of emotion recognition. The accuracy recognizing five different emotions exceeds 80% by using only MFCC coefficients with VQ model. This is a simple but efficient approach, the result is even much better than those obtained with the same database in human evaluations by listening and judging without returning permission nor comparisons between sentences (Inger Samso Engberg and Anya Varnich Hansen, 2001).
See-May Phoong, Georges Quénot, Eric Castelli
ISCC2
1996 From real-time emulation to ASIC integration for image processing applications
abstract
A methodology for automatically deriving image processing ASICs from their real-time emulation on the data-flow functional computer is presented. The aim of the method is to reduce the time and effort required to synthesize and validate ASICs after emulation. This is achieved by optimizing the architecture validated on the emulator and integrating the optimized resources. The paper presents the derivation of a 1100 MIPS defect detector.
Ivan C. Kraljic, Georges Quénot, Bertrand Y. Zavidovique
IEEE Trans. Very Large Scale Integr. Syst.2
1995 A methodology for rapid prototyping of real-time image processing VLSI systems
abstract
A methodology aimed at prototyping real-time image processing VLSI systems is presented. The prototyping methodology includes the emulation of the application in its target environment (i.e., in real time and on real-world scenes) and the automatic derivation of a VLSI chip-set implementing the application. Our methodology is based on a custom emulator called the Data-Flow Functional Computer (DFFC) which is dedicated to real-time image processing. Our method encompasses in a coherent environment: the validation of the high level specification; the simultaneous validation of an implementation of the design suitable for integration; and a method for integrating the validated emulator architecture as a VLSI system. The results of the successful prototyping of a defect detector algorithm are presented.
Ivan C. Kraljic, Georges Quénot, Bertrand Y. Zavidovique
RSP2
1993 A wavefront array processor for on the fly processing of digital video streams
abstract
The authors present a wavefront array processor architecture developed at ETCA and dedicated to real-time processing of digital video streams. The core of the architecture is a mesh-connected three-dimensional network of 1024 custom processing elements. Each processing element can perform up to 50 millions 8- or 16-bit operations per second, working with a 25 MHz clock frequency. Thus algorithms are facilitated by the routing capabilities of each processing element. The machine is fully data-driven and is a "pure" data-flow one since there are no address flows. Algorithms and architecture are described using a data-flow graphs formalism. Image processing applications are decomposed into elementary operators that correspond to physical processors in a one to one fashion. Several algorithms can be simultaneously mapped and independently executed on the processor network. Referring to the academic wavefront array paradigm, "exotic features" are exhibited. They are related to the wavefront propagation mode at run-time and to the heterogeneous nature of data-flow that are piped into or from the processor network. These features are shown to make the architecture well-suited for fast prototyping of low-level image processing automata.>
Georges Quénot, C. Coutelle, Jocelyn Sérot, Bertrand Y. Zavidovique
ASAP1
1993 Functional programming on a dataflow architecture: Applications in real-time image processing
Jocelyn Sérot, Georges Quénot, Bertrand Y. Zavidovique
Mach. Vis. Appl.2
1992 The 'orthogonal algorithm' for optical flow detection using dynamic programming
abstract
An algorithm for optical flow detection is introduced. It is based on an iterative search for a displacement field that minimizes the L/sub 1/ or L/sub 2/ distance between two images. Both changes are sliced into parallel and overlapping strips. Corresponding strips are aligned using dynamic programming. Two passes are performed using orthogonal slicing directions. This process is iterated in a pyramidal fashion by reducing the spacing and width of the strips. This algorithm provides very-high-quality matching for calibrated patterns as well as for human visual sensation. The results appear to be at least as good as those obtained with classical optical flow detection methods.>
Georges Quénot
ICASSP1
1992 The ETCA Data-Flow Functional Computer for Real-Time Image Processing
abstract
A data-flow computer, constituted of a large array of data-flow processors and programmed using a functional language, and its application to real-time image processing are presented. The approach integrates tightly and efficiently the data-flow architecture principle and the functional programming concept. It leads to a very simple and regular hardware. It also provides a very efficient and user-friendly software interface for application development. It remains applicable in systems that includes processors of different granularity. Results of an experimental system, running several image processing algorithm in real time at digital video speed, are reported.>
Georges Quénot, Bertrand Y. Zavidovique
ICCD1
1986 A dynamic time warp VLSI processor for continuous speech recognition
abstract
We present a VLSI processor designed to compute dynamic time warping algorithms for speech recognition with extreme rapidity. This processor works as a coprocessor in a classical system including a standard microprocessor and a digital signal processor. It uses its own local memory for reference utterances and intermediate results. It has been designed to give maximum efficiency on continuous speech recognition applications with or without syntax constraints. Its flexibility permits software optimisation and its use in a large number of different applications. We use a sequential approach for DTW computations and work along the time axis. All the calculations are carried out on each frame of the unknown utterance as soon as it arrives from the DSP and DTW computations therefore take place in real time. Response time is in hundredth of second; intermediate results are obtained before the end of the sentence. A system using this chip will be able to carry out continuous speech recognition in real timee on a vocabulary of 300 references. Many of those chips can be used in parallel on a single system.
Georges Quénot, Jean-Luc Gauvain, Jean-Jacques Gangolf, Joseph Mariani
ICASSP1