EDBT 2026 Demo / reviewers in the wild / expert
Guillaume Gravier
dblp:92/4096
· DBLP profile ↗
106ranked-venue papers
12as first author
10since 2021 · last 2026
0000-0002-2266-5682ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 78 · 10 first-author · 2 since 2021Artificial intelligence and machine learning · 56 · 7 first-author · 9 since 2021Databases, data management, data science and information retrieval · 12 · 4 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | n-Gram Injection into Transformers for Dynamic Language Model Adaptation in Handwritten Text RecognitionabstractTransformer-based encoder-decoder networks have recently achieved impressive results in handwritten text recognition, partly thanks to their auto-regressive decoder which implicitly learns a language model. However, such networks suffer from a large performance drop when evaluated on a target corpus whose language distribution is shifted from the source text seen during training. To retain recognition accuracy despite this language shift, we propose an external n-gram injection (NGI) for dynamic adaptation of the network's language modeling at inference time. Our method allows switching to an n-gram language model estimated on a corpus close to the target distribution, therefore mitigating bias without any extra training on target image-text pairs. We opt for an early injection of the n-gram into the transformer decoder so that the network learns to fully leverage text-only data at the low additional cost of n-gram inference. Experiments on three handwritten datasets demonstrate that the proposed NGI significantly reduces the performance gap between source and target corpora. Florent Meyer, Laurent Guichard, Yann Soullard, Denis Coquenet, Guillaume Gravier, Bertrand Coüasnon |
ICDAR (2) | 5 |
| 2026 | A Study on Building Efficient Zero-Shot Relation Extraction ModelsabstractInternational audience Hugo Thomas, Caio F. Corro, Guillaume Gravier, Pascale Sébillot |
LREC | 3 |
| 2025 | Relaxed Syntax Modeling in Transformers for Future-Proof License Plate Recognition
Florent Meyer, Laurent Guichard, Denis Coquenet, Guillaume Gravier, Yann Soullard, Bertrand Coüasnon |
ICDAR (4) | 4 |
| 2024 | One-shot relation retrieval in news archives: adapting N-way K-shot relation Classification for efficient knowledge extractionabstractOne-shot relation retrieval is the knowledge extraction task that consists in searching in a textual dataset for all occurrences of a relation of interest, named the source relation, characterized by a single example—a relation being a link between a pair of entities in an utterance. Performing this task on large datasets requires an intelligent system to automate the process, for instance when exploring news archives for press review or business intelligence. We propose a framework that leverages the representation learning capabilities of N-way K-shot models for few-shot relation Classification and extends these models to enable one-shot retrieval with a rejection class. At evaluation time, one-shot relation retrieval is performed in a N-way K-shot setting where 1 of the N ways (or relations) is the source relation and the N-1 others are distractors, i.e., relations modeling a rejection class. We benchmark this framework and investigate the influence of the number and the choice of distractors on the standard TACREV and FewRel datasets. Experimental results demonstrate the effectiveness of our approach to address this highly challenging task, however with high variability primarily induced by the type of the source relation. Experiments also highlight a sound strategy for the choice of distractors—a large number of distractors at an intermediate distance from the embedding of the source relation in the latent space learned by the model—, which provides a competing trade-of between recall and precision. This strategy is globally optimal but can however be surpassed on certain source relations by others, depending on the characteristics of the source relation, paving the way for future work. We finally show the substantial benefit of two-shot retrieval over one-shot retrieval, which sheds light on the design of actual intelligent applications leveraging one- or few-shot relation retrieval. Hugo Thomas, Guillaume Gravier, Pascale Sébillot |
KES | 2 |
| 2023 | Filtering Safe Temporal Motifs in Dynamic Graphs for Dissemination Purposes
Carolina Stephanie Jerônimo de Almeida, Simon Malinowski, Zenilton Kleber Gonçalves do Patrocínio Jr., Guillaume Gravier, Silvio Jamil Ferzoli Guimarães |
CIARP | 4 |
| 2023 | A Novel Method for Temporal Graph Classification based on Transitive ReductionabstractDomains such as bio-informatics, social network analysis, and computer vision, describe relations between entities and cannot be interpreted as vectors or fixed grids, instead, they are naturally represented by graphs. Often this kind of data evolves over time in a dynamic world, respecting a temporal order being known as temporal graphs. The latter became a challenge since subgraph patterns are very difficult to find and the distance between those patterns may change irregularly over time. While state-of-the-art methods are primarily designed for static graphs and may not capture temporal information, recent works have proposed mapping temporal graphs to static graphs to allow for the use of conventional static kernels and graph neural approaches. In this study, we compare the transitive reduction impact on these mappings in terms of accuracy and computational efficiency across different classification tasks. Furthermore, we introduce a novel mapping method using a transitive reduction approach that outperforms existing techniques in terms of classification accuracy. Our experimental results demonstrate the effectiveness of the proposed mapping method in improving the accuracy of supervised classification for temporal graphs while maintaining reasonable computational efficiency. Carolina Stephanie Jerônimo de Almeida, Zenilton Kleber Gonçalves do Patrocínio Jr., Simon Malinowski, Silvio Jamil Ferzoli Guimarães, Guillaume Gravier |
DSAA | 5 |
| 2023 | Regularization, Semi-supervision, and Supervision for a Plausible Attention-Based Explanation
Duc Hau Nguyen, Cyrielle Mallart, Guillaume Gravier, Pascale Sébillot |
NLDB | 3 |
| 2022 | Affect in Multimedia: Benchmarking Violent Scenes DetectionabstractIn this article, we report on the creation of a publicly available, common evaluation framework for Violent Scenes Detection (VSD) in Hollywood and YouTube videos. We propose a robust data set, the VSD96, with more than 96 hours of video of various genres, annotations at different levels of detail (e.g., shot-level, segment-level), annotations of mid-level concepts (e.g., blood, fire), various pre-computed multi-modal descriptors, and over 230 system output results as baselines. This is the most comprehensive data set available to this date tailored to the VSD task and was extensively validated during the MediaEval benchmarking campaigns. Furthermore, we provide an in-depth analysis of the crucial components of VSD algorithms, by reviewing the capabilities and the evolution of existing systems (e.g., overall trends and outliers, the influence of the employed features and fusion techniques, the influence of deep learning approaches). Finally, we discuss the possibility of going beyond state-of-the-art performance via an ad-hoc late fusion approach. Experimentation is carried out on the VSD96 data. We provide the most important lessons learned and gained insights. The increasing number of publications using the VSD96 data underline the importance of the topic. The presented and published resources are a practitioner's guide and also a strong baseline to overcome, which will help researchers for the coming years in analyzing aspects of audio-visual affect and violence detection in movies and videos. Mihai Gabriel Constantin, Liviu-Daniel Stefan, Bogdan Ionescu, Claire-Hélène Demarty, Mats Sjöberg, Markus Schedl, Guillaume Gravier |
IEEE Trans. Affect. Comput. | 7 |
| 2021 | A Study of the Plausibility of Attention between RNN Encoders in Natural Language InferenceabstractAttention maps in neural models for NLP are appealing to explain the decision made by a model, hopefully emphasizing words that justify the decision. While many empirical studies hint that attention maps can provide such justification from the analysis of sound examples, only a few assess the plausibility of explanations based on attention maps, i.e., the usefulness of attention maps for humans to understand the decision. These studies furthermore focus on text classification. In this paper, we report on a preliminary assessment of attention maps in a sentence comparison task, namely natural language inference. We compare the cross-attention weights between two RNN encoders with human-based and heuristic-based annotations on the eSNLI corpus. We show that the heuristic reasonably correlates with human annotations and can thus facilitate evaluation of plausible explanations in sentence comparison tasks. Raw attention weights however remain only loosely related to a plausible explanation. Duc Hau Nguyen, Guillaume Gravier, Pascale Sébillot |
ICMLA | 2 |
| 2021 | Hierarchical multi-label propagation using speaking face graphs for multimodal person discovery
Gabriel Barbosa da Fonseca, Gabriel Sargent, Ronan Sicre, Zenilton Kleber Gonçalves do Patrocínio Jr., Guillaume Gravier, Silvio Jamil Ferzoli Guimarães |
Multim. Tools Appl. | 5 |
| 2020 | Rethinking deep active learning: Using unlabeled data at model trainingabstractActive learning typically focuses on training a model on few labeled examples alone, while unlabeled ones are only used for acquisition. In this work we depart from this setting by using both labeled and unlabeled data during model training across active learning cycles. We do so by using unsupervised feature learning at the beginning of the active learning pipeline and semi-supervised learning at every active learning cycle, on all available data. The former has not been investigated before in active learning, while the study of latter in the context of deep learning is scarce and recent findings are not conclusive with respect to its benefit. Our idea is orthogonal to acquisition strategies by using more data, much like ensemble methods use more models. By systematically evaluating on a number of popular acquisition strategies and datasets, we find that the use of unlabeled data during model training brings a spectacular accuracy improvement in image classification, compared to the differences between acquisition strategies. We thus explore smaller label budgets, even one label per class. Oriane Siméoni, Mateusz Budnik, Yannis Avrithis, Guillaume Gravier |
ICPR | 4 |
| 2020 | A correlation-based entity embedding approach for robust entity linkingabstractEntity alignment is a crucial tool in knowledge discovery to reconcile knowledge from different sources. Recent state-of-the-art approaches leverage joint embedding of knowledge graphs (KGs) so that similar entities from different KGs are close in the embedded space. Whatever the joint embedding technique used, a seed set of aligned entities, often provided by (time-consuming) human expertise, is required to learn the joint KG embedding and/or a mapping between KG embeddings. In this context, a key issue is to limit the size and quality requirement for the seed. State-of-the-art methods usually learn the embedding by explicitly minimizing the distance between aligned entities from the seed and uniformly maximizing the distance for entities not in the seed. In contrast, we design a less restrictive optimization criterion that indirectly minimizes the distance between aligned entities in the seed by globally maximizing the dimension-wise correlation among all the embeddings of seed entities. Within an iterative entity alignment system, the correlation-based entity embedding function achieves state-of-the-art results and is shown to significantly increase robustness to the seed's size and accuracy. It ultimately enables fully unsupervised entity alignment using a seed automatically generated with a symbolic alignment method based on entities' names. Cheikh Brahim El Vaigh, François Torregrossa, Robin Allesiardo, Guillaume Gravier, Pascale Sébillot |
ICTAI | 4 |
| 2020 | On the Correlation of Word Embedding Evaluation MetricsabstractWord embeddings intervene in a wide range of natural language processing tasks. These geometrical representations are easy to manipulate for automatic systems. Therefore, they quickly invaded all areas of language processing. While they surpass all predecessors, it is still not straightforward why and how they do so. In this article, we propose to investigate all kind of evaluation metrics on various datasets in order to discover how they correlate with each other. Those correlations lead to 1) a fast solution to select the best word embeddings among many others, 2) a new criterion that may improve the current state of static Euclidean word embeddings, and 3) a way to create a set of complementary datasets, i.e. each dataset quantifies a different aspect of word embeddings. François Torregrossa, Vincent Claveau, Nihel Kooli, Guillaume Gravier, Robin Allesiardo |
LREC | 4 |
| 2020 | A Novel Path-Based Entity Relatedness Measure for Efficient Collective Entity Linking
Cheikh Brahim El Vaigh, François Goasdoué, Guillaume Gravier, Pascale Sébillot |
ISWC (1) | 3 |
| 2019 | Using Knowledge Base Semantics in Context-Aware Entity LinkingabstractEntity linking is a core task in textual document processing, which consists in identifying the entities of a knowledge base (KB) that are mentioned in a text. Approaches in the literature consider either independent linking of individual mentions or collective linking of all mentions. Regardless of this distinction, most approaches rely on the Wikipedia encyclopedic KB in order to improve the linking quality, by exploiting its entity descriptions (web pages) or its entity interconnections (hyperlink graph of web pages). In this paper, we devise a novel collective linking technique which departs from most approaches in the literature by relying on a structured RDF KB. This allows exploiting the semantics of the interrelationships that candidate entities may have at disambiguation time rather than relying on raw structural approximation based on Wikipedia's hyperlink graph. The few approaches that also use an RDF KB simply rely on the existence of a relation between the candidate entities to which mentions may be linked. Instead, we weight such relations based on the RDF KB structure and propose an efficient decoding strategy for collective linking. Experiments on standard benchmarks show significant improvement over the state of the art. Cheikh Brahim El Vaigh, François Goasdoué, Guillaume Gravier, Pascale Sébillot |
DocEng | 3 |
| 2018 | A Study on Multimodal Video Hyperlinking with Visual AggregationabstractVideo hyperlinking offers a way to explore a video collection, making use of links that connect segments having related content. Hyperlinking systems thus seek to automatically create links by connecting given anchor segments to relevant targets within the collection. In this paper, we further investigate multimodal representations of video segments in a hyper-linking system based on bidirectional deep neural networks, which achieved state-of-the-art results in the TRECVid 2016 evaluation. A systematic study of different input representations is done with a focus on the aggregation of the representation of multiple keyframes. This includes, in particular, the use of memory vectors as a novel aggregation technique, which provides a significant improvement over other aggregation methods on the final hyperlinking task. Additionally, the use of metadata is investigated leading to increased performance and lower computational requirements for the system. Mateusz Budnik, Mikail Demirdelen, Guillaume Gravier |
ICME | 3 |
| 2017 | Linking Multimedia Content for Efficient News BrowsingabstractAs the amount of news information available online grows, media are in need of advanced tools to explore the information surrounding specific events before writing their own piece of news, e.g., adding context and insight. While many tools exist to extract information from large datasets, they do not offer an easy way to gain insight from a news collection by browsing, going from article to article and viewing unaltered original content. Such browsing tools require the creation of rich underlying structures such as graph representations. These representations can be further enhanced by typing links that connect nodes, in order to inform the user on the nature of their relation. In this article, we introduce an efficient way to generate links between news items in order to obtain an easily navigable graph, and enrich this graph by automatically typing created links. User evaluations are conducted on real world data in order to assess for the interest of both the graph representation and link typing in a press reviewing task, showing a significant improvement compared to classical search engines. Rémi Bois, Guillaume Gravier, Eric Jamet, Emmanuel Morin, Maxime Robert, Pascale Sébillot |
ICMR | 2 |
| 2017 | Generative Adversarial Networks for Multimodal Representation Learning in Video HyperlinkingabstractContinuous multimodal representations suitable for multimodal information retrieval are usually obtained with methods that heavily rely on multimodal autoencoders. In video hyperlinking, a task that aims at retrieving video segments, the state of the art is a variation of two interlocked networks working in opposing directions. These systems provide good multimodal embeddings and are also capable of translating from one representation space to the other. Operating on representation spaces, these networks lack the ability to operate in the original spaces (text or image), which makes it difficult to visualize the crossmodal function, and do not generalize well to unseen data. Recently, generative adversarial networks have gained popularity and have been used for generating realistic synthetic data and for obtaining high-level, single-modal latent representation spaces. In this work, we evaluate the feasibility of using GANs to obtain multimodal representations. We show that GANs can be used for multimodal representation learning and that they provide multimodal representations that are superior to representations obtained with multimodal autoencoders. Additionally, we illustrate the ability of visualizing crossmodal translations that can provide human-interpretable insights on learned GAN-based video hyperlinking models. Vedran Vukotic, Christian Raymond, Guillaume Gravier |
ICMR | 3 |
| 2017 | NexGenTV: Providing Real-Time Insight during Political Debates in a Second Screen ApplicationabstractSecond screen applications are becoming key for broadcasters exploiting the convergence of TV and Internet. Authoring such applications however remains costly. In this paper, we present a second screen authoring application that leverages multimedia content analytics and social media monitoring. A back-office is dedicated to easy and fast content ingestion, segmentation, description and enrichment with links to entities and related content. From the back-end, broadcasters can push enriched content to front-end applications providing customers with highlights, entity and content links, overviews of social network, etc. The demonstration operates on political debates ingested during the 2017 French presidential election, enabling insights on the debates. Olfa Ben Ahmed, Gabriel Sargent, Florian Garnier, Benoit Huet, Vincent Claveau, Laurence Couturier, Raphaël Troncy, Guillaume Gravier, Philémon Bouzy, Fabrice Leménorel |
ACM Multimedia | 8 |
| 2017 | Exploiting Multimodality in Video Hyperlinking to Improve Target Diversity
Rémi Bois, Vedran Vukotic, Anca-Roxana Simon, Ronan Sicre, Christian Raymond, Pascale Sébillot, Guillaume Gravier |
MMM (2) | 7 |
| 2017 | Special Issue on Content Based Multimedia Indexing
Ioannis Kompatsiaris, Guillaume Gravier |
Multim. Tools Appl. | 2 |
| 2017 | Content-based unsupervised segmentation of recurrent TV programs using grammatical inference
Bingqing Qu, Félicien Vallet, Jean Carrive, Guillaume Gravier |
Multim. Tools Appl. | 4 |
| 2016 | Audio word similarity for clustering with zero resources based on iterative HMM classificationabstractRecent work on zero resource word discovery makes intensive use of audio fragment clustering to find repeating speech patterns. In the absence of acoustic models, the clustering step traditionally relies on dynamic time warping (DTW) to compare two samples and thus suffers from the known limitations of this technique. We propose a new sample comparison method, called similarity by iterative classification, that exploits the modeling capacities of hidden Markov models (HMM) with no supervision. The core idea relies on the use of HMMs trained on randomly labeled data and exploits the fact that similar samples are more likely to be classified together by a large number of random classifiers than dissimilar ones. The resulting similarity measure is compared to DTW on two tasks, namely nearest neighbor retrieval and clustering, showing that the generalization capabilities of probabilistic machine learning significantly benefit to audio word comparison and overcome many of the limitations of DTW-based comparison. Amelie Royer, Guillaume Gravier, Vincent Claveau |
ICASSP | 2 |
| 2016 | A Step Beyond Local Observations with a Dialog Aware Bidirectional GRU Network for Spoken Language UnderstandingabstractInternational audience Vedran Vukotic, Christian Raymond, Guillaume Gravier |
INTERSPEECH | 3 |
| 2016 | Bidirectional Joint Representation Learning with Symmetrical Deep Neural Networks for Multimodal and Crossmodal ApplicationsabstractCommon approaches to problems involving multiple modalities (classification, retrieval, hyperlinking, etc.) are early fusion of the initial modalities and crossmodal translation from one modality to the other. Recently, deep neural networks, especially deep autoencoders, have proven promising both for crossmodal translation and for early fusion via multimodal embedding. In this work, we propose a flexible crossmodal deep neural network architecture for multimodal and crossmodal representation. By tying the weights of two deep neural networks, symmetry is enforced in central hidden layers thus yielding a multimodal representation space common to the two original representation spaces. The proposed architecture is evaluated in multimodal query expansion and multimodal retrieval tasks within the context of video hyperlinking. Our method demonstrates improved crossmodal translation capabilities and produces a multimodal embedding that significantly outperforms multimodal embeddings obtained by deep autoencoders, resulting in an absolute increase of 14.14 in precision at 10 on a video hyperlinking task. Vedran Vukotic, Christian Raymond, Guillaume Gravier |
ICMR | 3 |
| 2016 | Shaping-Up Multimedia Analytics: Needs and Expectations of Media Professionals
Guillaume Gravier, Martin Ragot, Laurent Amsaleg, Rémi Bois, Grégoire Jadi, Eric Jamet, Laura Monceaux, Pascale Sébillot |
MMM (2) | 1 |
| 2016 | Partial least squares for face hashing
Cassio E. dos Santos, Ewa Kijak, Guillaume Gravier, William Robson Schwartz |
Neurocomputing | 3 |
| 2015 | Is it time to Switch to word embedding and recurrent neural networks for spoken language understanding?abstractRecently, word embedding representations have been investigated for slot filling in Spoken Language Understanding, along with the use of Neural Networks as classifiers.Neural Networks, especially Recurrent Neural Networks, that are specifically adapted to sequence labeling problems, have been applied successfully on the popular ATIS database.In this work, we make a comparison of this kind of models with the previously state-of-the-art Conditional Random Fields (CRF) classifier on a more challenging SLU database.We show that, despite efficient word representations used within these Neural Networks, their ability to process sequences is still significantly lower than for CRF, while also having a drawback of higher computational costs, and that the ability of CRF to model output label dependencies is crucial for SLU. Vedran Vukotic, Christian Raymond, Guillaume Gravier |
INTERSPEECH | 3 |
| 2015 | Overview of the 2015 Workshop on Speech, Language and Audio in MultimediaabstractThe Workshop on Speech, Language and Audio in Multimedia (SLAM) positions itself at at the crossroad of multiple scientific fields (music and audio processing, speech processing, natural language processing and multimedia) to discuss and stimulate research results, projects, datasets and benchmarks initiatives where audio, speech and language are applied to multimedia data. While the first two editions were collocated with major speech events, SLAM'15 is deeply rooted in the multimedia community, opening up to computer vision and multimodal fusion. To this end, the workshop emphasizes video hyperlinking as an showcase where computer vision meets speech and language. Such techniques provide a powerful illustration of how multimedia technologies incorporating speech, language and audio can make multimedia content collections better accessible, and thereby more useful, to users. Guillaume Gravier, Gareth J. F. Jones, Martha A. Larson, Roeland Ordelman |
ACM Multimedia | 1 |
| 2015 | Content-Based Discovery of Multiple Structures from Episodes of Recurrent TV Programs Based on Grammatical Inference
Bingqing Qu, Félicien Vallet, Jean Carrive, Guillaume Gravier |
MMM (1) | 4 |
| 2015 | VSD, a public dataset for the detection of violent scenes in movies: design, annotation, analysis and evaluation
Claire-Hélène Demarty, Cédric Penet, Mohammad Soleymani 0001, Guillaume Gravier |
Multim. Tools Appl. | 4 |
| 2015 | Variability modelling for audio events detection in movies
Cédric Penet, Claire-Hélène Demarty, Guillaume Gravier, Patrick Gros |
Multim. Tools Appl. | 3 |
| 2014 | Content-based inference of hierarchical structural grammar for recurrent TV programs using multiple sequence alignmentabstractRecently, unsupervised approaches were introduced to analyze the structure of TV programs, relying on the discovery of repeated elements within a program or across multiple episodes of the same program. These methods can discover key repeating elements, such as jingles and separators, however they cannot infer the entire structure of a program. In this paper, we propose a hierarchical use of grammatical inference to yield a temporal grammar of a program from a collection of episodes, discovering both the vocabulary of the grammar and the temporal organization of the words from the vocabulary. Using a set of basic event detectors and simple filtering techniques to detect repeating elements of interest, a symbolic representation of each episode is derived based on minimal domain knowledge. Grammatical inference based on multiple sequence alignment is then used in a hierarchical manner to provide a temporal grammar of the program at various levels of details. Experimental validation is performed on 3 distinct types of programs on 4 datasets. Qualitative analyses show that the grammars inferred at the different levels of the hierarchy are relevant and can be obtained from a fairly limited number of episodes. Bingqing Qu, Félicien Vallet, Jean Carrive, Guillaume Gravier |
ICME | 4 |
| 2014 | Audio thumbnails for spoken content without transcription based on a maximum motif coverage criterionabstractInternational audience Guillaume Gravier, Nathan Souviraà-Labastie, Sébastien Campion, Frédéric Bimbot |
INTERSPEECH | 1 |
| 2014 | The ETAPE speech processing evaluation
Olivier Galibert, Jérémy Leixa, Gilles Adda, Khalid Choukri, Guillaume Gravier |
LREC | 5 |
| 2014 | Bridging the gap between speech technology and natural language processing: an evaluation toolbox for term discovery systems
Bogdan Ludusan, Maarten Versteegh, Aren Jansen, Guillaume Gravier, Xuan-Nga Cao, Mark Johnson 0001, Emmanuel Dupoux |
LREC | 4 |
| 2014 | Language independent search in MediaEval's Spoken Web Search task
Florian Metze, Xavier Anguera Miró, Etienne Barnard, Marelie H. Davel, Guillaume Gravier |
Comput. Speech Lang. | 5 |
| 2014 | Classification-oriented structure learning in Bayesian networks for multimodal event detection in videos
Guillaume Gravier, Claire-Hélène Demarty, Siwar Baghdadi, Patrick Gros |
Multim. Tools Appl. | 1 |
| 2013 | Oriented pooling for dense and non-dense rotation-invariant featuresabstractInternational audience Wanlei Zhao, Guillaume Gravier, Hervé Jégou |
BMVC | 2 |
| 2013 | Leveraging Lexical Cohesion and Disruption for Topic SegmentationabstractTopic segmentation classically relies on one of two criteria, either finding areas with coherent vocabulary use or detecting discontinuities.In this paper, we propose a segmentation criterion combining both lexical cohesion and disruption, enabling a trade-off between the two.We provide the mathematical formulation of the criterion and an efficient graph based decoding algorithm for topic segmentation.Experimental results on standard textual data sets and on a more challenging corpus of automatically transcribed broadcast news shows demonstrate the benefit of such a combination.Gains were observed in all conditions, with segments of either regular or varying length and abrupt or smooth topic shifts.Long segments benefit more than short segments.However the algorithm has proven robust on automatic transcripts with short segments and limited vocabulary reoccurrences. Anca-Roxana Simon, Guillaume Gravier, Pascale Sébillot |
EMNLP | 2 |
| 2013 | The spoken web search task at MediaEval 2012abstractIn this paper, we describe the “Spoken Web Search” Task, which was held as part of the 2012 MediaEval benchmark evaluation campaign. The purpose of this task was to perform audio search with audio input in four languages, with very few resources being available. Continuing in the spirit of the 2011 SpokenWeb Search Task, which used speech from four Indian languages, the 2012 data was taken from the LWAZI corpus, to provide even more diversity and allow for a task that will allow both zero resource “pattern matching” approaches and “speech recognition” based approaches to participate. In this paper, we summarize the results from several independent systems, developed by nine teams, analyze their performance, and provide directions for future research. Florian Metze, Xavier Anguera Miró, Etienne Barnard, Marelie H. Davel, Guillaume Gravier |
ICASSP | 5 |
| 2013 | MODIS: an audio motif discovery software
Laurence Catanese, Nathan Souviraà-Labastie, Bingqing Qu, Sébastien Campion, Guillaume Gravier, Emmanuel Vincent 0001, Frédéric Bimbot |
INTERSPEECH | 5 |
| 2013 | Searching for Near-Duplicate Video Sequences from a Scalable Sequence AlignerabstractNear-duplicate video sequence identification consists in identifying real positions of a specific video clip in a video stream stored in a database. To address this problem, we propose a new approach based on a scalable sequence aligner borrowed from proteomics. Sequence alignment is performed on symbolic representations of features extracted from the input videos, based on an algorithm originally applied to bio-informatics. Experimental results demonstrate that our method performance achieved 94% recall with 100% precision, with an average searching time of about 1 second. Leonardo S. de Oliveira, Zenilton Kleber Gonçalves do Patrocínio Jr., Silvio Jamil Ferzoli Guimarães, Guillaume Gravier |
ISM | 4 |
| 2013 | Multimedia information seeking through search and hyperlinkingabstractSearching for relevant webpages and following hyperlinks to related content is a widely accepted and effective approach to information seeking on the textual web. Existing work on multimedia information retrieval has focused on search for individual relevant items or on content linking without specific attention to search results. We describe our research exploring integrated multimodal search and hyperlinking for multimedia data. Our investigation is based on the MediaEval 2012 Search and Hyperlinking task. This includes a known-item search task using the Blip10000 internet video collection, where automatically created hyperlinks link each relevant item to related items within the collection. The search test queries and link assessment for this task was generated using the Amazon Mechanical Turk crowdsourcing platform. Our investigation examines a range of alternative methods which seek to address the challenges of search and hyperlinking using multimodal approaches. The results of our experiments are used to propose a research agenda for developing effective techniques for search and hyperlinking of multimedia content. Maria Eskevich, Gareth J. F. Jones, Robin Aly, Roeland Ordelman, Danish Nadeem, Camille Guinaudeau, Guillaume Gravier, Pascale Sébillot, Tom De Nies, Pedro Debevere, Rik Van de Walle, Petra Galuscáková, Pavel Pecina, Martha A. Larson |
ICMR | 8 |
| 2013 | Retrieving geo-location of videos with a divide & conquer hierarchical multimodal approachabstractThis paper presents a strategy to identify the geographic location of videos. First, it relies on a multi-modal cascade pipeline that exploits the available sources of information, namely the user's upload history, his social network and a visual-based matching technique. Second, we present a novel divide & conquer strategy to better exploit the tags associated with the input video. It pre-selects one or several geographic area of interest of higher expected relevance and performs a deeper analysis inside the selected area(s) to return the coordinates most likely to be related to the input tags. The experiments were conducted as part of the MediaEval 2012 Placing Task. Our approach, which differs significantly from the other submitted techniques, achieves the best results on this benchmark when considering the same amount of external information, i.e. when not using any gazetteers nor any other kind of external information. Michele Trevisiol, Hervé Jégou, Jonathan Delhumeau, Guillaume Gravier |
ICMR | 4 |
| 2013 | Sim-min-hash: an efficient matching technique for linking large image collectionsabstractOne of the most successful method to link all similar images within a large collection is min-Hash, which is a way to significantly speed-up the comparison of images when the underlying image representation is bag-of-words. However, the quantization step of min-Hash introduces important information loss. In this paper, we propose a generalization of min-Hash, called Sim-min-Hash, to compare sets of real-valued vectors. We demonstrate the effectiveness of our approach when combined with the Hamming embedding similarity. Experiments on large-scale popular benchmarks demonstrate that Sim-min-Hash is more accurate and faster than min-Hash for similar image search. Linking a collection of one million images described by 2 billion local descriptors is done in 7 minutes on a single core machine. Wanlei Zhao, Hervé Jégou, Guillaume Gravier |
ACM Multimedia | 3 |
| 2013 | Dynamic Combination of Automatic Speech Recognition Systems by Driven DecodingabstractCombining automatic speech recognition (ASR) systems generally relies on the posterior merging of the outputs or on acoustic cross-adaptation. In this paper, we propose an integrated approach where outputs of secondary systems are integrated in the search algorithm of a primary one. In this driven decoding algorithm (DDA), the secondary systems are viewed as observation sources that should be evaluated and combined to others by a primary search algorithm. DDA is evaluated on a subset of the ESTER I corpus consisting of 4 hours of French radio broadcast news. Results demonstrate DDA significantly outperforms vote-based approaches: we obtain an improvement of 14.5% relative word error rate over the best single-systems, as opposed to the the 6.7% with a ROVER combination. An in-depth analysis of the DDA shows its ability to improve robustness (gains are greater in adverse conditions) and a relatively low dependency on the search algorithm. The application of DDA to both and beam-search-based decoder yields similar performances. Benjamin Lecouteux, Georges Linarès, Yannick Estève, Guillaume Gravier |
IEEE Trans. Speech Audio Process. | 4 |
| 2012 | BABAZ: A large scale audio search system for video copy detectionabstractThis paper presents BABAZ, an audio search system to search modified segments in large databases of music or video tracks. It is based on an efficient audio feature matching system which exploits the reciprocal nearest neighbors to produce a per-match similarity score. Temporal consistency is taken into account based on the audio matches, and boundary estimation allows the precise localization of the matching segments. The method is mainly intended for video retrieval based on their audio track, as typically evaluated in the copy detection task of TRECVID evaluation campaigns. The evaluation conducted on music retrieval shows that our system is comparable to a reference audio fingerprinting system for music retrieval, and significantly outperforms it on audio-based video retrieval, as shown by our experiments conducted on the dataset used in the copy detection task of TRECVID'2010 campaign. Hervé Jégou, Jonathan Delhumeau, Jiangbo Yuan, Guillaume Gravier, Patrick Gros |
ICASSP | 4 |
| 2012 | The Spoken Web Search Task at MediaEval 2011abstractIn this paper, we describe the “Spoken Web Search” Task, which was held as part of the 2011 MediaEval benchmark campaign. The purpose of this task was to perform audio search with audio input in four languages, with very few resources being available in each language. The data was taken from “spoken web” material collected over mobile phone connections by IBM India. We present results from several independent systems, developed by five teams and using different approaches, compare them, and provide analysis and directions for future research. Florian Metze, Nitendra Rajput, Xavier Anguera Miró, Marelie H. Davel, Guillaume Gravier, Charl Johannes van Heerden, Gautam Varma Mantena, Armando Muscariello, Kishore Prahallad, Igor Szöke, Javier Tejedor |
ICASSP | 5 |
| 2012 | Multimodal information fusion and temporal integration for violence detection in moviesabstractThis paper presents a violent shots detection system that studies several methods for introducing temporal and multimodal information in the framework. It also investigates different kinds of Bayesian network structure learning algorithms for modelling these problems. The system is trained and tested using the MediaEval 2011 Affect Task corpus, which comprises of 15 Hollywood movies. It is experimentally shown that both multimodality and temporality add interesting information into the system. Moreover, the analysis of the links between the variables of the resulting graphs yields important observations about the quality of the structure learning algorithms. Overall, our best system achieved 50% false alarms and 3% missed detection, which is among the best submissions in the MediaEval campaign. Cédric Penet, Claire-Hélène Demarty, Guillaume Gravier, Patrick Gros |
ICASSP | 3 |
| 2012 | Unsupervised Mining of Multiple Audiovisually Consistent Clusters for Video Structure AnalysisabstractWe address the problem of detecting multiple audiovisual events related to the edit structure of a video by incorporating an unsupervised cluster analysis technique into a cluster selection method designed to measure coherence between audio and visual segments. First, mutual information measure is used to select audio-visually consistent clusters from two dendrograms representing hierarchical clustering results respectively for the audio and visual modalities. A cluster analysis technique is then applied to define events from the audio-visual (AV) clusters with segments co-occurring frequently. Candidate events are then characterized by groups of AV clusters from which models are built by automatically selecting positive and negative examples. Experiments on the standard Canal9 data set demonstrates that our method is capable of discovering multiple audiovisual events in a totally unsupervised manner. Anh-Phuong Ta, Guillaume Gravier |
ICME | 2 |
| 2012 | Lexical-phonetic automata for spoken utterance indexing and retrievalabstractInternational audience Julien Fayolle, Murat Saraclar, Fabienne Moreau, Christian Raymond, Guillaume Gravier |
INTERSPEECH | 5 |
| 2012 | Integrating Stress Information in Large Vocabulary Continuous Speech RecognitionabstractIn this paper we propose a novel method for integrating stress information in the decoding step of a speech recognizer.A multiscale rhythm model was used to determine the stress scores for each syllable, which are further used to reinforce paths during search.Two strategies for integrating the stress were employed: the first one reinforces paths through all the syllables with a value proportional to the their stress score, while the second one enhances paths passing only through stressed syllables, but with a constant value.The former strategy slightly outperforms the later, bringing a relative improvement of more than 2% over the baseline.Furthermore, the stress information proved to be a robust feature, by performing well even for foreign-accented speech. Bogdan Ludusan, Stefan Ziegler, Guillaume Gravier |
INTERSPEECH | 3 |
| 2012 | Using broad phonetic classes to guide search in automatic speech recognitionabstractThis work presents a novel framework to guide the Viterbi decoding process of a hidden Markov model based speech recognition system by means of broad phonetic classes. In a first step, decision trees are employed, along with frame and segment based attributes, in order to detect broad phonetic classes in the speech signal. Then, the detected phonetic classes are used to reinforce paths in the search process, either at every frame or at phonetically significant landmarks. Results obtained on French broadcast news data show a relative improvement in word error rate of about 2% with respect to the baseline. Stefan Ziegler, Bogdan Ludusan, Guillaume Gravier |
INTERSPEECH | 3 |
| 2012 | The ETAPE corpus for the evaluation of speech-based TV content processing in the French language
Guillaume Gravier, Gilles Adda, Niklas Paulsson, Matthieu Carré 0003, Aude Giraudel, Olivier Galibert |
LREC | 1 |
| 2012 | Texmix: an automatically generated news navigation portalabstractThe Texmix demonstration presents an original interface to navigate a collection of broadcast news shows, exploiting speech transcription, natural language processing and image retrieval techniques. Navigation is performed through keywords search or through time or through maps, with links automatically created either between reports to follow the story or to the Web to know more about a story, a person or a fact. Image search technology is also integrated to find portions of the collection with similar images. We also present two original features to dynamically access videos, namely dynamic summary and geotagging. Morgan Bréhinier, Sébastien Campion, Guillaume Gravier |
ICMR | 3 |
| 2012 | Improving Cluster Selection and Event Modeling in Unsupervised Mining for Automatic Audiovisual Video Structuring
Anh-Phuong Ta, Mathieu Ben, Guillaume Gravier |
MMM | 3 |
| 2012 | Towards a new speech event detection approach for landmark-based speech recognitionabstractIn this work, we present a new approach for the classification and detection of speech units for the use in landmark or event-based speech recognition systems. We use segmentation to model any time-variable speech unit by a fixed-dimensional observation vector, in order to train a committee of boosted decision stumps on labeled training data. Given an unknown speech signal, the presence of a desired speech unit is estimated by searching for each time frame the corresponding segment, that provides the maximum classification score. This approach improves the accuracy of a phoneme classification task by 1.7%, compared to classification using HMMs. Applying this approach to the detection of broad phonetic landmarks inside a landmark-driven HMM-based speech recognizer significantly improves speech recognition. Stefan Ziegler, Bogdan Ludusan, Guillaume Gravier |
SLT | 3 |
| 2012 | Enhancing lexical cohesion measure with confidence measures, semantic relations and language model interpolation for multimedia spoken content topic segmentation
Camille Guinaudeau, Guillaume Gravier, Pascale Sébillot |
Comput. Speech Lang. | 2 |
| 2012 | Unsupervised Motif Acquisition in Speech via Seeded Discovery and Template Matching CombinationabstractThis paper describes and evaluates a computational architecture to discover and collect occurrences of speech repetitions, or motifs, in a totally unsupervised fashion, that is in the absence of acoustic, lexical or pronunciation modeling and training material. In the last few years, this task has known an increasing interest from the speech community because of a) its potential applicability in spoken document processing (as a preliminary step to summarization, topic clustering, etc.) and b) its novel methodology, that defines a new paradigm to speech processing that circumvents the issues common to all supervised, trained technologies. The contributions implied by the proposed system are two-fold: 1) the design of a discovery strategy that detects repetitions by extending matches of motif fragments, called seeds; 2) the implementation of template matching techniques to detect acoustically close segments, based on dynamic time warping (DTW) and self-similarity matrix (SSM) comparison of speech templates, in contrast to the decoding procedures of model-based recognition systems. The architecture is thoroughly evaluated on several hours of French broadcast news shows according to various parameter settings and acoustic features, namely mel-frequency cepstral coefficients (MFCCs) and different types of posteriorgrams: Gaussian mixture model (GMM)-based, and phone-based posteriors, in both language-matched and mismatched conditions. The evaluation highlights a) the improved robustness of the system that jointly employs DTW and SSM and b) the relevant impact of language-specific features to acoustic similarity detection based on template matching. Armando Muscariello, Guillaume Gravier, Frédéric Bimbot |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | Automatically finding semantically consistent n-grams to add new words in LVCSR systemsabstractThis paper presents a new method to automatically add re-grams containing out-of-vocabulary (OOV) words to a baseline language model (LM), where these re-grams are sought to be grammatically correct and to make sense according to the meaning of OOV words. First, this method consists in determining the word sequences, i.e., re-grams, in which the usage of a given OOV word is the most semantically consistent. Then, conditional probabilities of these re-grams have to be computed. To do this, semantic relations between words are used to assimilate each OOV word to several equivalent in vocabulary words. Based on these last words, n-grams from the baseline LM are re-used to find the word sequences to be added and to compute their probabilities. After augmenting the vocabulary and launching a recognition process, experiments show that our method results in WER improvements which are comparable to those obtained using a state-of-the-art open vocabulary LM. Gwénolé Lecorvé, Guillaume Gravier, Pascale Sébillot |
ICASSP | 2 |
| 2011 | Towards robust word discovery by self-similarity matrix comparisonabstractWord discovery is the task of discovering and collecting occurrences of repeating words in the absence of prior acoustic and linguistic knowledge, or training material. The capability of extracting such patterns (or motifs) represents a preliminary step towards automatic mining of contentful information in spoken documents. The absence of modelling and training data, forces the use of direct pattern matching of speech templates, which, in turn, is sensitive to speech variability, like the inter-speaker one, for instance. In the present work, a variability tolerant pattern recognition technique is proposed that relies on the comparison of self similarity matrices of speech sequences. The joint use of such technique and a dynamic time warping dissimilarity measure, is shown to account for more variability with respect to the DTW-based system alone, as demonstrated on several hours of broadcast news shows. Armando Muscariello, Guillaume Gravier, Frédéric Bimbot |
ICASSP | 2 |
| 2011 | Unsupervised mining of audiovisually consistent segments in videos with application to structure analysisabstractIn this paper, a multimodal event mining technique is proposed to discover repeating video segments exhibiting audio and visual consistency in a totally unsupervised manner. The mining strategy first exploits independent audio and visual cluster analysis to provide segments which are consistent in both their visual and audio modalities, thus likely corresponding to a unique underlying event. A subsequent modeling stage using discriminative models enables accurate detection of the underlying event throughout the video. Event mining is applied to unsupervised video structure analysis, using simple heuristics on occurrence patterns of the events discovered to select those relevant to the video structure. Results on TV programs ranging from news to talk shows and games, show that structurally relevant events are discovered with precisions ranging from 87% to 98% and recalls from 59% to 94 %. Mathieu Ben, Guillaume Gravier |
ICME | 2 |
| 2011 | A Study on Auditory Feature Spaces for Speech-Driven Lip AnimationabstractInternational audience Guylaine Le Jan, Yannick Benezeth, Guillaume Gravier, Frédéric Bimbot |
INTERSPEECH | 3 |
| 2011 | Zero-Resource Audio-Only Spoken Term Detection Based on a Combination of Template Matching Techniquesabstractspoken term detection, template matching, unsupervised learning, posterior features Armando Muscariello, Guillaume Gravier, Frédéric Bimbot |
INTERSPEECH | 2 |
| 2010 | CRF-based combination of contextual features to improve a posteriori word-level confidence measuresabstractInternational audience Julien Fayolle, Fabienne Moreau, Christian Raymond, Guillaume Gravier, Patrick Gros |
INTERSPEECH | 4 |
| 2010 | Improving ASR-based topic segmentation of TV programs with confidence measures and semantic relationsabstractInternational audience Camille Guinaudeau, Guillaume Gravier, Pascale Sébillot |
INTERSPEECH | 2 |
| 2010 | Morpho-syntactic post-processing of N-best lists for improved French automatic speech recognition
Stéphane Huet, Guillaume Gravier, Pascale Sébillot |
Comput. Speech Lang. | 2 |
| 2009 | Speaker adaptation by variable reference model subspace and application to large vocabulary speech recognitionabstractRecently, we presented a rapid speaker adaptation technique, reference model interpolation (RMI), which is based on the linear interpolation of speaker-dependent models and the a posteriori selection of reference models. The approach uses the a priori knowledge provided by a set of representative speakers to guide the estimation of a new speaker model in the speaker space. RMI achieved rapid supervised adaptation in phoneme decoding tasks. In this paper, we present two new results of RMI: firstly, we apply the RMI technique in a practical large vocabulary continuous speech recognition (LVCSR) system with unsupervised instantaneous adaptation. Secondly, we propose an evolutional subspace scenario which integrates the slow update of reference models with RMI rapid adaptation to achieve incremental adaptation. The unsupervised adaptation experiments carried out on broadcast news transcription task show encouraging results for both instantaneous and incremental adapatation. Wen Xuan Teng, Guillaume Gravier, Frédéric Bimbot, Frédéric Soufflet |
ICASSP | 2 |
| 2009 | The ester 2 evaluation campaign for the rich transcription of French radio broadcastsabstractThis paper reports on the final results of the ESTER 2evaluation campaign held from 2007 to April 2009. The aim of this campaign was to evaluate automatic radio broadcasts rich transcription systems for the French language. The evaluation tasks were divided into three main categories: audio event detection and tracking (e.g., speech vs. music, speaker tracking), orthographic transcription, and information extraction. The paper describes the data provided for the campaign, the task definitions and evaluation protocols as well as the results. 1. Sylvain Galliano, Guillaume Gravier, Laura Chaubard |
INTERSPEECH | 2 |
| 2009 | Constraint selection for topic-based MDI adaptation of language modelsabstractThis paper presents an unsupervised topic-based language model adaptation method which specializes the standard minimum information discrimination approach by identifying and combining topic-specific features.By acquiring a topic terminology from a thematically coherent corpus, language model adaptation is restrained to the sole probability re-estimation of n-grams ending with some topic-specific words, keeping other probabilities untouched.Experiments are carried out on a large set of spoken documents about various topics.Results show significant perplexity and recognition improvements which outperform results of classical adaptation techniques. Gwénolé Lecorvé, Guillaume Gravier, Pascale Sébillot |
INTERSPEECH | 2 |
| 2009 | Audio keyword extraction by unsupervised word discoveryabstractInternational audience Armando Muscariello, Guillaume Gravier, Frédéric Bimbot |
INTERSPEECH | 2 |
| 2009 | Can Automatic Speech Transcripts Be Used for Large Scale TV Stream Description and Structuring?abstractThe increasing quantity of TV material requires methods to help users navigate such data streams. Automatically associating a short textual description to each program in a stream, is a first stage to navigating or structuring tasks. Speech contained in TV broadcasts---accessible by means of automatic speech recognition systems in the absence of closed caption---is a highly valuable semantic clue that might be used to link existing textual description such as program guides, with video segments corresponding to program. However, high word error rates are to be expected on some programs, likely to jeopardize the usefulness of transcripts. The goal of this article is to determine to what extent automatic transcripts of TV streams, for various types of programs, can be used for structuring or navigating tasks. To this end, word-based and phonetic-based automatic association between video segments and program descriptions is used as a case study. We show that descriptions from a program guide can be associated with video segments with an accuracy of up to 65% and provide a valuable description to validate existing program labels. Such associations constitute a first stage for structuring task as they enable video segment textual characterization. Camille Guinaudeau, Guillaume Gravier, Pascale Sébillot |
ISM | 2 |
| 2009 | Variability Tolerant Audio Motif Discovery
Armando Muscariello, Guillaume Gravier, Frédéric Bimbot |
MMM | 2 |
| 2008 | An unsupervised web-based topic language model adaptation methodabstractThis paper focuses on a solution to better adapt ASR systems, whose language models (LM) are usually trained on topic-independent corpora, to new topics, in particular in the case of broadcast news. We propose a new complete and fully unsupervised technique that selects keywords from each segment using information retrieval methods, to build a thematically coherent adaptation corpus from the Internet. The LM used for the initial transcription is then adapted before rescoring word lattices. Experimental results demonstrate the validity of the proposed adaptation technique with a significant reduction of the perplexity after LM adaptation. Word error rates are also improved in some cases though to a lesser extent. Index Terms — Speech recognition, natural languages, Internet 1. Gwénolé Lecorvé, Guillaume Gravier, Pascale Sébillot |
ICASSP | 2 |
| 2008 | Generalized driven decoding for speech recognition system combinationabstractDriven decoding algorithm (DDA) is initially an integrated approach for the combination of 2 speech recognition (ASR) systems. It consists in guiding the search algorithm of a primary ASR system by the one-best hypothesis of an auxiliary system. In this paper, we generalize DDA to confusion-network driven decoding and we propose new combination schemes for multiple system combination. Since previous experiments involved 2 ASR systems on broadcast news data, the proposed extended DDA is evaluated using 3 ASR systems from different labs. Results show that generalized- DDA outperforms significantly ROVER method: we obtain a 15.7% relative word error rate improvement with respect to the best single system, as opposed to 8.5% with the ROVER combination. Benjamin Lecouteux, Georges Linarès, Yannick Estève, Guillaume Gravier |
ICASSP | 4 |
| 2008 | Structure learning in a Bayesian network-based video indexing frameworkabstractSeveral stochastic models provide an effective framework to identify the temporal structure of audiovisual data. Most of them need as input a first video structure, i.e. connections between features and video events. Provided that this structure is given as input, the parameters are then estimated from training data. Bayesian networks offer an additional feature, namely structure learning, which allows the automatic construction of the model structure from training data. Structure learning obviously leads to an increased generality of the model building process. This paper investigates the trade-off between the increase of generality and the quality of the results in video analysis. We model video data using dynamic Bayesian networks (DBNs) where the static part of the network accounts for the correlations between low-level features extracted from the raw data and between these features and the events considered. It is precisely this part of the network whose structure is automatically constructed from training data. Experimental results on a commercial detection case study application show that, even though the model structure is determined in a non supervised manner, the resulting model is effective for the detection of commercial segments in video data. Siwar Baghdadi, Guillaume Gravier, Claire-Hélène Demarty, Patrick Gros |
ICME | 2 |
| 2008 | Morphosyntactic Resources for Automatic Speech Recognition
Stéphane Huet, Guillaume Gravier, Pascale Sébillot |
LREC | 2 |
| 2008 | On the Use of Web Resources and Natural Language Processing Techniques to Improve Automatic Speech Recognition Systems
Gwénolé Lecorvé, Guillaume Gravier, Pascale Sébillot |
LREC | 2 |
| 2008 | Audiovisual integration with Segment Models for tennis video parsing
Manolis Delakis, Guillaume Gravier, Patrick Gros |
Comput. Vis. Image Underst. | 2 |
| 2007 | Morphosyntactic processing of n-best lists for improved recognition and confidence measure computationabstractInternational audience Stéphane Huet, Guillaume Gravier, Pascale Sébillot |
INTERSPEECH | 2 |
| 2007 | Rapid speaker adaptation by reference model interpolation
Wen Xuan Teng, Guillaume Gravier, Frédéric Bimbot, Frédéric Soufflet |
INTERSPEECH | 2 |
| 2006 | Corpus description of the ESTER Evaluation Campaign for the Rich Transcription of French Broadcast News
Sylvain Galliano, Edouard Geoffrois, Guillaume Gravier, Jean-François Bonastre, Djamel Mostefa, Khalid Choukri |
LREC | 3 |
| 2006 | Score oriented Viterbi search in sport video structuring using HMM and segment modelsabstractA key question in video indexing is the effective use of all the possible sources of information. Hidden Markov models (HMM) and segment models (SM) provide powerful frameworks for audiovisual integration and structure knowledge encoding. However, incorporating punctual symbolic information, such as inlaid score labels, with these models is not straightforward. We demonstrate that using the score labels as an additional feature is not efficient and propose a novel algorithm to efficiently incorporate the score information in the Viterbi decoding process. This score oriented Viterbi search guarantees an optimal solution consistent with the available score information. Experimental results demonstrate the effectiveness of the method and its robustness to inlaid score detection errors Manolis Delakis, Guillaume Gravier, Patrick Gros |
MMSP | 2 |
| 2006 | Audiovisual integration for tennis broadcast structuring
Ewa Kijak, Guillaume Gravier, Lionel Oisel, Patrick Gros |
Multim. Tools Appl. | 2 |
| 2006 | Experiments in audio source separation with one sensor for robust speech recognition
Elie-Laurent Benaroya, Frédéric Bimbot, Guillaume Gravier, Rémi Gribonval |
Speech Commun. | 3 |
| 2005 | Multimodal Segmental-Based Modeling of Tennis Video BroadcastsabstractEfficient multimodal fusion is a key feature of future video indexing systems. Hidden Markov models provide a powerful framework for video structure analysis but they require all video modalities to be strictly synchronous. Taking as a case study tennis broadcasts analysis, we introduce into video indexing segment models, a generalization of hidden Markov Models, where the fusion of different modalities can be performed with relaxed synchrony constraints. Segment models were experimentally proved to perform marginally better compared to hidden Markov models Manolis Delakis, Guillaume Gravier, Patrick Gros |
ICME | 2 |
| 2005 | A model space framework for efficient speaker detectionabstractIn this paper, we investigate the use of a distance between Gaussian mixture models for speaker detection. The proposed distance is derived from the KL divergences and is defined as an Euclidean distance in a particular model space. This distance is simply computable directly from the model parameters thus leading to a very efficient scoring process. This new framework for scoring is compared to the classical log likelihood ratio s-core approach on a speaker verification task of the NIST 2004 evaluation and on the speaker tracking task of the ESTER french evaluation. Results shows that the proposed approach is competitive and leads to computation times divided by a factor of more than 3. 1. Mathieu Ben, Guillaume Gravier, Frédéric Bimbot |
INTERSPEECH | 2 |
| 2005 | The ESTER phase II evaluation campaign for the rich transcription of French broadcast newsabstractThis paper gives the final results of the ESTER evaluation campaign which started in 2003 and ended in January 2005. The aim of this campaign was to evaluate automatic broadcast news rich transcription systems for the French language. The evaluation tasks were divided into three main categories: orthographic transcription, event detection and tracking (e.g. speech vs. music, speaker tracking), and information extraction. The last one, limited to named entity detection in this evaluation, was a preliminary test. The paper reports on protocols and gives the results obtained in the campaign. 1. Sylvain Galliano, Edouard Geoffrois, Djamel Mostefa, Khalid Choukri, Jean-François Bonastre, Guillaume Gravier |
INTERSPEECH | 6 |
| 2005 | Experiments on speaker tracking and segmentation in radio broadcast newsabstractIn this paper we describe the speaker tracking and clustering system that we implemented for the ESTER evaluation campaign. We present some experiments on normalization in speaker tracking, in particular concerning the use of t-norm for speaker tracking in broadcast news. Results show that the use of t-norm significantly improves the performance at low false alarm rates. In a second part of the paper, we study the possible interactions between speaker tracking and speaker segmentation (also known as speaker diarization). We show that speaker segmentation benefits from the use of speaker tracking as a prior information while the contrary is not true. Using speaker tracking before clustering can decrease the speaker segmentation error by 4 % absolute. 1. Daniel Moraru, Mathieu Ben, Guillaume Gravier |
INTERSPEECH | 3 |
| 2004 | Multiple events tracking in sound tracksabstractDetecting and tracking broad sound classes in audio documents is an important step toward structuration. In the case of complex audio scenes, such as TV broadcast sound tracks, one problem is that several audio events may occur simultaneously. We propose a two-step approach to detect superimposed events. The first step is a blind segmentation step, followed by an event detection step on each segment. In order to evaluate the quality of the system better, new performance measures have been introduced, more suited to the superimposed events detection task. We also extend the two-step approach with an equivalent Viterbi-based event detection approach. Michael Betser, Guillaume Gravier |
ICME | 2 |
| 2004 | Speaker diarization using bottom-up clustering based on a parameter-derived distance between adapted GMMs
Michael Betser, Frédéric Bimbot, Mathieu Ben, Guillaume Gravier |
INTERSPEECH | 4 |
| 2004 | The ESTER Evaluation Campaign for the Rich Transcription of French Broadcast News
Guillaume Gravier, Jean-François Bonastre, Edouard Geoffrois, Sylvain Galliano, Kevin McTait, Khalid Choukri |
LREC | 1 |
| 2004 | Tennis video abstraction from audio and visual cuesabstractWe propose a context-based model of video abstraction exploiting both audio and video features and applied to tennis TV programs. We can automatically produce different types of summary of a given video depending on the users' constraints or preferences. We have first designed an efficient and accurate temporal segmentation of the video into segments homogeneous w.r.t the camera motion. We introduce original visual descriptors related to the dominant and residual image motions. The different summary types are obtained by specifying adapted classification criteria which involve audio features to select the relevant segments to be included in the video abstract. The proposed scheme has been validated on 22 hours of tennis videos. François Coldefy, Patrick Bouthemy, Michael Betser, Guillaume Gravier |
MMSP | 4 |
| 2003 | HMM based structuring of tennis videos using visual and audio cuesabstractThis paper focuses on the use of hidden Markov models (HMMs) for structure analysis of videos, and demonstrates how they can be efficiently applied to merge audio and visual cues. Our approach is validated in the particular domain of tennis videos. The basic temporal unit is the video shot. Visual features describe the audio events within a video shot. The video structure parsing relies on the analysis of the temporal interleaving of video shots, with respect to prior information about tennis content and editing rules. As a result, typical tennis scenes are identified. In addition, each shot is assigned to a level in the hierarchy described in terms of point, game and set. Ewa Kijak, Guillaume Gravier, Patrick Gros, Lionel Oisel, Frédéric Bimbot |
ICME | 2 |
| 2003 | Recent advances in the automatic recognition of audiovisual speechabstractVisual speech information from the speaker's mouth region has been successfully shown to improve noise robustness of automatic speech recognizers, thus promising to extend their usability in the human computer interface. In this paper, we review the main components of audiovisual automatic speech recognition (ASR) and present novel contributions in two main areas: first, the visual front-end design, based on a cascade of linear image transforms of an appropriate video region of interest, and subsequently, audiovisual speech integration. On the latter topic, we discuss new work on feature and decision fusion combination, the modeling of audiovisual speech asynchrony, and incorporating modality reliability estimates to the bimodal recognition process. We also briefly touch upon the issue of audiovisual adaptation. We apply our algorithms to three multisubject bimodal databases, ranging from small- to large-vocabulary recognition tasks, recorded in both visually controlled and challenging environments. Our experiments demonstrate that the visual modality improves ASR over all conditions and data considered, though less so for visually challenging environments and large vocabulary tasks. Gerasimos Potamianos, Chalapathy Neti, Guillaume Gravier, Andrew W. Senior |
Proc. IEEE | 3 |
| 2002 | Maximum entropy and MCE based HMM stream weight estimation for audio-visual ASRabstractIn this paper, we propose a new fast and flexible algorithm based on the maximum entropy (MAXENT) criterion to estimate stream weights in a state-synchronous multi-stream HMM. The technique is compared to the minimum classification error (MCE) criterion and to a brute-force, grid-search optimization of the WER on both a small and a large vocabulary audio-visual continuous speech recognition task. When estimating global stream weights, the MAXENT approach gives comparable results to the grid-search and the MCE. Estimation of state dependent weights is also considered: We observe significant improvements in both the MAXENT and MCE criteria, which, however, do not result in significant WER gains. Guillaume Gravier, Scott Axelrod, Gerasimos Potamianos, Chalapathy Neti |
ICASSP | 1 |
| 2001 | Integrating contextual phonological rules in a large vocabulary decoderabstractInternational audience Guillaume Gravier, François Yvon, Bruno Jacob, Frédéric Bimbot |
INTERSPEECH | 1 |
| 2000 | A Markov random field based multi-band modelabstractAn extension of the multi-band model including inter-band control of time asynchrony is described. It is based on the framework of Markov random fields. The law of the speech process is given by a parametric Gibbs distribution and a maximum likelihood parameter estimation algorithm is developed. This random field model is applied to isolated word recognition. It is shown that similar performances are obtained with the new model and with standard HMM techniques in the mono-band case. In the multi-band case, it is shown that the recognition rate decreases when the number of bands is increased but that modeling inter-band synchrony limits the performance decrease. Guillaume Gravier, Marc Sigelle, Gérard Chollet |
ICASSP | 1 |
| 2000 | A Markov Random Field Model for Automatic Speech RecognitionabstractSpeech can be represented as a time/frequency distribution of energy using a multiband filter bank. A Markov random field model, which takes into account the possible time asynchrony across the bands, is estimated for each segmental units to be recognized. The law of the speech process is given by a parametric Gibbs distribution and a maximum likelihood parameter estimation algorithm is developed. Experiments are conducted on an isolated word recognition problem. It is shown that similar performances are obtained with the new model and with standard HMM techniques in the mono-band case. In the multiband case, it is shown that modeling interband synchrony is an interesting approach to increase the performance when the number of bands increases. Guillaume Gravier, Marc Sigelle, Gérard Chollet |
ICPR | 1 |
| 2000 | Speech modeling with state constrained Markov fields over frequency bands
Vincent Arsigny, Gérard Chollet, Guillaume Gravier, Marc Sigelle |
INTERSPEECH | 3 |
| 2000 | A further investigation on speech features for speaker characterizationabstractIn this article, we investigate on alternative speech features for speaker characterization. We study Line Spectrum Pairs features, Time-Frequency Principal Components and Discriminant Components of the Spectrum. These alternative features are tested and compared on a task of speaker verification. This task consists in verifying a claimed identity from a speech segment. Systems are evaluated on a subset of the evaluation data of the NIST 1999 speaker recognition campaign. The new speech features are also compared to the classical cepstral coefficients, which remain, in our experiments, the best performing features. Ivan Magrin-Chagnolleau, Guillaume Gravier, Mouhamadou Seck, Olivier Boëffard, Raphaël Blouet, Frédéric Bimbot |
INTERSPEECH | 2 |
| 1998 | Toward Markov random field modeling of speechabstractIn this paper, we present a new technique for statistical modeling of speech segments based on Markov random fields. Classical and multi-stream HMMs are particular cases of this more general family of models. However, the Random Field Model (RFM) proposed here can be seen as an extension of the multiband HMM in which interactions between the frequency bands have been added. In a first experiment, samples are drawn from different models and compared to real observations. This experiment shows that the RFM is able to produce realistic samples but a single HMM still performs better. Isolated word recognition experiments stress the fact that more work must be done on the RFM in order to reach the performances of classical hidden Markov modeling techniques. For the moment, the RFM parameters are estimated using a heuristic. We believe that a real maximum likelihood parameter estimation algorithm should improve the results. The main advantage of this new model is that it can easily be extended since a model is defined by some local interactions and the Gibbs potential functions associated to those interactions. 1 Guillaume Gravier, Marc Sigelle, Gérard Chollet |
ICSLP | 1 |
| 1997 | Model dependent spectral representations for speaker recognition
Guillaume Gravier, Chafic Mokbel, Gérard Chollet |
EUROSPEECH | 1 |
| 1997 | Optimal state dependent spectral representation for HMM modeling : a new theoretical framework
Chafic Mokbel, Guillaume Gravier, Gérard Chollet |
EUROSPEECH | 2 |
| 1996 | Combining methods to improve speaker verification decisionabstractThe aim of this paper is to describe how the combination of speaker verification algorithms with a priori decision thresholds can improve the overall robustness of a real application. The evaluation is performed in the context of a field application where each client is verified from a 7 digit pin code. This paper demonstrate that it is possible to increase the global performances of the system on combining the result of several algorithms. Dominique Genoud, Frédéric Bimbot, Guillaume Gravier, Gérard Chollet |
ICSLP | 3 |