Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Pascale Sébillot

dblp:90/3698 · DBLP profile ↗
← Back
32ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0002-5429-4302ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11Databases, data management, data science and information retrieval · 9 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Knowledge graphs · 60% Information retrieval · 40%
Artificial intelligence
1 paper
Information extraction and text analysis · 50% Knowledge representation and reasoning · 50%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge graphs
entity linking
1.012026
FAST-MEL: A Fast, Accurate, and Storage Efficient Solution for Multimodal Entity Linking · SIGIR 2026
Knowledge graphs › entity linking
multimodal entity linking
1.012026
FAST-MEL: A Fast, Accurate, and Storage Efficient Solution for Multimodal Entity Linking · SIGIR 2026
Information retrieval
multimodal retrieval
1.012026
FAST-MEL: A Fast, Accurate, and Storage Efficient Solution for Multimodal Entity Linking · SIGIR 2026
Information retrieval › text analysis
text segmentation
0.212013
Leveraging Lexical Cohesion and Disruption for Topic Segmentation · EMNLP 2013
Information retrieval › text analysis › text segmentation
topic segmentation
0.212013
Leveraging Lexical Cohesion and Disruption for Topic Segmentation · EMNLP 2013
Knowledge, reasoning and agents › Knowledge representation and reasoning › logic programming
inductive logic programming
0.012003
Learning Semantic Lexicons from a Part-of-Speech and Semantically Tagged Corpus Using Inductive Logic Programming · J. Mach. Learn. Res. 2003
Natural language and speech › Information extraction and text analysis › lexical semantics › lexicon learning
semantic lexicon learning
0.012003
Learning Semantic Lexicons from a Part-of-Speech and Semantically Tagged Corpus Using Inductive Logic Programming · J. Mach. Learn. Res. 2003

Methods — techniques the papers use, named apart from their topics

vectorized entity representation · 1.0encoder-based representation learning · 1.0graph-based decoding · 0.2inductive logic programming · 0.0
YearPublicationVenuePosition
2026 A Study on Building Efficient Zero-Shot Relation Extraction Models
abstract
International audience
Hugo Thomas, Caio F. Corro, Guillaume Gravier, Pascale Sébillot
LREC4
2026 FAST-MEL: A Fast, Accurate, and Storage Efficient Solution for Multimodal Entity Linking
abstract
Multimodal entity linking (MEL) is the task that consists of matching textual and visual mentions of entities in unstructured data to their corresponding entities in a knowledge base (KB). To be effective in large-scale practical settings, MEL systems must meet three objectives: high linking accuracy, computational efficiency, and storage efficiency, i.e., a compact yet efficient index of the KB. In this paper, we highlight that state-of-the-art systems fail to simultaneously satisfy these 3 requirements. To meet this three-fold objective, we propose FAST-MEL, a lightweight encoder-based MEL solution that relies on a novel and compact fixed-size vectorized representation of both the textual and visual information of each entity or mention. It matches the accuracy of the best systems but performs three orders of magnitude faster. It also consumes one order of magnitude less storage than the fastest systems.
Thomas Derrien, Laurent Amsaleg, Pascale Sébillot
SIGIR3
2024 One-shot relation retrieval in news archives: adapting N-way K-shot relation Classification for efficient knowledge extraction
abstract
One-shot relation retrieval is the knowledge extraction task that consists in searching in a textual dataset for all occurrences of a relation of interest, named the source relation, characterized by a single example—a relation being a link between a pair of entities in an utterance. Performing this task on large datasets requires an intelligent system to automate the process, for instance when exploring news archives for press review or business intelligence. We propose a framework that leverages the representation learning capabilities of N-way K-shot models for few-shot relation Classification and extends these models to enable one-shot retrieval with a rejection class. At evaluation time, one-shot relation retrieval is performed in a N-way K-shot setting where 1 of the N ways (or relations) is the source relation and the N-1 others are distractors, i.e., relations modeling a rejection class. We benchmark this framework and investigate the influence of the number and the choice of distractors on the standard TACREV and FewRel datasets. Experimental results demonstrate the effectiveness of our approach to address this highly challenging task, however with high variability primarily induced by the type of the source relation. Experiments also highlight a sound strategy for the choice of distractors—a large number of distractors at an intermediate distance from the embedding of the source relation in the latent space learned by the model—, which provides a competing trade-of between recall and precision. This strategy is globally optimal but can however be surpassed on certain source relations by others, depending on the characteristics of the source relation, paving the way for future work. We finally show the substantial benefit of two-shot retrieval over one-shot retrieval, which sheds light on the design of actual intelligent applications leveraging one- or few-shot relation retrieval.
Hugo Thomas, Guillaume Gravier, Pascale Sébillot
KES3
2023 Regularization, Semi-supervision, and Supervision for a Plausible Attention-Based Explanation
Duc Hau Nguyen, Cyrielle Mallart, Guillaume Gravier, Pascale Sébillot
NLDB4
2021 A Study of the Plausibility of Attention between RNN Encoders in Natural Language Inference
abstract
Attention maps in neural models for NLP are appealing to explain the decision made by a model, hopefully emphasizing words that justify the decision. While many empirical studies hint that attention maps can provide such justification from the analysis of sound examples, only a few assess the plausibility of explanations based on attention maps, i.e., the usefulness of attention maps for humans to understand the decision. These studies furthermore focus on text classification. In this paper, we report on a preliminary assessment of attention maps in a sentence comparison task, namely natural language inference. We compare the cross-attention weights between two RNN encoders with human-based and heuristic-based annotations on the eSNLI corpus. We show that the heuristic reasonably correlates with human annotations and can thus facilitate evaluation of plausible explanations in sentence comparison tasks. Raw attention weights however remain only loosely related to a plausible explanation.
Duc Hau Nguyen, Guillaume Gravier, Pascale Sébillot
ICMLA3
2020 A correlation-based entity embedding approach for robust entity linking
abstract
Entity alignment is a crucial tool in knowledge discovery to reconcile knowledge from different sources. Recent state-of-the-art approaches leverage joint embedding of knowledge graphs (KGs) so that similar entities from different KGs are close in the embedded space. Whatever the joint embedding technique used, a seed set of aligned entities, often provided by (time-consuming) human expertise, is required to learn the joint KG embedding and/or a mapping between KG embeddings. In this context, a key issue is to limit the size and quality requirement for the seed. State-of-the-art methods usually learn the embedding by explicitly minimizing the distance between aligned entities from the seed and uniformly maximizing the distance for entities not in the seed. In contrast, we design a less restrictive optimization criterion that indirectly minimizes the distance between aligned entities in the seed by globally maximizing the dimension-wise correlation among all the embeddings of seed entities. Within an iterative entity alignment system, the correlation-based entity embedding function achieves state-of-the-art results and is shown to significantly increase robustness to the seed's size and accuracy. It ultimately enables fully unsupervised entity alignment using a seed automatically generated with a symbolic alignment method based on entities' names.
Cheikh Brahim El Vaigh, François Torregrossa, Robin Allesiardo, Guillaume Gravier, Pascale Sébillot
ICTAI5
2020 A Novel Path-Based Entity Relatedness Measure for Efficient Collective Entity Linking
Cheikh Brahim El Vaigh, François Goasdoué, Guillaume Gravier, Pascale Sébillot
ISWC (1)4
2019 Using Knowledge Base Semantics in Context-Aware Entity Linking
abstract
Entity linking is a core task in textual document processing, which consists in identifying the entities of a knowledge base (KB) that are mentioned in a text. Approaches in the literature consider either independent linking of individual mentions or collective linking of all mentions. Regardless of this distinction, most approaches rely on the Wikipedia encyclopedic KB in order to improve the linking quality, by exploiting its entity descriptions (web pages) or its entity interconnections (hyperlink graph of web pages). In this paper, we devise a novel collective linking technique which departs from most approaches in the literature by relying on a structured RDF KB. This allows exploiting the semantics of the interrelationships that candidate entities may have at disambiguation time rather than relying on raw structural approximation based on Wikipedia's hyperlink graph. The few approaches that also use an RDF KB simply rely on the existence of a relation between the candidate entities to which mentions may be linked. Instead, we weight such relations based on the RDF KB structure and propose an efficient decoding strategy for collective linking. Experiments on standard benchmarks show significant improvement over the state of the art.
Cheikh Brahim El Vaigh, François Goasdoué, Guillaume Gravier, Pascale Sébillot
DocEng4
2017 Linking Multimedia Content for Efficient News Browsing
abstract
As the amount of news information available online grows, media are in need of advanced tools to explore the information surrounding specific events before writing their own piece of news, e.g., adding context and insight. While many tools exist to extract information from large datasets, they do not offer an easy way to gain insight from a news collection by browsing, going from article to article and viewing unaltered original content. Such browsing tools require the creation of rich underlying structures such as graph representations. These representations can be further enhanced by typing links that connect nodes, in order to inform the user on the nature of their relation. In this article, we introduce an efficient way to generate links between news items in order to obtain an easily navigable graph, and enrich this graph by automatically typing created links. User evaluations are conducted on real world data in order to assess for the interest of both the graph representation and link typing in a press reviewing task, showing a significant improvement compared to classical search engines.
Rémi Bois, Guillaume Gravier, Eric Jamet, Emmanuel Morin, Maxime Robert, Pascale Sébillot
ICMR6
2017 Exploiting Multimodality in Video Hyperlinking to Improve Target Diversity
Rémi Bois, Vedran Vukotic, Anca-Roxana Simon, Ronan Sicre, Christian Raymond, Pascale Sébillot, Guillaume Gravier
MMM (2)6
2016 Shaping-Up Multimedia Analytics: Needs and Expectations of Media Professionals
Guillaume Gravier, Martin Ragot, Laurent Amsaleg, Rémi Bois, Grégoire Jadi, Eric Jamet, Laura Monceaux, Pascale Sébillot
MMM (2)8
2014 Text recognition in multimedia documents: a study of two neural-based OCRs using and avoiding character segmentation
Khaoula Elagouni, Christophe Garcia, Franck Mamalet, Pascale Sébillot
Int. J. Document Anal. Recognit.4
2013 Leveraging Lexical Cohesion and Disruption for Topic Segmentation
abstract
Topic segmentation classically relies on one of two criteria, either finding areas with coherent vocabulary use or detecting discontinuities.In this paper, we propose a segmentation criterion combining both lexical cohesion and disruption, enabling a trade-off between the two.We provide the mathematical formulation of the criterion and an efficient graph based decoding algorithm for topic segmentation.Experimental results on standard textual data sets and on a more challenging corpus of automatically transcribed broadcast news shows demonstrate the benefit of such a combination.Gains were observed in all conditions, with segments of either regular or varying length and abrupt or smooth topic shifts.Long segments benefit more than short segments.However the algorithm has proven robust on automatic transcripts with short segments and limited vocabulary reoccurrences.
Anca-Roxana Simon, Guillaume Gravier, Pascale Sébillot
EMNLP3
2013 Multimedia information seeking through search and hyperlinking
abstract
Searching for relevant webpages and following hyperlinks to related content is a widely accepted and effective approach to information seeking on the textual web. Existing work on multimedia information retrieval has focused on search for individual relevant items or on content linking without specific attention to search results. We describe our research exploring integrated multimodal search and hyperlinking for multimedia data. Our investigation is based on the MediaEval 2012 Search and Hyperlinking task. This includes a known-item search task using the Blip10000 internet video collection, where automatically created hyperlinks link each relevant item to related items within the collection. The search test queries and link assessment for this task was generated using the Amazon Mechanical Turk crowdsourcing platform. Our investigation examines a range of alternative methods which seek to address the challenges of search and hyperlinking using multimodal approaches. The results of our experiments are used to propose a research agenda for developing effective techniques for search and hyperlinking of multimedia content.
Maria Eskevich, Gareth J. F. Jones, Robin Aly, Roeland Ordelman, Danish Nadeem, Camille Guinaudeau, Guillaume Gravier, Pascale Sébillot, Tom De Nies, Pedro Debevere, Rik Van de Walle, Petra Galuscáková, Pavel Pecina, Martha A. Larson
ICMR9
2012 Combining Multi-scale Character Recognition and Linguistic Knowledge for Natural Scene Text OCR
abstract
Understanding text captured in real-world scenes is a challenging problem in the field of visual pattern recognition and continues to generate a significant interest in the OCR (Optical Character Recognition) community. This paper proposes a novel method to recognize scene texts avoiding the conventional character segmentation step. The idea is to scan the text image with multi-scale windows and apply a robust recognition model, relying on a neural classification approach, to every window in order to recognize valid characters and identify non valid ones. Recognition results are represented as a graph model in order to determine the best sequence of characters. Some linguistic knowledge is also incorporated to remove errors due to recognition confusions. The designed method is evaluated on the ICDAR 2003 database of scene text images and outperforms state-of-the-art approaches.
Khaoula Elagouni, Christophe Garcia, Franck Mamalet, Pascale Sébillot
Document Analysis Systems4
2012 Text Recognition in Videos Using a Recurrent Connectionist Approach
Khaoula Elagouni, Christophe Garcia, Franck Mamalet, Pascale Sébillot
ICANN (2)4
2012 Enhancing lexical cohesion measure with confidence measures, semantic relations and language model interpolation for multimedia spoken content topic segmentation
Camille Guinaudeau, Guillaume Gravier, Pascale Sébillot
Comput. Speech Lang.3
2011 Automatically finding semantically consistent n-grams to add new words in LVCSR systems
abstract
This paper presents a new method to automatically add re-grams containing out-of-vocabulary (OOV) words to a baseline language model (LM), where these re-grams are sought to be grammatically correct and to make sense according to the meaning of OOV words. First, this method consists in determining the word sequences, i.e., re-grams, in which the usage of a given OOV word is the most semantically consistent. Then, conditional probabilities of these re-grams have to be computed. To do this, semantic relations between words are used to assimilate each OOV word to several equivalent in vocabulary words. Based on these last words, n-grams from the baseline LM are re-used to find the word sequences to be added and to compute their probabilities. After augmenting the vocabulary and launching a recognition process, experiments show that our method results in WER improvements which are comparable to those obtained using a state-of-the-art open vocabulary LM.
Gwénolé Lecorvé, Guillaume Gravier, Pascale Sébillot
ICASSP3
2011 A comprehensive neural-based approach for text recognition in videos using natural language processing
abstract
This work aims at helping multimedia content understanding by deriving benefit from textual clues embedded in digital videos. For this, we developed a complete video Optical Character Recognition system (OCR), specifically adapted to detect and recognize embedded texts in videos. Based on a neural approach, this new method outperforms related work, especially in terms of robustness to style and size variabilities, to background complexity and to low resolution of the image. A language model that drives several steps of the video OCR is also introduced in order to remove ambiguities due to a local letter by letter recognition and to reduce segmentation errors. This approach has been evaluated on a database of French TV news videos and achieves an outstanding character recognition rate of 95%, corresponding to 78% of words correctly recognized, which enables its incorporation into an automatic video indexing and retrieval system.
Khaoula Elagouni, Christophe Garcia, Pascale Sébillot
ICMR3
2010 Improving ASR-based topic segmentation of TV programs with confidence measures and semantic relations
abstract
International audience
Camille Guinaudeau, Guillaume Gravier, Pascale Sébillot
INTERSPEECH3
2010 Morpho-syntactic post-processing of N-best lists for improved French automatic speech recognition
Stéphane Huet, Guillaume Gravier, Pascale Sébillot
Comput. Speech Lang.3
2009 Constraint selection for topic-based MDI adaptation of language models
abstract
This paper presents an unsupervised topic-based language model adaptation method which specializes the standard minimum information discrimination approach by identifying and combining topic-specific features.By acquiring a topic terminology from a thematically coherent corpus, language model adaptation is restrained to the sole probability re-estimation of n-grams ending with some topic-specific words, keeping other probabilities untouched.Experiments are carried out on a large set of spoken documents about various topics.Results show significant perplexity and recognition improvements which outperform results of classical adaptation techniques.
Gwénolé Lecorvé, Guillaume Gravier, Pascale Sébillot
INTERSPEECH3
2009 Can Automatic Speech Transcripts Be Used for Large Scale TV Stream Description and Structuring?
abstract
The increasing quantity of TV material requires methods to help users navigate such data streams. Automatically associating a short textual description to each program in a stream, is a first stage to navigating or structuring tasks. Speech contained in TV broadcasts---accessible by means of automatic speech recognition systems in the absence of closed caption---is a highly valuable semantic clue that might be used to link existing textual description such as program guides, with video segments corresponding to program. However, high word error rates are to be expected on some programs, likely to jeopardize the usefulness of transcripts. The goal of this article is to determine to what extent automatic transcripts of TV streams, for various types of programs, can be used for structuring or navigating tasks. To this end, word-based and phonetic-based automatic association between video segments and program descriptions is used as a case study. We show that descriptions from a program guide can be associated with video segments with an accuracy of up to 65% and provide a valuable description to validate existing program labels. Such associations constitute a first stage for structuring task as they enable video segment textual characterization.
Camille Guinaudeau, Guillaume Gravier, Pascale Sébillot
ISM3
2008 An unsupervised web-based topic language model adaptation method
abstract
This paper focuses on a solution to better adapt ASR systems, whose language models (LM) are usually trained on topic-independent corpora, to new topics, in particular in the case of broadcast news. We propose a new complete and fully unsupervised technique that selects keywords from each segment using information retrieval methods, to build a thematically coherent adaptation corpus from the Internet. The LM used for the initial transcription is then adapted before rescoring word lattices. Experimental results demonstrate the validity of the proposed adaptation technique with a significant reduction of the perplexity after LM adaptation. Word error rates are also improved in some cases though to a lesser extent. Index Terms — Speech recognition, natural languages, Internet 1.
Gwénolé Lecorvé, Guillaume Gravier, Pascale Sébillot
ICASSP3
2008 Morphosyntactic Resources for Automatic Speech Recognition
Stéphane Huet, Guillaume Gravier, Pascale Sébillot
LREC3
2008 On the Use of Web Resources and Natural Language Processing Techniques to Improve Automatic Speech Recognition Systems
Gwénolé Lecorvé, Guillaume Gravier, Pascale Sébillot
LREC3
2007 Automatic Morphological Query Expansion Using Analogy-Based Machine Learning
Fabienne Moreau, Vincent Claveau, Pascale Sébillot
ECIR3
2007 Morphosyntactic processing of n-best lists for improved recognition and confidence measure computation
abstract
International audience
Stéphane Huet, Guillaume Gravier, Pascale Sébillot
INTERSPEECH3
2005 Combining statistical data analysis techniques to extract topical keyword classes from corpora
Mathias Rossignol, Pascale Sébillot
Intell. Data Anal.2
2004 From efficiency to portability: acquisition of semantic relations by semi-supervised machine learning
Vincent Claveau, Pascale Sébillot
COLING2
2003 Learning Semantic Lexicons from a Part-of-Speech and Semantically Tagged Corpus Using Inductive Logic Programming
Vincent Claveau, Pascale Sébillot, Cécile Fabre, Pierrette Bouillon
J. Mach. Learn. Res.2
2002 Acquisition of Qualia Elements from Corpora - Evaluation of a Symbolic Learning Method
Pierrette Bouillon, Vincent Claveau, Cécile Fabre, Pascale Sébillot
LREC4