Jérôme David

dblp:10/4191 · DBLP profile ↗
← Back
21ranked-venue papers
6as first author
4since 2021 · last 2025
0009-0007-0211-5003ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 11 · 4 first-author · 2 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorTheory of computation · 2 · 1 since 2021
YearPublicationVenuePosition
2025 Graph Embeddings Meet Link Keys Discovery for Entity Matching
abstract
Entity Matching (EM) automates the discovery of identity links between entities within different Knowledge Graphs (KGs). Link keys are crucial for EM, serving as rules allowing to identify identity links across different KGs, possibly described using different ontologies. However, the approach for extracting link keys struggles to scale on large KGs. While embedding-based EM methods efficiently handle large KGs they lack explainability. This paper proposes a novel hybrid EM approach to guarantee the scalability link key extraction approach and improve the explainability of embedding-based EM methods. First, embedding-based EM approaches are used to sample the KGs based on the identity links they generate, thereby reducing the search space to relevant sub-graphs for link key extraction. Second, rules (in the form of link keys) are extracted to explain the generation of identity links by the embedding-based methods. Experimental results demonstrate that the proposed approach allows link key extraction to scale on large KGs, preserving the quality of the extracted link keys. Additionally, it shows that link keys can improve the explainability of the identity links generated by embedding-methods, allowing for the regeneration of 77% of the identity links produced for a specific EM task, thereby providing an approximation of the reasons behind their generation.
Chloé Khadija Jradeh, Ensiyeh Raoufi, Jérôme David, Pierre Larmande, François Scharffe, Konstantin Todorov, Cássia Trojahn dos Santos
WWW3
2024 Discovering a Representative Set of Link Keys in RDF Datasets
Nacira Abbas, Alexandre Bazin, Jérôme David, Amedeo Napoli
EKAW3
2023 Discovery of link keys in resource description framework datasets based on pattern structures
Nacira Abbas, Alexandre Bazin, Jérôme David, Amedeo Napoli
Int. J. Approx. Reason.3
2021 Sandwich: An Algorithm for Discovering Relevant Link Keys in an LKPS Concept Lattice
Nacira Abbas, Alexandre Bazin, Jérôme David, Amedeo Napoli
ICFCA3
2020 Link key candidate extraction with relational concept analysis
Manuel Atencia, Jérôme David, Jérôme Euzenat, Amedeo Napoli, Jérémy Vizzini
Discret. Appl. Math.2
2019 Short-Term Temperature Forecasting on a Several Hours Horizon
Louis Desportes, Pierre Andry, Inbar Fijalkow, Jérôme David
ICANN (4)4
2019 Several Link Keys are Better than One, or Extracting Disjunctions of Link Key Candidates
abstract
Link keys express conditions under which instances of two classes of different RDF data sets may be considered as equal. As such, they can be used for data interlinking. There exist algorithms to extract link key candidates from RDF data sets and different measures have been defined to evaluate the quality of link key candidates individually. For certain data sets, however, it may be necessary to use more than one link key on a pair of classes to retrieve a more complete set of links. To this end, in this paper, we define disjunction of link keys, propose strategies to extract disjunctions of link key candidates from RDF data, and apply existing quality measures to evaluate them. We also report on experiments with these strategies.
Manuel Atencia, Jérôme David, Jérôme Euzenat
K-CAP2
2016 Uncertainty-Sensitive Reasoning for Inferring sameAs Facts in Linked Data
abstract
Discovering whether or not two URIs described in Linked Data—in the same or different RDF datasets—refer to the same real-world entity is crucial for building applications that exploit the cross-referencing of open data. A major challenge in data interlinking is to design tools that effectively deal with incomplete and noisy data, and exploit uncertain knowledge. In this paper, we model data interlinking as a reasoning problem with uncertainty. We introduce a probabilistic framework for modelling and reasoning over uncertain RDF facts and rules that is based on the semantics of probabilistic Datalog. We have designed an algorithm, ProbFR, based on this framework. Experiments on real-world datasets have shown the usefulness and effectiveness of our approach for data linkage and disambiguation.
Mustafa Al-Bakri, Manuel Atencia, Jérôme David, Steffen Lalande, Marie-Christine Rousset
ECAI3
2016 Cross-lingual RDF Thesauri Interlinking
Tatiana Lesnikova, Jérôme David, Jérôme Euzenat
LREC2
2015 What Is This Thing Called Linked Data?
abstract
The Linked Data initiative has made it possible for the web to evolve from being a global information space in which only documents are linked to one in which both documents and data are linked: a web of documents and data. This tutorial aims to give an overview of the principles, models and technologies underlying Linked Data.
Manuel Atencia, Jérôme David, Philippe Genoud
DocEng2
2015 Interlinking English and Chinese RDF Data Using BabelNet
abstract
Linked data technologies make it possible to publish and link structured data on the Web. Although RDF is not about text, many RDF data providers publish their data in their own language. Cross-lingual interlinking aims at discovering links between identical resources across knowledge bases in different languages. In this paper, we present a method for interlinking RDF resources described in English and Chinese using the BabelNet multilingual lexicon. Resources are represented as vectors of identifiers and then similarity between these resources is computed. The method achieves an F-measure of 88%. The results are also compared to a translation-based method.
Tatiana Lesnikova, Jérôme David, Jérôme Euzenat
DocEng2
2014 Data interlinking through robust linkkey extraction
abstract
Links are important for the publication of RDF data on the web. Yet, establishing links between data sets is not an easy task. We develop an approach for that purpose which extracts weak linkkeys. Linkkeys extend the notion of a key to the case of different data sets. They are made of a set of pairs of properties belonging to two different classes. A weak linkkey holds between two classes if any resources having common values for all of these properties are the same resources. An algorithm is proposed to generate a small set of candidate linkkeys. Depending on whether some of the, valid or invalid, links are known, we define supervised and non supervised measures for selecting the appropriate linkkeys. The supervised measures approximate precision and recall, while the non supervised measures are the ratio of pairs of entities a linkkey covers (coverage), and the ratio of entities from the same data set it identifies (discrimination). We have experimented these techniques on two data sets, showing the accuracy and robustness of both approaches.
Manuel Atencia, Jérôme David, Jérôme Euzenat
ECAI2
2013 Training sessions in a Master degree "Informatics as a Second Competence"
abstract
The ERAMIS acronym stands for European-Russian-Central Asian Network of Master's degrees “Informatics as a Second Competence”. The aim of this project is to create a Master degree “Computer Science as a Second Competence” in 9 beneficiary universities located in Kazakhstan, Kirghizstan and Russia. One crucial aspect of the project is training. In this contribution we describe how training sessions have been organized, the pedagogical issues that are involved and ways to address them.
Agathe Merceron, Jean-Michel Adam, Daniel Bardou, Jérôme David, Sergio Luján-Mora, Marek Milosz
EDUCON4
2012 Keys and Pseudo-Keys Detection for Web Datasets Cleansing and Interlinking
Manuel Atencia, Jérôme David, François Scharffe
EKAW2
2012 Experimenting with ontology distances in semantic social networks: Methodological remarks
abstract
Semantic social networks are social networks using ontologies for characterising resources shared within the network. It has been postulated that, in such networks, it is possible to discover social affinities between network members through measuring the similarity between the ontologies or part of ontologies they use. Using similar ontologies should reflect the cognitive disposition of the subjects. The main concern of this paper is the methodological aspect of experimenting in order to validate or invalidate such an hypothesis. Indeed, given the current lack of broad semantic social networks, it is difficult to rely on available data and experiments have to be designed from scratch. For that purpose, we first consider experimental settings that could be used and raise practical and methodological issues faced with analysing their results. We then describe a full experiments carried out according to some identified modalities and report the obtained results. The results obtained seem to invalidate the proposed hypothesis. We discuss why this may be so.
Jérôme David, Jérôme Euzenat, Jason J. Jung
SMC1
2010 Ontology Similarity in the Alignment Space
Jérôme David, Jérôme Euzenat, Ondrej Sváb-Zamazal
ISWC (1)1
2009 Detection and Transformation of Ontology Patterns
Ondrej Sváb-Zamazal, Vojtech Svátek, François Scharffe, Jérôme David
IC3K4
2008 Comparison between Ontology Distances (Preliminary Results)
Jérôme David, Jérôme Euzenat
ISWC1
2007 Association Rule Ontology Matching Approach
abstract
This article presents a hybrid, extensional and asymmetric matching approach designed to find out relations (equivalence and subsumption) between two textual hierarchies. By using the as-sociation rule paradigm and a statistical measure, this method relies on the following idea: ``An entity A will be more specific than or equivalent to an entity B if the vocabulary used to describe A and its instances tends to be included in that of B and its instances’’. This approach is divided into two parts: (1) The representation of each entity by a set of relevant terms and data; (2) The discovery of binary association rules between entities. The selection of rules uses two criteria for assessing the implication quality and reducing redundancy. The method is evaluated on two benchmarks. The first contains two hierarchies indexing textual documents and the second one is composed of OWL ontologies.
Jérôme David, Fabrice Guillet, Henri Briand
Int. J. Semantic Web Inf. Syst.1
2006 Matching directories and OWL ontologies with AROMA
abstract
This paper presents a simple and adaptable matching method dealing with web directories, catalogs and OWL ontologies. By using a well-known Knowledge Discovery in Databases model, such as the association rule paradigm, this method has the originality to be both extensional and asymmetric. It works at the terminological level (by selecting concept-relevant terms contained in documents) and permits to discover equivalence and also subsumption relations holding between entities (concepts and properties). This method relies on the implication intensity measure, a probabilistic model of deviation from independence. Selection of significant rules between concepts (or properties) is lead by two criteria permitting to assess respectively the implication quality and the generativity of the rule. Finally, the proposed method is evaluated on two benchmarks. The first contains two conceptual hierarchies populated with textual documents and the second one is composed of OWL ontologies.
Jérôme David, Fabrice Guillet, Henri Briand
CIKM1
2006 Conceptual Hierarchies Matching: An Approach Based on Discovery of Implication Rules Between Concepts
Jérôme David, Fabrice Guillet, Régis Gras, Henri Briand
ECAI1