Jordi Bernad

dblp:62/258 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
5since 2021 · last 2024
0000-0001-8531-353XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Data mining · 61% Distributed and cloud data management · 30% Knowledge graphs · 9%
Artificial intelligence
1 paper
Knowledge representation and reasoning · 77% Information extraction and text analysis · 23%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed and cloud data management
data partitioning
0.812024
Language-Model Based Informed Partition of Databases to Speed Up Pattern Mining · Proc. ACM Manag. Data 2024
Data mining › pattern mining › itemset mining
frequent itemset mining
0.812024
Language-Model Based Informed Partition of Databases to Speed Up Pattern Mining · Proc. ACM Manag. Data 2024
Data mining › pattern mining › itemset mining › frequent itemset mining
parallel frequent itemset mining
0.812024
Language-Model Based Informed Partition of Databases to Speed Up Pattern Mining · Proc. ACM Manag. Data 2024
Knowledge, reasoning and agents › Knowledge representation and reasoning › semantic representation › semantic relations
lexical relation classification
0.712023
No clues good clues: out of context Lexical Relation Classification · ACL (1) 2023
Knowledge graphs
knowledge graph mining
0.212024
Language-Model Based Informed Partition of Databases to Speed Up Pattern Mining · Proc. ACM Manag. Data 2024
Natural language and speech › Information extraction and text analysis
lexical semantics
0.212023
No clues good clues: out of context Lexical Relation Classification · ACL (1) 2023

Methods — techniques the papers use, named apart from their topics

word embeddings · 1.4clustering · 0.8distributional semantics · 0.7
YearPublicationVenuePosition
2024 MultiLexBATS: Multilingual Dataset of Lexical Semantic Relations
abstract
Understanding the relation between the meanings of words is an important part of comprehending natural language. Prior work has either focused on analysing lexical semantic relations in word embeddings or probing pretrained language models (PLMs), with some exceptions. Given the rarity of highly multilingual benchmarks, it is unclear to what extent PLMs capture relational knowledge and are able to transfer it across languages. To start addressing this question, we propose MultiLexBATS, a multilingual parallel dataset of lexical semantic relations adapted from BATS in 15 languages including low-resource languages, such as Bambara, Lithuanian, and Albanian. As experiment on cross-lingual transfer of relational knowledge, we test the PLMs’ ability to (1) capture analogies across languages, and (2) predict translation targets. We find considerable differences across relation types and languages with a clear preference for hypernymy and antonymy as well as romance languages.
Dagmar Gromann, Hugo Gonçalo Oliveira, Lucia Pitarch, Elena Apostol, Jordi Bernad, Eliot Bytyci, Chiara Cantone, Sara Carvalho, Francesca Frontini, Radovan Garabík, Jorge Gracia, Letizia Granata, Anas Fahad Khan, Timotej Knez, Penny Labropoulou, Chaya Liebeskind, Maria Pia di Buono, Ana Ostroski Anic, Sigita Rackeviciene, Ricardo Rodrigues 0001, Gilles Sérasset, Linas Selmistraitis, Mahammadou Sidibé, Purificação Silvano, Blerina Spahiu, Enriketa Sogutlu, Ranka Stankovic, Ciprian-Octavian Truica, Giedre Valunaite Oleskeviciene, Slavko Zitnik, Katerina Zdravkova
LREC/COLING5
2024 Building MUSCLE, a Dataset for MUltilingual Semantic Classification of Links between Entities
abstract
In this paper we introduce MUSCLE, a dataset for MUltilingual lexico-Semantic Classification of Links between Entities. The MUSCLE dataset was designed to train and evaluate Lexical Relation Classification (LRC) systems with 27K pairs of universal concepts selected from Wikidata, a large and highly multilingual factual Knowledge Graph (KG). Each pair of concepts includes its lexical forms in 25 languages and is labeled with up to five possible lexico-semantic relations between the concepts: hypernymy, hyponymy, meronymy, holonymy, and antonymy. Inspired by Semantic Map theory, the dataset bridges lexical and conceptual semantics, is more challenging and robust than previous datasets for LRC, avoids lexical memorization, is domain-balanced across entities, and enables enrichment and hierarchical information retrieval.
Lucia Pitarch, Carlos Bobed, David Abián, Jorge Gracia, Jordi Bernad
LREC/COLING5
2024 Language-Model Based Informed Partition of Databases to Speed Up Pattern Mining
abstract
Extracting interesting patterns from data is the main objective of Data Mining. In this context, Frequent Itemset Mining has shown its usefulness in providing insights from transactional databases, which, in turn, can be used to gain insights about the structure of Knowledge Graphs. While there have been a lot of advances in the field, due to the NP-hard nature of the problem, the main approaches still struggle when they are faced with large databases with large and sparse vocabularies, such as the ones obtained from graph propositionalizations. There have been efforts to propose parallel algorithms, but, so far, the goal has not been to tackle this source of complexity (i.e., vocabulary size), thus, in this paper, we propose to parallelize frequent itemset mining algorithms by partitioning the database horizontally (i.e., transaction-wise) while not neglecting all the possible vertical information (i.e., item-wise). Instead of relying on pure item co-appearance metrics, we advocate for the adoption of a different approach: modeling databases as documents, where each transaction is a sentence, and each item a word. In this way, we can apply recent language modeling techniques (i.e., word embeddings) to obtain a continuous representation of the database, clusterize it in different partitions, and apply any mining algorithm to them. We show how our proposal leads to informed partitions with a reduced vocabulary size and a reduced entropy (i.e., disorder). This enhances the scalability, allowing us to speed up mining even in very large databases with sparse vocabularies. We have carried out a thorough experimental evaluation over both synthetic and real datasets showing the benefits of our proposal.
Carlos Bobed, Jordi Bernad, Pierre Maillot
Proc. ACM Manag. Data2
2023 No clues good clues: out of context Lexical Relation Classification
abstract
Lucia Pitarch, Jordi Bernad, Lacramioara Dranca, Carlos Bobed Lisbona, Jorge Gracia. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Lucia Pitarch, Jordi Bernad, Lacramioara Dranca, Carlos Bobed, Jorge Gracia
ACL (1)2
2023 MEAN: Metaphoric Erroneous ANalogies dataset for PTLMs metaphor knowledge probing
Lucia Pitarch, Jordi Bernad, Jorge Gracia
LDK2
2008 Semantic Discovery of the User Intended Query in a Selectable Target Query Language
abstract
The syntactic approach of most of Web search engines still has the drawback of not considering the semantics of the keywords entered by the user. So, users usually have to browse many hits looking for the information they want. In this paper, we present a system that, given a set of keywords with well defined semantics, automatically generates a set of formal queries, in the query language of the user's choice, which attempt to capture what the user had in mind when she or he wrote those keywords. The system uses ontologies and a description logics reasoner to perform a semantic enrichment of user keywords to improve the discovering of possible user queries and to reject semantically inconsistent queries.
Carlos Bobed, Raquel Trillo Lado, Eduardo Mena, Jordi Bernad
Web Intelligence4