Zied Bouraoui

dblp:134/4606 · DBLP profile ↗
← Back
61ranked-venue papers
10as first author
35since 2021 · last 2026
0000-0002-1662-4163ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 50 · 10 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 6 first-author · 13 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 4 since 2021Theory of computation · 6 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Generalizing Analogical Inference from Boolean to Continuous Domains
abstract
Analogical reasoning is a powerful inductive mechanism, widely used in human cognition and increasingly applied in artificial intelligence. Formal frameworks for analogical inference have been developed for Boolean domains, where inference is provably sound for affine functions and approximately correct for functions close to affine. These results have informed the design of analogy-based classifiers. However, they do not extend to regression tasks or continuous domains. In this paper, we revisit analogical inference from a foundational perspective. We first present a counterexample showing that existing generalization bounds fail even in the Boolean setting. We then introduce a unified framework for analogical reasoning in real-valued domains based on parameterized analogies defined via generalized means. This model subsumes both Boolean classification and regression, and supports analogical inference over continuous functions. We characterize the class of analogy-preserving functions in this setting and derive both worst-case and average-case error bounds under smoothness assumptions. Our results offer a general theory of analogical inference across discrete and continuous domains.
Francisco Cunha, Yves Lepage, Miguel Couceiro, Zied Bouraoui
AAAI4
2026 Credal Concept Bottleneck Models for Epistemic-Aleatoric Uncertainty Decomposition
abstract
Concept Bottleneck Models (CBMs) predict through human-interpretable concepts, but they typically output point concept probabilities that conflate epistemic uncertainty (reducible model underspecification) with aleatoric uncertainty (irreducible input ambiguity).This makes concept-level uncertainty hard to interpret and, more importantly, hard to act upon.We introduce CREDENCE (Credal Ensemble Concept Estimation), a CBM framework that decomposes concept uncertainty by construction.CREDENCE represents each concept as a credal prediction (a probability interval), derives epistemic uncertainty from disagreement across diverse concept heads, and estimates aleatoric uncertainty via a dedicated ambiguity output trained to match annotator disagreement when available.The resulting signals support prescriptive decisions: automate low-uncertainty cases, prioritize data collection for high-epistemic cases, route high-aleatoric cases to human review, and abstain when both are high.Across several tasks, we show that epistemic uncertainty is positively associated with prediction errors, whereas aleatoric uncertainty closely tracks annotator disagreement, providing guidance beyond error correlation.Our implementation is available at the following link: https://github.com/Tankiit/ Credal_Sets/tree/ensemble-credal-cbm
Tanmoy Mukherjee, Thomas Bailleux, Pierre Marquis, Zied Bouraoui
ACL (1)4
2026 Learning Long-Document Embeddings via Chunk-Context Entailment
Waheed Ahmed Abro, Naïm Es-Sebbani, Zied Bouraoui
LREC3
2026 Coupling description-guided prototype and feature-focused memory replay for continual relation extraction
Na Li 0018, Yongji Dai, Dunlu Peng, Zied Bouraoui
Neurocomputing4
2025 Modeling Complex Semantics Relation with Contrastively Fine-Tuned Relational Encoders
abstract
Modeling relationships between concepts and entities is essential for many applications.While Large Language Models (LLMs) capture relational and commonsense knowledge effectively, they are computationally expensive and often underperform in tasks requiring efficient relational encoding, such as relation induction, extraction, and information retrieval.Despite advancements in learning relational embeddings, existing methods often fail to capture nuanced representations and the rich semantics needed for high-quality embeddings.In this work, we propose different relational encoders designed to capture diverse relational aspects and semantic properties of entity pairs.Although several datasets exist for training such encoders, they often rely on structured knowledge bases or predefined schemas, which primarily encode simple and static relations.To overcome this limitation, we also introduce a novel dataset generation method leveraging LLMs to create a diverse spectrum of relationships.Our experiments demonstrate the effectiveness of our proposed encoders and the benefits of our generated dataset.
Naïm Es-Sebbani, Esteban Marquer, Zied Bouraoui
ACL (1)3
2025 Enhancing DR Classification with Swin Transformer and Shifted Window Attention
Meher Boulaabi, Takwa Ben Aïcha Gader, Afef Kacem, Zied Bouraoui
AIME (2)4
2025 Grouping Entities with Shared Properties using Multi-Facet Prompting and Property Embeddings
abstract
Methods for learning taxonomies from data have been widely studied.We study a specific version of this task, called commonality identification, where only the set of entities is given and we need to find meaningful ways to group those entities.While LLMs should intuitively excel at this task, it is difficult to directly use such models in large domains.In this paper, we instead use LLMs to describe the different properties that are satisfied by each of the entities individually.We then use pretrained embeddings to cluster these properties, and finally group entities that have properties which belong to the same cluster.To achieve good results, it is paramount that the properties predicted by the LLM are sufficiently diverse.We find that this diversity can be improved by prompting the LLM to structure the predicted properties into different facets of knowledge.1
Amit Gajbhiye, Thomas Bailleux, Zied Bouraoui, Luis Espinosa Anke, Steven Schockaert
EMNLP3
2025 Frequency Domain Information Integrated Network for Low-Light Image Enhancement
abstract
Low-light images often suffer from significant noise and detail loss, making it challenging to effectively distinguish signals from noise when processed directly in the spatial domain. To this end, we introduce frequency domain information to better distinguish high-frequency details from low-frequency components, thereby suppressing noise while enhancing details. The proposed method converts sRGB images into raw-RGB images by reversing the Image Signal Processor (ISP) pipeline to avoid unnecessary effects. A frequency information interaction processing unit is then designed to enhance low-frequency textures using high-frequency information and correct high-frequency data through low-frequency information to reduce structural deformation during restoration. The signal-to-noise ratio prior is introduced in the hidden layer to further remove noise. Experimental results show that the proposed method outperforms existing color and structure enhancement methods.
Na Li 0018, Dunlu Peng, Zied Bouraoui
ICASSP4
2025 Grounding Agent Reasoning in Image Schemas: A Neurosymbolic Approach to Embodied Cognition
François Olivier, Zied Bouraoui
AAMAS2
2025 Similarity-based Prompt Optimization for Distilling Factual Knowledge From Language Models
abstract
Considerable attention has been devoted recently to the problem of exploring and distilling factual knowledge from Pre-trained Language Models (PLMs). While continuous prompts are currently considered the best approach, determining the number and positions of continuous prompt vectors remains challenging to obtain the optimal prompts. In this paper, we introduce a new similarity-based algorithm that automatically searches for the best positions and number of continuous prompt vectors based on the similarity score between prompt representation and the relation representation for the final optimal prompts. The prompt optimization process is done before model training, making it efficient and less time-consuming. Our experiments demonstrate that our method significantly outperforms state-of-the-art methods in probing factual knowledge.
Na Li 0018, Dunlu Peng, Zied Bouraoui
IJCNN4
2025 Towards a Neurosymbolic Reasoning System Grounded in Schematic Representations
abstract
Despite significant progress in natural language understanding, Large Language Models (LLMs) remain error-prone when performing logical reasoning, often lacking the robust mental representations that enable human-like comprehension. We introduce a prototype neurosymbolic system, Embodied-LM, that grounds understanding and logical reasoning in schematic representations based on image schemas—recurring patterns derived from sensorimotor experience that structure human cognition. Our system operationalizes the spatial foundations of these cognitive structures using declarative spatial reasoning within Answer Set Programming. Through evaluation on logical deduction problems, we demonstrate that LLMs can be guided to interpret scenarios through embodied cognitive structures, that these structures can be formalized as executable programs, and that the resulting representations support effective logical reasoning with enhanced interpretability. While our current implementation focuses on spatial primitives, it establishes the computational foundation for incorporating more complex and dynamic representations.
François Olivier, Zied Bouraoui
NeSy2
2025 A parallel network encoding dialog history template for end-to-end task-oriented dialog
Guisong Yang, Decao Ma, Na Li 0018, Zied Bouraoui
CCF Trans. Pervasive Comput. Interact.5
2025 Modeling Multi-modal Cross-interaction for Multi-label Few-shot Image Classification Based on Local Feature Selection
abstract
The aim of multi-label few-shot image classification (ML-FSIC) is to assign semantic labels to images, in settings where only a small number of training examples are available for each label. A key feature of the multi-label setting is that an image often has several labels, which typically refer to objects appearing in different regions of the image. When estimating label prototypes, in a metric-based setting, it is thus important to determine which regions are relevant for which labels, but the limited amount of training data and the noisy nature of local features make this highly challenging. As a solution, we propose a strategy in which label prototypes are gradually refined. First, we initialize the prototypes using word embeddings, which allows us to leverage prior knowledge about the meaning of the labels. Second, taking advantage of these initial prototypes, we then use a Loss Change Measurement (LCM) strategy to select the local features from the training images (i.e., the support set) that are most likely to be representative of a given label. Third, we construct the final prototype of the label by aggregating these representative local features using a multi-modal cross-interaction mechanism, which again relies on the initial word embedding-based prototypes. Experiments on COCO, PASCAL VOC, NUS-WIDE, and iMaterialist show that our model substantially improves the current state-of-the-art.
Kun Yan 0008, Zied Bouraoui, Fangyun Wei, Chang Xu 0002, Ping Wang 0003, Shoaib Jameel, Steven Schockaert
ACM Trans. Multim. Comput. Commun. Appl.2
2024 Vector Field Oriented Diffusion Model for Crystal Material Generation
abstract
Discovering crystal structures with specific chemical properties has become an increasingly important focus in material science. However, current models are limited in their ability to generate new crystal lattices, as they only consider atomic positions or chemical composition. To address this issue, we propose a probabilistic diffusion model that utilizes a geometrically equivariant GNN to consider atomic positions and crystal lattices jointly. To evaluate the effectiveness of our model, we introduce a new generation metric inspired by Frechet Inception Distance, but based on GNN energy prediction rather than InceptionV3 used in computer vision. In addition to commonly used metrics like validity, which assesses the plausibility of a structure, this new metric offers a more comprehensive evaluation of our model's capabilities. Our experiments on existing benchmarks show the significance of our diffusion model. We also show that our method can effectively learn meaningful representations.
Astrid Klipfel, Yaël Frégier, Adlane Sayede, Zied Bouraoui
AAAI4
2024 Self-supervised Segment Contrastive Learning for Medical Document Representation
Waheed Ahmed Abro, Hanane Kteich, Zied Bouraoui
AIME (1)3
2024 Can Language Models Learn Embeddings of Propositional Logic Assertions?
abstract
Natural language offers an appealing alternative to formal logics as a vehicle for representing knowledge. However, using natural language means that standard methods for automated reasoning can no longer be used. A popular solution is to use transformer-based language models (LMs) to directly reason about knowledge expressed in natural language, but this has two important limitations. First, the set of premises is often too large to be directly processed by the LM. This means that we need a retrieval strategy which can select the most relevant premises when trying to infer some conclusion. Second, LMs have been found to learn shortcuts and thus lack robustness, putting in doubt to what extent they actually understand the knowledge that is expressed. Given these limitations, we explore the following alternative: rather than using LMs to perform reasoning directly, we use them to learn embeddings of individual assertions. Reasoning is then carried out by manipulating the learned embeddings. We show that this strategy is feasible to some extent, while at the same time also highlighting the limitations of directly fine-tuning LMs to learn the required embeddings.
Nurul Fajrin Ariyani, Zied Bouraoui, Richard Booth 0001, Steven Schockaert
LREC/COLING2
2024 AMenDeD: Modelling Concepts by Aligning Mentions, Definitions and Decontextualised Embeddings
abstract
Contextualised Language Models (LM) improve on traditional word embeddings by encoding the meaning of words in context. However, such models have also made it possible to learn high-quality decontextualised concept embeddings. Three main strategies for learning such embeddings have thus far been considered: (i) fine-tuning the LM to directly predict concept embeddings from the name of the concept itself, (ii) averaging contextualised representations of mentions of the concept in a corpus, and (iii) encoding definitions of the concept. As these strategies have complementary strengths and weaknesses, we propose to learn a unified embedding space in which all three types of representations can be integrated. We show that this allows us to outperform existing approaches in tasks such as ontology completion, which heavily depends on access to high-quality concept embeddings. We furthermore find that mentions and definitions are well-aligned in the resulting space, enabling tasks such as target sense verification, even without the need for any fine-tuning.
Amit Gajbhiye, Zied Bouraoui, Luis Espinosa Anke, Steven Schockaert
LREC/COLING2
2024 REFINE-LM: Mitigating Language Model Stereotypes via Reinforcement Learning
abstract
With the introduction of (large) language models, there has been significant concern about the unintended bias such models may inherit from their training data. A number of studies have shown that such models propagate gender stereotypes, as well as geographical and racial bias, among other biases. While existing works tackle this issue by preprocessing data and debiasing embeddings, the proposed methods require a lot of computational resources and annotation effort while being limited to certain types of biases. To address these issues, we introduce REFINE-LM, a debiasing method that uses reinforcement learning to handle different types of biases without any fine-tuning. By training a simple model on top of the word probability distribution of a LM, our bias agnostic reinforcement learning method enables model debiasing without human annotations or significant computational resources. Experiments conducted on a wide range of models, including several LMs, show that our method (i) significantly reduces stereotypical biases while preserving LMs performance; (ii) is applicable to different types of biases, generalizing across contexts such as gender, ethnicity, religion, and nationality-based biases; and (iii) it is not expensive to train.
Rameez Qureshi, Naïm Es-Sebbani, Luis Galárraga, Yvette Graham, Miguel Couceiro, Zied Bouraoui
ECAI6
2024 Synergies between machine learning and reasoning - An introduction by the Kay R. Amel group
abstract
This paper proposes a tentative and original survey of meeting points between Knowledge Representation and Reasoning (KRR) and Machine Learning (ML), two areas which have been developed quite separately in the last four decades. First, some common concerns are identified and discussed such as the types of representation used, the roles of knowledge and data, the lack or the excess of information, or the need for explanations and causal understanding. Then, the survey is organised in seven sections covering most of the territory where KRR and ML meet. We start with a section dealing with prototypical approaches from the literature on learning and reasoning: Inductive Logic Programming, Statistical Relational Learning, and Neurosymbolic AI, where ideas from rule-based reasoning are combined with ML. Then we focus on the use of various forms of background knowledge in learning, ranging from additional regularisation terms in loss functions, to the problem of aligning symbolic and vector space representations, or the use of knowledge graphs for learning. Then, the next section describes how KRR notions may benefit to learning tasks. For instance, constraints can be used as in declarative data mining for influencing the learned patterns; or semantic features are exploited in low-shot learning to compensate for the lack of data; or yet we can take advantage of analogies for learning purposes. Conversely, another section investigates how ML methods may serve KRR goals. For instance, one may learn special kinds of rules such as default rules, fuzzy rules or threshold rules, or special types of information such as constraints, or preferences. The section also covers formal concept analysis and rough sets-based methods. Yet another section reviews various interactions between Automated Reasoning and ML, such as the use of ML methods in SAT solving to make reasoning faster. Then a section deals with works related to model accountability, including explainability and interpretability, fairness and robustness. Finally, a section covers works on handling imperfect or incomplete data, including the problem of learning from uncertain or coarse data, the use of belief functions for regression, a revision-based view of the EM algorithm, the use of possibility theory in statistics, or the learning of imprecise models. This paper thus aims at a better mutual understanding of research in KRR and ML, and how they can cooperate. The paper is completed by an abundant bibliography.
Ismaïl Baaj, Zied Bouraoui, Antoine Cornuéjols, Thierry Denoeux, Sébastien Destercke, Didier Dubois, Marie-Jeanne Lesot, João Marques-Silva 0001, Jérôme Mengin, Henri Prade, Steven Schockaert, Mathieu Serrurier, Olivier Strauss, Christel Vrain
Int. J. Approx. Reason.2
2023 Equivariant Message Passing Neural Network for Crystal Material Discovery
abstract
Automatic material discovery with desired properties is a fundamental challenge for material sciences. Considerable attention has recently been devoted to generating stable crystal structures. While existing work has shown impressive success on supervised tasks such as property prediction, the progress on unsupervised tasks such as material generation is still hampered by the limited extent to which the equivalent geometric representations of the same crystal are considered. To address this challenge, we propose EPGNN a periodic equivariant message-passing neural network that learns crystal lattice deformation in an unsupervised fashion. Our model equivalently acts on lattice according to the deformation action that must be performed, making it suitable for crystal generation, relaxation and optimisation. We present experimental evaluations that demonstrate the effectiveness of our approach.
Astrid Klipfel, Zied Bouraoui, Olivier Peltre, Yaël Frégier, Najwa Harrati, Adlane Sayede
AAAI2
2023 What do Deck Chairs and Sun Hats Have in Common? Uncovering Shared Properties in Large Concept Vocabularies
abstract
Concepts play a central role in many applications.This includes settings where concepts have to be modelled in the absence of sentence context.Previous work has therefore focused on distilling decontextualised concept embeddings from language models.But concepts can be modelled from different perspectives, whereas concept embeddings typically mostly capture taxonomic structure.To address this issue, we propose a strategy for identifying what different concepts, from a potentially large concept vocabulary, have in common with others.We then represent concepts in terms of the properties they share with the other concepts.To demonstrate the practical usefulness of this way of modelling concepts, we consider the task of ultra-fine entity typing, which is a challenging multi-label classification problem.We show that by augmenting the label set with shared properties, we can improve the performance of the state-of-the-art models for this task. 1
Amit Gajbhiye, Zied Bouraoui, Na Li 0018, Usashi Chatterjee, Luis Espinosa Anke, Steven Schockaert
EMNLP2
2023 An Application of Priority-Based Lightweight Ontology Merging
abstract
International audience
Rim Mohamed, Truong-Thanh Ma, Zied Bouraoui
ICAART (2)3
2023 Unified Model for Crystalline Material Generation
abstract
One of the greatest challenges facing our society is the discovery of new innovative crystal materials with specific properties. Recently, the problem of generating crystal materials has received increasing attention, however, it remains unclear to what extent, or in what way, we can develop generative models that consider both the periodicity and equivalence geometric of crystal structures. To alleviate this issue, we propose two unified models that act at the same time on crystal lattice and atomic positions using periodic equivariant architectures. Our models are capable to learn any arbitrary crystal lattice deformation by lowering the total energy to reach thermodynamic stability. Code and data are available at https://github.com/aklipf/GemsNet.
Astrid Klipfel, Yaël Frégier, Adlane Sayede, Zied Bouraoui
IJCAI4
2023 Optimized Crystallographic Graph Generation for Material Science
abstract
Graph neural networks are widely used in machine learning applied to chemistry, and in particular for material science discovery. For crystalline materials, however, generating graph-based representation from geometrical information for neural networks is not a trivial task. The periodicity of crystalline needs efficient implementations to be processed in real-time under a massively parallel environment. With the aim of training graph-based generative models of new material discovery, we propose an efficient tool to generate cutoff graphs and k-nearest-neighbours graphs of periodic structures within GPU optimization. We provide pyMatGraph a Pytorch-compatible framework to generate graphs in real-time during the training of neural network architecture. Our tool can update a graph of a structure, making generative models able to update the geometry and process the updated graph during the forward propagation on the GPU side. Our code is publicly available at https://github.com/aklipf/mat-graph.
Astrid Klipfel, Yaël Frégier, Adlane Sayede, Zied Bouraoui
IJCAI4
2023 Distilling Semantic Concept Embeddings from Contrastively Fine-Tuned Language Models
abstract
Learning vectors that capture the meaning of concepts remains a fundamental challenge. Somewhat surprisingly, perhaps, pre-trained language models have thus far only enabled modest improvements to the quality of such concept embeddings. Current strategies for using language models typically represent a concept by averaging the contextualised representations of its mentions in some corpus. This is potentially sub-optimal for at least two reasons. First, contextualised word vectors have an unusual geometry, which hampers downstream tasks. Second, concept embeddings should capture the semantic properties of concepts, whereas contextualised word vectors are also affected by other factors. To address these issues, we propose two contrastive learning strategies, based on the view that whenever two sentences reveal similar properties, the corresponding contextualised vectors should also be similar. One strategy is fully unsupervised, estimating the properties which are expressed in a sentence from the neighbourhood structure of the contextualised word embeddings. The second strategy instead relies on a distant supervision signal from ConceptNet. Our experimental results show that the resulting vectors substantially outperform existing concept embeddings in predicting the semantic properties of concepts, with the ConceptNet-based strategy achieving the best results. These findings are furthermore confirmed in a clustering task and in the downstream task of ontology completion.
Na Li 0018, Hanane Kteich, Zied Bouraoui, Steven Schockaert
SIGIR3
2023 Revision of prioritized Eℒ ontologies
Rim Mohamed, Zied Loukil, Faïez Gargouri, Zied Bouraoui
Appl. Intell.4
2022 Inferring Prototypes for Multi-Label Few-Shot Image Classification with Word Vector Guided Attention
abstract
Multi-label few-shot image classification (ML-FSIC) is the task of assigning descriptive labels to previously unseen images, based on a small number of training examples. A key feature of the multi-label setting is that images often have multiple labels, which typically refer to different regions of the image. When estimating prototypes, in a metric-based setting, it is thus important to determine which regions are relevant for which labels, but the limited amount of training data makes this highly challenging. As a solution, in this paper we propose to use word embeddings as a form of prior knowledge about the meaning of the labels. In particular, visual prototypes are obtained by aggregating the local feature maps of the support images, using an attention mechanism that relies on the label embeddings. As an important advantage, our model can infer prototypes for unseen labels without the need for fine-tuning any model parameters, which demonstrates its strong generalization abilities. Experiments on COCO and PASCAL VOC furthermore show that our model substantially improves the current state-of-the-art.
Kun Yan 0008, Chenbin Zhang, Ping Wang 0003, Zied Bouraoui, Shoaib Jameel, Steven Schockaert
AAAI5
2022 Evolution of Prioritized Eℒ Ontologies
Rim Mohamed, Zied Loukil, Faïez Gargouri, Zied Bouraoui
IEA/AIE4
2022 Representing Vietnamese Traditional Dances and Handling Inconsistent Information
Salem Benferhat, Zied Bouraoui, Truong-Thanh Ma, Karim Tabia
IPMU (2)2
2022 Region-Based Merging of Open-Domain Terminological Knowledge
Zied Bouraoui, Sébastien Konieczny, Thanh Ma, Nicolas Schwind, Ivan Varzinczak
KR1
2022 Tree Edit Distance Based Ontology Merging Evaluation Framework
Zied Bouraoui, Sébastien Konieczny, Thanh Ma, Ivan Varzinczak
KSEM (2)1
2022 Sentence Selection Strategies for Distilling Word Embeddings from BERT
abstract
Many applications crucially rely on the availability of high-quality word vectors. To learn such representations, several strategies based on language models have been proposed in recent years. While effective, these methods typically rely on a large number of contextualised vectors for each word, which makes them impractical. In this paper, we investigate whether similar results can be obtained when only a few contextualised representations of each word can be used. To this end, we analyse a range of strategies for selecting the most informative sentences. Our results show that with a careful selection strategy, high-quality word vectors can be learned from as few as 5 to 10 sentences.
Zied Bouraoui, Luis Espinosa Anke, Steven Schockaert
LREC2
2021 Few-Shot Image Classification with Multi-Facet Prototypes
abstract
The aim of few-shot learning (FSL) is to learn how to recognize image categories from a small number of training examples. A central challenge is that the available training examples are normally insufficient to determine which visual features are most characteristic of the considered categories. To address this challenge, we organise these visual features into facets, which intuitively group features of the same kind (e.g. features that are relevant to shape, color, or texture). This is motivated from the assumption that (i) the importance of each facet differs from category to category and (ii) it is possible to predict facet importance from a pre-trained embedding of the category names. In particular, we propose an adaptive similarity measure, relying on predicted facet importance weights for a given set of categories. This measure can be used in combination with a wide array of existing metric-based methods. Experiments on miniImageNet and CUB show that our approach improves the state-of-the-art in metric-based FSL.
Kun Yan 0008, Zied Bouraoui, Ping Wang 0003, Shoaib Jameel, Steven Schockaert
ICASSP2
2021 Modelling General Properties of Nouns by Selectively Averaging Contextualised Embeddings
abstract
While the success of pre-trained language models has largely eliminated the need for high-quality static word vectors in many NLP applications, static word vectors continue to play an important role in tasks where word meaning needs to be modelled in the absence of linguistic context. In this paper, we explore how the contextualised embeddings predicted by BERT can be used to produce high-quality word vectors for such domains, in particular related to knowledge base completion, where our focus is on capturing the semantic properties of nouns. We find that a simple strategy of averaging the contextualised embeddings of masked word mentions leads to vectors that outperform the static word vectors learned by BERT, as well as those from standard word embedding models, in property induction tasks. We notice in particular that masking target words is critical to achieve this strong performance, as the resulting vectors focus less on idiosyncratic properties and more on general semantic properties. Inspired by this view, we propose a filtering strategy which is aimed at removing the most idiosyncratic mention vectors, allowing us to obtain further performance gains in property induction.
Na Li 0018, Zied Bouraoui, José Camacho-Collados, Luis Espinosa Anke, Qing Gu 0001, Steven Schockaert
IJCAI2
2021 Aligning Visual Prototypes with BERT Embeddings for Few-Shot Learning
abstract
Few-shot learning (FSL) is the task of learning to recognize previously unseen categories of images from a small number of training examples. This is a challenging task, as the available examples may not be enough to unambiguously determine which visual features are most characteristic of the considered categories. To alleviate this issue, we propose a method that additionally takes into account the names of the image classes. While the use of class names has already been explored in previous work, our approach differs in two key aspects. First, while previous work has aimed to directly predict visual prototypes from word embeddings, we found that better results can be obtained by treating visual and text-based prototypes separately. Second, we propose a simple strategy for learning class name embeddings using the BERT language model, which we found to substantially outperform the GloVe vectors that were used in previous work. We furthermore propose a strategy for dealing with the high dimensionality of these vectors, inspired by models for aligning cross-lingual word embeddings. We provide experiments on miniImageNet, CUB and tieredImageNet, showing that our approach consistently improves the state-of-the-art in metric-based FSL.
Kun Yan 0008, Zied Bouraoui, Ping Wang 0003, Shoaib Jameel, Steven Schockaert
ICMR2
2020 Modelling Semantic Categories Using Conceptual Neighborhood
abstract
While many methods for learning vector space embeddings have been proposed in the field of Natural Language Processing, these methods typically do not distinguish between categories and individuals. Intuitively, if individuals are represented as vectors, we can think of categories as (soft) regions in the embedding space. Unfortunately, meaningful regions can be difficult to estimate, especially since we often have few examples of individuals that belong to a given category. To address this issue, we rely on the fact that different categories are often highly interdependent. In particular, categories often have conceptual neighbors, which are disjoint from but closely related to the given category (e.g. fruit and vegetable). Our hypothesis is that more accurate category representations can be learned by relying on the assumption that the regions representing such conceptual neighbors should be adjacent in the embedding space. We propose a simple method for identifying conceptual neighbors and then show that incorporating these conceptual neighbors indeed leads to more accurate region based representations.
Zied Bouraoui, José Camacho-Collados, Luis Espinosa Anke, Steven Schockaert
AAAI1
2020 Inducing Relational Knowledge from BERT
abstract
One of the most remarkable properties of word embeddings is the fact that they capture certain types of semantic and syntactic relationships. Recently, pre-trained language models such as BERT have achieved groundbreaking results across a wide range of Natural Language Processing tasks. However, it is unclear to what extent such models capture relational knowledge beyond what is already captured by standard word embeddings. To explore this question, we propose a methodology for distilling relational knowledge from a pre-trained language model. Starting from a few seed instances of a given relation, we first use a large text corpus to find sentences that are likely to express this relation. We then use a subset of these extracted sentences as templates. Finally, we fine-tune a language model to predict whether a given word pair is likely to be an instance of some relation, when given an instantiated template for that relation as input.
Zied Bouraoui, José Camacho-Collados, Steven Schockaert
AAAI1
2020 A Mixture-of-Experts Model for Learning Multi-Facet Entity Embeddings
abstract
Various methods have already been proposed for learning entity embeddings from text descriptions.Such embeddings are commonly used for inferring properties of entities, for recommendation and entity-oriented search, and for injecting background knowledge into neural architectures, among others.Entity embeddings essentially serve as a compact encoding of a similarity relation, but similarity is an inherently multi-faceted notion.By representing entities as single vectors, existing methods leave it to downstream applications to identify these different facets, and to select the most relevant ones.In this paper, we propose a model that instead learns several vectors for each entity, each of which intuitively captures a different aspect of the considered domain.We use a mixture-of-experts formulation to jointly learn these facet-specific embeddings.The individual entity embeddings are learned using a variant of the GloVe model, which has the advantage that we can easily identify which properties are modelled well in which of the learned embeddings.This is exploited by an associated gating network, which uses pre-trained word vectors to encourage the properties that are modelled by a given embedding to be semantically coherent, i.e. to encourage each of the individual embeddings to capture a meaningful facet.
Rana Alshaikh, Zied Bouraoui, Shelan Jeawak, Steven Schockaert
COLING2
2020 Consolidating Modal Knowledge Bases
Zied Bouraoui, Jean-Marie Lagniez, Pierre Marquis, Valentin Montmirail
ECAI1
2020 Model-based Merging of Open-Domain Ontologies
abstract
Conceptual knowledge, encoded in ontologies or knowledge graphs, plays an essential role in many areas, including Semantic Web, Information Retrieval, and Natural Language Processing. Considerable attention has recently been devoted to the problem of unifying and linking available ontologies. While the vast majority of existing work focuses on matching or aligning resources, in this paper, we investigate the application of belief merging theory to ontology merging to obtain a unique perspective. We consider the setting where different ontologies share the same terminology (i.e., assuming that they are already mapped to each other). However, they express knowledge in different and potentially conflicting ways. In order to get a unified view of the knowledge conveyed by the different ontologies, we start by providing a semantic-based merging model. Our method retrieves all the interpretations in which the outcome can be found. We support demonstrating the method's effectiveness by an experimental evaluation of the method on existing open-domain ontologies.
Zied Bouraoui, Sébastien Konieczny, Truong-Thanh Ma, Ivan Varzinczak
ICTAI1
2020 Hierarchical Linear Disentanglement of Data-Driven Conceptual Spaces
abstract
Conceptual spaces are geometric meaning representations in which similar entities are represented by similar vectors. They are widely used in cognitive science, but there has been relatively little work on learning such representations from data. In particular, while standard representation learning methods can be used to induce vector space embeddings from text corpora, these differ from conceptual spaces in two crucial ways. First, the dimensions of a conceptual space correspond to salient semantic features, known as quality dimensions, whereas the dimensions of learned vector space embeddings typically lack any clear interpretation. This has been partially addressed in previous work, which has shown that it is possible to identify directions in learned vector spaces which capture semantic features. Second, conceptual spaces are normally organised into a set of domains, each of which is associated with a separate vector space. In contrast, learned embeddings represent all entities in a single vector space. Our hypothesis in this paper is that such single-space representations are sub-optimal for learning quality dimensions, due to the fact that semantic features are often only relevant to a subset of the entities. We show that this issue can be mitigated by identifying features in a hierarchical fashion. Intuitively, the top-level features split the vector space into different domains, making it possible to subsequently identify domain-specific quality dimensions.
Rana Alshaikh, Zied Bouraoui, Steven Schockaert
IJCAI2
2019 Automated Rule Base Completion as Bayesian Concept Induction
abstract
Considerable attention has recently been devoted to the problem of automatically extending knowledge bases by applying some form of inductive reasoning. While the vast majority of existing work is centred around so-called knowledge graphs, in this paper we consider a setting where the input consists of a set of (existential) rules. To this end, we exploit a vector space representation of the considered concepts, which is partly induced from the rule base itself and partly from a pre-trained word embedding. Inspired by recent approaches to concept induction, we then model rule templates in this vector space embedding using Gaussian distributions. Unlike many existing approaches, we learn rules by directly exploiting regularities in the given rule base, and do not require that a database with concept and relation instances is given. As a result, our method can be applied to a wide variety of ontologies. We present experimental results that demonstrate the effectiveness of our method.
Zied Bouraoui, Steven Schockaert
AAAI1
2019 Learning Conceptual Spaces with Disentangled Facets
abstract
Conceptual spaces are geometric representations of meaning that were proposed by Gärdenfors (2000).They share many similarities with the vector space embeddings that are commonly used in natural language processing.However, rather than representing entities in a single vector space, conceptual spaces are usually decomposed into several facets, each of which is then modelled as a relatively lowdimensional vector space.Unfortunately, the problem of learning such conceptual spaces has thus far only received limited attention.To address this gap, we analyze how, and to what extent, a given vector space embedding can be decomposed into meaningful facets in an unsupervised fashion.While this problem is highly challenging, we show that useful facets can be discovered by relying on word embeddings to group semantically related features.
Rana Alshaikh, Zied Bouraoui, Steven Schockaert
CoNLL2
2019 An Automatic Extraction Tool for Ethnic Vietnamese Thai Dances Concepts
abstract
In recent year, preservation and promotion of the ICHs are one of the problems of interest. In this paper, we focus on modelling the traditional dance domain, particularly modelling traditional Vietnamese dances. To conserve significant characteristics of dances, we proposed an ontology to represent the significant movements features of Ethnic Vietnamese Thai Dances (EVTDs). Particularly, a detailed description of the movement schemas of EVTDs is presented in this paper. Additionally, we present how to build an automatic extraction tool to collect the fundamental movements data of EVTDs using machine learning. Finally, we represented explicitly how to store those extracted features from raw dance videos into prioritized Ontology-based proposed.
Truong-Thanh Ma, Salem Benferhat, Zied Bouraoui, Karim Tabia, Thanh-Nghi Do, Nguyen-Khang Pham
ICMLA3
2019 Ontology Completion Using Graph Convolutional Networks
Na Li 0018, Zied Bouraoui, Steven Schockaert
ISWC (1)2
2018 Unsupervised Learning of Distributional Relation Vectors
abstract
Word embedding models such as GloVe rely on co-occurrence statistics to learn vector representations of word meaning.While we may similarly expect that cooccurrence statistics can be used to capture rich information about the relationships between different words, existing approaches for modeling such relationships are based on manipulating pre-trained word vectors.In this paper, we introduce a novel method which directly learns relation vectors from co-occurrence statistics.To this end, we first introduce a variant of GloVe, in which there is an explicit connection between word vectors and PMI weighted co-occurrence vectors.We then show how relation vectors can be naturally embedded into the resulting vector space.
Shoaib Jameel, Zied Bouraoui, Steven Schockaert
ACL (1)2
2018 Relation Induction in Word Embeddings Revisited
abstract
Given a set of instances of some relation, the relation induction task is to predict which other word pairs are likely to be related in the same way. While it is natural to use word embeddings for this task, standard approaches based on vector translations turn out to perform poorly. To address this issue, we propose two probabilistic relation induction models. The first model is based on translations, but uses Gaussians to explicitly model the variability of these translations and to encode soft constraints on the source and target words that may be chosen. In the second model, we use Bayesian linear regression to encode the assumption that there is a linear relationship between the vector representations of related words, which is considerably weaker than the assumption underlying translation based models.
Zied Bouraoui, Shoaib Jameel, Steven Schockaert
COLING1
2018 Learning Conceptual Space Representations of Interrelated Concepts
abstract
Several recently proposed methods aim to learn conceptual space representations from large text collections. These learned representations associate each object from a given domain of interest with a point in a high-dimensional Euclidean space, but they do not model the concepts from this domain, and can thus not directly be used for categorization and related cognitive tasks. A natural solution is to represent concepts as Gaussians, learned from the representations of their instances, but this can only be reliably done if sufficiently many instances are given, which is often not the case. In this paper, we introduce a Bayesian model which addresses this problem by constructing informative priors from background knowledge about how the concepts of interest are interrelated with each other. We show that this leads to substantially better predictions in a knowledge base completion task.
Zied Bouraoui, Steven Schockaert
IJCAI1
2018 Qualitative-Based Possibilistic EL Ontology
Rim Mohamed, Zied Loukil, Zied Bouraoui
PRIMA3
2018 An Ontology-based Modelling of Vietnamese Traditional Dances (S)
abstract
Ontology is an essential resource to enhance the performance of information processing system as well as is an intelligent storage area served for management of largescale heterogeneous digital contents resulting.In this paper, we propose the initial steps for reconstructing a significant schema of Vietnamese traditional dances.Most of the typical dances of Vietnamese community are recorded in multimedia format, in raw videos.Accordingly, we concentrated on analyzing and collecting knowledge of the dance experts at art schools in Vietnam to classify and to determine the primary features that would be stored in the ontology.We propose an ontologybased modelling for the cultural heritage domain of Vietnamese traditional dance.
Truong-Thanh Ma, Salem Benferhat, Zied Bouraoui, Karim Tabia, Thanh-Nghi Do, Huu-Hoa Nguyen
SEKE3
2017 Inductive Reasoning about Ontologies Using Conceptual Spaces
abstract
Structured knowledge about concepts plays an increasingly important role in areas such as information retrieval. The available ontologies and knowledge graphs that encode such conceptual knowledge, however, are inevitably incomplete. This observation has led to a number of methods that aim to automatically complete existing knowledge bases. Unfortunately, most existing approaches rely on black box models, e.g. formulated as global optimization problems, which makes it difficult to support the underlying reasoning process with intuitive explanations. In this paper, we propose a new method for knowledge base completion, which uses interpretable conceptual space representations and an explicit model for inductive inference that is closer to human forms of commonsense reasoning. Moreover, by separating the task of representation learning from inductive reasoning, our method is easier to apply in a wider variety of contexts. Finally, unlike optimization based approaches, our method can naturally be applied in settings where various logical constraints between the extensions of concepts need to be taken into account.
Zied Bouraoui, Shoaib Jameel, Steven Schockaert
AAAI1
2017 A Polynomial Algorithm for Merging Lightweight Ontologies in Possibility Theory Under Incommensurability Assumption
Salem Benferhat, Zied Bouraoui, Ma Thi Chau, Sylvain Lagrue, Julien Rossit
ICAART (2)2
2017 MEmbER: Max-Margin Based Embeddings for Entity Retrieval
abstract
We propose a new class of methods for learning vector space embeddings of entities. While most existing methods focus on modelling similarity, our primary aim is to learn embeddings that are interpretable, in the sense that query terms have a direct geometric representation in the vector space. Intuitively, we want all entities that have some property (i.e. for which a given term is relevant) to be located in some well-defined region of the space. This is achieved by imposing max-margin constraints that are derived from a bag-of-words representation of the entities. The resulting vector spaces provide us with a natural vehicle for identifying entities that have a given property (or ranking them according to how much they have the property), and conversely, to describe what a given set of entities have in common. As we show in our experiments, our models lead to a substantially better performance in a range of entity-oriented search tasks, such as list completion and entity ranking.
Shoaib Jameel, Zied Bouraoui, Steven Schockaert
SIGIR2
2017 Min-based possibilistic DL-Lite
abstract
DL-Lite is one of the most important fragments of description logics that allows a flexible representation of knowledge with a tractable computational complexity of the reasoning process. This article investigates an extension of the main fragments of DL-Lite to deal with uncertainty associated with objects, concepts or relations using a possibility theory framework. Possibility theory offers a natural framework for representing uncertain and incomplete information. It is particularly useful for handling inconsistent knowledge. We first provide foundations of possibilistic DL-Lite , denoted by $$\\pi $$ - $$DL$$ - $$Lite$$ , by extending the $$DL$$ - $$Lit{e}_{core}$$ logic, the core fragment of all DL-Lite logics, within possibility theory setting. We present syntax and semantics of $$\\pi $$ - $$DL$$ - $$Lit{e}_{core}$$ , study the reasoning tasks and show how to compute the inconsistency degree of a $$\\pi $$ - $$DL$$ - $$Lit{e}_{core}$$ knowledge base. We then extend our possibilistic approach to $$DL$$ - $$Lit{e}_{F}$$ and $$DL$$ - $$Lit{e}_{R}$$ , two important fragments of DL-Lite family. Finally, we address the problem of query answering over a $$\\pi $$ - $$DL$$ - $$Lite$$ knowledge base. An important result of the article is that the extension of the expressive power of DL-Lite is done without additional extra-computational costs.
Salem Benferhat, Zied Bouraoui
J. Log. Comput.2
2016 Non-Objection Inference for Inconsistency-Tolerant Query Answering
Salem Benferhat, Zied Bouraoui, Madalina Croitoru, Odile Papini, Karim Tabia
IJCAI2
2016 Inconsistency-Tolerant Query Answering: Rationality Properties and Computational Complexity Analysis
Jean-François Baget, Salem Benferhat, Zied Bouraoui, Madalina Croitoru, Marie-Laure Mugnier, Odile Papini, Swan Rocher, Karim Tabia
JELIA3
2016 A General Modifier-Based Framework for Inconsistency-Tolerant Query Answering
Jean-François Baget, Salem Benferhat, Zied Bouraoui, Madalina Croitoru, Marie-Laure Mugnier, Odile Papini, Swan Rocher, Karim Tabia
KR3
2015 How to Select One Preferred Assertional-Based Repair from Inconsistent and Prioritized DL-Lite Knowledge Bases?
Salem Benferhat, Zied Bouraoui, Karim Tabia
IJCAI2
2014 Assertional-based Prioritized Removed Sets Revision of DL-LiteR Knowledge Bases
abstract
The paper proposes an extension of “Prioritized Removed Sets Revision” (PRSR) to DL-LiteRstratified knowledge bases. The revision strategy is based on inconsistency minimization and consists in determining smallest subsets of assertions to be dropped from the current DL-LiteRknowledge base, taking the stratification into account, in order to restore consistency and accept the input. We consider different forms of input: membership assertion, positive inclusion axiom or negative inclusion axiom. We show that according to the form of input and under some conditions PRSR can be achieved in polynomial time.
Salem Benferhat, Zied Bouraoui, Odile Papini, Éric Würbel
ECAI2
2014 A Prioritized Assertional-Based Revision for DL-Lite Knowledge Bases
Salem Benferhat, Zied Bouraoui, Odile Papini, Éric Würbel
JELIA2
2013 Min-Based Fusion of Possibilistic DL-Lite Knowledge Bases
abstract
DL-Lite is one of the most important tractable fragment of DLs that provides a powerful framework to compactly encode available knowledge with a low computational complexity of the reasoning process. In semantic web area, merging different and often conflicting sources of information, has been recognized as an important problem. Pieces of information to be combined are provided with uncertainty due for instance to the reliability of sources. Possibility theory offers an important tool for representing and reasoning with uncertain, partial and inconsistent pieces of information. This paper first presents possibilistic DL-Lite, denoted by π-DL-Lite as an extension of DL-Lite within a possibility theory setting. It then focuses on the use of a minimum-based (min-based) operator, well known as idempotent conjunctive operator to combine π-DLLite possibility distributions and it shows that the semantic fusion of π-DL-Lite possibility distributions has a natural syntactic counterpart when dealing with π-DL-Lite knowledge bases. The min-based fusion operator is recommended when distinct sources that provide information are dependent.
Salem Benferhat, Zied Bouraoui, Zied Loukil
Web Intelligence2