Engelbert Mephu Nguifo

dblp:10/3832 · DBLP profile ↗
← Back
61ranked-venue papers
6as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 37 · 4 first-author · 12 since 2021Databases, data management, data science and information retrieval · 11 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 6Systems, architecture and hardware · 3 · 1 since 2021Theory of computation · 3Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
YearPublicationVenuePosition
2026 OnlineBootKNN: An Unsupervised Framework for Detecting Anomalies in Spectral Data Streams
abstract
Monitoring the elemental composition of materials in order to detect abnormal conditions in real-time is essential for applications like manufacturing quality control, environmental monitoring, and space exploration. This is achieved using sensors that analyze the interaction of a material with electromagnetic radiation, producing spectral data streams or a sequence of instances where each represents an ordered set of wavelengths with an associated intensity. While many unsupervised anomaly detection methods exist for tabular streaming data, their applicability to spectral streams remains underexplored. To address this gap, we consider our spectra in a multivariate stream setting and benchmark the performance of state-of-the-art tabular anomaly detection methods on this data. Furthermore, we introduce OnlineBootKNN, a novel unsupervised framework that combines k-nearest neighbors with online bootstrapping and a z-score test to detect anomalies in real-time. We demonstrate the high performance and robustness of our method, as well as the efficacy of the autoencoder-based method, KitNet, on newly simulated real-world spectral datasets. In addition, we compare their efficiency against the other tested techniques. Finally, we highlight the inherent interpretability of OnlineBootKNN, which is crucial for identifying the specific wavelengths, and thus elements, responsible for a detected anomaly.
Nicolas Rojas Varela, Julien Ah-Pine, Engelbert Mephu Nguifo
AAAI3
2026 Accelerating Frequent Gradual Pattern Discovery Through Dimensionality Reduction
Herman Tcheneghon Motcheyo, Issam Falih, Lauraine Tiogning Kueti, Engelbert Mephu Nguifo
ISMIS4
2025 WeightedHGE: Weighted Heterogeneous Graph Embedding
abstract
In this paper, we introduce a novel method for embedding weighted graphs, designed to capture the significance of relationships in real-world weighted graph data. Unlike existing models that treat all relationships equally, our approach incorporates edge weights directly into the embedding process, enhancing the representation of highly weighted connections. By introducing modifications to both the scoring and loss functions, our method emphasizes the importance of weighted relationships during training. Theoretical analysis shows that our model maintains efficiency and scalability while significantly improving the capacity to represent weighted graph structures. Experimental results demonstrate that our approach consistently outperforms baseline methods, including TransE, in tasks involving weighted relationships, showcasing its robustness and applicability.
Khouloud Ammar, Wissem Inoubli, Sami Zghal, Engelbert Mephu Nguifo
KES4
2025 XLITE-Unet: Extremely Light and Efficient Deep learning architecture with selective atrous and axial depthwise convolution for image segmentation
Cezar Mbiethieu, Norbert Tsopzé, Engelbert Mephu Nguifo
Comput. Vis. Image Underst.3
2025 A degree centrality-enhanced computational approach for local network alignment leveraging knowledge graph embeddings
Warith Eddine Djeddi, Sadok Ben Yahia, Engelbert Mephu Nguifo
Expert Syst. Appl.3
2025 Navigating complexity: a comprehensive review of heterogeneous information networks and embedding techniques
Khouloud Ammar, Wissem Inoubli, Sami Zghal, Engelbert Mephu Nguifo
Knowl. Inf. Syst.4
2024 Scaling Knowledge Graph Embedding with Parallel TransE and Graph Partitioning
abstract
Knowledge graph embedding has emerged as a fundamental technique to represent entities and relationships in knowledge graphs within low-dimensional vector spaces. Among these methods, translation-based approaches stand out by treating relations as translations from head entities to tail entities, achieving state-of-the-art results. However, the training process of these methods can be prohibitively time-consuming, especially for large knowledge graphs, posing significant challenges in practical applications. As knowledge graphs grow in size and complexity, surpassing the capacities of existing systems, there is an urgent need for scalable solutions in knowledge representation learning. These graphs, comprising millions of nodes and billions of edges, serve as powerful data structures for representing and understanding complex networks of knowledge. Translation-based models, particularly TransE, have been instrumental in encoding structured information about entities and their relationships in low-dimensional embedding spaces. However, the current implementation of TransE is constrained to single-node machines, limiting its scalability and applicability. To address these limitations, this paper proposes leveraging parallel computing technique, such as parallel TransE with graph partitioning, to enhance scalability and efficiency in knowledge graph embedding.
Khouloud Ammar, Wissem Inoubli, Sami Zghal, Engelbert Mephu Nguifo
AICCSA4
2024 Revisiting Frequent (Closed) Gradual Itemsets Mining
abstract
The task of mining gradual itemsets holds significant importance in pattern mining, particularly when working with numerical data. It involves the discovery of covariations between attributes in the form of “The more/less X,…, the more/less Y,” referred to as gradual itemsets. However, discovering these itemsets remains challenging, partly due to the exponential combinatorial search space involved in large-scale data processing. Consequently, existing algorithms for gradual itemset mining encounter difficulties, such as slow processing speeds, and occasional failures to terminate due to the overwhelming number of candidate itemsets requiring exploration. A large number of candidates is generated, but a large proportion of them turns out to be infrequent once their supports are computed. This paper introduces an approach to streamline this process by efficiently reducing the number of candidates for which support needs to be computed through the introduction of a stricter upper-bound criterion. By circumventing the costly support computation for numerous candidate itemsets, our approach exhibits efficiency in terms of speed when applied to real databases, including large-scale databases that pose challenges for existing algorithms. Furthermore, we establish a connection in terms of pattern coverage between the two principal gradualness semantics commonly employed in the literature.
Jerry Lonlac, Bernoulli Fotsing Tchide, Alain Bertrand Bomgni, Arnaud Doniec, Engelbert Mephu Nguifo
ICTAI5
2024 Scalable and accurate subsequence transform for time series classification
Michael Franklin Mbouopda, Engelbert Mephu Nguifo
Pattern Recognit.2
2023 Extracting Frequent Gradual Patterns Based on SAT
abstract
International audience
Jerry Lonlac, Imen Ouled Dlala, Saïd Jabbour, Engelbert Mephu Nguifo, Badran Raddaoui, Lakhdar Sais
DATA4
2023 DGCN: Learning Graph Representations Via Dense Connections
abstract
In the last decades, learning over graph data has became one of the most challenging tasks in deep learning. The generally proposed Graph Neural Network (GNN) framework computes a hidden state for every node in the graph by applying nonlinear transformations to its neighborhood. Then updates the node of interest's hidden state. Nevertheless, the node features can contain discriminative information. That can get lost over GNN transformations. In this work, we present a new variant of GNN architecture where we combine node features and GNN activations to learn node representations in the graph. We conduct extensive experiments on two graph prediction tasks (node classification and link prediction) and show that our method can match and outperforms state-of-the-art results on five challenging datasets.
Khairi Abidi, Wissem Inoubli, Engelbert Mephu Nguifo
KES3
2023 Trans-Trip: Translation-based embedding with Triplets for Heterogeneous Graphs
abstract
Heterogeneous graphs (HG) are an effective way of abstracting complex systems, including social, biological, and economic systems. However, modeling these graphs is challenging due to their high dimensionality, sparsity, and heterogeneity. Traditional approaches designed for homogeneous graphs struggle to handle the diverse types of entities and relationships present in HG, leading to a loss of information and potentially inaccurate embeddings. To address these challenges, we introduce a novel method, Trans-Trip: Translation-based embedding with Triplets for Heterogeneous Graphs, that leverages the power of triplets of (entity, relation, entity) to accurately represent the various types of relationships among nodes and links. Trans-trip effectively captures the rich semantics embedded in HGs, overcoming the challenges of global coherence and entity projection. By leveraging the flexibility and interpretability of triplets, our method can handle the multi-types nodes/links present in HG and can capture complex higher-order structures. We demonstrate the effectiveness of our proposed method on several benchmark datasets, showing that it outperforms existing embedding methods. Trans-trip provides a more accurate and interpretable representation of HG, which can be used across various fields, such as biology, social networks, and e-commerce.
Khouloud Ammar, Wissem Inoubli, Sami Zghal, Amel Borji, Engelbert Mephu Nguifo
KES5
2022 A survey on unsupervised learning algorithms for detecting abnormal points in streaming data
abstract
One of the critical tasks of data stream analysis is anomaly detection. Various methods based on multiple assumptions have been reported in the literature. However, there is still a lack of experimental comparison of those methods, which makes it difficult to choose a specific one. In this paper, we compared unsupervised data stream abnormal point detection methods on various datasets with emphasis on their performance and runtime, as well as the presence of concept drift, seasonality, trend, and cycle as a characteristic of the dataset. Our experiments show that forecasting-based methods are the ones managing the best seasonality and trend, and lightweight models performing online gradient descent have a lower execution time. The details of our experiments are available online.
Anne Marthe Sophie Ngo Bibinbe, Michael Franklin Mbouopda, Gertrude Raissa Mbiadou Saleu, Engelbert Mephu Nguifo
IJCNN4
2022 A distributed and incremental algorithm for large-scale graph clustering
Wissem Inoubli, Sabeur Aridhi, Haithem Mezni, Mondher Maddouri, Engelbert Mephu Nguifo
Future Gener. Comput. Syst.5
2022 On the design of a similarity function for sparse binary data with application on protein function annotation
Marcelo B. A. Veras, Bishnu Sarker, Sabeur Aridhi, João Paulo Pordeus Gomes, José A. F. de Macêdo, Engelbert Mephu Nguifo, Marie-Dominique Devignes, Malika Smaïl-Tabbone
Knowl. Based Syst.6
2020 Categorical fuzzy entropy c-means
abstract
Hard and fuzzy clustering algorithms are part of the partition-based clustering family. They are widely used in real-world applications to cluster numerical and categorical data. While in hard clustering an object is assigned to a cluster with certainty, in fuzzy clustering an object can be assigned to different clusters given a membership degree. For both types of method an entropy can be incorporated into the objective function, mostly to avoid solutions raising too much uncertainties. In this paper, we present an extension of a fuzzy clustering method for categorical data using fuzzy centroids. The new algorithm, referred to as Categorical Fuzzy Entropy (CFE), integrates an entropy term in the objective function. This allows a better fuzzification of the cluster prototypes. Experiments on ten real-world data sets and statistical comparisons show that the new method can efficiently handle categorical data.
Abdoul Jalil Djiberou Mahamadou, Violaine Antoine, Engelbert Mephu Nguifo, Sylvain Moreno
FUZZ-IEEE3
2020 Special issue on "Advances on Large Evolving Graphs"
Sabeur Aridhi, José A. F. de Macêdo, Engelbert Mephu Nguifo, Karine Zeitouni
Future Gener. Comput. Syst.3
2020 A novel algorithm for searching frequent gradual patterns from an ordered data set
abstract
Mining frequent simultaneous attribute co-variations in numerical databases is also called frequent gradual pattern problem. Few efficient algorithms for automatically extracting such patterns have been reported in the literature. Their main difference resides in the variation semantics used. However in applications with temporal order relations, those algorithms fail to generate correct frequent gradual patterns as they do not take this temporal constraint into account in the mining process. In this paper, we propose an approach for extracting frequent gradual patterns for which the ordering of supporting objects matches the temporal order. This approach considerably reduces the number of gradual patterns within an ordered data set. The experimental results show the benefits of our approach.
Jerry Lonlac, Engelbert Mephu Nguifo
Intell. Data Anal.2
2020 Frobenius correlation based u-shapelets discovery for time series clustering
Vanel Steve Siyou Fotso, Engelbert Mephu Nguifo, Philippe Vaslin
Pattern Recognit.2
2019 A Structure Based Multiple Instance Learning Approach for Bacterial Ionizing Radiation Resistance Prediction
abstract
Ionizing-radiation-resistant bacteria (IRRB) could be used for bioremediation of radioactive wastes and in the therapeutic industry. Limited computational works are available for the prediction of bacterial ionizing radiation resistance (IRR). In this work, we present ABClass, an in silico approach that predicts if an unknown bacterium belongs to IRRB or ionizing-radiation-sensitive bacteria (IRSB). This approach is based on a multiple instance learning (MIL) formulation of the IRR prediction problem. It takes into account the relation between semantically related instances across bags. In ABClass, a preprocessing step is performed in order to extract substructures/motifs from each set of related sequences. These motifs are then used as attributes to construct a vector representation for each set of sequences. In order to compute partial prediction results, a discriminative classifier is applied to each sequence of the unknown bag and its correspondent related sequences in the learning dataset. Finally, an aggregation method is applied to generate the final result. The algorithm provides good overall accuracy rates. ABClass can be downloaded at the following link: http://homepages.loria.fr/SAridhi/software/MIL/.
Manel Zoghlami, Sabeur Aridhi, Mondher Maddouri, Engelbert Mephu Nguifo
KES4
2019 Corrections to "A Novel Computational Approach for Global Alignment for Multiple Biological Networks"
abstract
Presents corrections to the paper, A novel computational approach for global alignment for multiple biological networks,” (Djeddi, W.E., et al), Trans. Comput. Biol. Bioinf., vol. 15, no. 6, pp. 2060–2066, Nov./Dec. 2018.
Warith Eddine Djeddi, Sadok Ben Yahia, Engelbert Mephu Nguifo
IEEE ACM Trans. Comput. Biol. Bioinform.3
2018 An Approach for Extracting Frequent (Closed) Gradual Patterns Under Temporal Constraint
abstract
Gradual patterns that capture the order correlations of the form "The more/less X, then the more/less Y" play an important role in many real world applications. In this paper, we propose an approach for extracting (closed) frequent gradual patterns when the ordering of supporting objects matches the temporal order. This approach allows to reduce the quantity of mined patterns when the objects follow a temporal order relation. The experimental results obtained on the paleoecological data show the efficiency of our approach and the interpretation of those results bring new knowledge to paleoecological experts.
Jerry Lonlac, Yannick Miras, Aude Beauger, Vincent Mazenod, Jean-Luc Peiry, Engelbert Mephu Nguifo
FUZZ-IEEE6
2018 An experimental survey on big data frameworks
Wissem Inoubli, Sabeur Aridhi, Haithem Mezni, Mondher Maddouri, Engelbert Mephu Nguifo
Future Gener. Comput. Syst.5
2018 A Novel Computational Approach for Global Alignment for Multiple Biological Networks
abstract
Due to the rapid progress of biological networks for modeling biological systems, a lot of biomolecular networks have been producing more and more protein-protein interaction (PPI) data. Analyzing protein-protein interaction networks aims to find regions of topological and functional (dis)similarities between molecular networks of different species. The study of PPI networks has the potential to teach us as much about life process and diseases at the molecular level. Although few methods have been developed for multiple PPI network alignment and thus, new network alignment methods are of a compelling need. In this paper, we propose a novel algorithm for a global alignment of multiple protein-protein interaction networks called MAPPIN. The latter relies on information available for the proteins in the networks, such as sequence, function, and network topology. Our algorithm is perfectly designed to exploit current multi-core CPU architectures, and has been extensively tested on a real data (eight species). Our experimental results show that MAPPIN significantly outperforms NetCoffee in terms of coverage. Nevertheless, MAPPIN is handicapped by the time required to load the gene annotation file. An extensive comparison versus the pioneering PPI methods also show that MAPPIN is often efficient in terms of coverage, mean entropy, or mean normalized.
Warith Eddine Djeddi, Sadok Ben Yahia, Engelbert Mephu Nguifo
IEEE ACM Trans. Comput. Biol. Bioinform.3
2017 On Containment of Triclusters Collections Generated by Quantified Box Operators
Dmitrii Egurnov, Dmitry I. Ignatov, Engelbert Mephu Nguifo
ISMIS3
2017 MR-SimLab: Scalable subgraph selection with label similarity for big data
Wajdi Dhifli, Sabeur Aridhi, Engelbert Mephu Nguifo
Inf. Syst.3
2016 Boolean factors based Artificial Neural Network
abstract
Due to its ability to solve nonlinear problems, Artificial Neural Network (ANN) could be applied in several areas of life. However, defining its architecture for solving a given problem is not formalized and remains an open research problem. On the other hand the complexity of such a technique due to its “black box” aspect, makes its interpretation more tedious. Since optimal factors completely cover the data and therefore give an explanation to these data, we propose in this paper to build feedforward ANNs using the optimal factors obtained from the boolean context representing a data. We show through experiments and comparisons on the use datasets that this approach provides relatively better results than those existing in the literature.
Lauraine Tiogning Kueti, Norbert Tsopzé, Cezar Mbiethieu, Engelbert Mephu Nguifo, Laure Pauline Fotso
IJCNN4
2015 Density-based data partitioning strategy to approximate large-scale subgraph mining
Sabeur Aridhi, Laurent d'Orazio, Mondher Maddouri, Engelbert Mephu Nguifo
Inf. Syst.4
2013 Looking for a structural characterization of the sparseness measure of (frequent closed) itemset contexts
Tarek Hamrouni, Sadok Ben Yahia, Engelbert Mephu Nguifo
Inf. Sci.3
2012 Ranking and Selecting Association Rules Based on Dominance Relationship
abstract
The huge number of association rules represent the main hamper that a decision maker faces. In order to bypass this hamper, an efficient selection of rules has to be performed. Since selection is necessarily based on evaluation, many interestingness measures have been proposed. However, the abundance of these measures gave rise to a new problem, namely the heterogeneity of the evaluation results and this created confusion to the decision. In this respect, we propose a novel approach to discover interesting association rules without favoring or excluding any measure by adopting the notion of dominance between association rules. Our approach bypasses the problem of measure heterogeneity and unveils a compromise between their evaluations. Interestingly enough, the proposed approach also avoids another non-trivial problem which is the threshold value specification.
Slim Bouker, Rabie Saidi, Sadok Ben Yahia, Engelbert Mephu Nguifo
ICTAI4
2012 Multi-paradigm Generation of Tutoring Feedback in Robotic Arm Manipulation Training
Philippe Fournier-Viger, Roger Nkambou, André Mayers, Engelbert Mephu Nguifo, Usef Faghihi
ITS4
2012 CMRules: Mining sequential rules common to several sequences
Philippe Fournier-Viger, Usef Faghihi, Roger Nkambou, Engelbert Mephu Nguifo
Knowl. Based Syst.4
2011 Towards a generalization of decompositional approach of rule extraction from multilayer artificial neural network
abstract
The current development of knowledge discovery domain has pointed out a high number of applications where the need of explanation is at the heart of the process. Using neural networks for those applications requires to be able to provide a set of rules extracted from the trained neural networks, that can help the user to comprehend the learning process. The current literature reports two kinds of rules: `if condition then conclusion' (called if-then) and `if m of conditions then conclusion' (also called MofN). We propose a new method able to extract one intermediate structure (called generators list) from which it is possible to extract both forms of rules. The extracted structure is a generic representation that gives the possibility to the user to visualize each form of rules extracted from the multilayer artificial neural networks.
Norbert Tsopzé, Engelbert Mephu Nguifo, Gilbert Tindo
IJCNN2
2011 Learning task models in ill-defined domain using an hybrid knowledge discovery framework
Roger Nkambou, Philippe Fournier-Viger, Engelbert Mephu Nguifo
Knowl. Based Syst.3
2010 ITS in Ill-Defined Domains: Toward Hybrid Approaches
Philippe Fournier-Viger, Roger Nkambou, Engelbert Mephu Nguifo, André Mayers
Intelligent Tutoring Systems (2)3
2010 Protein sequences classification by means of feature extraction with substitution matrices
abstract
BACKGROUND: This paper deals with the preprocessing of protein sequences for supervised classification. Motif extraction is one way to address that task. It has been largely used to encode biological sequences into feature vectors to enable using well-known machine-learning classifiers which require this format. However, designing a suitable feature space, for a set of proteins, is not a trivial task. For this purpose, we propose a novel encoding method that uses amino-acid substitution matrices to define similarity between motifs during the extraction step. RESULTS: In order to demonstrate the efficiency of such approach, we compare several encoding methods using some machine learning classifiers. The experimental results showed that our encoding method outperforms other ones in terms of classification accuracy and number of generated attributes. We also compared the classifiers in term of accuracy. Results indicated that SVM generally outperforms the other classifiers with any encoding method. We showed that SVM, coupled with our encoding method, can be an efficient protein classification system. In addition, we studied the effect of the substitution matrices variation on the quality of our method and hence on the classification quality. We noticed that our method enables good classification accuracies with all the substitution matrices and that the variances of the obtained accuracies using various substitution matrices are slight. However, the number of generated features varies from a substitution matrix to another. Furthermore, the use of already published datasets allowed us to carry out a comparison with several related works. CONCLUSIONS: The outcomes of our comparative experiments confirm the efficiency of our encoding method to represent protein sequences in classification tasks.
Rabie Saidi, Mondher Maddouri, Engelbert Mephu Nguifo
BMC Bioinform.3
2009 Exploiting Partial Problem Spaces Learned from Users' Interactions to Provide Key Tutoring Services in Procedural and Ill-Defined Domains
abstract
In previous works, we showed how sequential pattern mining can be used to extract a partial problem space from logged user interactions for a procedural and ill-defined domain where classic domain knowledge acquisition approaches don't work well. In this paper, we describe in details how such a problem space can support important tutoring services such as (1) recognizing the plan of a learner, (2) providing hints and (3) estimating the profile of a learner including its expertise level and missing or misunderstandood skills.
Philippe Fournier-Viger, Roger Nkambou, Engelbert Mephu Nguifo
AIED3
2009 OACAS - Ontologies Alignment using Composition and Aggregation of Similarities
Sami Zghal, Marouen Kachroudi, Sadok Ben Yahia, Engelbert Mephu Nguifo
KEOD4
2009 Sweeping the disjunctive search space towards mining new exact concise representations of frequent itemsets
Tarek Hamrouni, Sadok Ben Yahia, Engelbert Mephu Nguifo
Data Knowl. Eng.3
2009 A new generic basis of "factual" and "implicative" association rules
abstract
The extremely large number of association rules that can be drawn from – even reasonably sized datasets, bootstrapped the development of more acute techniques or methods to reduce the size of the reported rule sets. In this context, the battery of re
Sadok Ben Yahia, Ghada Gasmi, Engelbert Mephu Nguifo
Intell. Data Anal.3
2008 M-CLANN: Multi-class Concept Lattice-Based Artificial Neural Network for Supervised Classification
Engelbert Mephu Nguifo, Norbert Tsopzé, Gilbert Tindo
ICANN (2)1
2008 Using Knowledge Discovery Techniques to Support Tutoring in an Ill-Defined Domain
Roger Nkambou, Engelbert Mephu Nguifo, Philippe Fournier-Viger
Intelligent Tutoring Systems2
2007 A Framework for Problem-Solving Knowledge Mining from Users' Actions
Roger Nkambou, Engelbert Mephu Nguifo, Olivier Couturier, Philippe Fournier-Viger
AIED2
2007 Extraction of Association Rules Based on Literalsets
Ghada Gasmi, Sadok Ben Yahia, Engelbert Mephu Nguifo, Slim Bouker
DaWaK3
2007 About the Lossless Reduction of the Minimal Generator Family of a Context
Tarek Hamrouni, Petko Valtchev, Sadok Ben Yahia, Engelbert Mephu Nguifo
ICFCA4
2007 A scalable association rule visualization towards displaying large amounts of knowledge
abstract
Providing efficient and easy-to-use graphical tools to users is a promising challenge of data mining (DM). These tools must be able to generate explicit knowledge and to restitute it. Visualization techniques have shown to be an efficient solution to achieve such goal. Even though considered as a key step in the mining process, the visualization step of association rules received much less attention than that paid to the extraction one. Nevertheless, some graphical tools have been developed to extract and visualize association rules. In those tools, various approaches are proposed to filter the huge number of association rules before the visualization step. However both DM steps (association rule extraction and visualization) are treated separately in a one way process. Our approach differs, and uses meta-knowledge to guide the user during the mining process. Standing at the crossroads of DM and Human-Computer Interaction (HCI), we present an integrated framework covering both steps of the DM process. Furthermore, our approach can easily integrate previous techniques of association rule visualization.
Olivier Couturier, Tarek Hamrouni, Sadok Ben Yahia, Engelbert Mephu Nguifo
IV4
2005 IGB: A New Informative Generic Base of Association Rules
Ghada Gasmi, Sadok Ben Yahia, Engelbert Mephu Nguifo, Yahya Slimani
PAKDD3
2004 Revisiting Generic Bases of Association Rules
Sadok Ben Yahia, Engelbert Mephu Nguifo
DaWaK2
2004 Contextual generic association rules visualization using hierarchical fuzzy meta-rules
abstract
Traditional framework for mining association rules has pointed out the derivation of many redundant rules, in order to be reliable in a decision making process, such discovered rules have to be concise and easily understandable for users or as well as an input to visualization tools. We present a 3D histograms-based visualization prototype for handling generic bases of association rules. An interesting feature of the prototype is that it provides a "contextual" exploration of such rule set. Such additional displayed knowledge, based on the construction of fuzzy meta-rules, enhances man-machine interaction by emulating a cooperative behavior.
Sadok Ben Yahia, Engelbert Mephu Nguifo
FUZZ-IEEE2
2004 A Comparative Study of FCA-Based Supervised Classification Algorithms
Huaiyu Fu, Huaiguo Fu, Patrick Njiwoua, Engelbert Mephu Nguifo
ICFCA4
2004 A Parallel Algorithm to Generate Formal Concepts for Large Data
Huaiguo Fu, Engelbert Mephu Nguifo
ICFCA2
2004 Mining frequent closed itemsets for large data
abstract
Mining frequent closed itemsets is one effective method to analyse frequent pattern, and further, to generate association rules. Several algorithms were proposed to generate frequent closed itemsets, including CLOSE, A-CLOSE, CLOSET, CHARM and CLOSET + etc. However it's still hard for these algorithms to deal with dense and very large data. In this paper, we analyze the search space of frequent closed itemsets and propose a new decomposition algorithm for mining frequent closed itemsets called PFC. PFC can dynamically generate non-overlapping partitions of the search space and mine frequent closed itemsets in each partition. Furthermore, each partition is independent and only shares the same source data with other partitions. So it is possible to implement PFC with multi-threads or parallel methods, and prune efficiently the search space of frequent closed itemsets. In this study, P FC is implemented in Java. We compare PFC with an author's C++ version of CLOSET + on some large VCI repository datasets and on the worst case. The preliminary experimental results demonstrate good performance of PFC for dealing with dense and very large data.
Huaiguo Fu, Engelbert Mephu Nguifo
ICMLA2
2004 Emulating a Cooperative Behavior in a Generic Association Rule Visualization Tool
abstract
Traditional framework for mining association rules has pointed out the derivation of many redundant rules. In order to be reliable in a decision making process, such discovered rules have to be both concise and easily understandable for users, and/or as an input to visualization tools [P. Adriaans et al. (1997)]. We present a graphical visualization prototype for handling generic bases of association rules. We discuss also the most adequate graphical visualization technique depending on the intrinsic structure of the generic bases of association rules. An interesting feature of the prototype is that it provides a "contextual" exploration of such rule set. Such exploration, based on the discovery of fuzzy meta-rules, enhances man-machine interaction by emulating a cooperative behavior.
Sadok Ben Yahia, Engelbert Mephu Nguifo
ICTAI2
2003 Partitioning Large Data to Scale up Lattice-Based Algorithm
abstract
Concept lattice is an effective tool and platform for data analysis and knowledge discovery such as classification or association rules mining. The lattice algorithm to build formal concepts and concept lattice plays an essential role in the application of concept lattice. We propose a new efficient scalable lattice-based algorithm: ScalingNextClosure to decompose the search space of any huge data in some partitions, and then generate independently concepts (or closed itemsets) in each partition. The experimental results show the efficiency of this algorithm.
Huaiguo Fu, Engelbert Mephu Nguifo
ICTAI2
2002 Concept lattice-based knowledge discovery in databases
abstract
(2002). Concept lattice-based knowledge discovery in databases. Journal of Experimental & Theoretical Artificial Intelligence: Vol. 14, No. 2-3, pp. 75-79.
Engelbert Mephu Nguifo, Vincent Duquenne, Michel Liquiere
J. Exp. Theor. Artif. Intell.1
2001 IGLUE: A lattice-based constructive induction system
Engelbert Mephu Nguifo, Patrick Njiwoua
Intell. Data Anal.1
1999 Exemplar-Based Prototype Selection for a Multi-Strategy Learning System
abstract
Multistrategy learning (MSL) consists of combining at least two different learning strategies to bring out a powerful system, where the drawbacks of the basic algorithms are avoided. In this scope, instance-based learning (IBL) techniques are often used as the basic component. However, one of the major drawbacks of IBL is the prototype selection problem which consists in selecting a subset of representative instances in order to reduce the classification process. This paper presents a novel approach which consists of three steps. The first one builds a set of lattice-based hypotheses that characterize the training data set. Given an unseen example, the second step selects a subset of training instances through the way they verify the same hypotheses as the unseen example. Finally the last step uses this subset of training instances as the prototypes for the classification of the unseen example. Results of experiments that we conducted show the effectiveness of our approach compared to standard ML techniques on different datasets.
Patrick Njiwoua, Engelbert Mephu Nguifo
ICTAI2
1998 Using Lattice-Based Framework as a Tool for Feature Extraction
Engelbert Mephu Nguifo, Patrick Njiwoua
ECML1
1997 IGLUE: An Instance-Based Learning System over Lattice Theory
abstract
Concept learning is one of the most studied areas in machine learning. A lot of work in this domain deals with decision trees. In this paper, we are concerned with a different kind of technique based on Galois lattices or concept lattices. We present a new semilattice based system, IGLUE, that uses the entropy function with a tap-down approach to select concepts during the lattice construction. Then IGLUE generates new relevant numerical features by transforming initial boolean features over these concepts. IGLUE uses the new features to redescribe examples. Finally, IGLUE applies the Mahanalobe's distance as a similarity measure between examples.
Patrick Njiwoua, Engelbert Mephu Nguifo
ICTAI2
1994 Galois Lattice: A Framework for Concept Learning-Design, Evaluation and Refinement
abstract
The previously-reported LEGAL system is an empirical machine learning system based on Galois Lattice. Its aim is first to produce a semi-lattice from a concept denoted by a set of objects which are described with binary attributes. Then using some selected attribute conjunctions in the semi-lattice and a majority vote principle, LEGAL predicts new examples from unseen objects. This paper describes a new version LEGAL-E and its application to two biological problems: the prediction of splice junctions sites and the promoter recognition. Results obtained are far better than those of some symbolic learning systems, and are as better as those of some best neural networks methods. Moreover some empirical properties shared by LEGAL-E and neural networks are described. Finally this paper shows how the semi-lattice can be used as a dynamic neural network architecture in order to combine both learning techniques for knowledge refinement.>
Engelbert Mephu Nguifo
ICTAI1
1993 Prediction of Primate Splice Junction Gene Sequences with a Cooperative Knowledge Acquisition System
Engelbert Mephu Nguifo, Jean Sallantin
ISMB1