Mehdi Kaytoue-Uberall

dblp:00/6831 · also Mehdi Kaytoue · DBLP profile ↗
← Back
38ranked-venue papers
8as first author
7since 2021 · last 2024
0000-0002-1569-5242ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 5 first-author · 2 since 2021Databases, data management, data science and information retrieval · 16 · 4 first-author · 3 since 2021Theory of computation · 14 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-authorSoftware engineering, systems software and programming languages · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2024 DeepLSH: Deep Locality-Sensitive Hash Learning for Fast and Efficient Near-Duplicate Crash Report Detection
abstract
Automatic crash bucketing is a crucial phase in the software development process for efficiently triaging bug reports. It generally consists in grouping similar reports through clustering techniques. However, with real-time streaming bug collection, systems are needed to quickly answer the question: What are the most similar bugs to a new one?, that is, efficiently find near-duplicates. It is thus natural to consider nearest neighbors search to tackle this problem and especially the well-known locality-sensitive hashing (LSH) to deal with large datasets due to its sublinear performance and theoretical guarantees on the similarity search accuracy. Surprisingly, LSH has not been considered in the crash bucketing literature. It is indeed not trivial to derive hash functions that satisfy the so-called locality-sensitive property for the most advanced crash bucketing metrics. Consequently, we study in this paper how to leverage LSH for this task. To be able to consider the most relevant metrics used in the literature, we introduce DeepLSH, a Siamese DNN architecture with an original loss function, that perfectly approximates the locality-sensitivity property even for Jaccard and Cosine metrics for which exact LSH solutions exist. We support this claim with a series of experiments on an original dataset, which we make available.
Youcef Remil, Ahmed Anes Bendimerad, Romain Mathonat, Chedy Raïssi, Mehdi Kaytoue-Uberall
ICSE5
2023 Three Views on Dependency Covers from an FCA Perspective
Jaume Baixeries, Víctor Codocedo, Mehdi Kaytoue-Uberall, Amedeo Napoli
ICFCA3
2023 On-Premise AIOps Infrastructure for a Software Editor SME: An Experience Report
abstract
Information Technology has become a critical component in various industries, leading to an increased focus on software maintenance and monitoring. With the complexities of modern software systems, traditional maintenance approaches have become insufficient. The concept of AIOps has emerged to enhance predictive maintenance using Big Data and Machine Learning capabilities. However, exploiting AIOps requires addressing several challenges related to the complexity of data and incident management. Commercial solutions exist, but they may not be suitable for certain companies due to high costs, data governance issues, and limitations in covering private software. This paper investigates the feasibility of implementing on-premise AIOps solutions by leveraging open-source tools. We introduce a comprehensive AIOps infrastructure that we have successfully deployed in our company, and we provide the rationale behind different choices that we made to build its various components. Particularly, we provide insights into our approach and criteria for selecting a data management system and we explain its integration. Our experience can be beneficial for companies seeking to internally manage their software maintenance processes with a modern AIOps approach.
Ahmed Anes Bendimerad, Youcef Remil, Romain Mathonat, Mehdi Kaytoue-Uberall
ESEC/SIGSOFT FSE4
2021 Anytime Subgroup Discovery in High Dimensional Numerical Data
abstract
Subgroup discovery (SD) enables one to elicit patterns that strongly discriminate a class label. When it comes to numerical data, most of the existing SD approaches perform data discretizations and thus suffer from information loss. A few algorithms avoid such a loss by considering the search space of every interval pattern built on the dataset numerical values and provide an “anytime” property: at any moment, they are able to provide a result that improves over time. Given a sufficient time/memory budget, they may eventually complete an exhaustive search. However, such approaches are often intractable when dealing with high-dimensional numerical data, for instance, when extracting features from real-life multivariate time series. To overcome such limitations, we propose MonteCloPi, an approach based on a bottom-up exploration of numerical patterns with a Monte Carlo Tree Search. It enables to have a better exploration-exploitation trade-off between exploration and exploitation when sampling huge search spaces. Our extensive set of experiments proves the efficiency of MonteCloPi on high-dimensional data with hundreds of attributes. We finally discuss the actionability of discovered subgroups when looking for skill analysis from Rocket League action logs.
Romain Mathonat, Diana Nurbakova, Jean-François Boulicaut, Mehdi Kaytoue-Uberall
DSAA4
2021 Interpretable Summaries of Black Box Incident Triaging with Subgroup Discovery
abstract
The need of predictive maintenance comes with an increasing number of incidents reported by monitoring systems and equipment/software users. In the front line, on-call engineers (OCEs) have to quickly assess the degree of severity of an incident and decide which service to contact for corrective actions. To automate these decisions, several predictive models have been proposed, but the most efficient models are opaque (say, black box), strongly limiting their adoption. In this paper, we propose an efficient black box model based on 170K incidents reported to our company over the last 7 years and emphasize on the need of automating triage when incidents are massively reported on thousands of servers running our product, an ERP. Recent developments in eXplainable Artificial Intelligence (XAI) help in providing global explanations to the model, but also, and most importantly, with local explanations for each model prediction/outcome. Sadly, providing a human with an explanation for each outcome is not conceivable when dealing with an important number of daily predictions. To address this problem, we propose an original data-mining method rooted in Subgroup Discovery, a pattern mining technique with the natural ability to group objects that share similar explanations of their black box predictions and provide a description for each group. We evaluate this approach and present our preliminary results which give us good hope towards an effective OCE's adoption. We believe that this approach provides a new way to address the problem of model agnostic outcome explanation.
Youcef Remil, Ahmed Anes Bendimerad, Marc Plantevit, Céline Robardet, Mehdi Kaytoue-Uberall
DSAA5
2021 "What makes my queries slow?": Subgroup Discovery for SQL Workload Analysis
abstract
Among daily tasks of database administrators (DBAs), the analysis of query workloads to identify schema issues and improving performances is crucial. Although DBAs can easily pinpoint queries repeatedly causing performance issues, it remains challenging to automatically identify subsets of queries that share some properties only (a pattern) and simultaneously foster some target measures, such as execution time. Patterns are defined on combinations of query clauses, environment variables, database alerts and metrics and help answer questions like what makes SQL queries slow? What makes I/O communications high? Automatically discovering these patterns in a huge search space and providing them as hypotheses for helping to localize issues and root-causes is important in the context of explainable AI. To tackle it, we introduce an original approach rooted on Subgroup Discovery. We show how to instantiate and develop this generic data-mining framework to identify potential causes of SQL workloads issues. We believe that such data-mining technique is not trivial to apply for DBAs. As such, we also provide a visualization tool for interactive knowledge discovery. We analyse a one week workload from hundreds of databases from our company, make both the dataset and source code available, and experimentally show that insightful hypotheses can be discovered.
Youcef Remil, Ahmed Anes Bendimerad, Romain Mathonat, Philippe Chaleat, Mehdi Kaytoue-Uberall
ASE5
2021 Anytime mining of sequential discriminative patterns in labeled sequences
Romain Mathonat, Diana Nurbakova, Jean-François Boulicaut, Mehdi Kaytoue-Uberall
Knowl. Inf. Syst.4
2020 A Behavioral Pattern Mining Approach to Model Player Skills in Rocket League
abstract
Competitive gaming, or esports, is now well-established and brought the game industry in a novel era. It comes with many challenges among which evaluating the level of a player, given the strategies and skills she masters. We are interested in automatically identifying the so called skillshots from game traces of Rocket League, a "soccer with rocket-powered cars" game. From a pure data point of view, each skill execution is unique and standard pattern matching may be insufficient. We propose a non trivial data-centric approach based on pattern mining and supervised learning techniques. We show through an extensive set of experiments that most of Rocket League skillshots can be efficiently detected and used for player modelling. It unveils applications for match making, supporting game commentators and learning systems among others.
Romain Mathonat, Jean-François Boulicaut, Mehdi Kaytoue-Uberall
CoG3
2019 FSSD - A Fast and Efficient Algorithm for Subgroup Set Discovery
abstract
Subgroup discovery (SD) is the task of discovering interpretable patterns in the data that stand out w.r.t. some property of interest. Discovering patterns that accurately discriminate a class from the others is one of the most common SD tasks. Standard approaches of the literature are based on local pattern discovery, which is known to provide an overwhelmingly large number of redundant patterns. To solve this issue, pattern set mining has been proposed: instead of evaluating the quality of patterns separately, one should consider the quality of a pattern set as a whole. The goal is to provide a small pattern set that is diverse and well-discriminant to~the target class. In this work, we introduce a novel formulation of the task of diverse subgroup set discovery where both discriminative power and diversity of the subgroup set are incorporated in the same quality measure. We propose an efficient and parameter-free algorithm dubbed FSSD and based on a greedy scheme. FSSD uses several optimization strategies that enable to efficiently provide a high quality pattern set in a short amount of time.
Adnene Belfodil, Aimene Belfodil, Ahmed Anes Bendimerad, Philippe Lamarre, Céline Robardet, Mehdi Kaytoue-Uberall, Marc Plantevit
DSAA6
2019 SeqScout: Using a Bandit Model to Discover Interesting Subgroups in Labeled Sequences
abstract
It is extremely useful to exploit labeled datasets not only to learn models but also to improve our understanding of a domain and its available targeted classes. The so-called subgroup discovery task has been considered for a long time. It concerns the discovery of patterns or descriptions, the set of supporting objects of which have interesting properties, e.g., they characterize or discriminate a given target class. Though many subgroup discovery algorithms have been proposed for transactional data, discovering subgroups within labeled sequential data and thus searching for descriptions as sequential patterns has been much less studied. In that context, exhaustive exploration strategies can not be used for real-life applications and we have to look for heuristic approaches. We propose the algorithm SeqScout to discover interesting subgroups (w.r.t. a chosen quality measure) from labeled sequences of itemsets. This is a new sampling algorithm that mines discriminant sequential patterns using a multi-armed bandit model. It is an anytime algorithm that, for a given budget, finds a collection of local optima in the search space of descriptions and thus subgroups. It requires a light configuration and it is independent from the quality measure used for pattern scoring. Furthermore, it is fairly simple to implement. We provide qualitative and quantitative experiments on several datasets to illustrate its added-value.
Romain Mathonat, Diana Nurbakova, Jean-François Boulicaut, Mehdi Kaytoue-Uberall
DSAA4
2019 Mining Formal Concepts Using Implications Between Items
Aimene Belfodil, Adnene Belfodil, Mehdi Kaytoue-Uberall
ICFCA3
2019 Sampling Representation Contexts with Attribute Exploration
Víctor Codocedo, Jaume Baixeries, Mehdi Kaytoue-Uberall, Amedeo Napoli
ICFCA3
2019 Chemical features mining provides new descriptive structure-odor relationships
abstract
An important goal in researching the biology of olfaction is to link the perception of smells to the chemistry of odorants. In other words, why do some odorants smell like fruits and others like flowers? While the so-called stimulus-percept issue was resolved in the field of color vision some time ago, the relationship between the chemistry and psycho-biology of odors remains unclear up to the present day. Although a series of investigations have demonstrated that this relationship exists, the descriptive and explicative aspects of the proposed models that are currently in use require greater sophistication. One reason for this is that the algorithms of current models do not consistently consider the possibility that multiple chemical rules can describe a single quality despite the fact that this is the case in reality, whereby two very different molecules can evoke a similar odor. Moreover, the available datasets are often large and heterogeneous, thus rendering the generation of multiple rules without any use of a computational approach overly complex. We considered these two issues in the present paper. First, we built a new database containing 1689 odorants characterized by physicochemical properties and olfactory qualities. Second, we developed a computational method based on a subgroup discovery algorithm that discriminated perceptual qualities of smells on the basis of physicochemical properties. Third, we ran a series of experiments on 74 distinct olfactory qualities and showed that the generation and validation of rules linking chemistry to odor perception was possible. Taken together, our findings provide significant new insights into the relationship between stimulus and percept in olfaction. In addition, by automatically extracting new knowledge linking chemistry of odorants and psychology of smells, our results provide a new computational framework of analysis enabling scientists in the field to test original hypotheses using descriptive or predictive modeling.
Carmen C. Licon, Guillaume Bosc, Mohammed Sabri, Marylou Mantel, Arnaud P. Fournel, Caroline Bushdid, Jérôme Golebiowski, Céline Robardet, Marc Plantevit, Mehdi Kaytoue-Uberall, Moustafa Bensafi
PLoS Comput. Biol.10
2018 Anytime Subgroup Discovery in Numerical Domains with Guarantees
Aimene Belfodil, Adnene Belfodil, Mehdi Kaytoue-Uberall
ECML/PKDD (2)3
2018 Characterizing approximate-matching dependencies in formal concept analysis with pattern structures
Jaume Baixeries, Víctor Codocedo, Mehdi Kaytoue-Uberall, Amedeo Napoli
Discret. Appl. Math.3
2018 Anytime discovery of a diverse set of patterns with Monte Carlo tree search
Guillaume Bosc, Jean-François Boulicaut, Chedy Raïssi, Mehdi Kaytoue-Uberall
Data Min. Knowl. Discov.4
2017 A Proposition for Sequence Mining Using Pattern Structures
Víctor Codocedo, Guillaume Bosc, Mehdi Kaytoue-Uberall, Jean-François Boulicaut, Amedeo Napoli
ICFCA3
2017 Mining the Lattice of Binary Classifiers for Identifying Duplicate Labels in Behavioral Data
Quentin Labernia, Víctor Codocedo, Céline Robardet, Mehdi Kaytoue-Uberall
IEA/AIE (2)4
2017 Mining Convex Polygon Patterns with Formal Concept Analysis
abstract
Pattern mining is an important task in AI for eliciting hypotheses from the data. When it comes to spatial data, the geo-coordinates are often considered independently as two different attributes. Consequently, rectangular patterns are searched for. Such an arbitrary form is not able to capture interesting regions in general. We thus introduce convex polygons, a good trade-off for capturing high density areas in any pattern mining task. Our contribution is threefold: (i) We formally introduce such patterns in Formal Concept Analysis (FCA), (ii) we give all the basic bricks for mining polygons with exhaustive search and pattern sampling, and (iii) we design several algorithms that we compare experimentally.
Aimene Belfodil, Sergei O. Kuznetsov, Céline Robardet, Mehdi Kaytoue-Uberall
IJCAI4
2017 Exceptional contextual subgraph mining
Mehdi Kaytoue-Uberall, Marc Plantevit, Albrecht Zimmermann, Ahmed Anes Bendimerad, Céline Robardet
Mach. Learn.1
2017 A Pattern Mining Approach to Study Strategy Balance in RTS Games
abstract
Whereas purest strategic games such as Go and Chess seem timeless, the lifetime of a video game is short, influenced by popular culture, trends, boredom, and technological innovations. Even the important budget and developments allocated by editors cannot guarantee a timeless success. Instead, novelties and corrections are proposed to extend an inevitably bounded lifetime. Novelties can unexpectedly break the balance of a game, as players can discover unbalanced strategies that developers did not take into account. In the new context of electronic sports, an important challenge is to be able to detect game balance issues. In this paper, we consider real-time strategy (RTS) games and present an efficient pattern mining algorithm as a basic tool for game balance designers that enables one to search for unbalanced strategies in historical data through a knowledge discovery in databases (KDD) process. We experiment with our algorithm on StarCraft II historical data, played professionally as an electronic sport.
Guillaume Bosc, Philip Tan, Jean-François Boulicaut, Chedy Raïssi, Mehdi Kaytoue-Uberall
IEEE Trans. Comput. Intell. AI Games5
2016 Local Subgroup Discovery for Eliciting and Understanding New Structure-Odor Relationships
Guillaume Bosc, Jérôme Golebiowski, Moustafa Bensafi, Céline Robardet, Marc Plantevit, Jean-François Boulicaut, Mehdi Kaytoue-Uberall
DS7
2016 What Did I Do Wrong in My MOBA Game? Mining Patterns Discriminating Deviant Behaviours
abstract
The success of electronic sports (eSports), where professional gamers participate in competitive leagues and tournaments, brings new challenges for the video game industry. Other than fun, games must be difficult and challenging for eSports professionals but still easy and enjoyable for amateurs. In this article, we consider Multi-player Online Battle Arena games (MOBA) and particularly, "Defense of the Ancients 2", commonly known simply as DOTA2. In this context, a challenge is to propose data analysis methods and metrics that help players to improve their skills. We design a data mining-based method that discovers strategic patterns from historical behavioral traces: Given a model encoding an expected way of playing (the norm), we are interested in patterns deviating from the norm that may explain a game outcome from which player can learn more efficient ways of playing. The method is formally introduced and shown to be adaptable to different scenarios. Finally, we provide an experimental evaluation over a dataset of 10 000 behavioral game traces.
Olivier Cavadenti, Víctor Codocedo, Jean-François Boulicaut, Mehdi Kaytoue-Uberall
DSAA4
2016 h(odor): Interactive Discovery of Hypotheses on the Structure-Odor Relationship in Neuroscience
Guillaume Bosc, Marc Plantevit, Jean-François Boulicaut, Moustafa Bensafi, Mehdi Kaytoue-Uberall
ECML/PKDD (3)5
2015 When cyberathletes conceal their game: Clustering confusion matrices to identify avatar aliases
abstract
Video game is a very lucrative industry, unleashed by the ubiquity of gaming devices, multi-player networks and live broadcasting platforms. Games generate large amounts of behavioural data which are valuable to face the new challenges of video game analytics such as detecting balance issues, bugs and cheaters. In electronic sports (e-sports), cyberathletes conceal their online training using different aliases or avatars (virtual identities), which allow them not being recognized by the opponents they may face in future competitions (with cash prices challenging already most of the traditional sports). It was recently suggested that behavioural data generated by the games allows predicting the avatar associated to a game play with high accuracy. However, when a player uses several avatars, accuracy drastically drops as prediction models cannot easily differentiate the player's different avatar aliases. Since mappings between players and avatars do not exist, we introduce the avatar aliases identification problem and propose an original approach for alias resolution based on supervised classification and Formal Concept Analysis. We thoroughly evaluate our method with the video game Starcraft 2 which has a very wide and active community with players from diverse cultures and nations. We show that under some circumstances, the avatars of a given player can easily be recognized as such. These results are valuable for e-sport structures (to help preparing tournaments), and game editors (detecting cheaters or usurpers).
Olivier Cavadenti, Víctor Codocedo, Jean-François Boulicaut, Mehdi Kaytoue-Uberall
DSAA4
2015 Gazouille: Detecting and Illustrating Local Events from Geolocalized Social Media Streams
Pierre Houdyer, Albrecht Zimmermann, Mehdi Kaytoue-Uberall, Marc Plantevit, Céline Robardet
ECML/PKDD (3)3
2015 Pattern Structures and Concept Lattices for Data Mining and Knowledge Processing
Mehdi Kaytoue-Uberall, Víctor Codocedo, Aleksey Buzmakov 0002, Jaume Baixeries, Sergei O. Kuznetsov, Amedeo Napoli
ECML/PKDD (3)1
2014 Triggering patterns of topology changes in dynamic graphs
abstract
To describe the dynamics taking place in networks that structurally change over time, we propose an approach to search for attributes whose value changes impact the topology of the graph. In several applications, it appears that the variations of a group of attributes are often followed by some structural changes in the graph that one may assume they generate. We formalize the triggering pattern discovery problem as a method jointly rooted in sequence mining and graph analysis. We apply our approach on three real-world dynamic graphs of different natures - a co-authoring network, an airline network, and a social bookmarking system - assessing the relevancy of the triggering pattern mining approach.
Mehdi Kaytoue-Uberall, Yoann Pitarch, Marc Plantevit, Céline Robardet
ASONAM1
2014 Mining Balanced Sequential Patterns in RTS Games
abstract
The video game industry has grown enormously over the last twenty years, bringing new challenges to the artificial intelligence and data analysis communities. We tackle here the problem of automatic discovery of strategies in real-time strategy games through pattern mining. Such patterns are the basic units for many tasks such as automated agent design, but also to build tools for the professionally played video games in the electronic sports scene. Our formalization relies on a sequential pattern mining approach and a novel measure, the balance measure, telling how a strategy is likely to win. We experiment our methodology on a real-time strategy game that is professionally played in the electronic sport community.
Guillaume Bosc, Mehdi Kaytoue-Uberall, Chedy Raïssi, Jean-François Boulicaut, Philip Tan
ECAI2
2013 Mining Statistically Significant Sequential Patterns
abstract
Recent developments in the frequent pattern mining framework uses additional measures of interest to reduce the set of discovered patterns. We introduce a rigorous and efficient approach to mine statistically significant, unexpected patterns in sequences of item sets. The proposed methodology is based on a null model for sequences and on a multiple testing procedure to extract patterns of interest. Experiments on sequences of replays of a video game demonstrate the scalability and the efficiency of the method to discover unexpected game strategies.
Cécile Low-Kam, Chedy Raïssi, Mehdi Kaytoue-Uberall, Jian Pei 0001
ICDM3
2013 Using Pattern Structures for Analyzing Ontology-Based Annotations of Biomedical Data
Adrien Coulet, Florent Domenach, Mehdi Kaytoue-Uberall, Amedeo Napoli
ICFCA3
2011 Biclustering Numerical Data in Formal Concept Analysis
Mehdi Kaytoue-Uberall, Sergei O. Kuznetsov, Amedeo Napoli
ICFCA1
2011 Numerical Information Fusion: Lattice of Answers with Supporting Arguments
abstract
The problem addressed in this paper is the merging of numerical information provided by several sources. Merging conflicting pieces of information into an interpretable and useful format is a tricky task even when an information fusion method is chosen. The use of formal concept analysis and pattern structures enables us to associate subsets of sources to combination results obtainable from consistent subsets of pieces of information. This provides a lattice of arguments where the reliability of sources can be taken into account. Instead of providing a unique fusion result, the method yields a structured view of partial results labelled by subsets of sources and allows us to argue about the most appropriate evaluation. The approach is illustrated with an experiment on a real-world application to decision aid in agricultural practices.
Zainab Assaghir, Amedeo Napoli, Mehdi Kaytoue-Uberall, Didier Dubois, Henri Prade
ICTAI3
2011 Revisiting Numerical Pattern Mining with Formal Concept Analysis
abstract
HAL is a multi-disciplinary open access archive for the deposit and dissemination of sci-entific research documents, whether they are pub-lished or not. The documents may come from teaching and research institutions in France or abroad, or from public or private research centers. L’archive ouverte pluridisciplinaire HAL, est destinée au dépôt et a ̀ la diffusion de documents scientifiques de niveau recherche, publiés ou non, émanant des établissements d’enseignement et de recherche français ou étrangers, des laboratoires publics ou privés.
Mehdi Kaytoue-Uberall, Sergei O. Kuznetsov, Amedeo Napoli
IJCAI1
2011 Mining gene expression data with pattern structures in formal concept analysis
Mehdi Kaytoue-Uberall, Sergei O. Kuznetsov, Amedeo Napoli, Sébastien Duplessis
Inf. Sci.1
2010 Embedding tolerance relations in formal concept analysis: an application in information fusion
abstract
This paper shows how to embed a similarity relation between complex descriptions in concept lattices. We formalize similarity by a tolerance relation: objects are grouped within a same concept when having similar descriptions, extending the ability of FCA to deal with complex data. We propose two different approaches.~A first classical manner defines a discretization procedure. A second way consists in representing data by pattern structures, from which a pattern concept lattice can be constructed directly. In this case, considering a tolerance relation can be mathematically defined by a projection in a meet-semi-lattice. This allows to use concept lattices for their knowledge representation and reasoning abilities without transforming data. We show finally that resulting lattices are useful for solving information fusion problems.
Mehdi Kaytoue-Uberall, Zainab Assaghir, Amedeo Napoli, Sergei O. Kuznetsov
CIKM1
2010 Managing Information Fusion with Formal Concept Analysis
Zainab Assaghir, Mehdi Kaytoue-Uberall, Amedeo Napoli, Henri Prade
MDAI2
2009 Two FCA-Based Methods for Mining Gene Expression Data
Mehdi Kaytoue-Uberall, Sébastien Duplessis, Sergei O. Kuznetsov, Amedeo Napoli
ICFCA1