David Ing

dblp:01/11025 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 2 since 2021Theory of computation · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 On Integrating Logical Analysis of Data into Random Forests
abstract
Random Forests (RFs) are one of the most popular classifiers in machine learning. RF is an ensemble learning method that combines multiple Decision Trees (DTs), providing a more robust and accurate model than a single DT. However, one of the main step of RFs is the random selection of many different features during the construction phase of DTs, resulting in a forest with various features, which makes it difficult to extract short and concise explanations. In this paper, we propose integrating Logical Analysis of Data (LAD) into RFs. LAD is a pattern learning framework that combines optimization, Boolean functions, and combinatorial theory. One of its main goals is to generate minimal support sets (MSSes) that discriminate between different groups of data. More precisely, we show how to enhance the classical RF algorithm by randomly choosing MSSes rather than randomly choosing feature subsets that potentially contain irrelevant features for constructing DTs. Experiments on benchmark datasets reveal that integrating LAD into classical RFs using MSSes can maintain similar performance in terms of accuracy, produce forests of similar size, reduce the set of used features, and enable the extraction of significantly shorter explanations compared to classical RFs.
David Ing, Saïd Jabbour, Lakhdar Sais
IJCAI1
2025 Text Mining from Migration Narratives
David Ing, Fabien Delorme, Saïd Jabbour, Nelly Robin, Lakhdar Sais
ECML/PKDD (8)1
2024 LAD-based Feature Selection for Optimal Decision Trees and Other Classifiers
abstract
The curse of dimensionality presents a significant challenge in data mining, pattern recognition, computer vision, and machine learning applications. Feature selection is a primary approach to address this challenge. It aims to eliminate irrelevant and redundant features while preserving the relevant ones to reduce computation time, improve prediction performance, and enhance the understanding of data. In this paper, we introduce a new feature selection (FS) technique based on the Logical Analysis of Data (LAD), a pattern learning framework that combines optimization, Boolean functions, and combinatorial theory. One of its main objectives is to generate minimal support sets of features (subsets of features) that discriminate between different groups of data. To generate such subsets, we first reduce the complexity of the LAD optimization task by transforming it into the problem of enumerating minimal hitting sets in a hypergraph, for which efficient implementations exist. Those feature subsets are then ranked based on a scoring method before selecting the highest quality one. Moreover, we explore the relationship between optimal Decision Trees (DTs) and LAD-based FS, introducing new optimality criteria, namely DTs involving a minimum number of features. Finally, we conduct comparative evaluations of LAD-based approach against several state-of-the-art (SOTA) FS methods on benchmark datasets, including two-class binary datasets and numerical datasets with two and multiple classes. Experiments reveal that our approach is competitive with SOTA methods, selecting high-quality feature subsets that maintain or enhance the performance of DTs and other classifiers like SVM, KNN, and Naive Bayes.
David Ing, Saïd Jabbour, Lakhdar Sais, Fabien Delorme
KR1
2023 Classification with Explanation for Human Trafficking Networks
abstract
On a worldwide scale, an increasing number of victims of human trafficking were observed these last years, covering a majority of countries and territories. Among them, a large portion of women and girls are recruited primarily for sexual exploitation. United Nations Office on Drugs and Crime (UNODC) highlights the difficulties of access to justice which deprive victims of protection, a central issue behind our work. Our contribution is part of an emerging research trend, combining Artificial Intelligence (AI), Humanities and Social Sciences (HSS). It makes an original use of legal database to identify Human Trafficking Networks (HTNs), involving both sexual abuse victims and exploiters. First, a reformulation of the legal database as a numerical database is proposed, using new features expressing relationships between people involved in the same court case, likely to better reveal HTNs. Secondly, six machine learning algorithms, including Decision Tree, Random Forest, Gradient Boosting, Logistic Regression, Support Vector Machine (SVM) and K-Nearest Neighbors (KNN) are used to train on numerical database and learn to classify the input court case into one of the three classes: Not suspicious, Suspicious, or Probably suspicious. We in details discuss knowledge-based feature engineering, dataset balancing, parameters tuning, and best models selection. The comparative empirical evaluations between those classification algorithms have been conducted in order to highlights the relevance of our HTNs detection approach. To help the end-users, to better understand the displayed HTNs, for Decision Tree and Random Forest, we also provide explanations of why such court case can be classified. Those results were finally discussed with experts in the field of human trafficking, providing us with interesting feedback shedding light to this multidimensional form of modern-day slavery problem.
David Ing, Fabien Delorme, Saïd Jabbour, Nelly Robin, Lakhdar Sais
DSAA1
2012 Declarative web application development: encapsulating dynamic JavaScript widgets (abstract only)
abstract
The development of modern, highly interactive AJAX Web applications that enable dynamic visualization of data requires writing a great deal of tedious "plumbing code" to interface data between browser-based DOM and AJAX components, the application server, and the SQL database. Worse, each of these layers utilizes a different language. Further, much code is needed to keep the page and application states in sync using an imperative paradigm, which hurts simplicity. These factors result in a frustrating experience for today's Web developer. The FORWARD Project aims to alleviate this frustration by enabling pages that are "rendered views", in the SQL sense of "view". Our work in the project has led to a highly declarative approach whereby JavaScript/AJAX UI widgets automatically render views over the application state (database + session data + page data) without requiring the developer to tediously code how changes to the application state lead to invocation of the components' update methods.
Robert Bolton, David Ing, Christopher Rebert, Kristina Lam Thai
SIGMOD Conference2