Thi-Bich-Hanh Dao

dblp:00/1932 · DBLP profile ↗
← Back
29ranked-venue papers
9as first author
12since 2021 · last 2026
0000-0002-2740-6954ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 8 first-author · 12 since 2021Software engineering, systems software and programming languages · 7 · 3 first-author · 1 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-authorTheory of computation · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Heterogeneous Pattern Sampling According to Frequency
Rayane Lachache, Djawad Bekkoucha, Abdelkader Ouali, Bruno Crémilleux, Thi-Bich-Hanh Dao, Christel Vrain
IDA5
2026 WISTERIA: Weak Implicit Signal-based Temporal Relation Extraction with Attention
Duy Dao Do, Anaïs Lefeuvre-Halftermeyer, Thi-Bich-Hanh Dao
LREC3
2026 Enhancing concept-based image classification via knowledge reasoning
abstract
Deep learning has made great progress in supervised image classification. However, a persistent challenge in this field is the need for better explainability. This is crucial for building trust, troubleshooting, and ensuring regulatory compliance, ultimately leading to more responsible and effective AI applications. To tackle this challenge, we have developed a deep learning framework based on concepts, a recent method aimed at enhancing explainability, albeit with less emphasis on image classification performance. Our innovation lies in integrating a specific knowledge graph as a reasoning tool. This approach leverages the interdependence between concepts and classes to improve the performance of the concept-based model, achieving a balance between explainability and efficiency. Results from both synthetic and real data validate the effectiveness of our approach in balancing efficiency and explainability. Additionally, our model supports human test-time intervention to update its final prediction after incorporating new expert feedback. Our experiments show significant improvements in both classification and concept efficiencies.
Franck Anaël Mbiaya, Frédéric Ros, Christel Vrain, Thi-Bich-Hanh Dao, Yves Lucas
Knowl. Based Syst.4
2025 A Constrained Declarative Based Approach for Explainable Clustering
Mathieu Guilbert, Christel Vrain, Thi-Bich-Hanh Dao
IDA3
2024 Rule-Based Constraint Elicitation For Active Constraint-Incremental Clustering
abstract
Constrained clustering algorithms integrate user knowledge as constraints in the clustering process to guide it towards a desired outcome. When interacting with users, it is essential to quickly ask simple questions to identify informative constraints that will efficiently enhance an initial partition. We propose a new active query strategy for incremental clustering that translates user feedback into interpretable decision rules and identifies relevant points for queries using rule-based heuristics Experiments on benchmark datasets highlight the benefits of our new approach, making it suitable for real-world applications.
Aymeric Beauchamp, Thi-Bich-Hanh Dao, Samir Loudni, Christel Vrain
ICTAI2
2024 Knowledge graph-based image classification
abstract
This paper introduces a deep learning method for image classification that leverages knowledge formalised as a graph created from information represented by pairs attribute/value. The proposed method investigates a loss function that adaptively combines the classical cross-entropy commonly used in deep learning with a novel penalty function. The novel loss function is derived from the representation of nodes after embedding the knowledge graph and incorporates the proximity between class and image nodes. Its formulation enables the model to focus on identifying the boundary between the most challenging classes to distinguish. Experimental results on several image databases demonstrate improved performance compared to state-of-the-art methods, including classical deep learning algorithms and recent algorithms that incorporate knowledge represented by a graph.
Franck Anaël Mbiaya, Christel Vrain, Frédéric Ros, Thi-Bich-Hanh Dao, Yves Lucas
Data Knowl. Eng.4
2024 A review on declarative approaches for constrained clustering
abstract
Clustering is an important Machine Learning task, which aims at discovering the implicit structure of data. Applying a clustering algorithm is easy but since clustering is an unsupervised task, tuning it so that the results is appropriate to the expert expectations is much less obvious. To overcome this, expert knowledge can be integrated into a clustering process; this is generally formalized as constraints on the desired output, thus leading to constrained clustering. There are two lines of research for clustering: distance based clustering, where data are grouped into clusters according to their dissimilarity and conceptual clustering, where a cluster must be a concept that is a set of objects and a set of properties that describe them. This second approach relies on Formal Concept Analysis and benefits from advances in Pattern Mining. [69] has shown the interest of declarative approaches for pattern mining and has led to a new research direction for clustering that is interested in the use of declarative frameworks, such as Integer Linear Programming, Constraint Programming or SAT for clustering. This has several advantages: finding a global optimum, integrating different kinds of constraints, even complex ones in a clustering process and even combining conceptual and distance-based clustering. In this paper we present an inventory of constraints and a survey of declarative methods for constrained clustering.
Thi-Bich-Hanh Dao, Christel Vrain
Int. J. Approx. Reason.1
2023 Incremental Constrained Clustering by Minimal Weighted Modification
abstract
Clustering is a well-known task in Data Mining that aims at grouping data instances according to their similarity. It is an exploratory and unsupervised task whose results depend on many parameters, often requiring the expert to iterate several times before satisfaction. Constrained clustering has been introduced for better modeling the expectations of the expert. Nevertheless constrained clustering is not yet sufficient since it usually requires the constraints to be given before the clustering process. In this paper we address a more general problem that aims at modeling the exploratory clustering process, through a sequence of clustering modifications where expert constraints are added on the fly. We present an incremental constrained clustering framework integrating active query strategies and a Constraint Programming model to fit the expert expectations while preserving the stability of the partition, so that the expert can understand the process and apprehend its impact. Our model supports instance and group-level constraints, which can be relaxed. Experiments on reference datasets and a case study related to the analysis of satellite image time series show the relevance of our framework.
Aymeric Beauchamp, Thi-Bich-Hanh Dao, Samir Loudni, Christel Vrain
CP2
2023 Identifying relevant descriptors for tweet sets
abstract
Twitter is a media where information flows in vast volumes. Even considering a particular topic, the collected tweets come in a wide variety of forms. This makes the data description problem complex. In this paper, given a set of tweets, we consider the problem of producing a set of relevant descriptors to characterize the tweets. Messages and words are considered in an embedding space learned with Doc2Vec, a model well suited for short documents. We propose to leverage this model to identify text units that may span over several words. We then propose to measure the impact of the units on their document representation through ablation. The most important units are selected as descriptors. Experiments have been conducted in the context of a tweet cluster description problem on two datasets. One is about the storm Alex, which struck France in October 2020, and the other about the beginning of the Russia-Ukraine war in February 2022. The results show the interest of our method compared to existing work.
Olivier Gracianne, Anaïs Lefeuvre-Halftermeyer, Thi-Bich-Hanh Dao
ICTAI3
2022 Presenting an event through the description of related tweets clusters
abstract
Twitter data are a mirror to real world events and information stakeholders willing to benefit from this rich resource are numerous. Nonetheless, the events conveyed through this media are large-scale and heterogeneous in their manifestations. This implies the need to detect smaller components, and to provide a description of these finer-grained units, which we call sub-events. In this work we propose a method to identify and to describe such sub-events. Word and document embeddings have proven efficient in capturing information about semantic relations among text data. Therefore we exploit this to build vectors from tweets in a representation space where the similarity between tweet vectors expresses a form of topical coherence between tweets. We leverage this tweet vector representation to build tweet groupings with a clustering method, where each group contains elements with close topics. These clusters are then described on the basis of a set of description tags. We present three methods to pick words from tweets in order to obtain this set. With this tag set, we propose to build a description for each cluster using a declarative Integer Linear Programming model. We develop an end-to-end procedure to implement this approach and we experiment it on two real datasets. One was collected with a request on the storm Alex, which struck France in October 2020, and the other with a request on the Ukraine war which started in February 2022. We propose two evaluation metrics and their results in order to show the interest of our method to select words to describe their cluster.
Olivier Gracianne, Anaïs Lefeuvre-Halftermeyer, Thi-Bich-Hanh Dao
ICTAI3
2022 Anchored Constrained Clustering Ensemble
abstract
In the context of semi-supervised learning and clustering ensemble methods, we introduce a novel strategy to consider not only pairwise constraints, but also triplet constraints. As far as we are aware, the latter have not been addressed in the literature of semi-supervised clustering ensembles. The strategy consists of a post-processing applied once a consensus partition has been built. Taking into account the fact that the clusters of the consensus partition are usually not spherical, in order to maintain their complex shapes, we first generate anchors, which are data points judged representative of the clusters, in such a way that every point in the cluster has an anchor close to it. These anchors are then used to create a data structure that we call an allocation matrix, which measures the assignment score of each point to each cluster in the consensus partition. Such a matrix is provided to an Integer Linear Programing (ILP) model to find the partition which is closest to the consensus partition, while satisfying the constraints. The experimental results show that our method, given an initial consensus partition and a set of constraints, allows to satisfy all the given constraints, modifying the initial partition, but without deteriorating considerably its quality.
Mathieu Guilbert, Christel Vrain, Thi-Bich-Hanh Dao, Marcílio Carlos Pereira de Souto
IJCNN3
2022 Knowledge Integration in Deep Clustering
Nguyen-Viet-Dung Nghiem, Christel Vrain, Thi-Bich-Hanh Dao
ECML/PKDD (1)3
2020 Constrained Clustering via Post-processing
Nguyen-Viet-Dung Nghiem, Christel Vrain, Thi-Bich-Hanh Dao, Ian Davidson
DS3
2018 Descriptive Clustering: ILP and CP Formulations with Applications
abstract
In many settings just finding a good clustering is insufficient and an explanation of the clustering is required. If the features used to perform the clustering are interpretable then methods such as conceptual clustering can be used. However, in many applications this is not the case particularly for image, graph and other complex data. Here we explore the setting where a set of interpretable discrete tags for each instance is available. We formulate the descriptive clustering problem as a bi-objective optimization to simultaneously find compact clusters using the features and to describe them using the tags. We present our formulation in a declarative platform and show it can be integrated into a standard iterative algorithm to find all Pareto optimal solutions to the two objectives. Preliminary results demonstrate the utility of our approach on real data sets for images and electronic health care records and that it outperforms single objective and multi-view clustering baselines.
Thi-Bich-Hanh Dao, Chia-Tung Kuo, S. S. Ravi, Christel Vrain, Ian Davidson
IJCAI1
2018 Constrained distance based clustering for time-series: a comparative and experimental study
Thomas Andrew Lampert, Thi-Bich-Hanh Dao, Baptiste Lafabregue, Nicolas Serrette, Germain Forestier, Bruno Crémilleux, Christel Vrain, Pierre Gançarski
Data Min. Knowl. Discov.2
2017 A Framework for Minimal Clustering Modification via Constraint Programming
abstract
Consider the situation where your favorite clustering algorithm applied to a data set returns a good clustering but there are a few undesirable properties. One adhoc way to fix this is to re-run the clustering algorithm and hope to find a better variation. Instead, we propose to not run the algorithm again but minimally modify the existing clustering to remove the undesirable properties. We formulate the minimal clustering modification problem where we are given an initial clustering produced from any algorithm. The clustering is then modified to: i) remove the undesirable properties and ii) be minimally different to the given clustering. We show the underlying feasibility sub-problem can be intractable and demonstrate the flexibility of our constraint programming formulation. We empirically validate its usefulness through experiments on social network and medical imaging data sets.
Chia-Tung Kuo, S. S. Ravi, Thi-Bich-Hanh Dao, Christel Vrain, Ian Davidson
AAAI3
2017 Constrained clustering by constraint programming
Thi-Bich-Hanh Dao, Khanh-Chuong Duong, Christel Vrain
Artif. Intell.1
2016 A Framework for Actionable Clustering Using Constraint Programming
abstract
Consider if you wish to cluster your ego network in Facebook so as to find several useful groups each of which you can invite to a different dinner party. You may require that each cluster must contain equal number of males and females, that the width of a cluster in terms of age is at most 10 and that each person in a cluster should have at least r other people with the same hobby. These are examples of cardinality, geometric and density requirements/constraints respectfully that can make the clustering useful for a given purpose. However existing formulations of constrained clustering were not designed to handle these constraints since they typically deal with low-level, instance-level constraints. We formulate a constraint programming (CP) languages formulation of clustering with these cluster-level styles of constraints which we call actionable clustering. Experimental results show the potential uses of this work to make clustering more actionable. We also show that these constraints can be used to improve the accuracy of semi-supervised clustering.
Thi-Bich-Hanh Dao, Christel Vrain, Khanh-Chuong Duong, Ian Davidson
ECAI1
2016 Repetitive Branch-and-Bound Using Constraint Programming for Constrained Minimum Sum-of-Squares Clustering
abstract
Minimum sum-of-squares clustering (MSSC) is a widely studied task and numerous approximate as well as a number of exact algorithms have been developed for it. Recently the interest of integrating prior knowledge in data mining has been shown, and much attention has gone into incorporating user constraints into clustering algorithms in a generic way.
Tias Guns, Thi-Bich-Hanh Dao, Christel Vrain, Khanh-Chuong Duong
ECAI2
2015 Constrained Minimum Sum of Squares Clustering by Constraint Programming
Thi-Bich-Hanh Dao, Khanh-Chuong Duong, Christel Vrain
CP1
2014 Model-theory and implementation of property grammars with features
abstract
Property Grammar (PG) is a formalism introduced by Blache [1], which aims at describing syntax in terms of local constraints that can be independently violated. A promising feature of this formalism lies in its ability to account for ungrammatical utterances, thus departing from classical formalisms of generative-enumerative syntax. In this article, we present a model-theoretic description of PG that improves on previous work by providing support for properties augmented with feature constraints. (e.g. the requirement and agreement properties). While providing a formal definition of the semantics of feature-based PG, we illustrate various uses of features within this formalism and give a general framework to interpret them. In a second part, we show how this formalization of PG can be turned into a Constraint Optimization Problem to implement a PG parser that supports the computation of both syntactic trees (for grammatical sentences) and quasi-syntactic trees (i.e. linguistically motivated syntactic structure for ungrammatical utterances). Finally, we briefly report on the implementation of such a parser using the Gecode library for Constraint Programming.
Denys Duchier, Thi-Bich-Hanh Dao, Yannick Parmentier 0001
J. Log. Comput.2
2013 A Filtering Algorithm for Constrained Clustering with Within-Cluster Sum of Dissimilarities Criterion
abstract
Constrained clustering is an important task in Data Mining. In the last ten years, many works have been done to extend classical clustering algorithms to handle user-defined constraints, but restricted to handle one kind of user-constraints. In a previous work [1], we have proposed a declarative and generic framework, based on Constraint Programming, which enables to design a clustering task by specifying an optimization criterion and different kinds of user-constraints. One of the criteria is the within-cluster sum of dissimilarities, which is represented by a sum constraint and reified equality constraints V=Σ1≤i<;j≤n(G[i]==G[j])aij· A direct implementation using predefined constraints is not effective as the propagation of theses constraints is weak. In this paper, we consider this criterion as a global constraint and develop a filtering algorithm for it. This filtering helps to improve significantly the model performance. Experiments on classical databases show the interest of our approach.
Thi-Bich-Hanh Dao, Khanh-Chuong Duong, Christel Vrain
ICTAI1
2013 A Declarative Framework for Constrained Clustering
Thi-Bich-Hanh Dao, Khanh-Chuong Duong, Christel Vrain
ECML/PKDD (3)1
2008 Theory of finite or infinite trees revisited
abstract
Abstract We present in this paper a first-order axiomatization of an extended theory T of finite or infinite trees, built on a signature containing an infinite set of function symbols and a relation finite(t), which enables to distinguish between finite and infinite trees. We show that T has at least one model and prove its completeness by giving not only a decision procedure, but a full first-order constraint solver that gives clear and explicit solutions for any first-order constraint satisfaction problem in T. The solver is given in the form of 16 rewriting rules that transform any first-order constraint ϕ into an equivalent disjunction φ of simple formulas such that φ is either the formula true or the formula false or a formula having at least one free variable, being equivalent neither to true nor to false and where the solutions of the free variables are expressed in a clear and explicit way. The correctness of our rules implies the completeness of T. We also describe an implementation of our algorithm in CHR (Constraint Handling Rules) and compare the performance with an implementation in C++ and that of a recent decision procedure for decomposable theories.
Khalil Djelloul, Thi-Bich-Hanh Dao, Thom W. Frühwirth
Theory Pract. Log. Program.2
2006 Solving First-Order Constraints in the Theory of the Evaluated Trees
Thi-Bich-Hanh Dao, Khalil Djelloul
ICLP1
2003 Intermediate (Learned) Consistencies
Arnaud Lallouet, Andrei Legtchenko, Thi-Bich-Hanh Dao, AbdelAli Ed-Dbali
CP3
2003 Finite Domain Constraint Solver Learning
Arnaud Lallouet, Thi-Bich-Hanh Dao, Andrei Legtchenko, AbdelAli Ed-Dbali
IJCAI2
2002 Indexical-Based Solver Learning
Thi-Bich-Hanh Dao, Arnaud Lallouet, Andrei Legtchenko, Lionel Martin
CP1
2000 Expressiveness of Full First Order Constraints in the Algebra of Finite or Infinite Trees
Alain Colmerauer, Thi-Bich-Hanh Dao
CP2