EDBT 2026 Demo / reviewers in the wild / expert
Tai Dinh
dblp:171/2667 · also Duy-Tai Dinh
· DBLP profile ↗
15ranked-venue papers
7as first author
8since 2021 · last 2026
0000-0001-7597-4262ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 6 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An upper bound on the silhouette evaluation metric for clusteringabstract• Derives a global, data-dependent upper bound for average silhouette width (ASW). • Gives pointwise silhouette-width bounds, and aggregates them to the ASW ceiling. • Computes bounds from pairwise dissimilarities in O(n2 log n) time overall only. • Experiments show bounds aid ASW interpretation; usefulness is datasetdependent. • Extends the framework to an upper bound for the macro-averaged silhouette score. The silhouette coefficient quantifies, for each observation, the balance between within-cluster cohesion and between-cluster separation, taking values in the range [ − 1 , 1 ] . The average silhouette width ( ASW ) is a widely used internal measure of clustering quality, with higher values indicating more cohesive and well-separated clusters. However, the dataset-specific maximum of ASW is typically unknown, and the standard upper limit of 1 is rarely attainable. In this work, we derive for each data point a sharp upper bound on its silhouette width and aggregate these to obtain a canonical upper bound of the ASW . This bound—often substantially below 1—enhances the interpretability of empirical ASW values by providing guidance on how close a given clustering result is to the best possible outcome for that dataset. We evaluate the usefulness of the upper bound on a variety of datasets and conclude that it can meaningfully enrich cluster quality evaluation; however, its practical relevance depends on the specific dataset. Finally, we extend the framework to establish an upper bound of the macro-averaged silhouette. Hugo Sträng, Tai Dinh |
Pattern Recognit. | 2 |
| 2025 | Sentiment Analysis of Travel Reviews in Kyoto Using LLMs
Shehan Liyanaarachchi, Tai Dinh, Wuyi Yue, Trung Vo, Minh Le Nguyen 0001 |
ADMA (3) | 2 |
| 2025 | A Mixed-Methods Framework for Analyzing Travel Reviews with LLMs and MAXQDA: A Case Study in Osaka
Daud Santoso, Tai Dinh, Shehan Liyanaarachchi, Wuyi Yue |
MEDI | 2 |
| 2025 | An efficient fusion-based deep learning framework for land use and land cover image clustering
Tai Dinh, Zdena Dobesová, Huynh Van Hong, Daniil Lisik, Rameesh Khan |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | Categorical data clustering: 25 years beyond K-modes
Tai Dinh, Wong Hauchi, Philippe Fournier-Viger, Daniil Lisik, Minh-Quyet Ha, Hieu-Chi Dam, Van-Nam Huynh |
Expert Syst. Appl. | 1 |
| 2023 | Understanding Mobile Game Reviews Through Sentiment Analysis: A Case Study of PUBGm
Yang Yu 0046, Tai Dinh, Fangyu Yu, Van-Nam Huynh |
MEDI | 2 |
| 2022 | Multi-core parallel algorithms for hiding high-utility sequential patterns
Ut Huynh, Bac Le, Tai Dinh, Hamido Fujita |
Knowl. Based Syst. | 3 |
| 2021 | Clustering mixed numerical and categorical data with missing valuesabstractThis paper proposes a novel framework for clustering mixed numerical and categorical data with missing values. It integrates the imputation and clustering steps into a single process, which results in an algorithm named Clustering Mixed Numerical and Categorical Data with Missing Values (k-CMM). The algorithm consists of three phases. The initialization phase splits the input dataset into two parts based on missing values in objects and attributes types. The imputation phase uses the decision-tree-based method to find the set of correlated data objects. The clustering phase uses the mean and kernel-based methods to form cluster centers at numerical and categorical attributes, respectively. The algorithm also uses the squared Euclidean and information-theoretic-based dissimilarity measure to compute the distances between objects and cluster centers. An extensive experimental evaluation was conducted on real-life datasets to compare the clustering quality of k-CMM with state-of-the-art clustering algorithms. The execution time, memory usage, and scalability of k-CMM for various numbers of clusters or data sizes were also evaluated. Experimental results show that k-CMM can efficiently cluster missing mixed datasets as well as outperform other algorithms when the number of missing values increases in the datasets. Tai Dinh, Van-Nam Huynh, Songsak Sriboonchitta |
Inf. Sci. | 1 |
| 2020 | k-PbC: an improved cluster center initialization for categorical data clustering
Tai Dinh, Van-Nam Huynh |
Appl. Intell. | 1 |
| 2018 | k-CCM: A Center-Based Algorithm for Clustering Categorical Data with Missing Values
Tai Dinh, Van-Nam Huynh |
MDAI | 1 |
| 2018 | A New Context-Based Clustering Framework for Categorical Data
Thanh-Phu Nguyen, Tai Dinh, Van-Nam Huynh |
PRICAI (1) | 2 |
| 2018 | An efficient algorithm for mining periodic high-utility sequential patterns
Tai Dinh, Bac Le, Philippe Fournier-Viger, Van-Nam Huynh |
Appl. Intell. | 1 |
| 2018 | A pure array structure and parallel strategy for high-utility sequential pattern mining
Bac Le, Ut Huynh, Tai Dinh |
Expert Syst. Appl. | 3 |
| 2018 | An efficient algorithm for Hiding High Utility Sequential Patterns
Bac Le, Tai Dinh, Van-Nam Huynh, Quang Minh Nguyen, Philippe Fournier-Viger |
Int. J. Approx. Reason. | 2 |
| 2017 | Mining Periodic High Utility Sequential Patterns
Tai Dinh, Van-Nam Huynh, Bac Le |
ACIIDS (1) | 1 |