Tai Dinh

dblp:171/2667 · also Duy-Tai Dinh · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
8since 2021 · last 2026
0000-0001-7597-4262ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 6 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 An upper bound on the silhouette evaluation metric for clustering
abstract
• Derives a global, data-dependent upper bound for average silhouette width (ASW). • Gives pointwise silhouette-width bounds, and aggregates them to the ASW ceiling. • Computes bounds from pairwise dissimilarities in O(n2 log n) time overall only. • Experiments show bounds aid ASW interpretation; usefulness is datasetdependent. • Extends the framework to an upper bound for the macro-averaged silhouette score. The silhouette coefficient quantifies, for each observation, the balance between within-cluster cohesion and between-cluster separation, taking values in the range [ − 1 , 1 ] . The average silhouette width ( ASW ) is a widely used internal measure of clustering quality, with higher values indicating more cohesive and well-separated clusters. However, the dataset-specific maximum of ASW is typically unknown, and the standard upper limit of 1 is rarely attainable. In this work, we derive for each data point a sharp upper bound on its silhouette width and aggregate these to obtain a canonical upper bound of the ASW . This bound—often substantially below 1—enhances the interpretability of empirical ASW values by providing guidance on how close a given clustering result is to the best possible outcome for that dataset. We evaluate the usefulness of the upper bound on a variety of datasets and conclude that it can meaningfully enrich cluster quality evaluation; however, its practical relevance depends on the specific dataset. Finally, we extend the framework to establish an upper bound of the macro-averaged silhouette.
Hugo Sträng, Tai Dinh
Pattern Recognit.2
2025 Sentiment Analysis of Travel Reviews in Kyoto Using LLMs
Shehan Liyanaarachchi, Tai Dinh, Wuyi Yue, Trung Vo, Minh Le Nguyen 0001
ADMA (3)2
2025 A Mixed-Methods Framework for Analyzing Travel Reviews with LLMs and MAXQDA: A Case Study in Osaka
Daud Santoso, Tai Dinh, Shehan Liyanaarachchi, Wuyi Yue
MEDI2
2025 An efficient fusion-based deep learning framework for land use and land cover image clustering
Tai Dinh, Zdena Dobesová, Huynh Van Hong, Daniil Lisik, Rameesh Khan
Eng. Appl. Artif. Intell.1
2025 Categorical data clustering: 25 years beyond K-modes
Tai Dinh, Wong Hauchi, Philippe Fournier-Viger, Daniil Lisik, Minh-Quyet Ha, Hieu-Chi Dam, Van-Nam Huynh
Expert Syst. Appl.1
2023 Understanding Mobile Game Reviews Through Sentiment Analysis: A Case Study of PUBGm
Yang Yu 0046, Tai Dinh, Fangyu Yu, Van-Nam Huynh
MEDI2
2022 Multi-core parallel algorithms for hiding high-utility sequential patterns
Ut Huynh, Bac Le, Tai Dinh, Hamido Fujita
Knowl. Based Syst.3
2021 Clustering mixed numerical and categorical data with missing values
abstract
This paper proposes a novel framework for clustering mixed numerical and categorical data with missing values. It integrates the imputation and clustering steps into a single process, which results in an algorithm named Clustering Mixed Numerical and Categorical Data with Missing Values (k-CMM). The algorithm consists of three phases. The initialization phase splits the input dataset into two parts based on missing values in objects and attributes types. The imputation phase uses the decision-tree-based method to find the set of correlated data objects. The clustering phase uses the mean and kernel-based methods to form cluster centers at numerical and categorical attributes, respectively. The algorithm also uses the squared Euclidean and information-theoretic-based dissimilarity measure to compute the distances between objects and cluster centers. An extensive experimental evaluation was conducted on real-life datasets to compare the clustering quality of k-CMM with state-of-the-art clustering algorithms. The execution time, memory usage, and scalability of k-CMM for various numbers of clusters or data sizes were also evaluated. Experimental results show that k-CMM can efficiently cluster missing mixed datasets as well as outperform other algorithms when the number of missing values increases in the datasets.
Tai Dinh, Van-Nam Huynh, Songsak Sriboonchitta
Inf. Sci.1
2020 k-PbC: an improved cluster center initialization for categorical data clustering
Tai Dinh, Van-Nam Huynh
Appl. Intell.1
2018 k-CCM: A Center-Based Algorithm for Clustering Categorical Data with Missing Values
Tai Dinh, Van-Nam Huynh
MDAI1
2018 A New Context-Based Clustering Framework for Categorical Data
Thanh-Phu Nguyen, Tai Dinh, Van-Nam Huynh
PRICAI (1)2
2018 An efficient algorithm for mining periodic high-utility sequential patterns
Tai Dinh, Bac Le, Philippe Fournier-Viger, Van-Nam Huynh
Appl. Intell.1
2018 A pure array structure and parallel strategy for high-utility sequential pattern mining
Bac Le, Ut Huynh, Tai Dinh
Expert Syst. Appl.3
2018 An efficient algorithm for Hiding High Utility Sequential Patterns
Bac Le, Tai Dinh, Van-Nam Huynh, Quang Minh Nguyen, Philippe Fournier-Viger
Int. J. Approx. Reason.2
2017 Mining Periodic High Utility Sequential Patterns
Tai Dinh, Van-Nam Huynh, Bac Le
ACIIDS (1)1