Thu Nguyen 0001

dblp:47/3996-1 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
16since 2021 · last 2026
0000-0001-7044-1731ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Low-Dimension Representation Estimation in Principal Component Analysis Under Missing Data
Thanh Tu Do, Van Hua, Uyen Dang, Thu Nguyen 0001, Steven Alexander Hicks, Pål Halvorsen, Michael Riegler 0001, Binh T. Nguyen 0001
MMM (2)4
2026 DPERC: Direct Parameter Estimation for Mixed Data with Random Missingness
Tuan L. Vo, Uyen Dang, Thu Nguyen 0001, Pål Halvorsen, Michael Riegler 0001, Binh T. Nguyen 0001
MMM (2)3
2026 PICA: Interpretable imputation for randomly missing data
Tuan L. Vo, Uyen Dang, Van Hua, Xuan Hoang Nguyen, Thu Nguyen 0001, Bao Huynh
Inf. Sci.5
2025 Data Imputation for Noisy Time-Series Data in Healthcare
Lien P. Le, Thi Xuan-Hien Nguyen, Thu Nguyen 0001, Michael Riegler 0001, Pål Halvorsen, Binh T. Nguyen 0001
ICCCI (2)3
2025 Reproducibility Companion Paper: NIF: A Fast Implicit Image Compression with Bottleneck Layers and Modulated Sinusoidal Activations
abstract
In this companion paper, we reproduce the experiments presented in our work titled ''NIF: A Fast Implicit Image Compression with Bottleneck Layers and Modulated Sinusoidal Activations'' [2], presented at ACM Multimedia 2023. In this study, we present the architecture and the technical details of our implementation and provide instructions to reproduce the main results, the ablation study, the plots and the figures presented in the paper. All the material described in this paper is released on GitHub [3], featuring the full results, a reference software implementation and a generic environment setup that works on any system, even without a GPU.
Lorenzo Catania, Dario Allegra, Luigi Capogrosso, Thu Nguyen 0001
ACM Multimedia4
2024 Blockwise Principal Component Analysis for monotone missing data imputation and dimensionality reduction
abstract
Monotone missing data is a common problem in data analysis. However, imputation combined with dimensionality reduction can be computationally expensive, especially with the increasing size of datasets. We propose a Blockwise Principal Component Analysis Imputation (BPI) framework for dimensionality reduction and imputation of monotone missing data to address this issue. The framework conducts Principal Component Analysis on the observed part of each monotone block of the data and then imputes on merging the obtained principal components using a chosen imputation technique. BPI can work with various imputation techniques and can significantly reduce imputation time compared to conducting dimensionality reduction after imputation. This makes it a practical and efficient approach for large datasets with monotone missing data. Our experiments validate the improvement in speed while achieving an accuracy that is comparable to the common strategy of imputation prior to dimensional reduction.
Tu T. Do, Anh M. Vu, Tuan L. Vo, Hoang Thien Ly, Thu Nguyen 0001, Steven Alexander Hicks, Pål Halvorsen, Michael Riegler 0001, Binh T. Nguyen 0001
IJCNN5
2024 Correlation Visualization Under Missing Values: A Comparison Between Imputation and Direct Parameter Estimation Methods
Nhat-Hao Pham, Khanh-Linh Vo, Mai Anh Vu, Thu Nguyen 0001, Michael Riegler 0001, Pål Halvorsen, Binh T. Nguyen 0001
MMM (4)4
2023 Faster Imputation Using Singular Value Decomposition for Sparse Data
Linh G. H. Tran, Bao H. Le, Thuong H. T. Nguyen, Thu Nguyen 0001, Hien D. Nguyen 0002, Binh T. Nguyen 0001
ACIIDS (1)5
2023 Principal Components Analysis Based Imputation for Logistic Regression
Thuong H. T. Nguyen, Bao Le, Linh G. H. Tran, Thu Nguyen 0001, Binh T. Nguyen 0001
IEA/AIE (1)5
2023 Combining datasets to improve model fitting
abstract
For many use cases, combining information from different datasets can be of interest to improve a machine learning model's performance, especially when the number of samples from at least one of the datasets is small. An additional challenge in such cases is that the features from these datasets are not identical, even though there are some commonly shared features among the datasets. To tackle this, we propose a novel framework called Combine datasets based on Imputation (ComImp). In addition, we propose PCA-ComImp, a variant of ComImp that utilizes Principle Component Analysis (PCA), where dimension reduction is conducted before combining datasets. This is useful when the datasets have a large number of features that are not shared across them. Furthermore, our framework can also be utilized for data preprocessing by imputing missing data, i.e., filling in the missing entries while combining different datasets. To illustrate the performance and practicability of the proposed methods and their potential usages, we conduct experiments for various tasks (regression, classification) and for different data types (tabular data, time series data) when the datasets to be combined have missing data. We also investigate how the devised methods can be used with transfer learning to provide even further model training improvement. Our results indicate that can provide extra improvement when being used in combination with transfer learning.
Thu Nguyen 0001, Rabindra Khadka, Nhan Phan, Anis Yazidi, Pål Halvorsen, Michael Riegler 0001
IJCNN1
2023 Multimedia Datasets: Challenges and Future Possibilities
Thu Nguyen 0001, Andrea M. Storås, Vajira Thambawita, Steven Alexander Hicks, Pål Halvorsen, Michael Riegler 0001
MMM (2)1
2022 Unequal Covariance Awareness for Fisher Discriminant Analysis and Its Variants in Classification
abstract
Fisher Discriminant Analysis (FDA) is one of the essential tools for feature extraction and classification. In addition, it motivates the development of many improved techniques based on the FDA to adapt to different problems or data types. However, none of these approaches make use of the fact that the assumption of equal covariance matrices in FDA is usually not satisfied in practical situations. Therefore, we propose a novel classification rule for the FDA that accounts for this fact, mitigating the effect of unequal covariance matrices in the FDA. Furthermore, since we only modify the classification rule, the same can be applied to many FDA variants, improving these algorithms further. Theoretical analysis reveals that the new classification rule allows the implicit use of the class covariance matrices while increasing the number of parameters to be estimated by a small amount compared to going from FDA to Quadratic Discriminant Analysis. We illustrate our idea via experiments, which shows the superior performance of the modified algorithms based on our new classification rule compared to the original ones.
Thu Nguyen 0001, Quang M. Le, Son N. T. Tu, Binh T. Nguyen 0001
IJCNN1
2022 Parallel feature selection based on the trace ratio criterion
abstract
The growth of data today poses a challenge in management and inference. While feature extraction methods are capable of reducing the size of the data for inference, they do not help in minimizing the cost of data storage. On the other hand, feature selection helps to remove the redundant features and therefore is helpful not only in inference but also in reducing management costs. This work presents a novel parallel feature selection approach for classification, namely Parallel Feature Selection using Trace criterion (PFST), which scales up to very large datasets. Our method uses trace criterion, a measure of class separability used in Fisher's Discriminant Analysis, to evaluate feature usefulness. We analyzed the criterion's desirable properties theoretically. Based on the criterion, PFST rapidly finds important features out of a set of features for big datasets by first making a forward selection with early removal of seemingly redundant features parallelly. After the most important features are included in the model, we check back their contribution for possible interaction that may improve the fit. Lastly, we make a backward selection to check back possible redundant added by the forward steps. We evaluate our methods via various experiments using Linear Discriminant Analysis as the classifier on selected features. The experiments show that our method can produce a small set of features in a fraction of the amount of time by the other methods under comparison. In addition, the classifier trained on the features selected by PFST not only achieves better accuracy than the ones chosen by other approaches, but can also achieve better accuracy than the classification on all available features.
Thu Nguyen 0001, Thanh Nhan Phan, Van Nhuong Nguyen, Binh T. Nguyen 0001, Pål Halvorsen, Michael Riegler 0001
IJCNN1
2022 ASMCNN: An efficient brain extraction using active shape model and convolutional neural networks
Duy M. H. Nguyen, Duy M. Nguyen, Truong Thanh Nhat Mai, Thu Nguyen 0001, Khanh T. Tran, Anh Triet Nguyen, Bao T. Pham, Binh T. Nguyen 0001
Inf. Sci.4
2022 DPER: Direct Parameter Estimation for Randomly missing data
Thu Nguyen 0001, Khoi Minh Nguyen-Duy, Duy Ho Minh Nguyen, Binh T. Nguyen 0001, Bruce A. Wade
Knowl. Based Syst.1
2021 EPEM: Efficient Parameter Estimation for Multiple Class Monotone Missing Data
Thu Nguyen 0001, Duy M. H. Nguyen, Binh T. Nguyen 0001, Bruce A. Wade
Inf. Sci.1
2020 Deep Matrix Tri-Factorization: Mining Vertex-wise Interactions in Multi-Space Attributed Graphs
abstract
Mining vertex-wise interactions in graphs helps reveal useful information in real-world applications, such as bioinformatics networks BioGRID and DrugBank and academic networks DBLP and Arxiv. A main challenge in developing a general learning method for this setting is that each vertex may be associated with features from heterogeneous feature spaces, representing very disparate information. Moreover, features could be raw and low-level, leading to sparse representations. Some solutions in this area treat all feature spaces as equally important and concatenate features from heterogeneous feature spaces into a single feature vector. Others harmonize different feature spaces by respecting their relative significance in mining vertex-wise interactions but requiring construct specialized harmonizing function and/or handcrafting expressive features, both of which entail expert knowledge. Motivated by this observation, we propose a new learning paradigm named Deep Matrix Tri-Factorization (DM3F), which draws insights from deep models: (i) DM3F replaces the linear combination with a neural architecture that can learn an arbitrary harmonizing function from data; and (ii) DM3F allows raw feature inputs and automatically extracts high-level feature representations via a layer-by-layer learning mechanism. These two characteristics of DM3F make it accessible for users without expert knowledge. DM3F includes two orthogonal and complementary models, allowing an ensemble mechanism to optimize its performance during both training and predicting. A theoretical analysis of DM3F reveals that it possesses several desirable properties, including that it strictly generalizes matrix factorization models. We demonstrate the performance of DM3F on two real-world datasets.
Yi He 0007, Sheng Chen 0008, Thu Nguyen 0001, Bruce A. Wade, Xindong Wu 0001
SDM3