VLDB 2026 Research / reviewers in the wild / expert
Binh P. Nguyen
dblp:88/8221
· DBLP profile ↗
40ranked-venue papers
5as first author
28since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 1 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 7 · 7 since 2021Systems, architecture and hardware · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TTVAE: Transformer-Based Generative Modeling for Tabular Data Generation (Abstract Reprint)abstractTabular data synthesis presents unique challenges, with Transformer models remaining underexplored despite the applications of Variational Autoencoders and Generative Adversarial Networks. To address this gap, we propose the Transformer-based Tabular Variational AutoEncoder (TTVAE), leveraging the attention mechanism for capturing complex data distributions. The inclusion of the attention mechanism enables our model to understand complex relationships among heterogeneous features, a task often difficult for traditional methods. TTVAE facilitates the integration of interpolation within the latent space during the data generation process. Specifically, TTVAE is trained once, establishing a low-dimensional representation of real data, and then various latent interpolation methods can efficiently generate synthetic latent points. Through extensive experiments on diverse datasets, TTVAE consistently achieves state-of-the-art performance, highlighting its adaptability across different feature types and data sizes. This innovative approach, empowered by the attention mechanism and the integration of interpolation, addresses the complex challenges of tabular data synthesis, establishing TTVAE as a powerful solution. Alex X. Wang, Binh P. Nguyen |
AAAI | 2 |
| 2026 | Predicting Fibromyalgia Pain From Self-reported Health Data Using Time-Series and Deep Learning Models
Delnia Alipour, Olga Perepelkina, Alessandro Vinciarelli, Tahir Janmohamed, Binh P. Nguyen, Simone Stumpf |
AIME (1) | 5 |
| 2026 | Edge-updating graph neural networks for modeling feature interactions in tabular dataabstractWe proposed a message-passing graph neural network (GNN) based on graph isomorphism network (GIN) for learning on tabular data. Fully-connected, unweighted feature graphs were constructed from tabular data using contextual feature encodings for numerical and categorical features. A classification node was added to the feature graph to represent the entire graph during inference. Feature interactions were modeled using the proposed architecture, which used a neural network to learn edge attributes while leveraging residual connections for both node and edge updates to alleviate oversmoothing, a commonly-faced problem in GNNs. Our model was evaluated on 12 publicly available datasets and achieved the best mean rank among 6 tabular deep learning and GNN models. Furthermore, since gradient-boosted decision trees are considered to be state-of-the-art models for tabular data, we compared our model to XGBoost and CatBoost and found that our model outperformed both models in all datasets when using default hyperparameters and achieved the best results in 8 datasets when using tuned hyperparameters. A further comparison was made against 5 commonly used or recently-proposed GNNs to investigate the effectiveness of our model, in which our model achieved the top result in all datasets. Pimwipa Charuthamrong, Colin R. Simpson, Binh P. Nguyen |
Neural Networks | 3 |
| 2025 | Differential Evolutionary for Label Ordering in Multi-label Classification
Bach Hoai Nguyen, Binh P. Nguyen, Vinh Truong Hoang |
ADMA (4) | 2 |
| 2025 | iACP-KAN: Identifying Anticancer Peptides using Kolmogorov-Arnold NetworkabstractAnticancer peptides (ACPs) represent a promising therapeutic approach for cancer treatment, offering potential advantages over traditional methods by providing more selective targeting and reduced toxicity. However, the clinical implementation of ACPs has been hindered by insufficient effective prediction methods. This study introduces iACP-KAN, a novel computational model for identifying anticancer peptides using a Kolmogorov-Arnold Network (KAN) with innovative feature integration. The proposed methodology combines four handcrafted features: Dipeptide Composition, Binary Profile Features, Pseudo Amino Acid Composition, and K-mer Composition, with sequence embedding features learned through a Long Short-Term Memory. Utilizing two independent datasets, the model demonstrated superior performance across multiple evaluation metrics, consistently outperforming seven state-of-the-art methods. For Dataset 1, the proposed approach achieved the highest Area Under the Receiver Operating Characteristic Curve (AUROC) of 0.8713 and Area Under the Precision-Recall Curve (AUPRC) of 0.5026 on the test set. For Dataset 2, it obtained an AUROC of 0.8551 and an AUPRC of 0.8713 on the test set. The results proved its potential to advance ACP prediction and support drug discovery efforts. Thanh-Hoang Nguyen-Vo, Quang H. Trinh, Huu-Thanh Duong, Trang T. T. Do 0001, Binh P. Nguyen |
CIBCB | 6 |
| 2025 | GSDeep-DTA: A Hybrid Graph and Sequence-based Deep Learning Framework for Robust Drug-Target Affinity PredictionabstractRobust prediction of drug-target binding affinity (DTA) is essential for accelerating drug discovery pipelines, as it reduces experimental failures in cold-start scenarios where novel compounds or targets are involved. Although existing DTA models primarily rely on sequence-based or graph-based representations, a limited number of studies have explored the integration of both approaches. However, effectively encoding and integrating the diverse features of drugs and proteins while enhancing predictive performance remains a challenging task. This work proposes a hybrid graph- and sequence-based framework for robust DTA prediction, GSDeep-DTA. The model integrates graph neural networks to represent drug molecular structures and protein contact maps, and cascades Convolutional Neural Network - Bidirectional Long Short-Term Memory modules with transformer embeddings to hierarchically capture sequential dependencies. We adopt a weighted sum fusion mechanism to integrate these heterogeneous features, which balances effectiveness and simplicity compared to more complex fusion techniques. Experimental results on the Davis benchmark dataset show that our model outperforms state-of-the-art DTA prediction models. In addition, we evaluate its ability to generalize in cold-start scenarios, assessing its performance on novel drugs and/or proteins. Our findings highlight the potential of hybrid graph-and sequence-based deep learning models for improved DTA prediction. W. G. D. M. Samankula, Joanne Harvey, Binh P. Nguyen |
CIBCB | 3 |
| 2025 | MolHFCNet: Enhancing Molecular Graph Representations with Hierarchical Feature Combining and Hybrid PretrainingabstractEfficient molecular property prediction is crucial in bioinformatics and cheminformatics, with applications in drug discovery, materials science, and chemical engineering. This paper introduces MolHFCNet, a graph neural network designed to enhance molecular representation learning. At its core, the n-Hierarchical Features Combining (n-HFC) module aggregates information across multiple hierarchical feature spaces, effectively capturing both local and global graph structures. Unlike conventional models, n-HFC maintains computational complexity comparable to a single full-dimensional graph layer while supporting either 2D or 3D molecular graphs, ensuring flexibility across tasks. Furthermore, we propose a novel graph pretraining strategy that integrates predictive and contrastive learning, enabling the model to capture local chemical interactions and global molecular contexts for robust embeddings. Experimental results on benchmark datasets demonstrate MolHFCNet’s superior accuracy and efficiency compared to state-of-the-art methods, highlighting the potential of high-order hierarchical feature learning for advancing molecular graph analysis. Our code is available at https://github.com/ndlongvn/MolHFCNet. Duy-Long Nguyen, Ho Viet Duc Luong, Anh-Thu Ngo-Tran, Quang H. Nguyen 0001, Binh P. Nguyen |
IJCAI | 5 |
| 2025 | Unleashing SAM for Few-Shot Medical Image Segmentation with Dual-Encoder and Automated Prompting
Cuong M. Pham, Phi-Le Nguyen, Thanh Trung Nguyen, Vu Minh Hieu Phan, Binh P. Nguyen |
MICCAI (6) | 5 |
| 2025 | Generative AI for Tabular Data Synthesis
Alex X. Wang, Binh P. Nguyen, Colin R. Simpson |
PAKDD (4) | 2 |
| 2025 | TTVAE: Transformer-based generative modeling for tabular data generation
Alex X. Wang, Binh P. Nguyen |
Artif. Intell. | 2 |
| 2025 | Blending is all you need: Data-centric ensemble synthetic data
Alex X. Wang, Colin R. Simpson, Binh P. Nguyen |
Inf. Sci. | 3 |
| 2025 | CTVAE: Contrastive Tabular Variational Autoencoder for imbalance dataabstractAbstract Class imbalance, where datasets often lack sufficient samples for minority classes, is a persistent challenge in machine learning. Existing solutions often generate synthetic data to mitigate this issue, but they typically struggle with complex data distributions, primarily because they focus on oversampling the minority class while neglecting the relationships with the majority class. To overcome these limitations, we propose the Contrastive Tabular Variational Autoencoder (CTVAE), which integrates conditional Variational Autoencoders with contrastive learning techniques. CTVAE excels at generating high-quality synthetic samples that capture the intricate data distributions of both minority and majority classes. Additionally, it can be seamlessly integrated with variants of the Synthetic Minority Oversampling Technique (SMOTE) for enhanced effectiveness. Experimental results demonstrate that CTVAE substantially improves classification performance on imbalanced datasets, offering a more robust and holistic solution to the class imbalance problem. Alex X. Wang, Minh Quang Le, Huu-Thanh Duong, Bay Nguyen Van, Binh P. Nguyen |
Knowl. Inf. Syst. | 5 |
| 2025 | Deterministic Autoencoder using Wasserstein loss for tabular data generationabstractTabular data generation is a complex task due to its distinctive characteristics and inherent complexities. While Variational Autoencoders have been adapted from the computer vision domain for tabular data synthesis, their reliance on non-deterministic latent space regularization introduces limitations. The stochastic nature of Variational Autoencoders can contribute to collapsed posteriors, yielding suboptimal outcomes and limiting control over the latent space. This characteristic also constrains the exploration of latent space interpolation. To address these challenges, we present the Tabular Wasserstein Autoencoder (TWAE), leveraging the deterministic encoding mechanism of Wasserstein Autoencoders. This characteristic facilitates a deterministic mapping of inputs to latent codes, enhancing the stability and expressiveness of our model's latent space. This, in turn, enables seamless integration with shallow interpolation mechanisms like the synthetic minority over-sampling technique (SMOTE) within the data generation process via deep learning. Specifically, TWAE is trained once to establish a low-dimensional representation of real data, and various latent interpolation methods efficiently generate synthetic latent points, achieving a balance between accuracy and efficiency. Extensive experiments consistently demonstrate TWAE's superiority, showcasing its versatility across diverse feature types and dataset sizes. This innovative approach, combining WAE principles with shallow interpolation, effectively leverages SMOTE's advantages, establishing TWAE as a robust solution for complex tabular data synthesis. Alex X. Wang, Binh P. Nguyen |
Neural Networks | 2 |
| 2024 | Hierarchical Federated Learning in MEC Networks with Knowledge DistillationabstractModern automobiles are equipped with advanced computing capabilities, allowing them to become powerful computing units capable of processing a large amount of data and training machine learning models. However, machine learning algorithms typically require a large centralized dataset, raising concerns about users’ privacy. Federated Learning (FL) is a distributed machine learning paradigm that tackles this problem, allowing intelligent vehicles to collaboratively train machine learning models locally without having to compromise their private data. Multiple works have concentrated on applying Federated Learning on Mobile Edge Computing (MEC) networks with a 3-tier architecture consisting of mobile clients, edge servers, and cloud servers, where the edge server aggregates its local set of clients, and the cloud server aggregates edge servers to learn a global model. This approach helps reduce the expensive communication costs to the far-away cloud server. However, this 3-tier paradigm faces several challenges, a notable one being clients’ constant mobility, leading to regional edges having a fluctuating set of participating clients at each round, which we refer to as distribution drift. This phenomenon introduces instability to the local training process, leading to suboptimal accuracy and convergence. As a solution, we propose a local training process based on the knowledge distillation mechanism. Specifically, we employ the global model and an ensemble of historical regional models from the edge servers as sources of knowledge to guide the local training process, preventing the local models from drifting away from the global knowledge and preserving information from clients that left the region. Experimental results showed that the proposed method helps achieve better performance compared to other baselines. Tuan Dung Nguyen, Ngoc Anh Tong, Binh P. Nguyen, Nguyen Quoc Viet Hung, Phi-Le Nguyen |
IJCNN | 3 |
| 2024 | A Contrastive Learning and Graph-based Approach for Missing Modalities in Multimodal Federated LearningabstractFederated Learning has emerged as a decentralized method for training machine learning models using distributed data sources. It ensures privacy by allowing clients to collaboratively learn a shared global model while keeping their data stored locally. However, a significant challenge arises when dealing with missing modalities in clients’ datasets, where certain features or modalities are unavailable or incomplete, leading to heterogeneous data distribution. Previous studies have addressed this issue, but they fall short in addressing generalizability across diverse, unobserved individuals. This study introduces MIFL, a novel framework for handling modality-missing clients in Multimodal Federated Learning. MIFL utilizes a unique approach, optimizing a client’s local model with existing modalities while incorporating absent modalities from clients. Aggregation of these models is performed through a graph-based attentive aggregation method, maintaining generalized characteristics by updating a global model averaged across clients. Our experimental results demonstrate the effectiveness of MIFL across various client configurations with statistical heterogeneity, showcasing its potential for addressing the challenge of missing modalities in Federated Learning. Thu Hang Phung, Binh P. Nguyen, Thanh-Hung Nguyen, Nguyen Quoc Viet Hung, Phi-Le Nguyen |
IJCNN | 2 |
| 2024 | i6mA-CNN: A Web-based System to Identify DNA N6-Methyladenine Sites in Mouse GenomesabstractN6-methyladenine (6mA) is one of the most common epigenetic modifications of DNA sequences found in both eukaryotes and prokaryotes. In prokaryotes, 6mA is closely associated with various biochemical processes such as DNA replication, repair, transcription, and cellular defense. In eukaryotes, the biological roles and behaviors of this methylation type have not been fully understood. Therefore, gaining more knowledge about 6mA sites is important and contributes to uncovering the characteristics and unexplored biological functions of 6mA. In our study, we propose an effective computational method called i6mA-CNN using convolutional neural networks with a fusion of multiple receptive fields. The 6mA sequences of Mus musculus (mice) were retrieved from the MethSMRT database and then refined to create a benchmark dataset. To fairly evaluate the performance of the model, we conducted multiple experiments and compared i6mA-CNN with other methods on the same independent test set. The results indicated that i6mACNN achieved better performance, with a value of 0.98 for both the area under the receiver operating characteristic curve and the area under the precision-recall curve. Thanh-Hoang Nguyen-Vo, Susanto Rahardja, Binh P. Nguyen |
ISCAS | 3 |
| 2024 | Identifying Nephrotoxicity of Small Molecules Using Machine LearningabstractNephrotoxicity is a severe condition characterized by kidney damage resulting from exposure to harmful substances such as drugs, diagnostic agents, chemicals, or environmental toxins. The potential for nephrotoxicity in drug molecules remains significant, often leading to severe consequences for patients. Despite existing computational methods for identifying nephrotoxic molecules, these approaches fail to provide stable and reliable performance due to biased modeling (e.g., small sample sizes, imbalanced classes, and data leakage). In this study, we offer a refined dataset for nephrotoxicity prediction tasks. Our dataset was collected from existing studies, rebalanced, and carefully curated to improve the quality of data for Quantitative Structure-Activity Relationship modeling. Additionally, we implemented a series of 32 prediction models using eight machine learning algorithms in combination with three types of molecular representations. The implementation of these machine learning models serves as a preliminary survey of how conventional methods perform on the refined dataset. Our findings indicated that all implemented models achieved satisfactory performance. Our dataset could serve as a valuable resource for developing more advanced prediction methods in the future. Thanh-Hoang Nguyen-Vo, Linh Bui, Trang T. T. Do 0001, Susanto Rahardja, Binh P. Nguyen |
TENCON | 5 |
| 2024 | Comparative Analysis of Oversampling Techniques and Deep Learning for Imbalanced Tabular Data
Alex X. Wang, Colin R. Simpson, Binh P. Nguyen |
TENCON | 3 |
| 2024 | Enhancing public research on citizen data: An empirical investigation of data synthesis using Statistics New Zealand's Integrated Data InfrastructureabstractThe Integrated Data Infrastructure (IDI) in New Zealand is a critical asset that integrates citizen data from various public and private organizations for population-level analyses. However, access restrictions within the IDI environment present challenges for fully utilizing its potential. This study examines synthetic data as a potential solution, offering a comprehensive framework for generating customizable and easily implementable synthetic data. The evaluation of multiple data synthesis algorithms considers statistical similarity, machine learning utility, and privacy concerns. The findings reveal that distance-based algorithms, like SMOTE, strike a balance between accuracy and computational cost, making them suitable for IDI. The study also identifies the need for a clear release guide for micro-level synthetic data and proposes exploring a fully automatic data evaluation pipeline in future research. Additionally, the study highlights opportunities enabled by synthetic data, such as familiarization with administrative datasets, reproducibility of studies, pilot analyses, and enhanced cross-domain collaboration. Overall, the proposed framework and findings offer valuable insights and guidance for synthetic data projects within the IDI, advancing synthetic data privacy research and facilitating reproducibility, collaboration, and data sharing in the IDI ecosystem. Alex X. Wang, Stefanka S. Chukova, Andrew Sporle, Barry J. Milne, Colin R. Simpson, Binh P. Nguyen |
Inf. Process. Manag. | 6 |
| 2023 | Digital Phenotype Representation by Statistical, Information Theory, Data-Driven Approach with Digital Health DataabstractDigital phenotyping (DP) is a multidisciplinary field of science that quantifies the individual level phenotype through active and passive data. Although DP is a multidisciplinary field, there lacks a technical and a systematic approach to representing DP. This work proposes the development of digital phenotype profile (DPP) to represent a user’s physical and behavioural health baseline through systematic investigations with an emphasis on robustness and explainability. To achieve this, a Statistical, Information Theory, and Data-driven (SID) pipeline will develop the foundation of the DPP. SID evaluates the non-linearity of the signal to offer inference for domain-specific feature extraction, evaluates the information theory to rank the DPP parameters, and imputes missing data for robust analysis, respectively. SID was applied to a 24-hr Multi-Level dataset and was able to represent individual DPPs. The respective DPPs were visualized and clusters of awake and asleep were used for individual specific modelling. Binh P. Nguyen, Michael Nigro, Alice Rueda, Venkat Bhat, Sridhar Krishnan 0001 |
ICASSP | 1 |
| 2023 | Correction: A lightweight classification of adaptor proteins using transformer networks
Sylwan Rahardja, Mou Wang, Binh P. Nguyen, Pasi Fränti, Susanto Rahardja |
BMC Bioinform. | 3 |
| 2023 | Ensemble k-nearest neighbors based on centroid displacementabstractk-nearest neighbors (k-NN) is a well-known classification algorithm that is widely used in different domains. Despite its simplicity, effectiveness and robustness, k-NN is limited by the use of the Euclidean distance as the similarity metric, the arbitrarily selected neighborhood size k, the computational challenge of high-dimensional data, and the use of the simple majority voting rule in class determination. We sought to address the last issue and proposed the Centroid Displacement-based k-NN algorithm, where centroid displacement is used for class determination. This paper presents a simple yet efficient variant of our previous work, named Ensemble Centroid Displacement-based k-NN, which leverages the homogeneity of the nearest neighbors of test instances. Extensive experiments on various real and synthetic datasets were conducted to show the effectiveness and robustness of the proposed algorithm. Our experimental results demonstrate that the proposed algorithm is able to enhance the classification performance of the standard k-NN algorithm and its variants and also improve the computational efficiency. The performance of our algorithm was consistent and robust for both balanced and imbalanced datasets. Alex X. Wang, Stefanka S. Chukova, Binh P. Nguyen |
Inf. Sci. | 3 |
| 2022 | Implementation and Analysis of Centroid Displacement-Based k-Nearest Neighbors
Alex X. Wang, Stefanka S. Chukova, Binh P. Nguyen |
ADMA (1) | 3 |
| 2022 | Multi-level Community-awareness Graph Neural Networks for Neural Machine TranslationabstractNeural Machine Translation (NMT) aims to translate the source- to the target-language while preserving the original meaning. Linguistic information such as morphology, syntactic, and semantics shall be grasped in token embeddings to produce a high-quality translation. Recent works have leveraged the powerful Graph Neural Networks (GNNs) to encode such language knowledge into token embeddings. Specifically, they use a trained parser to construct semantic graphs given sentences and then apply GNNs. However, most semantic graphs are tree-shaped and too sparse for GNNs which cause the over-smoothing problem. To alleviate this problem, we propose a novel Multi-level Community-awareness Graph Neural Network (MC-GNN) layer to jointly model local and global relationships between words and their linguistic roles in multiple communities. Intuitively, the MC-GNN layer substitutes a self-attention layer at the encoder side of a transformer-based machine translation model. Extensive experiments on four language-pair datasets with common evaluation metrics show the remarkable improvements of our method while reducing the time complexity in very long sentences. Binh P. Nguyen, Long H. B. Nguyen, Dinh Dien |
COLING | 1 |
| 2022 | A lightweight classification of adaptor proteins using transformer networksabstractBACKGROUND: Adaptor proteins play a key role in intercellular signal transduction, and dysfunctional adaptor proteins result in diseases. Understanding its structure is the first step to tackling the associated conditions, spurring ongoing interest in research into adaptor proteins with bioinformatics and computational biology. Our study aims to introduce a small, new, and superior model for protein classification, pushing the boundaries with new machine learning algorithms. RESULTS: We propose a novel transformer based model which includes convolutional block and fully connected layer. We input protein sequences from a database, extract PSSM features, then process it via our deep learning model. The proposed model is efficient and highly compact, achieving state-of-the-art performance in terms of area under the receiver operating characteristic curve, Matthew's Correlation Coefficient and Receiver Operating Characteristics curve. Despite merely 20 hidden nodes translating to approximately 1% of the complexity of previous best known methods, the proposed model is still superior in results and computational efficiency. CONCLUSIONS: The proposed model is the first transformer model used for recognizing adaptor protein, and outperforms all existing methods, having PSSM profiles as inputs that comprises convolutional blocks, transformer and fully connected layers for the use of classifying adaptor proteins. Sylwan Rahardja, Mou Wang, Binh P. Nguyen, Pasi Fränti, Susanto Rahardja |
BMC Bioinform. | 3 |
| 2022 | Bone age assessment and sex determination using transfer learning
Quang H. Nguyen 0001, Binh P. Nguyen, Minh T. Nguyen, Matthew Chua 0001, Trang T. T. Do 0001, Nhung Nghiem |
Expert Syst. Appl. | 2 |
| 2021 | Robo-advisor using genetic algorithm and BERT sentiments from tweets for hybrid portfolio optimisation
Edmund Kwong Wei Leow, Binh P. Nguyen, Matthew Chua 0001 |
Expert Syst. Appl. | 2 |
| 2021 | Prediction of FMN Binding Sites in Electron Transport Chains Based on 2-D CNN and PSSM ProfilesabstractFlavin mono-nucleotides (FMNs) are cofactors that hold responsibility for carrying and transferring electrons in the electron transport chain stage of cellular respiration. Without being facilitated by FMNs, energy production is stagnant due to the interruption in most of the cellular processes. Investigation on FMN's functions, therefore, can gain holistic understanding about human diseases and molecular information on drug targets. We proposed a deep learning model using a two-dimensional convolutional neural network and position specific scoring matrices that could identify FMN interacting residues with the sensitivity of 83.7 percent, specificity of 99.2 percent, accuracy of 98.2 percent, and Matthews correlation coefficients of 0.85 for an independent dataset containing 141 FMN binding sites and 1,920 non-FMN binding sites. The proposed method outperformed other previous studies using similar evaluation metrics. Our positive outcome can also promote the utilization of deep learning in dealing with various problems in bioinformatics and computational biology. Nguyen-Quoc-Khanh Le, Binh P. Nguyen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2020 | Automated Grading in Diabetic Retinopathy Using Image Processing and Modified EfficientNet
Hung N. Pham, Renjie Tan, Yu Tian Cai, Shahril Mustafa, Ngan Chong Yeo, Hui Juin Lim, Trang T. T. Do 0001, Binh P. Nguyen, Matthew Chua 0001 |
ICCCI | 8 |
| 2020 | High-content image generation for drug discovery using generative adversarial networks
Shaista Hussain, Ayesha Anees, Ankit Das, Binh P. Nguyen, Mardiana Marzuki, Shuping Lin, Graham Wright, Amit Singhal 0003 |
Neural Networks | 4 |
| 2019 | DeLHCA: Deep transfer learning for high-content analysis of the effects of drugs on immune cellsabstractAnalysis of high-content screening (HCS) data mostly relies on supervised machine learning based approaches employing user-defined image features. This strategy has limited applications due to the requirement of a priori knowledge of expected cellular phenotypes / perturbations and the time-consuming process of manually annotating these phenotypes. To address these issues, we propose a machine learning based unsupervised framework for high-content analysis. The framework performs anomaly detection using features transferred from natural images to the cellular images by deep learning models. We applied this framework to detect anomalous effects of FDA approved drugs on human monocytic cells. Drug anomaly detection based on image features derived using three deep learning architectures, DenseNet-121, ResNet-50 and VGG-16, is compared with the anomaly scores computed from user-defined features extracted from individually segmented cells. The drug anomaly scores of automatically extracted deep features and user-defined features were found to be comparable. Our method has broad implications for faster and reliable analysis of high-content data with limited human interaction which can provide new biological insights and identification of drug candidates for repurposing of FDA approved drugs for new clinical conditions. Shaista Hussain, Ankit Das, Binh P. Nguyen, Mardiana Marzuki, Shuping Lin, Arun Kumar 0006, Graham Wright, Amit Singhal 0003 |
TENCON | 3 |
| 2019 | iProDNA-CapsNet: identifying protein-DNA binding residues using capsule neural networksabstractBACKGROUND: Since protein-DNA interactions are highly essential to diverse biological events, accurately positioning the location of the DNA-binding residues is necessary. This biological issue, however, is currently a challenging task in the age of post-genomic where data on protein sequences have expanded very fast. In this study, we propose iProDNA-CapsNet - a new prediction model identifying protein-DNA binding residues using an ensemble of capsule neural networks (CapsNets) on position specific scoring matrix (PSMM) profiles. The use of CapsNets promises an innovative approach to determine the location of DNA-binding residues. In this study, the benchmark datasets introduced by Hu et al. (2017), i.e., PDNA-543 and PDNA-TEST, were used to train and evaluate the model, respectively. To fairly assess the model performance, comparative analysis between iProDNA-CapsNet and existing state-of-the-art methods was done. RESULTS: Under the decision threshold corresponding to false positive rate (FPR) ≈ 5%, the accuracy, sensitivity, precision, and Matthews's correlation coefficient (MCC) of our model is increased by about 2.0%, 2.0%, 14.0%, and 5.0% with respect to TargetDNA (Hu et al., 2017) and 1.0%, 75.0%, 45.0%, and 77.0% with respect to BindN+ (Wang et al., 2010), respectively. With regards to other methods not reporting their threshold settings, iProDNA-CapsNet also shows a significant improvement in performance based on most of the evaluation metrics. Even with different patterns of change among the models, iProDNA-CapsNets remains to be the best model having top performance in most of the metrics, especially MCC which is boosted from about 8.0% to 220.0%. CONCLUSIONS: According to all evaluation metrics under various decision thresholds, iProDNA-CapsNet shows better performance compared to the two current best models (BindN and TargetDNA). Our proposed approach also shows that CapsNet can potentially be used and adopted in other biological applications. Binh P. Nguyen, Quang H. Nguyen 0001, Giang-Nam Doan-Ngoc, Thanh-Hoang Nguyen-Vo, Susanto Rahardja |
BMC Bioinform. | 1 |
| 2017 | A two-level clustering approach for multidimensional transfer function specification in volume visualization
Lile Cai, Binh P. Nguyen, Chee-Kong Chui, Sim Heng Ong |
Vis. Comput. | 2 |
| 2016 | Automated brain tumor segmentation using kernel dictionary learning and superpixel-level featuresabstractBrain tumor segmentation, an essential but challenging task, has long attracted much attention from the medical imaging community. Recently, successful applications of sparse coding and dictionary learning has emerged in various vision problems including image segmentation. In this paper, a superpixel-based framework for automated brain tumor segmentation is introduced. The kernel trick is adopted in dictionary learning to transform superpixel-level features to a high-dimensional feature space where their nonlinear similarities are considered to generate discriminative sparse codes. A graph is constructed from the approximation errors given by dictionaries modeling different brain tumor structures so that superpixels belonging to particular tumor regions can be efficiently identified. The proposed framework is evaluated on brain magnetic resonance images of high-grade glioma (HGG) patients provided by the multi-modal Brain Tumor Segmentation (BRATS) Benchmark. Results show that the proposed framework achieves competitive performance when compared with the state-of-the-art methods. Binh P. Nguyen, Chee-Kong Chui, Sim Heng Ong |
SMC | 2 |
| 2016 | Comments on and corrections to 'Hardware-software co-design architecture for joint photo expert graphic XR encoder'abstractIn the aforementioned paper, Tseng and Lai presented a hardware‐sharing design of elementary transform operations for the photo overlap transform and the photo core transform in JPEG XR. In this letter, we point out some errors in their work and suggest the corresponding corrections. Trang T. T. Do 0001, Binh P. Nguyen |
IET Image Process. | 2 |
| 2015 | Rule-Enhanced Transfer Function Generation for Medical Volume VisualizationabstractAbstract In volume visualization, transfer functions are used to classify the volumetric data and assign optical properties to the voxels. In general, transfer functions are generated in a transfer function space, which is the feature space constructed by data values and properties derived from the data. If volumetric objects have the same or overlapping data values, it would be difficult to separate them in the transfer function space. In this paper, we present a rule‐enhanced transfer function design method that allows important structures of the volume to be more effectively separated and highlighted. We define a set of rules based on the local frequency distribution of volume attributes. A rule‐selection method based on a genetic algorithm is proposed to learn the set of rules that can distinguish the user‐specified target tissue from other tissues. In the rendering stage, voxels satisfying these rules are rendered with higher opacities in order to highlight the target tissue. The proposed method was tested on various volumetric datasets to enhance the visualization of important structures that are difficult to be visualized by traditional transfer function design methods. The results demonstrate the effectiveness of the proposed method. Lile Cai, Binh P. Nguyen, Chee-Kong Chui, Sim Heng Ong |
Comput. Graph. Forum | 2 |
| 2015 | Robust Biometric Recognition From Palm Depth Images for Gloved HandsabstractBiometric recognition can be used to improve gesture-based interfaces by automatically identifying operators. Traditional palm biometric recognition techniques depend on palm appearance features, but these features are not available in an operating theater where gloves are worn. We propose a depth-based solution for palm biometric recognition. Based on the depth image, our system automatically segments the user's palm and extracts finger dimensions. The finger dimensions are further scaled according to the sensed depth to obtain the true finger dimensions, which are then used as features to characterize the palm. Finally, a modified$k$-nearest neighbors algorithm that assigns class labels based on the centroid displacement of each class in the neighboring points is applied to recognize the palm based on the geometric features. An accuracy of 96.24% was achieved for the biometric recognition of 4057 gloved palm samples captured at different angles and depths from 27 users. This accuracy is comparable with those of other state-of-the-art classification algorithms and demonstrates that biometric recognition may be viable for settings with gloved hands such as surgery. Binh P. Nguyen, Wei-Liang Tay, Chee-Kong Chui |
IEEE Trans. Hum. Mach. Syst. | 1 |
| 2013 | Near-Field Communication Transceiver System Modeling and Analysis Using SystemC/SystemC-AMS With the Consideration of Noise IssuesabstractSystemC, as a C++-based hardware description language, is used for system architecture design, large digital hardware, software, and their interaction. Its extension, SystemC-AMS, provides the capability of abstract modeling to deliver analog system-level simulation of “real-time” application scenarios. SystemC and SystemC-AMS help designers to analyze a whole mixed-signal system and further guide the circuit design to reduce the design cost. This paper presents SystemC (2.2.0) and SystemC-AMS (1.0 Beta2) modeling of a near-field communication (NFC) system working in passive mode, based on the proximity contactless identification cards ISO/IEC 14443 international standard. The NFC transceiver system includes reader and card analog blocks, digital blocks, and antennas. Problems caused by realistic imperfections are considered, simulated, and then solved by modifying the design at a system level, which is significant to high-level modeling. Systematic simulation is given to prove SystemC/SystemC-AMS is an accurate and efficient tool to model a heterogeneous mixed-signal system in an early-design stage. Dian Zhou, Minghua Li, Binh P. Nguyen, Xuan Zeng 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2012 | A clustering-based system to automate transfer function design for medical image visualization
Binh P. Nguyen, Wei-Liang Tay, Chee-Kong Chui, Sim Heng Ong |
Vis. Comput. | 1 |
| 2010 | An efficient clustering method for fast rendering of time-varying volumetric medical data
Zhenlan Wang, Binh P. Nguyen, Chee-Kong Chui, Harry Qin, Chuan-Heng Ang, Sim Heng Ong |
Vis. Comput. | 2 |