Viet Cuong Nguyen

dblp:36/9125 · also Cuong V. Nguyen, Nguyen Viet Cuong · DBLP profile ↗
← Back
27ranked-venue papers
16as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 11 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Theory of computation · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DEGRE: Dynamic Gating Ensembles for Trust-Aware Rejection in Medical Image Diagnostics
abstract
For artificial intelligence to be safely deployed in high-risk domains, it must reliably know its limits. Selective prediction, or learning with a reject option, addresses this by enabling a model to abstain from prediction on inputs it deems unreliable, deferring them to a human expert. While deep ensembles have emerged as a leading approach for uncertainty estimation, their potential is often squandered by rejection methods that rely on static thresholds applied to the mean prediction. In this paper, we propose to learn a dynamic rejection policy directly from the rich behavioral signals of the ensemble itself. Our framework, DEGRE (Dynamic Ensembles Gating for REjection), is a novel meta-learning approach that trains a lightweight gating network on the ensemble’s consensus confidence and its internal disagreement (variance)— to explicitly discriminate between correct and incorrect predictions. Through rigorous evaluation across twelve diverse medical imaging benchmarks (MRI, X-ray, CT), DEGRE significantly advances selective prediction, achieving an average risk-coverage (AURC) reduction of 68.2% compared to the standard ensemble baseline. By providing a more reliable method for a model to recognize its own limitations, this learned, adaptive rejection mechanism paves the way for safer and more responsible integration of AI into critical clinical workflows.
Hong-Hai Nguyen, Duong Bach, Nam Phan, Viet Cuong Nguyen, Cuong Do 0001
AAAI4
2025 Supporters and Skeptics: LLM-Based Analysis of Engagement with Mental Health (Mis)Information Content on Video-Sharing Platforms
abstract
Over one in five adults in the US lives with a mental illness. In the face of a shortage of mental health professionals and offline resources, online short-form video content has grown to serve as a crucial conduit for disseminating mental health help and resources. However, the ease of content creation and access also contributes to the spread of misinformation, posing risks to accurate diagnosis and treatment. Detecting and understanding engagement with such content is crucial to mitigating their harmful effects on public health. We perform the first quantitative study of the phenomenon using YouTube Shorts and Bitchute as the sites of study. We contribute MentalMisinfo, a novel labeled mental health misinformation (MHMisinfo) dataset of 739 videos (639 from Youtube and 100 from Bitchute) and 135372 comments in total, using an expert-driven annotation schema. We first found that few-shot in-context learning with large language models (LLMs) are effective in detecting MHMisinfo videos. Next, we discover distinct and potentially alarming linguistic patterns in how audiences engage with MHMisinfo videos through commentary on both video-sharing platforms. Across the two platforms, comments could exacerbate prevailing stigma with some groups showing heightened susceptibility to and alignment with MHMisinfo. We discuss technical and public health-driven adaptive solutions to tackling the "epidemic" of mental health misinformation online.
Viet Cuong Nguyen, Mini Jain, Abhijat Chauhan, Heather Jaime Soled, Santiago Alvarez Lesmes, Michael L. Birnbaum, Sunny X. Tang, Srijan Kumar, Munmun De Choudhury
ICWSM1
2025 Fake advertisements detection using automated multimodal learning: a case study for Vietnamese real estate data
abstract
Abstract The popularity of e-commerce has given rise to fake advertisements that can expose users to financial and data risks while damaging the reputation of these e-commerce platforms. For these reasons, detecting and removing such fake advertisements are important for the success of e-commerce websites. In this paper, we propose FADAML, a novel end-to-end machine learning system to detect and filter out fake online advertisements. Our system combines techniques in multimodal machine learning and automated machine learning to achieve a high detection rate. As a case study, we apply FADAML to detect fake advertisements on popular Vietnamese real estate websites. Our experiments show that we can achieve 91.5% detection accuracy, which significantly outperforms three different state-of-the-art fake news detection systems.
Trung T. Nguyen, Viet Cuong Nguyen
Appl. Intell.3
2024 Patient Perspectives on AI-Driven Predictions of Schizophrenia Relapses: Understanding Concerns and Opportunities for Self-Care and Treatment
abstract
Early detection and intervention for relapse is important in the treatment of schizophrenia spectrum disorders. Researchers have developed AI models to predict relapse from patient-contributed data like social media. However, these models face challenges, including misalignment with practice and ethical issues related to transparency, accountability, and potential harm. Furthermore, how patients who have recovered from schizophrenia view these AI models has been underexplored. To address this gap, we first conducted semi-structured interviews with 28 patients and reflexive thematic analysis, which revealed a disconnect between AI predictions and patient experience, and the importance of the social aspect of relapse detection. In response, we developed a prototype that used patients' Facebook data to predict relapse. Feedback from seven patients highlighted the potential for AI to foster collaboration between patients and their support systems, and to encourage self-reflection. Our work provides insights into human-AI interaction and suggests ways to empower people with schizophrenia.
Dong Whi Yoo, Hayoung Woo, Viet Cuong Nguyen, Michael L. Birnbaum, Kaylee Payne Kruzan, Jennifer G. Kim, Gregory D. Abowd, Munmun De Choudhury
CHI3
2024 Hamiltonian Monte Carlo on ReLU Neural Networks is Inefficient
abstract
We analyze the error rates of the Hamiltonian Monte Carlo algorithm with leapfrog integrator for Bayesian neural network inference. We show that due to the non-differentiability of activation functions in the ReLU family, leapfrog HMC for networks with these activation functions has a large local error rate of $\Omega(\epsilon)$ rather than the classical error rate of $\mathcal{O}(\epsilon^3)$. This leads to a higher rejection rate of the proposals, making the method inefficient. We then verify our theoretical findings through empirical simulations as well as experiments on a real-world dataset that highlight the inefficiency of HMC inference on ReLU-based neural networks compared to analytical networks.
Vu C. Dinh, Lam S. Ho, Viet Cuong Nguyen
NeurIPS3
2024 The Matter of Captchas: An Analysis of a Brittle Security Feature on the Modern Web
abstract
The web ecosystem is a fast-paced environment. In this dynamic landscape, new security features are offered one after another to enhance the security and robustness of web applications and the operations they handle. This paper focuses on a fragile but still in-use security feature, text-based CAPTCHAs, that had been wildly used by web applications in the past to protect against automated attacks such as credential stuffing and account hijacking. The paper first investigates what it takes to develop automated scanners that can solve previously unseen text-based CAPTCHAs. We evaluated the possibility of developing and integrating a pre-trained CAPTCHA solver in the automated web scanning process without using a significantly large training dataset. We also perform an analysis of the impact of such autonomous scanners on CAPTCHA-enabled websites. Our analysis shows that solvable text-based CAPTCHAs on login, contact, and comment pages of websites are not uncommon. In particular, we identified over 3,100 text-based CAPTCHA websites in critical sectors such as finance, government, and health with hundreds of thousands of users. We showed that a web scanner with a pre-trained solver could solve more than 20% of previously unseen CAPTCHAs in just one single attempt. This result is worrisome considering the substantial potential to autonomously run the operation across thousands of websites on a daily basis with minimal training. The findings suggest that the integration of autonomous scanning with pre-training and local optimization of models can significantly increase adversaries' asymmetric power to launch their attacks cheaper and faster.
Behzad Ousat, Esteban Schafir, Duc C. Hoang, Mohammad Ali Tofighi, Viet Cuong Nguyen, Sajjad Arshad, A. Selcuk Uluagac, Amin Kharraz
WWW5
2023 Simple Transferability Estimation for Regression Tasks
abstract
We consider transferability estimation, the problem of estimating how well deep learning models transfer from a source to a target task. We focus on regression tasks, which received little previous attention, and propose two simple and computationally efficient approaches that estimate transferability based on the negative regularized mean squared error of a linear regression model. We prove novel theoretical results connecting our approaches to the actual transferability of the optimal target models obtained from the transfer learning process. Despite their simplicity, our approaches significantly outperform existing state-of-the-art regression transferability estimators in both accuracy and efficiency. On two large-scale keypoint regression benchmarks, our approaches yield 12% to 36% better results on average while being at least 27% faster than previous state-of-the-art methods.
Cuong N. Nguyen, Lam Si Tung Ho, Vu C. Dinh, Anh T. Tran, Tal Hassner, Viet Cuong Nguyen
UAI7
2022 Generalization Bounds for Deep Transfer Learning Using Majority Predictor Accuracy
Cuong N. Nguyen, Lam Si Tung Ho, Vu C. Dinh, Tal Hassner, Viet Cuong Nguyen
ISITA5
2022 Bayesian active learning with abstention feedbacks
Viet Cuong Nguyen, Lam Si Tung Ho, Huan Xu 0001, Vu C. Dinh, Binh T. Nguyen 0001
Neurocomputing1
2022 Learning for amalgamation: A multi-source transfer learning framework for sentiment classification
Viet Cuong Nguyen, Khiem H. Le, Anh M. Tran, Quang Hong Pham, Binh T. Nguyen 0001
Inf. Sci.1
2021 Multimodal Machine Learning for Credit Modeling
abstract
Credit ratings are traditionally generated using models that use financial statement data and market data, which is tabular (numeric and categorical). Practitioner and academic models do not include text data. Using an automated approach to combine long-form text from SEC filings with the tabular data, we show how multimodal machine learning using stack ensembling and bagging can generate more accurate rating predictions. This paper demonstrates a methodology to use big data to extend tabular data models, which have been used by the ratings industry for decades, to the class of multimodal machine learning models.
Viet Cuong Nguyen, Sanjiv R. Das, John He, Shenghua Yue, Vinay Hanumaiah, Xavier Ragot
COMPSAC1
2021 A Novel Approach for Enhancing Vietnamese Sentiment Classification
Viet Cuong Nguyen, Khiem H. Le, Binh T. Nguyen 0001
IEA/AIE (2)1
2021 DRL-Based Intelligent Resource Allocation for Diverse QoS in 5G and toward 6G Vehicular Networks: A Comprehensive Survey
abstract
The vehicular network is taking great attention from both academia and industry to enable the intelligent transportation system (ITS), autonomous driving, and smart cities. The system provides extremely dynamic features due to the fast mobile characteristics. While the number of different applications in the vehicular network is growing fast, the quality of service (QoS) in the 5G vehicular network becomes diverse. One of the most stringent requirements in the vehicular network is a safety‐critical real‐time system. To guarantee low‐latency and other diverse QoS requirements, wireless network resources should be effectively utilized and allocated among vehicles, such as computation power in cloud, fog, and edge servers; spectrum at roadside units (RSUs); and base stations (BSs). Historically, optimization problems have mostly been investigated to formulate resource allocation and are solved by mathematical computation methods. However, the optimization problems are usually nonconvex and hard to be solved. Recently, machine learning (ML) is a powerful technique to cope with the complexity in computation and has capability to cope with big data and data analysis in the heterogeneous vehicular network. In this paper, an overview of resource allocation in the 5G vehicular network is represented with the support of traditional optimization and advanced ML approaches, especially a deep reinforcement learning (DRL) method. In addition, a federated deep reinforcement learning‐ (FDRL‐) based vehicular communication is proposed. The challenges, open issues, and future research directions for 5G and toward 6G vehicular networks, are discussed. A multiaccess edge computing assisted by network slicing and a distributed federated learning (FL) technique is analyzed. A FDRL‐based UAV‐assisted vehicular communication is discussed to point out the future research directions for the networks.
Hoa Tt. Nguyen, Hai T. Do, Hoang T. Hua, Viet Cuong Nguyen
Wirel. Commun. Mob. Comput.5
2020 LEEP: A New Measure to Evaluate Transferability of Learned Representations
abstract
We introduce a new measure to evaluate the transferability of representations learned by classifiers. Our measure, the Log Expected Empirical Prediction (LEEP), is simple and easy to compute: when given a classifier trained on a source data set, it only requires running the target data set through this classifier once. We analyze the properties of LEEP theoretically and demonstrate its effectiveness empirically. Our analysis shows that LEEP can predict the performance and convergence speed of both transfer and meta-transfer learning methods, even for small or imbalanced data. Moreover, LEEP outperforms recently proposed transferability measures such as negative conditional entropy and H scores. Notably, when transferring from ImageNet to CIFAR100, LEEP can achieve up to 30% improvement compared to the best competing method in terms of the correlations with actual transfer accuracy.
Viet Cuong Nguyen, Tal Hassner, Matthias W. Seeger, Cédric Archambeau
ICML1
2020 An Efficient Framework for Vietnamese Sentiment Classification
abstract
With the booming development of E-commerce platforms in many counties, there is a massive amount of customers’ review data in different products and services. Understanding customers’ feedbacks in both current and new products can give online retailers the possibility to improve the product quality, meet customers’ expectations, and increase the corresponding revenue. In this paper, we investigate the Vietnamese sentiment classification problem on two datasets containing Vietnamese customers’ reviews. We propose eight different approaches, including Bi-LSTM, Bi-LSTM + Attention, Bi-GRU, Bi-GRU + Attention, Recurrent CNN, Residual CNN, Transformer, and PhoBERT, and conduct all experiments on two datasets, AIVIVN 2019 and our dataset self-collected from multiple Vietnamese e-commerce websites. The experimental results show that all our proposed methods outperform the winning solution of the competition “AIVIVN 2019 Sentiment Champion” with a significant margin. Especially, Recurrent CNN has the best performance in comparison with other algorithms in terms of both AUC (98.48%) and F1-score (93.42%) in this competition dataset and also surpasses other techniques in our dataset collected. Finally, we aim to publish our codes, and these two data-sets later to contribute to the current research community related to the field of sentiment analysis.
Viet Cuong Nguyen, Khiem H. Le, Anh M. Tran, Binh T. Nguyen 0001
SoMeT1
2020 The Hybrid Solar-RF Energy for Base Transceiver Stations
abstract
The base transceiver stations (BTS) are telecom infrastructures that facilitate wireless communication between the subscriber device and the telecom operator networks. They are deployed in suitable places having a lot of freely propagating ambient radio frequency (RF) and solar energies. This paper is aimed at converting received ambient environmental energy into usable electricity to power the stations. We proposed a hybrid energy harvesting system that can collect energy from RF and solar energies at the same time. The sources are combined to provide to a significant amount, to contribute to operational expenditures that reduce energy costs, and to improve the energy efficiency of the base station sites in rural areas from the most common renewable resources since the base stations are major consumers of cellular networks. The hybrid systems are designed with circuits, simulated, and compared to show their good performance to the base stations. PSIM, PROTEUS, and MATLAB software are used to simulate for evaluating the voltage and the current output of the hybrid systems that meet the power requirements. The design and simulation results show the feasibility of our proposed method with the battery storage that can be deployed not only in real base stations but also for other electrical operated systems.
Viet Cuong Nguyen, Toan V. Quyen, Anh M. Le, Linh H. Truong
Wirel. Commun. Mob. Comput.1
2019 Transferability and Hardness of Supervised Classification Tasks
abstract
We propose a novel approach for estimating the difficulty and transferability of supervised classification tasks. Unlike previous work, our approach is solution agnostic and does not require or assume trained models. Instead, we estimate these values using an information theoretic approach: treating training labels as random variables and exploring their statistics. When transferring from a source to a target task, we consider the conditional entropy between two such variables (i.e., label assignments of the two tasks). We show analytically and empirically that this value is related to the loss of the transferred model. We further show how to use this value to estimate task hardness. We test our claims extensively on three large scale data sets - CelebA (40 tasks), Animals with Attributes 2 (85 tasks), and Caltech-UCSD Birds 200 (312 tasks) - together representing 437 classification tasks. We provide results showing that our hardness and transferability estimates are strongly correlated with empirical hardness and transferability. As a case study, we transfer a learned face recognition model to CelebA attribute classification tasks, showing state of the art accuracy for tasks estimated to be highly transferable.
Anh Tuan Tran 0001, Viet Cuong Nguyen, Tal Hassner
ICCV2
2018 Variational Continual Learning
Viet Cuong Nguyen, Yingzhen Li, Thang D. Bui, Richard E. Turner
ICLR (Poster)1
2017 Streaming Sparse Gaussian Process Approximations
abstract
Sparse pseudo-point approximations for Gaussian process (GP) models provide a suite of methods that support deployment of GPs in the large data regime and enable analytic intractabilities to be sidestepped. However, the field lacks a principled method to handle streaming data in which both the posterior distribution over function values and the hyperparameter estimates are updated in an online fashion. The small number of existing approaches either use suboptimal hand-crafted heuristics for hyperparameter learning, or suffer from catastrophic forgetting or slow updating when new data arrive. This paper develops a new principled framework for deploying Gaussian process probabilistic models in the streaming setting, providing methods for learning hyperparameters and optimising pseudo-input locations. The proposed framework is assessed using synthetic and real-world datasets.
Thang D. Bui, Viet Cuong Nguyen, Richard E. Turner
NIPS2
2016 Robustness of Bayesian Pool-Based Active Learning Against Prior Misspecification
abstract
We study the robustness of active learning (AL) algorithms against prior misspecification: whether an algorithm achieves similar performance using a perturbed prior as compared to using the true prior. In both the average and worst cases of the maximum coverage setting, we prove that all alpha-approximate algorithms are robust (i.e., near alpha-approximate) if the utility is Lipschitz continuous in the prior. We further show that robustness may not be achieved if the utility is non-Lipschitz. This suggests we should use a Lipschitz utility for AL if robustness is required. For the minimum cost setting, we can also obtain a robustness result for approximate AL algorithms. Our results imply that many commonly used AL algorithms are robust against perturbed priors. We then propose the use of a mixture prior to alleviate the problem of prior misspecification. We analyze the robustness of the uniform mixture prior and show experimentally that it performs reasonably well in practice.
Viet Cuong Nguyen, Wee Sun Lee
AAAI1
2015 Learning from Non-iid Data: Fast Rates for the One-vs-All Multiclass Plug-in Classifiers
Vu C. Dinh, Lam Si Tung Ho, Viet Cuong Nguyen, Duy M. H. Nguyen, Binh T. Nguyen 0001
TAMC3
2014 Near-optimal Adaptive Pool-based Active Learning with General Loss
Viet Cuong Nguyen, Wee Sun Lee
UAI1
2014 Conditional random field with high-order dependencies for sequence labeling and segmentation
Viet Cuong Nguyen, Wee Sun Lee, Hai Leong Chieu
J. Mach. Learn. Res.1
2013 Generalization and Robustness of Batched Weighted Average Algorithm with V-Geometrically Ergodic Markov Data
Viet Cuong Nguyen, Lam Si Tung Ho, Vu C. Dinh
ALT1
2013 Active Learning for Probabilistic Hypotheses Using the Maximum Gibbs Error Criterion
abstract
We introduce a new objective function for pool-based Bayesian active learning with probabilistic hypotheses. This objective function, called the policy Gibbs error, is the expected error rate of a random classifier drawn from the prior distribution on the examples adaptively selected by the active learning policy. Exact maximization of the policy Gibbs error is hard, so we propose a greedy strategy that maximizes the Gibbs error at each iteration, where the Gibbs error on an instance is the expected error of a random classifier selected from the posterior label distribution on that instance. We apply this maximum Gibbs error criterion to three active learning scenarios: non-adaptive, adaptive, and batch active learning. In each scenario, we prove that the criterion achieves near-maximal policy Gibbs error when constrained to a fixed budget. For practical implementations, we provide approximations to the maximum Gibbs error criterion for Bayesian conditional random fields and transductive Naive Bayes. Our experimental results on a named entity recognition task and a text classification task show that the maximum Gibbs error criterion is an effective active learning criterion for noisy models.
Viet Cuong Nguyen, Wee Sun Lee, Kian Ming A. Chai, Hai Leong Chieu
NIPS1
2012 Mel-frequency Cepstral Coefficients for Eye Movement Identification
abstract
Human identification is an important task for various activities in society. In this paper, we consider the problem of human identification using eye movement information. This problem, which is usually called the eye movement identification problem, can be solved by training a multiclass classification model to predict a person's identity from his or her eye movements. In this work, we propose using Mel-frequency cepstral coefficients (MFCCs) to encode various features for the classification model. Our experiments show that using MFCCs to represent useful features such as eye position, eye difference, and eye velocity would result in a much better accuracy than using Fourier transform, cepstrum, or raw representations. We also compare various classification models for the task. From our experiments, linear-kernel SVMs achieve the best accuracy with 93.56% and 91.08% accuracy on the small and large datasets respectively. Besides, we conduct experiments to study how the movements of each eye contribute to the final classification accuracy.
Viet Cuong Nguyen, Vu C. Dinh, Lam Si Tung Ho
ICTAI1
2011 Improving Text Segmentation with Non-systematic Semantic Relation
Viet Cuong Nguyen, Minh Le Nguyen 0001, Akira Shimazu
CICLing (1)1