VLDB 2026 Research / reviewers in the wild / expert
Tuan Dung Nguyen
dblp:45/3372
· DBLP profile ↗
12ranked-venue papers
3as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Efficient and distributed learning · 52% Vision and language · 17% Information extraction and text analysis · 15% | |
| Theoretical computer science
2 papers |
Mathematical optimization · 100% | |
| Computer networks
2 papers |
Internet of things and sensor networks · 45% Network measurement and analytics · 45% Edge and fog computing · 10% | |
| Network and information security
1 paper |
Security and privacy of machine learning · 50% Privacy and data protection · 50% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational social science and digital humanities · 100% |
Topics — the 21 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
federated learning |
1.2 | 3 | 2024 | DONE: Distributed Approximate Newton-type Method for Federated Edge Learning · IEEE Trans. Parallel Distributed Syst. 2022 Personalized Federated Learning with Moreau Envelopes · NeurIPS 2020 Federated PCA on Grassmann Manifold for IoT Anomaly Detection · IEEE/ACM Trans. Netw. 2024 |
Computer vision › Vision and language
medical report generation |
1.0 | 1 | 2026 | Region-Grounded Report Generation for 3D Medical Imaging: A Fine-Grained Dataset and Graph-Enhanced Framework · ACL (1) 2026 |
Natural language and speech › Information extraction and text analysis
text classification |
0.9 | 1 | 2025 | MoVa: Towards Generalizable Classification of Human Morals and Values · EMNLP 2025 |
Network measurement and analytics
anomaly detection |
0.8 | 1 | 2024 | Federated PCA on Grassmann Manifold for IoT Anomaly Detection · IEEE/ACM Trans. Netw. 2024 |
Internet of things and sensor networks
iot security |
0.8 | 1 | 2024 | Federated PCA on Grassmann Manifold for IoT Anomaly Detection · IEEE/ACM Trans. Netw. 2024 |
Security and privacy of machine learning
federated learning |
0.8 | 1 | 2024 | Federated PCA on Grassmann Manifold for IoT Anomaly Detection · IEEE/ACM Trans. Netw. 2024 |
Privacy and data protection › privacy-preserving data analysis
privacy-preserving anomaly detection |
0.8 | 1 | 2024 | Federated PCA on Grassmann Manifold for IoT Anomaly Detection · IEEE/ACM Trans. Netw. 2024 |
Mathematical optimization › continuous optimization › convex optimization
first-order methods |
0.8 | 1 | 2024 | On Partial Optimal Transport: Revising the Infeasibility of Sinkhorn and Efficient Gradient Methods · AAAI 2024 |
Mathematical optimization
gradient descent |
0.8 | 1 | 2024 | On Partial Optimal Transport: Revising the Infeasibility of Sinkhorn and Efficient Gradient Methods · AAAI 2024 |
Mathematical optimization
optimal transport |
0.8 | 1 | 2024 | On Partial Optimal Transport: Revising the Infeasibility of Sinkhorn and Efficient Gradient Methods · AAAI 2024 |
Mathematical optimization › optimal transport
partial optimal transport |
0.8 | 1 | 2024 | On Partial Optimal Transport: Revising the Infeasibility of Sinkhorn and Efficient Gradient Methods · AAAI 2024 |
Machine learning › Efficient and distributed learning › federated learning
communication-efficient federated learning |
0.6 | 1 | 2022 | DONE: Distributed Approximate Newton-type Method for Federated Edge Learning · IEEE Trans. Parallel Distributed Syst. 2022 |
Machine learning › Efficient and distributed learning › federated learning
federated edge learning |
0.6 | 1 | 2022 | DONE: Distributed Approximate Newton-type Method for Federated Edge Learning · IEEE Trans. Parallel Distributed Syst. 2022 |
Mathematical optimization › distributed optimization
distributed newton-type methods |
0.6 | 1 | 2022 | DONE: Distributed Approximate Newton-type Method for Federated Edge Learning · IEEE Trans. Parallel Distributed Syst. 2022 |
Mathematical optimization
distributed optimization |
0.6 | 1 | 2022 | DONE: Distributed Approximate Newton-type Method for Federated Edge Learning · IEEE Trans. Parallel Distributed Syst. 2022 |
Machine learning › Optimization for machine learning
convergence analysis |
0.4 | 1 | 2020 | Personalized Federated Learning with Moreau Envelopes · NeurIPS 2020 |
Machine learning › Efficient and distributed learning › federated learning
personalized federated learning |
0.4 | 1 | 2020 | Personalized Federated Learning with Moreau Envelopes · NeurIPS 2020 |
Computer vision › 3D vision
3d scene understanding |
0.3 | 1 | 2026 | Region-Grounded Report Generation for 3D Medical Imaging: A Fine-Grained Dataset and Graph-Enhanced Framework · ACL (1) 2026 |
Machine learning › Efficient and distributed learning › distributed statistical learning
distributed PCA |
0.2 | 1 | 2024 | Federated PCA on Grassmann Manifold for IoT Anomaly Detection · IEEE/ACM Trans. Netw. 2024 |
Machine learning › Transfer learning and domain adaptation
domain adaptation |
0.2 | 1 | 2024 | On Partial Optimal Transport: Revising the Infeasibility of Sinkhorn and Efficient Gradient Methods · AAAI 2024 |
Edge and fog computing › distributed learning › federated learning
federated edge learning |
0.2 | 1 | 2022 | DONE: Distributed Approximate Newton-type Method for Federated Edge Learning · IEEE Trans. Parallel Distributed Syst. 2022 |
Methods — techniques the papers use, named apart from their topics
principal component analysis · 2.3grassmann manifold · 2.3alternating direction method of multipliers · 2.3richardson iteration · 1.7approximate newton method · 1.7sinkhorn algorithm · 1.5rounding algorithm · 1.5primal-dual accelerated gradient descent · 1.5dual extrapolation · 1.5graph neural network · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Region-Grounded Report Generation for 3D Medical Imaging: A Fine-Grained Dataset and Graph-Enhanced FrameworkabstractCong Huy Nguyen, Son Dinh Nguyen, Guanlin Li, Tuan Dung Nguyen, Aditya Narayan Sankaran, Mai Huy Thong, Thanh Trung Nguyen, Mai Hong Son, Reza Farahbakhsh, Phi Le Nguyen, Noel Crespi. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Cong Huy Nguyen, Son Dinh Nguyen, Tuan Dung Nguyen, Aditya Narayan Sankaran, Mai Huy Thong, Thanh Trung Nguyen, Mai Hong Son, Reza Farahbakhsh, Phi-Le Nguyen, Noël Crespi |
ACL (1) | 4 |
| 2025 | MoVa: Towards Generalizable Classification of Human Morals and ValuesabstractZiyu Chen, Junfei Sun, Chenxi Li, Tuan Dung Nguyen, Jing Yao, Xiaoyuan Yi, Xing Xie, Chenhao Tan, Lexing Xie. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Junfei Sun, Tuan Dung Nguyen, Jing Yao 0003, Xiaoyuan Yi, Xing Xie 0001, Chenhao Tan, Lexing Xie |
EMNLP | 4 |
| 2025 | Sign Language Recognition: A Large-scale Multi-view Dataset and Comprehensive EvaluationabstractVision-based sign language recognition is an extensively researched problem aimed at advancing communication be-tween deaf and hearing individuals. Numerous Sign Lan-guage Recognition (SLR) datasets have been introduced to promote research in this field, spanning multiple languages, vocabulary sizes, and signers. However, most existing pop-ular datasets focus predominantly on the frontal view of signers, neglecting visual information from other perspec-tives. In practice, many sign languages contain words that have similar hand movements and expressions, making it challenging to differentiate between them from a single frontal view. Although a few studies have proposed sign language datasets using multi-view data, these datasets remain limited in vocabulary size and scale, hindering their gener-alizability and practicality. To address this issue, we in-troduce a new large-scale, multi-view sign language recog-nition dataset spanning 1,000 glosses and 30 signers, re-sulting in over 84,000 multi-view videos. To the best of our knowledge, this is the first multi-view sign language recognition dataset of this scale. In conjunction with of-fering a comprehensive dataset, we perform extensive ex-periments to assess the performance of state-of-the-art Sign Language Recognition models utilizing on our dataset. The findings indicate that utilizing multi-view data substantially enhances model accuracy across all models, with a maxi-mum performance improvement of up to 19.75% compared to models trained on single-view data. Our dataset and baseline models are publicly accessible on GitHub11Available at https://github.com/Etdihatthoc/Multi-VSL_WACV_2025. Nguyen Son Dinh, Tuan Dung Nguyen, Duc Tri Tran, Nguyen Dang Huy Pham, Thuan Hieu Tran, Ngoc Anh Tong, Quang Huy Hoang, Phi-Le Nguyen |
WACV | 2 |
| 2024 | On Partial Optimal Transport: Revising the Infeasibility of Sinkhorn and Efficient Gradient MethodsabstractThis paper studies the Partial Optimal Transport (POT) problem between two unbalanced measures with at most n supports and its applications in various AI tasks such as color transfer or domain adaptation. There is hence a need for fast approximations of POT with increasingly large problem sizes in arising applications. We first theoretically and experimentally investigate the infeasibility of the state-of-the-art Sinkhorn algorithm for POT, which consequently degrades its qualitative performance in real world applications like point-cloud registration. To this end, we propose a novel rounding algorithm for POT, and then provide a feasible Sinkhorn procedure with a revised computation complexity of O(n^2/epsilon^4). Our rounding algorithm also permits the development of two first-order methods to approximate the POT problem. The first algorithm, Adaptive Primal-Dual Accelerated Gradient Descent (APDAGD), finds an epsilon-approximate solution to the POT problem in O(n^2.5/epsilon). The second method, Dual Extrapolation, achieves the computation complexity of O(n^2/epsilon), thereby being the best in the literature. We further demonstrate the flexibility of POT compared to standard OT as well as the practicality of our algorithms on real applications where two marginal distributions are unbalanced. Tuan Dung Nguyen, Quang Minh Nguyen, Hoang Nguyen 0010, Lam M. Nguyen, Kim-Chuan Toh |
AAAI | 2 |
| 2024 | Federated Deep Equilibrium Learning: Harnessing Compact Global Representations to Enhance Personalization
Tuan Dung Nguyen, Tung-Anh Nguyen, Choong Seon Hong, Suranga Seneviratne, Wei Bao 0001, Nguyen Hoang Tran |
CIKM | 2 |
| 2024 | Measuring Moral Dimensions in Social Media with MformerabstractThe ever-growing textual records of contemporary social issues, often discussed online with moral rhetoric, present both an opportunity and a challenge for studying how moral concerns are debated in real life. Moral foundations theory is a taxonomy of intuitions widely used in data-driven analyses of online content, but current computational tools to detect moral foundations suffer from the incompleteness and fragility of their lexicons and from poor generalization across data domains. In this paper, we fine-tune a large language model to measure moral foundations in text based on datasets covering news media and long- and short-form online discussions. The resulting model, called Mformer, outperforms existing approaches on the same domains by 4–12% in AUC and further generalizes well to four commonly used moral text datasets, improving by up to 17% in AUC. We present case studies using Mformer to analyze everyday moral dilemmas on Reddit and controversies on Twitter, showing that moral foundations can meaningfully describe people’s stance on social issues and such variations are topic-dependent. Pretrained model and datasets are released publicly. We posit that Mformer will help the research community quantify moral dimensions for a range of tasks and data domains, and eventually contribute to the understanding of moral situations faced by humans and machines. Tuan Dung Nguyen, Nicholas George Carroll, Alasdair Tran, Colin Klein, Lexing Xie |
ICWSM | 1 |
| 2024 | Hierarchical Federated Learning in MEC Networks with Knowledge DistillationabstractModern automobiles are equipped with advanced computing capabilities, allowing them to become powerful computing units capable of processing a large amount of data and training machine learning models. However, machine learning algorithms typically require a large centralized dataset, raising concerns about users’ privacy. Federated Learning (FL) is a distributed machine learning paradigm that tackles this problem, allowing intelligent vehicles to collaboratively train machine learning models locally without having to compromise their private data. Multiple works have concentrated on applying Federated Learning on Mobile Edge Computing (MEC) networks with a 3-tier architecture consisting of mobile clients, edge servers, and cloud servers, where the edge server aggregates its local set of clients, and the cloud server aggregates edge servers to learn a global model. This approach helps reduce the expensive communication costs to the far-away cloud server. However, this 3-tier paradigm faces several challenges, a notable one being clients’ constant mobility, leading to regional edges having a fluctuating set of participating clients at each round, which we refer to as distribution drift. This phenomenon introduces instability to the local training process, leading to suboptimal accuracy and convergence. As a solution, we propose a local training process based on the knowledge distillation mechanism. Specifically, we employ the global model and an ensemble of historical regional models from the edge servers as sources of knowledge to guide the local training process, preventing the local models from drifting away from the global knowledge and preserving information from clients that left the region. Experimental results showed that the proposed method helps achieve better performance compared to other baselines. Tuan Dung Nguyen, Ngoc Anh Tong, Binh P. Nguyen, Nguyen Quoc Viet Hung, Phi-Le Nguyen |
IJCNN | 1 |
| 2024 | Federated PCA on Grassmann Manifold for IoT Anomaly DetectionabstractWith the proliferation of the Internet of Things (IoT) and the rising interconnectedness of devices, network security faces significant challenges, especially from anomalous activities. While traditional machine learning-based intrusion detection systems (ML-IDS) effectively employ supervised learning methods, they possess limitations such as the requirement for labeled data and challenges with high dimensionality. Recent unsupervised ML-IDS approaches such as AutoEncoders and Generative Adversarial Networks (GAN) offer alternative solutions but pose challenges in deployment onto resource-constrained IoT devices and in interpretability. To address these concerns, this paper proposes a novel federated unsupervised anomaly detection framework – FedPCA – that leverages Principal Component Analysis (PCA) and the Alternating Directions Method Multipliers (ADMM) to learn common representations of distributed non-i.i.d. datasets. Building on the FedPCA framework, we propose two algorithms, FedPE in Euclidean space and FedPG on Grassmann manifolds. Our approach enables real-time threat detection and mitigation at the device level, enhancing network resilience while ensuring privacy. Moreover, the proposed algorithms are accompanied by theoretical convergence rates even under a sub-sampling scheme, a novel result. Experimental results on the UNSW-NB15 and TON-IoT datasets show that our proposed methods offer performance in anomaly detection comparable to non-linear baselines, while providing significant improvements in communication and memory efficiency, underscoring their potential for securing IoT networks. Tung-Anh Nguyen, Tuan Dung Nguyen, Wei Bao 0001, Suranga Seneviratne, Choong Seon Hong, Nguyen Hoang Tran |
IEEE/ACM Trans. Netw. | 3 |
| 2022 | Mapping Topics in 100, 000 Real-Life Moral Dilemmas
Tuan Dung Nguyen, Georgiana Lyall, Alasdair Tran, Minjeong Shin, Nicholas George Carroll, Colin Klein, Lexing Xie |
ICWSM | 1 |
| 2022 | DONE: Distributed Approximate Newton-type Method for Federated Edge LearningabstractThere is growing interest in applying distributed machine learning to edge computing, formingfederated edge learning. Federated edge learning faces non-i.i.d. and heterogeneous data, and the communication between edge workers, possibly through distant locations and with unstable wireless networks, is more costly than their local computational overhead. In this work, we propose${{\sf DONE}}$, a distributed approximate Newton-type algorithm with fast convergence rate for communication-efficient federated edge learning. First, with strongly convex and smooth loss functions,${{\sf DONE}}$approximates the Newton direction in a distributed manner using the classical Richardson iteration on each edge worker. Second, we prove that${{\sf DONE}}$has linear-quadratic convergence and analyze its communication complexities. Finally, the experimental results with non-i.i.d. and heterogeneous data show that${{\sf DONE}}$attains a comparable performance to Newton's method. Notably,${{\sf DONE}}$requires fewer communication iterations compared to distributed gradient descent and outperforms DANE, FEDL, and GIANT, state-of-the-art approaches, in the case of non-quadratic loss functions. Canh T. Dinh, Nguyen Hoang Tran, Tuan Dung Nguyen, Wei Bao 0001, Amir Rezaei Balef, Bing Bing Zhou, Albert Y. Zomaya |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2020 | Federated Learning with Proximal Stochastic Variance Reduced Gradient AlgorithmsabstractFederated Learning (FL) is a fast-developing distributed machine learning technique involving the participation of a massive number of user devices. While FL has benefits of data privacy and the abundance of user-generated data, its challenges of heterogeneity across users’ data and devices complicate algorithm design and convergence analysis. To tackle these challenges, we propose an algorithm that exploits proximal stochastic variance reduced gradient methods for non-convex FL. The proposed algorithm consists of two nested loops, which allow user devices to update their local models approximately up to an accuracy threshold (inner loop) before sending these local models to the server for global model update (outer loop). We characterize the convergence conditions for both local and global model updates and extract various insights from these conditions via the algorithm’s parameter control. We also propose how to optimize these parameters such that the training time of FL is minimized. Experimental results not only validate the theoretical convergence but also show that the proposed algorithm outperforms existing Stochastic Gradient Descent-based methods in terms of convergence speed in FL setting. Canh T. Dinh, Nguyen Hoang Tran, Tuan Dung Nguyen, Wei Bao 0001, Albert Y. Zomaya, Bing Bing Zhou |
ICPP | 3 |
| 2020 | Personalized Federated Learning with Moreau EnvelopesabstractFederated learning (FL) is a decentralized and privacy-preserving machine learning technique in which a group of clients collaborate with a server to learn a global model without sharing clients' data. One challenge associated with FL is statistical diversity among clients, which restricts the global model from delivering good performance on each client's task. To address this, we propose an algorithm for personalized FL (pFedMe) using Moreau envelopes as clients' regularized loss functions, which help decouple personalized model optimization from the global model learning in a bi-level problem stylized for personalized FL. Theoretically, we show that pFedMe convergence rate is state-of-the-art: achieving quadratic speedup for strongly convex and sublinear speedup of order 2/3 for smooth nonconvex objectives. Experimentally, we verify that pFedMe excels at empirical performance compared with the vanilla FedAvg and Per-FedAvg, a meta-learning based personalized FL algorithm. Canh T. Dinh, Nguyen Hoang Tran, Tuan Dung Nguyen |
NeurIPS | 3 |