Hoang D. Nguyen

dblp:146/2079 · DBLP profile ↗
← Back
21ranked-venue papers
2as first author
20since 2021 · last 2026
0000-0003-2541-3269ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 2 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Sound-AI: A Pedagogical Tool for Exploring AI in Audio and Bioacoustic Research
abstract
Artificial intelligence offers powerful methods for audio processing and analysis. Still, complex workflows and the required programming skills often limit access for students and domain experts, such as marine bioacousticians and soundscape ecologists. We present "AI EcoSound Tutor", a code-free and interactive tool that lowers these barriers by allowing users to construct and explore a complete AI pipeline for audio data analysis. Starting from raw recordings, users can choose from various feature extraction techniques (MFCC, OpenL3), apply dimensionality reduction methods (PCA, t-SNE, UMAP), and optionally perform unsupervised clustering (K-Means, GMM,HDBSCAN). The results are displayed with an interactive 2D visualisation where the user can compare multiple plots by employing various techniques, including PCA and t-SNE. Interactive plots enable the selection of points or clusters of interest, allowing exploration of spectrograms within the desired frequency range, and playing an audio clip corresponding to the selected points. An integrated "Help" feature provides explanations of each method (i.e., what it is, how it works, and its practical use in different domains, such as bioacoustics), fostering both conceptual understanding and useful skill acquisition as learning outcomes. For precomputed features or embeddings, this tool also supports training and evaluating a variety of machine learning models, providing visual feedback on the results. By merging accessibility, interactivity, pedagogy, and domain relevance, our application demystifies AI methods for interdisciplinary education and supporting research in audio analysis.
Hoang D. Nguyen, Rosane Minghim
AAAI2
2026 Reasoning Transfer for an Extremely Low-Resource and Endangered Language: Bridging Languages Through Sample-Efficient Language Understanding
abstract
Recent advances have enabled Large Language Models (LLMs) to tackle reasoning tasks by generating chain-of-thought (CoT) rationales, yet these gains have largely applied to high-resource languages, leaving low-resource languages underperformed. In this work, we first investigate CoT techniques in extremely low-resource scenarios through previous prompting, model editing, and fine-tuning approaches. We introduce \emph{English-Pivoted CoT Training}, leveraging the insight that LLMs internally operate in a latent space aligned toward the dominant language. Given input in a low-resource language, we perform supervised fine-tuning to generate CoT in English and output the final response in the target language. Across mathematical reasoning benchmarks, our approach outperforms other baselines with up to 28.33% improvement in low-resource scenarios. Our analyses and additional experiments, including Mixed-Language CoT and Two-Stage Training, show that explicitly separating language understanding from reasoning enhances crosslingual reasoning abilities. To facilitate future work, we also release LC2024, the first benchmark for mathematical task in Irish, an extremely low-resource and endangered language. Our results and resources highlight a practical pathway to multilingual reasoning without extensive retraining in every extremely low-resource language, despite data scarcity.
Khanh-Tung Tran, Barry O'Sullivan, Hoang D. Nguyen
AAAI3
2026 IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation
abstract
Recent advances in Large Language Models (LLMs) have demonstrated promising capabilities, yet their performance in multilingual and low-resource settings remains modest. Existing benchmarks often exhibit cultural bias, restrict evaluation to text-only, rely on multiple-choice formats, and, more importantly, are ineffectual for extremely low-resource languages. To address these gaps, we introduce IRLBench, presented in parallel English and Irish, which is considered definitely endangered by UNESCO. Our benchmark consists of 12 representative subjects developed from the 2024 Irish Leaving Certificate exam, enabling fine-grained analysis of model capabilities across domains. By framing the task as long-form generation and leveraging the official marking scheme, it supports not only a comprehensive evaluation of correctness but also language fidelity. Our extensive experiments of leading closed-source and open-source LLMs reveal a persistent performance gap between English and Irish, in which models produce valid Irish responses less than 80% of the time, and answer correctly 55.8% of the time compared to 76.2% in English for the best-performing model. With Irish as the case study, our work exposes systemic weaknesses in today's multilingual LLMs and provides a rigorous benchmark for evaluating true multilingual capabilities. We release IRLBench and an accompanying evaluation codebase to enable future research on robust, culturally aware multilingual AI development.
Khanh-Tung Tran, Barry O'Sullivan, Hoang D. Nguyen
KDD (1)4
2026 Empowering Multimodal Learning Analytics using Agentic AI: A Comprehensive Platform for Simulation-based Clinical Training with Intelligent Assessment
abstract
Simulation-based clinical training generates rich multimodal data that remains underused due to fragmented modalities, annotation bottlenecks, weak provenance, and tools misaligned with educator workflows. We introduce ClinVision, an educator-in-the-loop platform that operationalizes end-to-end multimodal learning analytics: synchronized multi-camera review, ISBAR-aligned scoring (a validated clinical communication framework), and templated reports with jump-to-evidence provenance. Agentic AI - systems that act on behalf of users while preserving human authority - assists with phrasing under explicit control (accept/edit/reject) and visible provenance, supporting accountable use rather than prescriptive automation. An in-learning deployment with five educators revealed full ISBAR coverage and time-efficient workflows, though AI suggestions were used selectively. We surface three transferable design tensions (assistance vs. authority, structure vs. flexibility, evidence vs. overload) and demonstrate that workflow integration, temporal primitives, and background AI assistance may better support high-stakes assessment than analytics or automation alone.
Kinza Salim, David Power, Murray Connolly, Ahmed Hamdy, Maya Contreras, Tai Tan Mai, George Shorten, Barry O'Sullivan, Vijayakumar Nanjappan, Hoang D. Nguyen
LAK11
2026 VIVID: A Culturally Grounded Benchmark Exposing the Figurative Language Gap in Vietnamese NLP
Tu Tran Do, Nhat Ngoc Nguyen, Tung Khanh Tran, Hoang D. Nguyen, Tu Minh Phuong, Long Hoang Dang
LREC4
2026 Irish-BLiMP: A Linguistic Benchmark for Evaluating Human and Language Model Performance in a Low-Resource Setting
Josh McGiff, Khanh-Tung Tran, William Mulcahy, Dáibhidh Ó Luinín, Jake Dalzell, Róisín Ní Bhroin, Adam Burke 0003, Barry O'Sullivan, Hoang D. Nguyen, Nikola S. Nikolov
LREC9
2026 Questionnaire Meets LLM: A Benchmark and Empirical Study of Structural Skills for Understanding Questions and Responses
Vijayakumar Nanjappan, Barry O'Sullivan, Hoang D. Nguyen
LREC4
2026 Designing AI-based decision support systems for carbon stock estimation and planning: A design science research
abstract
Mitigating climate change presents significant challenges, particularly in effective carbon stock estimation and planning. Decision-making based on quantified carbon stock is essential for enabling carbon sequestration efforts in natural ecosystems, yet traditional field measurements are labor-intensive, geographically constrained, and difficult to be adopted at scale. Motivated by this practical challenge, we undertook a design science research (DSR) project to develop AI-based decision support systems that enhance the efficiency, accessibility, and scalability for carbon stock estimation and planning. Guided by technology affordance theory, our iterative design, development, demonstration, and evaluation process has led to the conceptualization of four affordance-based design principles that build on and contribute to the growing design knowledge for environmental sustainability. This research advances our understanding of the critical problem and solution domains of carbon stock estimation and planning. It also facilitates the emerging discourse on balancing AI and human intelligence in designing systems for complex decision-making challenges for sustainability. This study provides an exemplar of DSR for supporting Sustainable Development Goal 13 by demonstrating how to design systems that facilitate carbon sequestration and societal transition toward environmental sustainability.
Amelia S. Li, Shan Ling Pan, Yenni Tim, Hoang D. Nguyen
Decis. Support Syst.4
2025 Sum Rate Maximization in Downlink HAP-RSMA-based THz Systems: A Generative Diffusion Model enabled RL Approach
abstract
This paper investigates the maximization of the achievable rate for users served by a high-altitude platform (HAP) acting as a flying base station in the downlink of rate-splitting multiple access (RSMA)-based terahertz (THz) communication systems. Considering the dynamic and uncertain environment caused by user mobility and molecular absorption effects, we propose a generative diffusion model (DM)-based deep reinforcement learning approach to address this challenge. The problem is formulated as a Markov decision process, aiming to maximize the long-term achievable rate for all users by jointly optimizing power allocation and the common rate splitting ratio. Moreover, the generative DM significantly improves the decision-making capabilities of a deep reinforcement learning algorithm, namely the deep deterministic policy gradient (DDPG). Experimental simulations demonstrate the effectiveness of the proposed DM-DDPG algorithm compared to alternative schemes.
Mai Le, Quoc-Viet Pham, Barry O'Sullivan, Hoang D. Nguyen
GLOBECOM4
2025 GraphLOMO: LOcating multiple objects without visual annotations
abstract
Abstract Locating multiple objects has become an important task in multimedia research and applications due to the common nature of real-world images. Object localization requires a large number of visual annotations, such as bounding boxes or segmentation, but the annotation process is labour-intensive and sometimes inextricable for human experts in complex domains such as manufacturing and medical fields. Moving beyond single object localization, this paper presents a weakly semi-supervised learning framework based on Graph Transformer Networks using Class Activation Maps to LOcate Multiple Objects (GraphLOMO) in images without visual annotations. Our method overcomes the computational challenges of gradient-based CAM while integrating topological information and prior knowledge into object localization. Moreover, we investigate the higher order of object inter-dependencies with the use of 3D adjacency matrix for better performance. Extensive empirical experiments are conducted on MS-COCO and Pascal VOC to establish a suitable performance measure and baselines, as well as a state-of-the-art for weakly semi-supervised multi-object localization.
Alex To, Joseph G. Davis, Hoang D. Nguyen
Mach. Vis. Appl.3
2024 UCCIX: Irish-eXcellence Large Language Model
abstract
The development of Large Language Models (LLMs) has predominantly focused on high-resource languages, leaving extremely low-resource languages like Irish with limited representation. This work presents UCCIX, a pioneering effort on the development of an open-source Irish-based LLM. We propose a novel framework for continued pre-training of LLMs specifically adapted for extremely low-resource languages, requiring only a fraction of the textual data typically needed for training LLMs according to scaling laws. Our model, based on Llama 2-13B [23], outperforms much larger models on Irish language tasks with up to 12% performance improvement, showcasing the effectiveness and efficiency of our approach. We also contribute comprehensive Irish benchmarking datasets, including IrishQA, a question-answering dataset, and Irish version of MT-bench [28]. These datasets enable rigorous evaluation and facilitate future research in Irish LLM systems. Our work aims to preserve and promote the Irish language, knowledge, and culture of Ireland in the digital era while providing a framework for adapting LLMs to other indigenous languages.
Khanh-Tung Tran, Barry O'Sullivan, Hoang D. Nguyen
ECAI3
2024 Collaborative Fair-is-Better Filtering for Implicit Feedback
abstract
With the wide adoption of recommender systems, fairness has increasingly become a critical topic in many applications, such as e-commerce, job search, and online entertainment. Collaborative filtering is susceptible to unfair recommendations for users from sensitive groups due to the non-negligible presence of biases. Whereas recent work mostly concerned fairness with the explicit ratings, we propose fairness-awareness recommendation models, referred to as CFBF, for implicit feedback (e.g., clicks, views, or purchases), which are ubiquitous in today’s context. This paper considers sensitive attributes, such as gender and age, in both disjoined and combined manners to investigate the models’ unfairness. We discuss several fairness metrics for implicit feedback recommendations based on Spearman’s rank correlation and Kendal Tau. Comprehensive experiments on Movielens and LastFM show that CFBF significantly improves user groups’ fairness with comparable and even better ranking performance.
Hoang V. Dong, Huu-Quang Nguyen, Hoang D. Nguyen, Duc-Trong Le
KES3
2024 Autofusion: Fusing Multi-Modalities with Interactions
abstract
In the context of escalating data flows through diverse channels, multimodal machine learning holds the potential to simultaneously process varied data formats from multiple sources, offering robust solutions for uncertainty in various applications. The study highlights the often under-explored correlation information among data modalities and emphasizes the importance of disentangling enriched interactions for more informed decision-making. This paper navigates the burgeoning field of multimodal artificial intelligence (AI) by proposing Autofusion, a pioneering framework addressing representation learning and fusion challenges. Our proposed approach integrates autoencoder structures to address overfitting issues in unimodal machine learning, simultaneously tackling information balancing challenges. The framework’s application in Alzheimer’s disease detection, using DementiaBank’s Pitt corpus, demonstrates promising results, outperforming unimodal methods and showcasing a substantial advantage over traditional fusion techniques. This research significantly contributes by introducing Autofusion as a comprehensive multimodal machine learning solution, demonstrating its efficacy through DementiaBank’s Pitt corpus to detect Alzheimer’s disease, and shedding light on the influential role of cross-modality interaction for enhanced performance in complex applications.
Thuy-Trinh Nguyen, Fani Deligianni, Hoang D. Nguyen
KES3
2024 NeuProNet: neural profiling networks for sound classification
abstract
Abstract Real-world sound signals exhibit various aspects of grouping and profiling behaviors, such as being recorded from identical sources, having similar environmental settings, or encountering related background noises. In this work, we propose novel neural profiling networks (NeuProNet) capable of learning and extracting high-level unique profile representations from sounds. An end-to-end framework is developed so that any backbone architectures can be plugged in and trained, achieving better performance in any downstream sound classification tasks. We introduce an in-batch profile grouping mechanism based on profile awareness and attention pooling to produce reliable and robust features with contrastive learning. Furthermore, extensive experiments are conducted on multiple benchmark datasets and tasks to show that neural computing models under the guidance of our framework gain significant performance gaps across all evaluation tasks. Particularly, the integration of NeuProNet surpasses recent state-of-the-art (SoTA) approaches on UrbanSound8K and VocalSound datasets with statistically significant improvements in benchmarking metrics, up to 5.92% in accuracy compared to the previous SoTA method and up to 20.19% compared to baselines. Our work provides a strong foundation for utilizing neural profiling for machine learning tasks.
Khanh-Tung Tran, Xuan-Son Vu, Khuong Nguyen, Hoang D. Nguyen
Neural Comput. Appl.4
2023 Personalization for Robust Voice Pathology Detection in Sound Waves
abstract
Automatic voice pathology detection is promising for noninvasive screening and early intervention using sound signals. Nevertheless, existing methods are susceptible to covariate shifts due to background noises, human voice variations, and data selection biases leading to severe performance degradation in real-world scenarios. Hence, we propose a non-invasive framework that contrastively learns personalization from sound waves as a pre-train and predicts latent-spaced profile features through semi-supervised learning. It allows all subjects from various distributions (e.g., regionality, gender, age) to benefit from personalized predictions for robust voice pathology in a privacy-fulfilled manner. We extensively evaluate the framework on four real-world respiratory illnesses datasets, including Coswara, COUGHVID, ICBHI, and our private dataset - ASound under multiple covariate shift settings (i.e., cross-dataset), improving up to 4.12% in overall performance.
Khanh-Tung Tran, Truong Hoang, Duy Khuong Nguyen, Hoang D. Nguyen, Xuan-Son Vu
INTERSPEECH4
2023 Multimodal Machine Learning for Mental Disorder Detection: A Scoping Review
abstract
Recent advancements in machine learning and multimedia technologies have paved new ways for automatic medical diagnosis. In mental health, multimodal inputs such as visual and audible sensing data are promising to investigate the underlying mechanisms of many conditions, such as depression and bipolar disorders. With the increasing burden on healthcare systems, timely diagnosis of mental diseases using multiple modalities might benefit millions of people worldwide. This scoping review provides an exploratory overview of recent multimodal machine learning approaches for mental disorder screening. We also discuss a generalised end-to-end multimodal machine learning pipeline for future research and development of multimodal disease detection.
Thuy-Trinh Nguyen, Viet Hoang-Quoc Pham, Duc-Trong Le, Xuan-Son Vu, Fani Deligianni, Hoang D. Nguyen
KES6
2023 Reliable Sound Re-labeling Approach for Bird Sound Classification
abstract
Passive acoustic monitoring has become an effective and economically scalable solution for wildlife monitoring in recent years, especially for population monitoring of endangered birds in hardly accessible and high-elevation areas. As sound recorders are deployed on the field for extended periods (months), continuous sound streams are often complex in nature, with noises and intermittent signals. Generally, it is prohibitively expensive to label bird sounds with the exact onset and offset time, thus, during training, the data is often provided with only presence/absence labeling (weak labeling) that states which bird species are present in each long recording without temporal information. During test time, it is, however, desirable to provide fine-grained detection in short audio segments. To bridge the gap between the difference of weak labels used for training and strong labels used for testing, we propose a re-labeling approach with two stages: (1) we train a model with weak labeling; and (2) using the model obtained from stage 1, we generate labels for short audio segments and retrain the model on short audio segments with the newly generated labels. We applied our approach in both classification and sound event detection and achieved consistently good performance across multiple random seeds. In BirdCLEF 2022, our model ranked in the top 1.1%, the 9th best entry out of the total 807 entries in the competition.
Hoang Van Truong, Nghia NVN, Alex To, Duc-Trong Le, Hoang D. Nguyen
KES5
2023 Seismic fragility analysis of steel moment frames using machine learning models
Hoang D. Nguyen, Young-Joo Lee, James M. LaFave, Myoungsu Shin
Eng. Appl. Artif. Intell.1
2023 Comparative study on the performance of different machine learning techniques to predict the shear strength of RC deep beams: Model selection and industry implications
Khuong Le Nguyen, Hoa Thi Trinh, Thanh T. Nguyen, Hoang D. Nguyen
Expert Syst. Appl.4
2021 Modular Graph Transformer Networks for Multi-Label Image Classification
abstract
With the recent advances in graph neural networks, there is a rising number of studies on graph-based multi-label classification with the consideration of object dependencies within visual data. Nevertheless, graph representations can become indistinguishable due to the complex nature of label relationships. We propose a multi-label image classification framework based on graph transformer networks to fully exploit inter-label interactions. The paper presents a modular learning scheme to enhance the classification performance by segregating the computational graph into multiple sub-graphs based on modularity. The proposed approach, named Modular Graph Transformer Networks (MGTN), is capable of employing multiple backbones for better information propagation over different sub-graphs guided by graph transformers and convolutions. We validate our framework on MS-COCO and Fashion550K datasets to demonstrate improvements for multi-label image classification. The source code is available at https://github.com/ReML-AI/MGTN.
Hoang D. Nguyen, Xuan-Son Vu, Duc-Trong Le
AAAI1
2020 Privacy-Preserving Visual Content Tagging using Graph Transformer Networks
abstract
With the rapid growth of Internet media, content tagging has become an important topic with many multimedia understanding applications, including efficient organisation and search. Nevertheless, existing visual tagging approaches are susceptible to inherent privacy risks in which private information may be exposed unintentionally. The use of anonymisation and privacy-protection methods is desirable, but with the expense of task performance. Therefore, this paper proposes an end-to-end framework (SGTN) using Graph Transformer and Convolutional Networks to significantly improve classification and privacy preservation of visual data. Especially, we employ several mechanisms such as differential privacy based graph construction and noise-induced graph transformation to protect the privacy of knowledge graphs. Our approach unveils new state-of-the-art on MS-COCO dataset in various semi-supervised settings. In addition, we showcase a real experiment in the education domain to address the automation of sensitive document tagging. Experimental results show that our approach achieves an excellent balance of model accuracy and privacy preservation on both public and private datasets.
Xuan-Son Vu, Duc-Trong Le, Christoffer Edlund, Lili Jiang 0002, Hoang D. Nguyen
ACM Multimedia5