Hung Cao

dblp:52/680 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Multi-agent systems · 41% Knowledge representation and reasoning · 36% Vision and language · 12%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 100%

Topics — the 6 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Multi-agent systems
formal verification of multi-agent systems
0.912025
Multimedia Verification Through Multi-Agent Deep Research Multimodal Large Language Models · ACM Multimedia 2025
Multimedia analysis and retrieval › harmful content detection
misinformation detection
0.912025
Multimedia Verification Through Multi-Agent Deep Research Multimodal Large Language Models · ACM Multimedia 2025
Multimedia analysis and retrieval › multimedia analysis › multimedia forensics
multimedia verification
0.912025
Multimedia Verification Through Multi-Agent Deep Research Multimodal Large Language Models · ACM Multimedia 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
explanation generation
0.812024
LangXAI: Integrating Large Vision Models for Generating Textual Explanations to Enhance Explainability in Visual Perception Tasks · IJCAI 2024
Computer vision › Vision and language › vision-language model
multimodal large language model
0.312025
Multimedia Verification Through Multi-Agent Deep Research Multimodal Large Language Models · ACM Multimedia 2025
Machine learning › Trustworthy machine learning
interpretability
0.212024
LangXAI: Integrating Large Vision Models for Generating Textual Explanations to Enhance Explainability in Visual Perception Tasks · IJCAI 2024

Methods — techniques the papers use, named apart from their topics

reverse image search · 1.7multimodal large language model · 1.7metadata analysis · 1.7fact-checking · 1.7large vision models · 0.8
YearPublicationVenuePosition
2026 Contestable Multi-Agent Debate with Arena-based Argumentative Computation for Multimedia Verification
abstract
Multimedia verification requires not only accurate conclusions but also transparent and contestable reasoning. We propose a contestable multi-agent framework that integrates multimodal large language models, external verification tools, and arena-based quantitative bipolar argumentation (A-QBAF) as a submission to the ICMR 2026 Grand Challenge on Multimedia Verification. Our method decomposes each case into claim-centered sections, retrieves targeted evidence, and converts evidence into structured support and attack arguments with provenance and strength scores. These arguments are resolved through small local argument graphs with selective clash resolution and uncertainty-aware escalation. The resulting system generates section-wise verification reports that are transparent, editable, and computationally practical for real-world multimedia verification. Our implementation is public at: https://github.com/Analytics-Everywhere-Lab/MV2026_the_liems.
Hung Truong Thanh Nguyen, Vo Thanh Khang Nguyen, Hoang-Loc Cao, Phuc Ho, Van Pham, Hung Cao
ICMR6
2025 Multimedia Verification Through Multi-Agent Deep Research Multimodal Large Language Models
abstract
This paper presents our submission to the ACMMM25 - Grand Challenge on Multimedia Verification. We developed a multi-agent verification system that combines Multimodal Large Language Models (MLLMs) with specialized verification tools to detect multimedia misinformation. Our system operates through six stages: raw data processing, planning, information extraction, deep research, evidence collection, and report generation. The core Deep Researcher Agent employs four tools: reverse image search, metadata analysis, fact-checking databases, and verified news processing that extracts spatial, temporal, attribution, and motivational context. We demonstrate our approach on a challenge dataset sample involving complex multimedia content. Our system successfully verified content authenticity, extracted precise geolocation and timing information, and traced source attribution across multiple platforms, effectively addressing real-world multimedia verification scenarios.
Huy Hoan Le, Van Sy Thinh Nguyen, Thi Le Chi Dang, Vo Thanh Khang Nguyen, Truong Thanh Hung Nguyen, Hung Cao
ACM Multimedia6
2025 Human-Centered Explainable Psychiatric Disorder Diagnosis System Using Wearable ECG Monitors
Truong Thanh Hung Nguyen, Alireza Rahimi, Veronica Whitford, Hélène Fournier, Irina Kondratova, René Richard, Hung Cao
PAKDD (2)7
2025 Multilingual Phishing Email Detection Using Lightweight Federated Learning
abstract
Given the escalating global threat of phishing emails, it is imperative to develop effective solutions to mitigate their potentially devastating impacts on society. This study endeavours to construct a federated multilingual spam detection system employing logistic regression, specifically targeting English, French, and Russian emails. This is the first work to the best of our knowledge which considers a non-deep learning setting for federated learning, and combines federated learning with multilingual phishing detection. Evaluation of the models is based on accuracy metrics which are compared with a most frequent class baseline. Our findings indicate that an optimal configuration comprises 10 clients undergoing 100 epochs of training with 100 rounds of federated learning, resulting in superior performance. Notably, this approach significantly outperforms the baseline, achieving an accuracy of $89.46 \%$ compared to $70 \%$.
Dakota Staples, Hung Cao, Saqib Hakak, Paul Cook
PST2
2024 Fetal QRS Detection from Single-Channel Abdominal ECG by Adaptive Improved Clustering
abstract
Assessment of fetal development and wellness throughout pregnancy is essential in the early detection of pregnancy anomalies. Recent advances in wearable technology have led to the development of home pregnancy monitoring systems, which hold potentials for the continuous monitoring of expectant mothers and their unborn children through fe-tal/maternal electrocardiogram (f/mECG) technology. Given the complexity of extracting useful information from f/mECG data, these devices are often equipped with multiple channels and sophisticated signal processing algorithms. Thus, the implementation of f/mECG monitoring in daily life remains unfeasible. In this study, we propose an algorithm for extracting the fetal electrocardiogram (fECG) using a single channel. The algorithm detects the peaks of the maternal ECG (mECG) through a clustering method, followed by peak correction. The mECG signal is then reconstructed with Principal Component Analysis (PCA) before applying the Template Subtraction (TS) method to obtain fECG. Rigorous testing with various online f/mECG databases demonstrates the potential to accurately extract fECG using only a single channel. The evaluation results based on the online dataset achieved an accuracy of 99.60% and an F1 score of 99.30%. This advancement also opens opportunities to optimize home-based f/mECG monitoring, making it more compact and reliable.
Thinh Nguyen-Quang, Tai Le, Duc Nguyen Minh, Hung Cao, Huy-Dung Han
BSN4
2024 LangXAI: Integrating Large Vision Models for Generating Textual Explanations to Enhance Explainability in Visual Perception Tasks
Truong Thanh Hung Nguyen, Tobias Clement, Phuc Truong Loc Nguyen, Nils Kemmerzell, Van Binh Truong, Vo Thanh Khang Nguyen, Mohamed Abdelaal 0001, Hung Cao
IJCAI8
2022 Composing Graphical Models with Generative Adversarial Networks for EEG Signal Modeling
abstract
Neural oscillations in the form of electroencephalogram (EEG) can reveal underlying brain functions, such as cognition, memory, perception, and consciousness. A comprehensive EEG computational model provides not only a stochastic procedure that directly generates data but also insights to further understand the neurological mechanisms. Here, we propose a generative and inference approach that combines the complementary benefits of probabilistic graphical models and generative adversarial networks (GANs) for EEG signal modeling. We investigate the method’s ability to jointly learn coherent generation and inverse inference models on the CHI-MIT epilepsy multi-channel EEG dataset. We further study the efficacy of the learned representations in epilepsy seizure detection formulated as an unsupervised learning problem. Quantitative and qualitative experimental results demonstrate the effectiveness and efficiency of our approach.
Khuong Vo, Manoj Vishwanath, Ramesh Srinivasan, Nikil Dutt, Hung Cao
ICASSP5
2017 Developing an edge computing platform for real-time descriptive analytics
abstract
The Internet of Mobile Things encompasses stream data being generated by sensors, network communications that pull and push these data streams, as well as running processing and analytics that can effectively leverage actionable information for transportation planning, management, and business advantage. Edge computing emerges as a new paradigm that decentralizes the communication, computation, control and storage resources from the cloud to the edge of the network. This paper proposes an edge computing platform where mobile edge nodes are physical devices deployed on a transit bus where descriptive analytics is used to uncover meaningful patterns from real-time transit data streams. An application experiment is used to evaluate the advantages and disadvantages of our proposed platform to support descriptive analytics at a mobile edge node and generate actionable information to transit managers.
Hung Cao, Monica Wachowicz, Sangwhan Cha
IEEE BigData1