VLDB 2026 Research / reviewers in the wild / expert
Xuan-Son Vu
dblp:151/8673
· DBLP profile ↗
19ranked-venue papers
9as first author
10since 2021 · last 2026
0000-0001-8820-2405ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 6 first-author · 9 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reason-to-Learn (R2L): Multi-Agent Knowledge Distillation for Lightweight LLMs in Sentiment Analysis
Le-Huy Tu, Vincent Nguyen 0001, Johanna Björklund, Xuan-Son Vu |
LREC | 5 |
| 2024 | Pseudonymization Categories across Domain BoundariesabstractLinguistic data, a component critical not only for research in a variety of fields but also for the development of various Natural Language Processing (NLP) applications, can contain personal information. As a result, its accessibility is limited, both from a legal and an ethical standpoint. One of the solutions is the pseudonymization of the data. Key stages of this process include the identification of sensitive elements and the generation of suitable surrogates in a way that the data is still useful for the intended task. Within this paper, we conduct an analysis of tagsets that have previously been utilized in anonymization and pseudonymization. We also investigate what kinds of Personally Identifiable Information (PII) appear in various domains. These reveal that none of the analyzed tagsets account for all of the PII types present cross-domain at the level of detailedness seemingly required for pseudonymization. We advocate for a universal system of tags for categorizing PIIs leading up to their replacement. Such categorization could facilitate the generation of grammatically, semantically, and sociolinguistically appropriate surrogates for the kinds of information that are considered sensitive in a given domain, resulting in a system that would enable dynamic pseudonymization while keeping the texts readable and useful for future research in various fields. Maria Irena Szawerna, Simon Dobnik, Therese Lindström Tiedemann, Ricardo Muñoz Sánchez, Xuan-Son Vu, Elena Volodina |
LREC/COLING | 5 |
| 2024 | NeuProNet: neural profiling networks for sound classificationabstractAbstract Real-world sound signals exhibit various aspects of grouping and profiling behaviors, such as being recorded from identical sources, having similar environmental settings, or encountering related background noises. In this work, we propose novel neural profiling networks (NeuProNet) capable of learning and extracting high-level unique profile representations from sounds. An end-to-end framework is developed so that any backbone architectures can be plugged in and trained, achieving better performance in any downstream sound classification tasks. We introduce an in-batch profile grouping mechanism based on profile awareness and attention pooling to produce reliable and robust features with contrastive learning. Furthermore, extensive experiments are conducted on multiple benchmark datasets and tasks to show that neural computing models under the guidance of our framework gain significant performance gaps across all evaluation tasks. Particularly, the integration of NeuProNet surpasses recent state-of-the-art (SoTA) approaches on UrbanSound8K and VocalSound datasets with statistically significant improvements in benchmarking metrics, up to 5.92% in accuracy compared to the previous SoTA method and up to 20.19% compared to baselines. Our work provides a strong foundation for utilizing neural profiling for machine learning tasks. Khanh-Tung Tran, Xuan-Son Vu, Khuong Nguyen, Hoang D. Nguyen |
Neural Comput. Appl. | 2 |
| 2023 | Personalization for Robust Voice Pathology Detection in Sound WavesabstractAutomatic voice pathology detection is promising for noninvasive screening and early intervention using sound signals. Nevertheless, existing methods are susceptible to covariate shifts due to background noises, human voice variations, and data selection biases leading to severe performance degradation in real-world scenarios. Hence, we propose a non-invasive framework that contrastively learns personalization from sound waves as a pre-train and predicts latent-spaced profile features through semi-supervised learning. It allows all subjects from various distributions (e.g., regionality, gender, age) to benefit from personalized predictions for robust voice pathology in a privacy-fulfilled manner. We extensively evaluate the framework on four real-world respiratory illnesses datasets, including Coswara, COUGHVID, ICBHI, and our private dataset - ASound under multiple covariate shift settings (i.e., cross-dataset), improving up to 4.12% in overall performance. Khanh-Tung Tran, Truong Hoang, Duy Khuong Nguyen, Hoang D. Nguyen, Xuan-Son Vu |
INTERSPEECH | 5 |
| 2023 | Multimodal Machine Learning for Mental Disorder Detection: A Scoping ReviewabstractRecent advancements in machine learning and multimedia technologies have paved new ways for automatic medical diagnosis. In mental health, multimodal inputs such as visual and audible sensing data are promising to investigate the underlying mechanisms of many conditions, such as depression and bipolar disorders. With the increasing burden on healthcare systems, timely diagnosis of mental diseases using multiple modalities might benefit millions of people worldwide. This scoping review provides an exploratory overview of recent multimodal machine learning approaches for mental disorder screening. We also discuss a generalised end-to-end multimodal machine learning pipeline for future research and development of multimodal disease detection. Thuy-Trinh Nguyen, Viet Hoang-Quoc Pham, Duc-Trong Le, Xuan-Son Vu, Fani Deligianni, Hoang D. Nguyen |
KES | 4 |
| 2023 | MetaVSID: A Robust Meta-Reinforced Learning Approach for VSI-DDoS Detection on the EdgeabstractThe explosive growth of end devices that generate massive amounts of data requires close-proximity computing resources for processing at the network’s edge. Having geographic distributions and limited resources of edge nodes or servers opens several doors for attackers to exploit them primarily to the detriment of deployed services; one of the recent attacks is Very Short Intermittent Distributed Denial of Services (VSI-DDoS). Deep learning-based models have been developed to detect and mitigate such attacks but cause the degrading quality of models due to covariate shifts when deployed in real-world environments. Therefore, we propose a new approach, called MetaVSID, to detect VSI-DDoS attacks in edge clouds using meta-reinforcement learning followed by ensemble learning to increase the robustness of the model in detecting VSI-DDoS attacks early. The proposed model can capture dynamic patterns of VSI-DDoS attacks, from which it identifies manipulated services and increase service availability when covariate shifts at deployment time. We carry out extensive experiments to validate the MetaVSID using both testbed and benchmark datasets. Via the meta-reinforced downsampling process, the proposed method improves sample efficiency, leading to cost-effective policies. Moreover, the optimized policies are generalized to adapt to dynamic changes in the training distribution. Our experimental results demonstrate that MetaVSID stably achieves better performance in multiple evaluation settings with the difference from baseline models from 1.5% to 7.5% in terms of AUC for both VSI-DDoS and DDoS detection, especially under covariate shift settings. Xuan-Son Vu, Maode Ma, Monowar Bhuyan |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2022 | Optimized and Adaptive Federated Learning for Straggler-Resilient Device SelectionabstractFederated Learning (FL) has evolved as a promising distributed learning paradigm in which data samples are disseminated over massively connected devices in an IID (Identical and Independent Distribution) or non-IID manner. FL follows a collaborative training approach where each device uses local training data to train local models, and the server generates a global model by combining the local model's parameters. However, FL is vulnerable to system heterogeneity when local devices have varying computational, storage, and communication capabilities over time. The presence of stragglers or low-performing devices in the learning process severely impacts the scalability of FL algorithms and significantly delays convergence. To mitigate this problem, we propose Fed-MOODS, a Multi-Objective Optimization-based Device Selection approach to reduce the effect of stragglers in the FL process. The primary criteria for optimization are to maximize: (i) the availability of the processing capacity of each device, (ii) the availability of the memory in devices, and (iii) the bandwidth capacity of the participating devices. The multi-objective optimization prioritizes devices from fast to slow. The approach involves faster devices in early global rounds and gradually incorporating slower devices from the Pareto fronts to improve the model's accuracy. The overall training time of Fed-MOODS is 1.8× and 1.48× faster than the baseline model (FedAvg) with random device selection for MNIST and FMNIST non-IID data, respectively. Fed-MOODS is extensively evaluated under multiple experimental settings, and the results show that Fed-MOODS has significantly improved model's convergence and performance. Fed-MOODS maintains fairness in the prioritized participation of devices and the model for both IID and non-IID settings. Sourasekhar Banerjee, Xuan-Son Vu, Monowar Bhuyan |
IJCNN | 2 |
| 2021 | Modular Graph Transformer Networks for Multi-Label Image ClassificationabstractWith the recent advances in graph neural networks, there is a rising number of studies on graph-based multi-label classification with the consideration of object dependencies within visual data. Nevertheless, graph representations can become indistinguishable due to the complex nature of label relationships. We propose a multi-label image classification framework based on graph transformer networks to fully exploit inter-label interactions. The paper presents a modular learning scheme to enhance the classification performance by segregating the computational graph into multiple sub-graphs based on modularity. The proposed approach, named Modular Graph Transformer Networks (MGTN), is capable of employing multiple backbones for better information propagation over different sub-graphs guided by graph transformers and convolutions. We validate our framework on MS-COCO and Fashion550K datasets to demonstrate improvements for multi-label image classification. The source code is available at https://github.com/ReML-AI/MGTN. Hoang D. Nguyen, Xuan-Son Vu, Duc-Trong Le |
AAAI | 2 |
| 2021 | Cformer: Semi-Supervised Text Clustering Based on Pseudo LabelingabstractWe propose a semi-supervised learning method called Cformer for automatic clustering of text documents in cases where clusters are described by a small number of labeled examples, while the majority of training examples are unlabeled. We motivate this setting with an application in contextual programmatic advertising, a type of content placement on news pages that does not exploit personal information about visitors but relies on the availability of a high-quality clustering computed on the basis of a small number of labeled samples. Arezoo Hatefi, Xuan-Son Vu, Monowar Bhuyan, Frank Drewes |
CIKM | 2 |
| 2021 | ICDAR 2021 Competition on Multimodal Emotion Recognition on Comics Scenes
Vincent Nguyen 0001, Xuan-Son Vu, Christophe Rigaud, Lili Jiang 0002, Jean-Christophe Burie |
ICDAR (4) | 2 |
| 2020 | Multimodal Review Generation with Privacy and Fairness AwarenessabstractUsers express their opinions towards entities (e.g., restaurants) via online reviews which can be in diverse forms such as text, ratings, and images.Modeling reviews are advantageous for user behavior understanding which, in turn, supports various user-oriented tasks such as recommendation, sentiment analysis, and review generation.In this paper, we propose MG-PriFair, a multimodal neural-based framework, which generates personalized reviews with privacy and fairness awareness.Motivated by the fact that reviews might contain personal information and sentiment bias, we propose a novel differentially private (dp)-embedding model for training privacy guaranteed embeddings and an evaluation approach for sentiment fairness in the food-review domain.Experiments on our novel review dataset show that MG-PriFair is capable of generating plausibly long reviews while controlling the amount of exploited user data and using the least sentimentbiased word embeddings.To the best of our knowledge, we are the first to bring user privacy and sentiment fairness into the review generation task.The dataset and source codes are available at https Xuan-Son Vu, Thanh-Son Nguyen 0001, Duc-Trong Le, Lili Jiang 0002 |
COLING | 1 |
| 2020 | Privacy-Preserving Visual Content Tagging using Graph Transformer NetworksabstractWith the rapid growth of Internet media, content tagging has become an important topic with many multimedia understanding applications, including efficient organisation and search. Nevertheless, existing visual tagging approaches are susceptible to inherent privacy risks in which private information may be exposed unintentionally. The use of anonymisation and privacy-protection methods is desirable, but with the expense of task performance. Therefore, this paper proposes an end-to-end framework (SGTN) using Graph Transformer and Convolutional Networks to significantly improve classification and privacy preservation of visual data. Especially, we employ several mechanisms such as differential privacy based graph construction and noise-induced graph transformation to protect the privacy of knowledge graphs. Our approach unveils new state-of-the-art on MS-COCO dataset in various semi-supervised settings. In addition, we showcase a real experiment in the education domain to address the automation of sensitive document tagging. Experimental results show that our approach achieves an excellent balance of model accuracy and privacy preservation on both public and private datasets. Xuan-Son Vu, Duc-Trong Le, Christoffer Edlund, Lili Jiang 0002, Hoang D. Nguyen |
ACM Multimedia | 1 |
| 2019 | dpUGC: Learn Differentially Private Representation for User Generated Contents (Best Paper Award, Third Place, Shared)
Xuan-Son Vu, Son N. Tran, Lili Jiang 0002 |
CICLing (1) | 1 |
| 2019 | Graph-based Interactive Data Federation System for Heterogeneous Data Retrieval and AnalyticsabstractGiven the increasing number of heterogeneous data stored in relational databases, file systems or cloud environment, it needs to be easily accessed and semantically connected for further data analytic. The potential of data federation is largely untapped, this paper presents an interactive data federation system (https://vimeo.com/319473546) by applying large-scale techniques including heterogeneous data federation, natural language processing, association rules and semantic web to perform data retrieval and analytics on social network data. The system first creates a Virtual Database (VDB) to virtually integrate data from multiple data sources. Next, a RDF generator is built to unify data, together with SPARQL queries, to support semantic data search over the processed text data by natural language processing (NLP). Association rule analysis is used to discover the patterns and recognize the most important co-occurrences of variables from multiple data sources. The system demonstrates how it facilitates interactive data analytic towards different application scenarios (e.g., sentiment analysis, privacy-concern analysis, community detection). Xuan-Son Vu, Addi Ait-Mlouk, Erik Elmroth, Lili Jiang 0002 |
WWW | 1 |
| 2018 | Self-adaptive Privacy Concern Detection for User-Generated Content
Xuan-Son Vu, Lili Jiang 0002 |
CICLing (1) | 1 |
| 2018 | Improving Recurrent Neural Networks with Predictive Propagation for Sequence Labelling
Son N. Tran, Qing Zhang 0001, Anthony N. Nguyen, Xuan-Son Vu, Ngo Tung Son |
ICONIP (1) | 4 |
| 2018 | Lexical-semantic resources: yet powerful resources for automatic personality classificationabstractIn this paper, we aim to reveal the impact of lexical-semantic resources, used in particular for word sense disambiguation and sense-level semantic categorization, on automatic personality classification task.While stylistic features (e.g., part-of-speech counts) have been shown their power in this task, the impact of semantics beyond targeted word lists is relatively unexplored.We propose and extract three types of lexical-semantic features, which capture high-level concepts and emotions, overcoming the lexical gap of word n-grams.Our experimental results are comparable to state-of-the-art methods, while no personality-specific resources are required. Xuan-Son Vu, Lucie Flek, Lili Jiang 0002, Iryna Gurevych |
GWC | 1 |
| 2017 | Personality-based Knowledge Extraction for Privacy-preserving Data AnalysisabstractIn this paper, we present a differential privacy preserving approach, which extracts personality-based knowledge to serve privacy guarantee data analysis on personal sensitive data. Based on the approach, we further implement an end-to-end privacy guarantee system, KaPPA, to provide researchers iterative data analysis on sensitive data. The key challenge for differential privacy is determining a reasonable amount of privacy budget to balance privacy preserving and data utility. Most of the previous work applies unified privacy budget to all individual data, which leads to insufficient privacy protection for some individuals while over-protecting others. In KaPPA, the proposed personality-based privacy preserving approach automatically calculates privacy budget for each individual. Our experimental evaluations show a significant trade-off of sufficient privacy protection and data utility. Xuan-Son Vu, Lili Jiang 0002, Anders Brändström, Erik Elmroth |
K-CAP | 1 |
| 2014 | Building a Vietnamese SentiWordNet Using Vietnamese Electronic Dictionary and String Kernel
Xuan-Son Vu, Hyun-Je Song, Seong-Bae Park |
PKAW | 1 |