Son T. Luu

dblp:256/5429 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
16since 2021 · last 2027
0000-0002-1231-5865ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 10 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2027 M-RAV: Multimodal retrieve-augment-verify framework for boosting zero-shot fact verification system with large language models
Son T. Luu, Trung Vo, Minh Le Nguyen 0001
Inf. Process. Manag.1
2026 Natural language processing and computational linguistics for Vietnamese: A comprehensive review
Khiem Vinh Tran, Triet Minh Thai, Son T. Luu, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen
Expert Syst. Appl.3
2026 USCRaKE: Unsupervised semantic chunk retrieval and knowledge-enhanced reasoning for multiple-choice question answering
Trung Vo, Son T. Luu, Minh Le Nguyen 0001
Neurocomputing2
2026 Abnormal job postings detection with explanation on Vietnamese job hunter platforms
Vuong-Quoc Huynh Nguyen, Bao-Khanh Phuoc Truong, Anh Hoang Tran, Duong Tien Pham, Son T. Luu, Trong-Hop Do
Neural Comput. Appl.5
2025 DeepSIX at ACM MM 2025 Grand Challenge: Enhancing Context Text Processing for Multimodal Hallucination Detection and Fact Verification
abstract
Significant advancements have been achieved in both fields of Natural Language Processing (NLP) and Computer Vision (CV) with the advent of Multimodal Large Language Models (MLLMs), sometimes referred to as large vision-language models (LVMs). MLLMs show promising ability in multimodal tasks, such as image captioning, visual question answering, etc. However, there is a concerning trend associated with the advancement in MLLMs. These models exhibit an inclination to generate hallucinations and misleading facts, resulting in seemingly plausible yet factually spurious content. To address these challenges, our team, DeepSIX, leverages recent advances in MLLMs to enhance the ability to detect hallucination and verify factual information within the scope of the ACM MM 2025 grand challenge 8: Truthful and Responsible Multimodal Learning (ResMM). We participated in both tasks: Multimodal Hallucination Detection (Task 1) and Multimodal Fact Checking (Task 2). Our approach leverages the interpretive power of the vision and language components of vision language models (VLMs) to analyze and summarize insights from text and images. It performs contextual reasoning by uncovering semantic relationships among entities in the text and objects in the images. By employing diverse prompting techniques, our method deconstructs critical entities in the text, effectively uncovers implicit relationships between text and images, and identifies hallucinations and false facts. Experimental results demonstrate the strength of our approach: it achieved second place in the Hallucination Detection task and third place in the Fact Verification task, confirming the potential of LLM-based methods in MLLMs. We open-source our code at https://github.com/JAIST-DeepSIX/ACMMM25
Hoang Chu, Huy Chu, Tan-Minh Nguyen, Son T. Luu, Cuong Hoang, Hiep Nguyen, Minh Le Nguyen 0001
ACM Multimedia4
2025 ViMRHP: A Vietnamese Benchmark Dataset for Multimodal Review Helpfulness Prediction via Human-AI Collaborative Annotation
Truc Mai-Thanh Nguyen, Dat Minh Nguyen, Son T. Luu, Kiet Van Nguyen
NLDB (1)3
2025 K-Bloom: unleashing the power of pre-trained language models in extracting knowledge graph with predefined relations
Trung Vo, Son T. Luu, Minh Le Nguyen 0001
Knowl. Inf. Syst.2
2025 MCVE: multimodal claim verification and explanation framework for fact-checking system
Son T. Luu, Trung Vo, Minh Le Nguyen 0001
Multim. Syst.1
2024 VlogQA: Task, Dataset, and Baseline Models for Vietnamese Spoken-Based Machine Reading Comprehension
abstract
This paper presents the development process of a Vietnamese spoken language corpus for machine reading comprehension (MRC) tasks and provides insights into the challenges and opportunities associated with using realworld data for machine reading comprehension tasks.The existing MRC corpora in Vietnamese mainly focus on formal written documents such as Wikipedia articles, online newspapers, or textbooks.In contrast, the VlogQA consists of 10,076 question-answer pairs based on 1,230 transcript documents sourced from YouTube -an extensive source of user-uploaded content, covering the topics of food and travel.By capturing the spoken language of native Vietnamese speakers in natural settings, an obscure corner overlooked in Vietnamese research, the corpus provides a valuable resource for future research in reading comprehension tasks for the Vietnamese language.Regarding performance evaluation, our deep-learning models achieved the highest F1 score of 75.34% on the test set, indicating significant progress in machine reading comprehension for Vietnamese spoken language data.In terms of EM, the highest score we accomplished is 53.97%, which reflects the challenge in processing spoken-based content and highlights the need for further improvement.
Thinh Phuoc Ngo, Khoa Tran Anh Dang, Son T. Luu, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen
EACL (1)3
2024 ZeFaV: Boosting Large Language Models for Zero-Shot Fact Verification
Son T. Luu, Hiep Nguyen, Trung Vo, Minh Le Nguyen 0001
PRICAI (2)1
2024 An approach of data augmentation to improve the performance of BERTology models for Vietnamese hate speech detection
Son T. Luu, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen
Multim. Tools Appl.1
2023 A Text-based Approach For Link Prediction on Wikipedia Articles
abstract
This paper present our work in the DSAA 2023 Challenge about Link Prediction for Wikipedia Articles. We use traditional machine learning models with POS tags (part-of-speech tags) features extracted from text to train the classification model for predicting whether two nodes has the link. Then, we use these tags to test on various machine learning models. We obtained the results by F1 score at 0.99999 and got $7^{th}$ place in the competition. Our source code is publicly available at this link: https://github.com/Tam1032/DSAA2023-Challenge-Link-prediction-DS-UIT_SAT.
Anh Hoang Tran, Tam Minh Nguyen, Son T. Luu
DSAA3
2022 Detecting Spam Reviews on Vietnamese E-Commerce Websites
Co Van Dinh, Son T. Luu, Anh Gia-Tuan Nguyen
ACIIDS (1)2
2022 UIT-ViCoV19QA: A Dataset for COVID-19 Community-based Question Answering on Vietnamese Language
Triet Minh Thai, Ngan Ha-Thao Chu, Anh Tuan Vo, Son T. Luu
PACLIC4
2021 A Large-Scale Dataset for Hate Speech Detection on Vietnamese Social Media Texts
Son T. Luu, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen
IEA/AIE (1)1
2021 Multi-Level Sentiment Analysis of Product Reviews Based on Grammar Rules
abstract
Vietnamese is a tonal and isolated language. Its highly ambiguity makes the designing of methods for sentiment analysis being difficult. For getting the most effectiveness, the designed method has to analyze sentiment of sentences based on combining the grammar and syllable structures of Vietnamese. In this paper, a method to build a Vietnamese dataset of product reviews with many sentiment levels, including very negative, negative, neutral, positive and very positive, is proposed. This method can be scaled to a large dataset using for analyzing sentiment of product reviews. Moreover, a solution to add more grammar rules of Vietnamese into the pre-processing of sentiment analysis is also constructed. Those rules simulate the sentiment recognition of humans and help to increase the accuracy of sentiment determination. The combination of grammar rules and some methods for sentiment analysis are experimented on the Vietnamese dataset of product reviews to classify them into sentiment-levels. The testing results show that their accuracy and F-measure are improved and suitable to apply in the practical business analyzing of customer behaviors.
Hien D. Nguyen 0002, Khiem Vinh Tran, Son T. Luu, Suong N. Hoang, Hieu T. Phan
SoMeT4
2020 Empirical Study of Text Augmentation on Social Media Text in Vietnamese
Son T. Luu, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen
PACLIC1
2020 Measure of the Content Creation Score on Social Network Using Sentiment Score and Passion Point
abstract
Social network is one of efficient tools for spreading information. The evaluation of the content creation of a user is a useful feature to improve the ability of information propagation on social network. In this paper, the measures for evaluating the user’s content creation are proposed. They include the passion point of a user with a brand and the quality of the user’s posts. The passion point is computed based on the sentiment score of posting and the activity of the user. The quality of the user’s posts is computed through the analyzing of the post’s content. Those measures are combined to analyze the interesting of posts. The proposed method has been tested and get the positive experimental results.
Hien D. Nguyen 0002, Tai Huynh, Son T. Luu, Suong N. Hoang, Vuong T. Pham, Ivan Zelinka
SoMeT3
2019 Knowledge representation for designing an Intelligent Tutoring System in learning of courses about Algorithms
abstract
Nowadays building the intelligent systems for education plays an important role to construct the smart city. Knowledge engineering provides technology to build the Intelligent Tutoring System (ITS) supporting for learning. In this paper, a model for representing the knowledge of courses about algorithms is presented. This knowledge model is the integrating between ontology and frames, called Integ-model. Based on the structure of Integ-model, problems about querying on the knowledge and visualizing of algorithms have been proposed and solved. The Integ-model is applied to organize the knowledge base of the intelligent system for learning Data Structures and Algorithms. This system can illustrate algorithms's process and the working of data structures in the course. It also helps learners to query the content of lessons in the course.
Trong T. Le, Son T. Luu, Hien D. Nguyen 0002, Nhon V. Do
APCC2