VLDB 2026 Research / reviewers in the wild / expert
Minh Nguyen 0001
dblp:83/2833-1
· DBLP profile ↗
29ranked-venue papers
5as first author
23since 2021 · last 2025
0000-0002-2757-8350ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 5 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Legibility vs. Extractability: Crafting Visual Defenses Against Automated OCR
Minh Nguyen 0001, Kien Tran, The Han Huynh |
ACIVS | 1 |
| 2025 | LLM-BotGuard: A Novel Framework for Detecting LLM-Driven Bots With Mixture of Experts and Graph Neural NetworksabstractDetecting social media bots has become increasingly critical due to their detrimental impact on online environments. With the emergence of sophisticated large language models (LLM) such as ChatGPT, bot detection faces new challenges. These bots based on LLMs exhibit human-like behaviors, and it is difficult for traditional detection approaches to identify them effectively. Such conventional methods struggle with the advanced features associated with LLM-driven bots, which possess contextual understanding and mimic human interaction patterns. The significance of detecting LLM-driven bots lies in their increased difficulty of detection and their potential to inflict more covert harm compared with traditional bots. To address these challenges, we propose LLM-BotGuard, a novel detection model that is capable of capturing the unique features of LLM-driven bots alongside other bot characteristics through three key modules, i.e., pattern-informed feature extraction module, mixture of experts module, and graph module with graph sample and aggregation networks. Extensive experiments have been conducted to evaluate the performance of the proposed LLM-BotGuard. The results demonstrate that LLM-BotGuard significantly outperforms baseline methods in detecting LLM-driven bots. The proposed LLM-BotGuard offers a robust solution for identifying sophisticated LLM-driven bots in online social networks. Jinglong Duan, Weihua Li 0007, Quan Bai 0001, Minh Nguyen 0001, Jianhua Jiang |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2025 | Enhancing Emotional Well-Being With IoT Data Solutions for Depression: A Systematic ReviewabstractEffectively caring for adults with depression is challenging. While technology offers potential improvements in emotional well-being through better monitoring, standardised methods to gather and analyse relevant data are highly fragmented. This Systematic Literature Review (SLR) explores using Internet of Things (IoT) based data collection and analysis to enhance emotional well-being and manage depression effectively. Following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) framework, we report in-depth findings from 42 studies, which were selected from an initial set of 559 published works. We find that current literature extensively covers important topics like IoT for detecting, analysing, and monitoring emotions, therapeutic interventions for emotional well-being, and predicting, detecting, and managing depression. IoT-based data collection and analysis solutions predominantly employ sensors and AI, respectively. The literature review identifies a gap in prioritising active systems that engage users, highlighting the need to address key aspects such as privacy and security. Sanaz Zamani, Roopak Sinha, Minh Nguyen 0001, Samaneh Madanian |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | Intent-Spectrum BotTracker: Tackling LLM-Based Social Media Bots Through an Enhanced BotRGCN Model with Intention and Entropy Measurement
Jinglong Duan, Weihua Li 0007, Quan Bai 0001, Minh Nguyen 0001 |
PKAW | 6 |
| 2024 | Enhancing Remote Sensing Image Retrieval: A Hierarchical Approach Integrating Visual and Semantic SimilaritiesabstractThe heightened revisiting frequency and expanded observation capabilities of satellites lead to the daily generation of a substantial volume of remote sensing images. Retrieving relevant data accurately from this extensive archive holds significant importance. Deep learning Content-Based Image Retrieval (CBIR) utilizes a feature extraction network pre-trained on image classification tasks to derive image-level features. Subsequently, a similarity measure is applied on these features to identify the archive images most closely resembling the query image. While image-level labels facilitate CBIR in retrieving images from the same category as the query image, they do not empower CBIR to differentiate between implicit sub-categories. For instance, although CBIR can discern between broader categories like “residential” and “forest”, it lacks the necessary semantic statistical information to distinguish more nuanced distinctions such as “high-density residential” from “medium-density residential”. To enhance image retrieval for greater similarity, we propose a Hierarchical Image Retrieval (HIR) approach that combines visual similarity with semantic statistics. In the first stage, visually similar images are identified using CBIR, while the second stage refines the selection based on semantic similarity derived from land-cover classification. The experimental results indicate that HIR achieves 20% higher retrieval accuracy for “residential” sub-categories and over 1% increase in retrieval accuracy across all classes. Wen Lu 0001, Minh Nguyen 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | CISO: Co-iteration semi-supervised learning for visual object detectionabstractAbstract Semi-supervised learning offers a solution to the high cost and limited availability of manually labeled samples in supervised learning. In semi-supervised visual object detection, the use of unlabeled data can significantly enhance the performance of deep learning models. In this paper, we introduce an end-to-end framework, named CISO (Co-Iteration Semi-Supervised Learning for Object Detection), which integrates a knowledge distillation approach and a collaborative, iterative semi-supervised learning strategy. To maximize the utilization of pseudo-label data and address the scarcity of pseudo-label data due to high threshold settings, we propose a mean iteration approach where all unlabeled data is applied to each training iteration. Pseudo-label data with high confidence is extracted based on an ever-changing threshold (average intersection over union of all pseudo-labeled data). This strategy not only ensures the accuracy of the pseudo-label but also optimizes the use of unlabeled data. Subsequently, we apply a weak-strong data augmentation strategy to update the model. Lastly, we evaluate CISO using Swin Transformer model and conduct comprehensive experiments on MS-COCO. Our framework showcases impressive results, outperforms the state-of-the-art methods by 2.16 mAP and 1.54 mAP with 10% and 5% labeled data, respectively. Jianchun Qi, Minh Nguyen 0001, Wei Qi Yan 0001 |
Multim. Tools Appl. | 2 |
| 2024 | NUNI-Waste: novel semi-supervised semantic segmentation waste classification with non-uniform data augmentationabstractAbstract Waste categorization and recycling are critical approaches for converting waste into valuable and functional materials, thereby significantly aiding in land preservation, reducing pollution, and optimizing resource usages. However, real-world classification and identification of recyclable waste face substantial hurdles due to the intricate and unpredictable nature of wastes, as well as the limited availability of comprehensive waste datasets. These factors limit efficacy of the existing research work in the domain of waste management. In this paper, we utilize semantic segmentation at individual pixel level and introduce a semi-supervised metod for authentic waste classification scenarios, leveraging the Zerowaste dataset. We devise a non-standard data augmentation strategy that mimics the ever-changing conditions of real-world waste environments. Additionally, we introduce an adaptive weighted loss function and dynamically adjust the ratio of positive to negative samples through a masking method, ensuring the model learns from relevant samples. Lastly, to maintain consistency between predictions made on data-augmented images and the original counterparts, we remove input perturbations. Our method proves to be effective, as verified by an array of standard experiments and ablation studies, achieved an accuracy improvement of 3.74% over the baseline Zerowaste method. Jianchun Qi, Minh Nguyen 0001, Wei Qi Yan 0001 |
Multim. Tools Appl. | 2 |
| 2024 | Apple ripeness identification from digital images using transformersabstractAbstract We describe a non-destructive test of apple ripeness using digital images of multiple types of apples. In this paper, fruit images are treated as data samples, artificial intelligence models are employed to implement the classification of fruits and the identification of maturity levels. In order to obtain the ripeness classifications of fruits, we make use of deep learning models to conduct our experiments; we evaluate the test results of our proposed models. In order to ensure the accuracy of our experimental results, we created our own dataset, and obtained the best accuracy of fruit classification by comparing Transformer model and YOLO model in deep learning, thereby attaining the best accuracy of fruit maturity recognition. At the same time, we also combined YOLO model with attention module and gave the fast object detection by using the improved YOLO model. Bingjie Xiao, Minh Nguyen 0001, Wei Qi Yan 0001 |
Multim. Tools Appl. | 2 |
| 2024 | Fruit ripeness identification using YOLOv8 modelabstractAbstract Deep learning-based visual object detection is a fundamental aspect of computer vision. These models not only locate and classify multiple objects within an image, but they also identify bounding boxes. The focus of this paper's research work is to classify fruits as ripe or overripe using digital images. Our proposed model extracts visual features from fruit images and analyzes fruit peel characteristics to predict the fruit's class. We utilize our own datasets to train two "anchor-free" models: YOLOv8 and CenterNet, aiming to produce accurate predictions. The CenterNet network primarily incorporates ResNet-50 and employs the deconvolution module DeConv for feature map upsampling. The final three branches of convolutional neural networks are applied to predict the heatmap. The YOLOv8 model leverages CSP and C2f modules for lightweight processing. After analyzing and comparing the two models, we found that the C2f module of the YOLOv8 model significantly enhances classification results, achieving an impressive accuracy rate of 99.5%. Bingjie Xiao, Minh Nguyen 0001, Wei Qi Yan 0001 |
Multim. Tools Appl. | 2 |
| 2024 | A Lightweight Transformer With Multigranularity Tokens and Connected Component Loss for Land Cover ClassificationabstractOnboard land cover classification provides ever-updating land cover information, supporting various intelligent satellite applications that demand timely autonomous decision-making based on current and continuous land cover data. However, due to space, weight, and power constraints, satellites possess limited computational resources, rendering them unable to execute conventional land cover classification networks. In response to this challenge, we have designed a lightweight network for land cover classification featuring two efficient Transformer attention mechanisms enhanced by Multi-Granularity Tokens. Diverging from traditional Transformer attention mechanisms that solely capture token-to-token correlations at a single granularity, our approach splits the tokens into four segments and utilizes Atrous Convolutions across various dilation rates to aggregate token segments from diverse receptive fields, forming token segments combinations that encompass not just point information but also information from patches of varying sizes. These Multi-Granularity Tokens are subsequently processed through the Windowed Squeeze Axial Transformer Attention (WSATA) and Multi-Granularity Bi-level Routing Attention (MGBRA) for feature enhancement. In another aspect, empirical observations reveal that prediction errors are more prone to manifest on land covers of small extent, however, conventional methods treat all pixels uniformly. This realization motivates us to propose a novel network-agnostic loss named Connected Component Loss, which specifically targets small-scale land covers and their boundaries. Quantitative metrics and visual interpretations from comprehensive experiments confirm that our method attains state-of-the-art accuracy on two land cover classification datasets while exhibits significantly faster inference speed than other lightweight networks, underscoring the practical potential of our method on embedded systems. Wen Lu 0001, Minh Nguyen 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | A Lightweight CNN-Transformer Network With Laplacian Loss for Low-Altitude UAV Imagery Semantic SegmentationabstractSemantic segmentation is crucial for enabling autonomous flight and landing of low-altitude Unmanned Aerial Vehicles (UAVs) and is indispensable for various intelligent applications. However, real-time semantic segmentation is a computationally intensive task because it involves pixel-wise classification, which renders conventional semantic segmentation networks impractical for deployment on embedded systems of limited hardware resources. Moreover, variations in flight height and object appearance increase the likelihood of misjudgment in segmentation results. To address these challenges, we propose an efficient approach consisting of a CNN-Transformer network and an auxiliary loss. The encoder of the network integrates a newly designed module, which equally handles objects with varying scales. The decoder is composed of the innovative Query-Value Squeeze Axial Transformer Attention, which reduces computational complexity from quadratic in terms of image size toO(2C(H2+W2)), linear in terms of image size. By incorporating Laplacian operator convolution, the novel network-agnostic loss effectively captures intricate patterns, boundaries, and small objects. This enables extra penalization of misjudgments in these areas and compels the network to focus on objects that are challenging to distinguish. Our approach attains impressive accuracy when processing 4K resolution images in real-time (15 FPS) on a mobile GPU. It demonstrates over 2x faster speed compared to representative lightweight networks, underscoring its suitability for onboard deployment. Wen Lu 0001, Minh Nguyen 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Enhanced Color QR Codes with Resilient Error Correction for Dirt-Prone Surfaces
Minh Nguyen 0001 |
ACIVS | 1 |
| 2023 | Exploring the Potential of Image Overlay in Self-supervised Learning: A Study on SimSiam Networks and Strategies for Preventing Model Collapse
Weihua Li 0007, Quan Bai 0001, Minh Nguyen 0001 |
PKAW | 4 |
| 2023 | A High-Accuracy Deformable Model for Human Face Mask Detection
Xinyi Gao 0002, Minh Nguyen 0001, Wei Qi Yan 0001 |
PSIVT | 2 |
| 2023 | Enhancement of Human Face Mask Detection Performance by Using Ensemble Learning Models
Xinyi Gao 0002, Minh Nguyen 0001, Wei Qi Yan 0001 |
PSIVT | 2 |
| 2023 | Enhancing Safety During Surgical Procedures with Computer Vision, Artificial Intelligence, and Natural Language Processing
Okeke Stephen, Minh Nguyen 0001 |
PSIVT | 2 |
| 2023 | Multiscale Kiwifruit Detection from Digital Images
Minh Nguyen 0001, Raymond Lutui, Wei Qi Yan 0001 |
PSIVT | 2 |
| 2023 | Computational Analysis of Table Tennis Matches from Real-Time Videos Using Deep Learning
Minh Nguyen 0001, Wei Qi Yan 0001 |
PSIVT | 2 |
| 2023 | Fruit ripeness identification using transformersabstractAbstract Pattern classification has always been essential in computer vision. Transformer paradigm having attention mechanism with global receptive field in computer vision improves the efficiency and effectiveness of visual object detection and recognition. The primary purpose of this article is to achieve the accurate ripeness classification of various types of fruits. We create fruit datasets to train, test, and evaluate multiple Transformer models. Transformers are fundamentally composed of encoding and decoding procedures. The encoder is to stack the blocks, like convolutional neural networks (CNN or ConvNet). Vision Transformer (ViT), Swin Transformer, and multilayer perceptron (MLP) are considered in this paper. We examine the advantages of these three models for accurately analyzing fruit ripeness. We find that Swin Transformer achieves more significant outcomes than ViT Transformer for both pears and apples from our dataset. Bingjie Xiao, Minh Nguyen 0001, Wei Qi Yan 0001 |
Appl. Intell. | 2 |
| 2023 | Sign language recognition from digital videos using feature pyramid network with detection transformerabstractAbstract Sign language recognition is one of the fundamental ways to assist deaf people to communicate with others. An accurate vision-based sign language recognition system using deep learning is a fundamental goal for many researchers. Deep convolutional neural networks have been extensively considered in the last few years, and a slew of architectures have been proposed. Recently, Vision Transformer and other Transformers have shown apparent advantages in object recognition compared to traditional computer vision models such as Faster R-CNN, YOLO, SSD, and other deep learning models. In this paper, we propose a Vision Transformer-based sign language recognition method called DETR (Detection Transformer), aiming to improve the current state-of-the-art sign language recognition accuracy. The DETR method proposed in this paper is able to recognize sign language from digital videos with a high accuracy using a new deep learning model ResNet152 + FPN (i.e., Feature Pyramid Network), which is based on Detection Transformer. Our experiments show that the method has excellent potential for improving sign language recognition accuracy. For instance, our newly proposed net ResNet152 + FPN is able to enhance the detection accuracy up to 1.70% on the test dataset of sign language compared to the standard Detection Transformer models. Besides, an overall accuracy 96.45% was attained by using the proposed method. Parma Nand, Md. Akbar Hossain, Minh Nguyen 0001, Wei Qi Yan 0001 |
Multim. Tools Appl. | 4 |
| 2022 | A Method for Face Image Inpainting Based on Autoencoder and Generative Adversarial Network
Xinyi Gao 0002, Minh Nguyen 0001, Wei Qi Yan 0001 |
PSIVT | 2 |
| 2022 | Waste Classification from Digital Images Using ConvNeXt
Jianchun Qi, Minh Nguyen 0001, Wei Qi Yan 0001 |
PSIVT | 2 |
| 2022 | Traffic Sign Recognition from Digital Images by Using Deep Learning
Jiawei Xing, Ziyuan Luo, Minh Nguyen 0001, Wei Qi Yan 0001 |
PSIVT | 3 |
| 2020 | Red-Green-Blue Augmented Reality Tags for Retail Stores
Minh Nguyen 0001, Wei Qi Yan 0001 |
ACIVS | 1 |
| 2019 | A Web-based Augmented Reality Plat-form using Pictorial QR Code for Educational Purposes and BeyondabstractAugmented Reality (AR) provides the capability to overlay virtual 3D information onto a 2D printed flat surface; for example, displaying a 3D model on a single flat card that accompanies with the diagram shown in a learning text-book. The student can zoom in and out, rotate, and perceive the animation of the figure in real-time. This will make the educational theory more attractive; hence, motivates students to learn. AR is a great tool; however, the setup and display are not straight-forward (there are many different AR markers with different encryption, decryption methods, and displaying flat-forms). In this paper, we proposed a portable browser-based platform which uses the advantages of AR along with scan-able QR Code on mobile phones to enhance instant 3D visualisation. The user only needs a smart-phone (Apple iPhone or Android) with Internet-enabled; no specific Apps are needed to install. The user scans the QR Code embedded in a colour image, the code will link to a public website, and the website will produce AR Experience right on top of the browser. As a result, it provides a stress-free, low-cost, portable, and promising solution for not only educational purposes but also many other fields such as gaming, property selling, e-commerce, reporting. The set up is convenient: the user uploads a picture (e.g. a racing car), and what actions to be related to it (a 3D model to display, or a movie to play). The system will add on the picture one small colour QR code (to redirect to an online URL) and a thin black border. The user also uploads the 3D model (GLTF files) that he wants to display on top of the card to finish the set-up. At the display, the user can print the AR card, point their smart-phone towards the card, and pre-setup AR models or actions will appear on it. To students, these 3D graphics or animations will allow them to learn and understand the lessons in a much more intuitive way. Minh Nguyen 0001, Minh Phu Lai, Wei Qi Yan 0001 |
VRST | 1 |
| 2018 | Human Behaviour Recognition Using Deep LearningabstractTraditional human behaviour recognition is mostly based on global features of digital images. Nowadays, with the increase of computing power and processing capacity, deep neural networks (DNNs) acquire a high possibility to detect any objects, which have effectively led to a new era of machine learning. In this paper, we investigated a human behaviour recognition using deep learning based on YOLOv3 model. After a number of experiments conducted, our YOLOv3 model had shown to achieve 80.20% of accuracy in human behaviour recognition with the speed of approximate 15 fps using GPU acceleration. Our direct contributions are: (1) data augment and collection, (2) adjusting deep neural network structures, and (3) superior performance in evaluations for our proposed deep learning model. Wei Qi Yan 0001, Minh Nguyen 0001 |
AVSS | 3 |
| 2018 | Enhancing Visualisation of Anatomical Presentation and Education Using Marker-based Augmented Reality Technology on Web-based PlatformabstractThe domain of teaching medicine involves the mastery of many complex skills that almost always need to be performed in real life situations following very high professional standards. However, this training is not always possible for various reasons such as ethics, safety and costs. Virtual reality (VR) and augmented reality (AR) have started to be widely used as alternative medical teaching practices. Our proposed AR system works online, so no installation is required on a users device; it only requires a generic colour web-cam to track a pictorial AR tag which contains a hidden QR code. These QR codes contain data, such as the ID of a three-dimensional (3D) model, and merge it with text so that it can be used as both an identity and tag pattern in the AR marker. The system can then show the corresponding computer-generated 3D anatomical models of organs; relevant text information about the subject is displayed above the AR Tag. The tag is numerically encrypted and decrypted and detectable by shape and orientation. Different from other similar techniques, our AR Tag is both a bar-code and a template marker, QR code is used to load a previously setup website, and then the detail of that QR code is used as a template to identify the border and orientation of the marker. The system is thus faster and more robust that allows users to control and navigate the 3D environment by zooming in and out and rotating left and right. It is hoped that this virtual environment will help reduce the need for real-life surgical practice, instead of increasing intuition, the direct 3D perception of the human body and other 3D medical imaging data (mimesis). This system could even be further developed to present the framework of a patient's anatomy. Minh Nguyen 0001, Hui Le, Wei Qi Yan 0001, Steffan Hooper |
AVSS | 2 |
| 2017 | Enhancing Textbook Study Experiences with Pictorial Bar-Codes and Augmented Reality
Minh Nguyen 0001 |
CAIP (2) | 2 |
| 2017 | A tile based colour picture with hidden QR code for augmented reality and beyondabstractMost existing Augmented Reality (AR) applications use either template (picture) markers or bar-code markers to overlay computer-generated graphics on the real world surfaces. The use of template markers is computationally expensive and unreliable. On the other hand, bar-code markers display only black and white blocks; thus, they look uninteresting and uninformative. In this short paper, we describe a new way to optically hide a QR code inside a tile based colour picture. Each AR marker is built from hundreds of small tiles (just like tiling a bathroom), and the unique gaps between the tiles are used to determine the elements of the hidden QR Code. This novel type of AR marker presents not only a realistic-looking colour picture but also contains self-Correcting information (stored in QR code). In this article, we demonstrate that this tile based colour picture with hidden QR code is relatively robust under various conditions and scaling. We believe many nowadays' AR challenges could be solved with this type of marker. AR-enabled medias could then be easily generated. For instance, it would be capable of storing and displaying virtual figures of an entire book or magazine. Thus, it provides a promising AR approach to be used in many different AR applications; and beyond, it may even replace the barcodes and QR Codes in some cases. Minh Nguyen 0001, Wei Qi Yan 0001 |
VRST | 1 |