Vincent Nguyen 0001

dblp:20/2617 · also Nhu-Van Nguyen · DBLP profile ↗
← Back
20ranked-venue papers
12as first author
9since 2021 · last 2026
0000-0003-2271-6918ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 9 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 3 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Reason-to-Learn (R2L): Multi-Agent Knowledge Distillation for Lightweight LLMs in Sentiment Analysis
Le-Huy Tu, Vincent Nguyen 0001, Johanna Björklund, Xuan-Son Vu
LREC3
2026 SAMix: Calibrated and Accurate Continual Learning via Sphere-Adaptive Mixup and Neural Collapse
abstract
Abstract While most continual learning methods focus on mitigating forgetting and improving accuracy, they often overlook the critical aspect of network calibration, despite its importance. Neural collapse, a phenomenon where last-layer features collapse to their class means, has demonstrated advantages in continual learning by reducing feature-classifier misalignment. Few works aim to improve the calibration of continual models for more reliable predictions. Our work goes a step further by proposing a novel method that not only enhances calibration but also improves performance by reducing overconfidence, mitigating forgetting, and increasing accuracy. We introduce Sphere-Adaptive Mixup (SAMix), an adaptive mixup strategy tailored for neural collapse-based methods. SAMix adapts the mixing process to the geometric properties of feature spaces under neural collapse, ensuring more robust regularization and alignment. Experiments show that SAMix significantly boosts performance, surpassing SOTA methods in continual learning while also improving model calibration. SAMix enhances both across-task accuracy and the broader reliability of predictions, making it a promising advancement for robust continual learning systems.
Trung-Anh Dang, Vincent Nguyen 0001, Ngoc-Son Vu, Christel Vrain
Mach. Learn.2
2026 Beyond labels: Semi-supervised underwater habitat detection with image enhancement as augmentation
abstract
Abstract Recent progress in deep learning for object detection has largely relied on big, well-labeled datasets. But when it comes to underwater imagery, annotation becomes a real challenge. It typically requires domain experts and a lot of manual effort. As a result, building large-scale labeled datasets in this field is not only time-consuming but also often impractical. In this paper, we present an in-depth investigation of semi-supervised learning models for marine habitat detection, with the goal of reducing the need for large amounts of labeled data while maintaining strong performance in the complex conditions found underwater. We evaluate these models using the Deepfish and UTDAC2020 datasets. By extensive experimentation, we explore how various factors influence performance: the amount of labeled training data, the role of contrastive learning, and the application of underwater image enhancement as an augmentation technique. Overall, this work provides a detailed analysis and highlights the potential of these techniques in marine habitat detection.
Rim Rahali, Vincent Nguyen 0001
Multim. Tools Appl.2
2025 Memory-efficient Continual Learning with Neural Collapse Contrastive
abstract
Contrastive learning has significantly improved representation quality, enhancing knowledge transfer across tasks in continual learning (CL). However, catastrophic forgetting remains a key challenge, as contrastive based methods primarily focus on “soft relationships” or “softness” between samples, which shift with changing data distributions and lead to representation overlap across tasks. Recently, the newly identified Neural Collapse phenomenon has shown promise in CL by focusing on ““““hard relationships” or “hardness” between samples and fixed proto-types. However, this approach overlooks “softness”, crucial for capturing intra-class variability, and this rigid focus can also pull old class representations toward current ones, increasing forgetting. Building on these insights, we propose Focal Neural Collapse Contrastive$(FNC^{2})$, a novel representation learning loss that effectively balances both soft and hard relationships. Additionally, we introduce the Hardness-Softness Distillation (HSD) loss to progressively preserve the knowledge gained from these relationships across tasks. Our method outperforms state-of-the-art approaches, particularly in minimizing memory reliance. Remarkably, even without the use of memory, our approach rivals rehearsal-based methods, offering a compelling solution for data privacy concerns.
Trung-Anh Dang, Vincent Nguyen 0001, Ngoc-Son Vu, Christel Vrain
WACV2
2025 Accumulating global channel-wise patterns via deformed-bottleneck recalibration for image classification
Thanh Tuan Nguyen 0001, Thanh Phuong Nguyen 0001, Vincent Nguyen 0001
Pattern Anal. Appl.3
2023 False Positive Reduction of Pulmonary Nodule on CT image using Attention-based Multiple Instance Learning
abstract
Diagnosis and treatment of multiple pulmonary nodules are clinically essential but challenging. Pulmonary nodule detection is important in early lung cancer detection and diagnosis. False positive reduction (FPR) is a significant stage of pulmonary nodule detection systems. Prior investigations on nodule candidate classification use solitary-nodule approaches, which ignore the relations between nodules. In this study, we propose to use Attention-based Deep Multiple Instance Learning, a variation of supervised learning to recognize true pulmonary nodules among a large group of candidates proposed from the detection stage. By treating the multiple nodules from a different patient, critical relational information between solitary-nodule is extracted and empirically proves the benefit of learning the relations between multiple nodules. An attention layer trained with CNN to replace typical pooling-based aggregation in multiple instance learning (MIL). Experiments of lung nodule FPR on the public LUNA16 dataset validate the effectiveness of the proposed method. The proposed method achieved an accuracy of 99.6%, specificity of 100%, recall of 99.92%, and F1 score of 99.6%. The experimental results reveal that our method can achieve satisfactory performance in FPR.
Chi Cuong Nguyen, Giang Son Tran, Vincent Nguyen 0001
CBMI3
2021 ICDAR 2021 Competition on Historical Map Segmentation
Joseph Chazalon, Edwin Carlinet, Yizi Chen, Julien Perret, Bertrand Dumenieu, Clément Mallet, Thierry Géraud, Vincent Nguyen 0001, Josef Baloun, Ladislav Lenc, Pavel Král
ICDAR (4)8
2021 Manga-MMTL: Multimodal Multitask Transfer Learning for Manga Character Analysis
Vincent Nguyen 0001, Christophe Rigaud, Arnaud Revel, Jean-Christophe Burie
ICDAR (2)1
2021 ICDAR 2021 Competition on Multimodal Emotion Recognition on Comics Scenes
Vincent Nguyen 0001, Xuan-Son Vu, Christophe Rigaud, Lili Jiang 0002, Jean-Christophe Burie
ICDAR (4)1
2020 An adaptive document recognition system for lettrines
Vincent Nguyen 0001, Mickaël Coustaty, Jean-Marc Ogier
Int. J. Document Anal. Recognit.1
2020 A learning approach with incomplete pixel-level labels for deep neural networks
Vincent Nguyen 0001, Christophe Rigaud, Arnaud Revel, Jean-Christophe Burie
Neural Networks1
2019 Post-OCR Error Detection by Generating Plausible Candidates
abstract
The accuracy of Optical Character Recognition (OCR) technologies considerably impacts the way digital documents are indexed, accessed and exploited. Post-processing approaches detect and correct remaining errors to improve the quality of OCR texts. However, state-of-the-art approaches still need to be improved. Most of the existing post-OCR techniques use predefined error position lists or apply simple techniques to detect errors. In this paper, we describe a novel error detector using different features from character-level (including character noisy channel, index of peculiarity) to word-level (such as frequencies of n-grams, skip-grams, part-of-speech) Experimental results show that our approach outperforms the best performing techniques in the ICDAR 2017 Competition on Post-OCR text correction.
Thi-Tuyet-Hai Nguyen, Adam Jatowt, Mickaël Coustaty, Vincent Nguyen 0001, Antoine Doucet
ICDAR4
2019 Multi-task Model for Comic Book Image Analysis
Vincent Nguyen 0001, Christophe Rigaud, Jean-Christophe Burie
MMM (2)1
2019 Comic MTL: optimized multi-task learning for comic book image analysis
Vincent Nguyen 0001, Christophe Rigaud, Jean-Christophe Burie
Int. J. Document Anal. Recognit.1
2015 Keyword Visual Representation for Image Retrieval and Image Annotation
abstract
Keyword-based image retrieval is more comfortable for users than content-based image retrieval. Because of the lack of semantic description of images, image annotation is often used a priori by learning the association between the semantic concepts (keywords) and the images (or image regions). This association issue is particularly difficult but interesting because it can be used for annotating images but also for multimodal image retrieval. However, most of the association models are unidirectional, from image to keywords. In addition to that, existing models rely on a fixed image database and prior knowledge. In this paper, we propose an original association model, which provides image-keyword bidirectional transformation. Based on the state-of-the-art Bag of Words model dealing with image representation, including a strategy of interactive incremental learning, our model works well with a zero-or-weak-knowledge image database and evolving from it. Some objective quantitative and qualitative evaluations of the model are proposed, in order to highlight the relevance of the method.
Vincent Nguyen 0001, Alain Boucher, Jean-Marc Ogier
Int. J. Pattern Recognit. Artif. Intell.1
2014 Multi-modal and Cross-Modal for Lecture Videos Retrieval
abstract
The problem of multi-modal and cross-modal lecture videos retrieval is studied in this paper, on the basis of the use of document analysis techniques. In the context of this paper, a lecture video is represented by a set of subjects, in which a subject is represented by a Bag of mixed words -visual words and textual words-, each of them coming from speech recognition and OCR engines. Our work relies on two assumptions 1) a video may contain multiple subjects, 2) multiple modalities exist in the same lecture video document. We propose in this research a combination of technologies issuing from image document analysis and text mining. Visual words and textual words in images of lecture slides are extracted based on text detection and graphics localization computed on the sequences captured with a camera. Assuming that a subject in the video composes of a set of slides, lecture slides are clustered in different groups representing different possible subjects by using mixed words extracted. Multimodal and cross-modal lecture video retrieval are realized by the Bag of Subjects model. We discuss the proposed indexing and retrieval approach for lecture videos and report a quantitative evaluation on lecture videos of our University. It is shown that using Bag of Subjects for lecture video retrieval improves the retrieval accuracy.
Vincent Nguyen 0001, Mickaël Coustaty, Jean-Marc Ogier
ICPR1
2013 Bag of subjects: lecture videos multimodal indexing
abstract
In this paper, we address multimodal indexing and retrieval for videos of lectures or seminars. This paper proposes a combination of technologies respectively issuing from image document analysis and text mining. Based on visual information and textual information extracted from slide images, we investigate a Bag of mixed Words (visual words and textual words) model to represent lecture slide's contents. Lecture videos are indexed and retrieved by using extended Bag of Words model. In this model, it is assumed that a video may contain multiple subjects; and this model discovers the visual representation of these subjects automatically and indexes the video accordingly. We discuss the mixed text/image query and proposed indexing approach for retrieval lecture videos and report a quantitative evaluation on lecture videos of our Lab.
Vincent Nguyen 0001, Jean-Marc Ogier, Franck Charneau
ACM Symposium on Document Engineering1
2013 Interactive Knowledge Learning for Ancient Images
abstract
This paper deals with cultural heritage preservation and ancient document indexing. In the management of historical documents, ancient images are described using semantic information, often manually annotated by historians. In this paper, we propose an approach to interactively propagate the historians' knowledge to a database of drop caps images manually populated by historians with drop caps image annotations. Based on a novel document indexing processing scheme which combines the use of the Zipf law and the use of bag of patterns, our approach extends the Bag of Words model to represent the knowledge by visual features through relevance feedback. Then annotation propagation is automatically performed to propagate knowledge to the drop caps image database. In this article, our approach is presented together with preliminary experimental results and an illustrative example.
Vincent Nguyen 0001, Mickaël Coustaty, Alain Boucher, Jean-Marc Ogier
ICDAR1
2012 PEDIVHANDI: Multimodal Indexation and Retrieval System for Lecture Videos
Vincent Nguyen 0001, Jean-Marc Ogier, Franck Charneau
ACCV (2)1
2009 Region-Based Semi-automatic Annotation Using the Bag of Words Representation of the Keywords
abstract
Automatic Image Annotation (AIA) tries to minimize the manual effort for image annotation. However, the performance of the AIA approaches is not satisfactory. The interaction of user is needed to solve this problem. The annotation is refined during the interaction by using semantic-based relevance feedback. This approach has a limit as only the annotations of found images during the interaction are updated. In this paper we introduce a novel method of semi-automatic annotation. The method is using visual feature representations of keywords which are improved during the region-based relevance feedback. The experiments show that this method gives good results and can be used to update the annotations for all images.
Vincent Nguyen 0001, Alain Boucher, Jean-Marc Ogier, Salvatore Tabbone
ICIG1