Quoc-Huy Trinh

dblp:293/9766 · DBLP profile ↗
← Back
13ranked-venue papers
7as first author
13since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Validating polyp and instrument segmentation methods in colonoscopy through Medico 2020 and MedAI 2021 Challenges
abstract
Automatic analysis of colonoscopy images has been an active field of research motivated by the importance of early detection of precancerous polyps. However, detecting polyps during the live examination can be challenging due to various factors such as variation of skills and experience among the endoscopists, lack of attentiveness, and fatigue leading to a high polyp miss-rate. Therefore, there is a need for an automated system that can flag missed polyps during the examination and improve patient care. Deep learning has emerged as a promising solution to this challenge as it can assist endoscopists in detecting and classifying overlooked polyps and abnormalities in real time, improving the accuracy of diagnosis and enhancing treatment. In addition to the algorithm’s accuracy, transparency and interpretability are crucial to explaining the whys and hows of the algorithm’s prediction. Further, conclusions based on incorrect decisions may be fatal, especially in medicine. Despite these pitfalls, most algorithms are developed in private data, closed source, or proprietary software, and methods lack reproducibility. Therefore, to promote the development of efficient and transparent methods, we have organized the “Medico automatic polyp segmentation (Medico 2020)” and “MedAI: Transparency in Medical Image Segmentation (MedAI 2021)” competitions. The Medico 2020 challenge received submissions from 17 teams, while the MedAI 2021 challenge also gathered submissions from another 17 distinct teams in the following year. We present a comprehensive summary and analyze each contribution, highlight the strength of the best-performing methods, and discuss the possibility of clinical translations of such methods into the clinic. Our analysis revealed that the participants improved dice coefficient metrics from 0.8607 in 2020 to 0.8993 in 2021 despite adding diverse and challenging frames (containing irregular, smaller, sessile, or flat polyps), which are frequently missed during a routine clinical examination. For the instrument segmentation task, the best team obtained a mean Intersection over union metric of 0.9364. For the transparency task, a multi-disciplinary team, including expert gastroenterologists, accessed each submission and evaluated the team based on open-source practices, failure case analysis, ablation studies, usability and understandability of evaluations to gain a deeper understanding of the models’ credibility for clinical deployment. The best team obtained a final transparency score of 21 out of 25. Through the comprehensive analysis of the challenge, we not only highlight the advancements in polyp and surgical instrument segmentation but also encourage subjective evaluation for building more transparent and understandable AI-based colonoscopy systems. Moreover, we discuss the need for multi-center and out-of-distribution testing to address the current limitations of the methods to reduce the cancer burden and improve patient care. • We present a detailed analysis of the Medico 2020 and MedAI 2021 challenges that are aimed at advancing automated polyp and instrument segmentation in colonoscopy for early colorectal cancer diagnosis by using novel deep learning methods. • To the best of our knowledge, MedAI 2021 is the first challenge to evaluate the transparency in both GI endoscopy and colonoscopy. Through the challenge, we invited the participants to list package dependencies and architecture code (with instructions for building, compiling, and training) and share trained model weights in a standardized format. Additionally, we invited participants to include the code for model evaluation and provide repository licensing information to enable others to use the code and the trained model responsibly. Moreover, we asked the participants to explain model predictions using intermediate heatmaps, perform ablation studies, conduct a thorough failure analysis, and share their code for reproducing the results. Finally, we performed a subjective evaluation by including an expert gastroenterologist in the group and gave the final transparency score based on the usefulness and understandability of the results. Our initiative aims to promote transparency in AI research and foster the development of reliable, interpretable, and trustworthy algorithms for use in medical image segmentation. • We provide a comparative analysis of the 34 proposed methods in both challenges (3 subtasks), covering small details of each team in the form of Tables, qualitative and quantitative results (failure analysis), and an in-depth analysis of the findings. • We explore trust, safety, interpretability, transparency, and generalizability issues and provide future strategies to overcome the current limitations of developed algorithms.
Debesh Jha, Vanshali Sharma, Debapriya Banik, Debayan Bhattacharya, Kaushiki Roy, Steven Alexander Hicks, Nikhil Kumar Tomar, Vajira Thambawita, Adrian Krenzer, Ge-Peng Ji, Sahadev Poudel, George Batchkala, Saruar Alam, Awadelrahman M. A. Ahmed, Quoc-Huy Trinh, Zeshan Khan, Tien-Phat Nguyen, Shruti Shrestha, Sabari Nathan, Jeonghwan Gwak, Ritika Kumari Jha, Zheyuan Zhang 0001, Alexander Schlaefer, Debotosh Bhattacharjee, Manas Kamal Bhuyan, Pradip K. Das, Deng-Ping Fan, Sravanthi Parasa, Sharib Ali, Michael Riegler 0001, Pål Halvorsen, Thomas de Lange, Ulas Bagci
Medical Image Anal.15
2025 CMATalk: Cross modality alignment for talking head generation
Xuan-Nam Cao, Quoc-Huy Trinh, Minh-Triet Tran
Multim. Tools Appl.2
2024 PGDS: Pose-Guidance Deep Supervision for Mitigating Clothes-Changing in Person Re-Identification
abstract
Person Re-Identification (Re-ID) task seeks to enhance the tracking of multiple individuals by surveillance cameras. It supports multimodal tasks, including text-based person retrieval and human matching. One of the most significant challenges faced in Re-ID is clothes-changing, where the same person may appear in different outfits. While previous methods have made notable progress in maintaining clothing data consistency and handling clothing change data, they still rely excessively on clothing information, which can limit performance due to the dynamic nature of human appearances. To mitigate this challenge, we propose the Pose-Guidance Deep Supervision (PGDS), an effective framework for learning pose guidance within the Re-ID task. It consists of three modules: a human encoder, a pose encoder, and a Pose-to-Human Projection module(PHP). Our framework guides the human encoder, i.e., the main re-identification model, with pose information from the pose encoder through multiple layers via the knowledge transfer mechanism from the PHP module, helping the human encoder learn body parts information without increasing computation resources in the inference stage. Through extensive experiments, our method surpasses the performance of current state-of-the-art methods, demonstrating its robustness and effectiveness for real-world applications. Our code is available at https://github.com/huyquoctrinh/PGDS.
Quoc-Huy Trinh, Nhat-Tan Bui, Dinh-Hieu Hoang, Phuoc-Thao Vo Thi, Hai-Dang Nguyen, Debesh Jha, Ulas Bagci, T. Hoang Ngan Le, Minh-Triet Tran
AVSS1
2024 SAM-EG: Segment Anything Model with Egde Guidance framework for efficient Polyp Segmentation
Quoc-Huy Trinh, Hai-Dang Nguyen, Bao-Tram Nguyen Ngoc, Debesh Jha, Ulas Bagci, Minh-Triet Tran
BMVC1
2024 EAPC: Emotion and Audio Prior Control Framework for the Emotional and Temporal Talking Face Generation
Xuan-Nam Cao, Quoc-Huy Trinh, Quoc-Anh Do-Nguyen, Van-Son Ho, Hoai-Thuong Dang, Minh-Triet Tran
ICAART (2)2
2024 KDAS: Knowledge Distillation via Attention Supervision Framework for Polyp Segmentation
abstract
Polyp segmentation, a contentious issue in medical imaging, has seen numerous proposed methods aimed at improving the quality of segmented masks. While current state-of-the-art techniques yield impressive results, the size and computational cost of these models create challenges for practical industry applications. To address this challenge, we present KDAS, a Knowledge Distillation framework that incorporates attention supervision, and our proposed Symmetrical Guiding Module. This framework is designed to facilitate a compact student model with fewer parameters, allowing it to learn the strengths of the teacher model and mitigate the inconsistency between teacher features and student features, a common challenge in Knowledge Distillation, via the Symmetrical Guiding Module. Through extensive experiments, our compact models demonstrate their strength by achieving competitive results with state-of-the-art methods, offering a promising approach to creating compact models with high accuracy for polyp segmentation and in the medical imaging field. The implementation is available on https://github.com/huyquoctrinh/KDAS.
Quoc-Huy Trinh, Minh-Van Nguyen, Phuoc-Thao Vo Thi
ICME1
2024 Pose-to-Human (P2H): A pose guidance framework via Gram matrix for Occluded Person Re-identification
abstract
The person identification task aims to generate robust human representation embeddings. Traditional methods have achieved competitive results, but they face occlusion challenges, making it difficult to distinguish between humans occluded by objects or other humans. This paper proposes Pose-to-Human (P2H), a pose guidance framework via Gram matrix and knowledge transfer mechanism to tackle occlusion problems in Person Re-identification to address this issue. Our framework comprises four modules: Human Encoder, Pose Encoder, Compact Modules (CPM), and the Gram Matrix Guiding (GMG). The Pose Encoder generates pose information used to guide the human embedding models, facilitating learning of other parts of the human anatomy, which enables the model to focus on the unoccluded parts of the body, thus mitigating the reliance on the data for part-to-part matching, which is the first limitation of previous works. Additionally, it can reduce the demand for general appearance information from the human encoder model. Our method achieves competitive results through extensive experiments on five datasets compared to state-of-the-art methods, making it a promising framework for addressing the Re-Identification task.
Quoc-Huy Trinh, Phuoc-Thao Vo Thi, Minh-Triet Tran, Hai-Dang Nguyen
KES1
2023 Meta-Polyp: A Baseline for Efficient Polyp Segmentation
abstract
In recent years, polyp segmentation has gained significant importance, and many methods have been developed using CNN, Vision Transformer, and Transformer techniques to achieve competitive results. However, these methods often face difficulties when dealing with out-of-distribution datasets, missing boundaries, and small polyps. In 2022, Meta-Former was introduced as a new baseline for vision, which not only improved the performance of multi-task computer vision but also addressed the limitations of the Vision Transformer and CNN family backbones. To further enhance segmentation, we propose a fusion of Meta-Former with UNet, along with the introduction of a Multi-scale Upsampling block with a level-up combination in the decoder stage to enhance the texture, also we propose the Convformer block base on the idea of the Meta-former to enhance the crucial information of the local feature. These blocks enable the combination of global information, such as the overall shape of the polyp, with local information and boundary information, which is crucial for the decision of the medical segmentation. Our proposed approach achieved competitive performance and obtained the top result in the State of the Art on the CVC-300 dataset, Kvasir, and CVC-ColonDB dataset. Apart from Kvasir-SEG, others are out-of-distribution datasets.
Quoc-Huy Trinh
CBMS1
2023 Graph for Transformer Feature: A New Approach for Face Anti-Spoofing
abstract
Face recognition is popular nowadays, however, Face antispoofing (FAS) poses a significant challenge for recognition systems due to the threat of external attacks.While many deep learning methods have been proposed to address this issue, they often face challenges in industry settings.Experiments found that patch extraction modules, such as the Vision Transformer and Swin Transformer, are effective for FAS in single images and perform well in industrial environments.From this point, we propose a model that leverages Transformer features and Graph Neural Networks to learn global information and identify correlations between patch features, which are critical for FAS.
Quoc-Huy Trinh, Xuan-Mao Nguyen, Hai-Dang Nguyen
ESANN1
2023 PEFNet: Positional Embedding Feature for Polyp Segmentation
Trong-Hieu Nguyen Mau, Quoc-Huy Trinh, Nhat-Tan Bui, Phuoc-Thao Vo Thi, Minh-Van Nguyen, Xuan-Nam Cao, Minh-Triet Tran, Hai-Dang Nguyen
MMM (2)2
2023 SpeechSyncNet: Speech to Talking Landmark via the fusion of prior frame landmark and the audio
abstract
Accurately generating talking landmarks from audio is critical for creating authentic talking head animations. This is a significant issue in the field of landmark generation from audio and has potential applications in areas such as virtual assistants, education, and entertainment. Previous methods evaluate the effectiveness to generate the talking landmark from audio, however, these methods have limitations such as bias and inconsistencies in landmark generation due to the lack of visual information from the previous frames. In this research, we propose SpeechSyncNet, a baseline that integrates a fusion of the prior landmark information and the audio feature to provide the lack of visual information from audio, also this method aims to improve the consistency of the landmark motion. Our proposed method yields competitive results compared to existing state-of-the-art methods and helps to improve the quality of talking face landmark generation.
Xuan-Nam Cao, Quoc-Huy Trinh, Van-Son Ho, Minh-Triet Tran
VCIP2
2022 Tiny convolution contextual neural network: a lightweight model for skin lesion detection
abstract
Skin Lesion is a controversial disease all over the world, particularly Melanoma which is a kind of skin cancer. In recent years, there are several methods of using the Convolutional Neural Network and Vision Transformer model have been proposed for the detection and classification of skin images and have achieved competitive results. In this paper, we introduce and demonstrate the efficiency of the Tiny Convolution Contextual Neural Network (TCC Neural Network) which is a tiny model with light-weight architecture than the previous architecture, and have fewer parameter than the popular model for classification of nine lesions from skin images. Our proposal achieves 0.75 on accuracy and 0.55 on F1 score with 5.6 million parameters in the Skin Lesion classification task.
Quoc-Huy Trinh, Trong-Hieu Nguyen Mau, Phuoc-Thao Vo Thi, Hai-Dang Nguyen
ICMV1
2021 SHREC 2021: Retrieval of cultural heritage objects
Ivan Sipiran, Patrick Lazo, Cristian López 0001, Milagritos Jimenez, Nihar Bagewadi, Benjamin Bustos, Hieu Dao, Shankar Gangisetty, Martin Hanik, Ngoc-Phuong Ho-Thi, Mike Holenderski, Dmitri Jarnikov, Arniel Labrada, Stefan Lengauer, Roxane Licandro, Dinh-Huan Nguyen, Thang-Long Nguyen-Ho, Luis A. Pérez Rey, Bang-Dang Pham, Reinhold Preiner, Tobias Schreck, Quoc-Huy Trinh, Loek Tonnaer, Christoph von Tycowicz, The-Anh Vu-Le
Comput. Graph.22