VLDB 2026 Research / reviewers in the wild / expert
T. Hoang Ngan Le
dblp:37/245 · also Ngan Le, Ngan T. H. Le, Thi Hoang Ngan Le
· DBLP profile ↗
83ranked-venue papers
16as first author
58since 2021 · last 2026
0000-0003-2571-0511ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 51 · 10 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 51 · 10 first-author · 36 since 2021Systems, architecture and hardware · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 5 since 2021Security and privacy · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric PerspectiveabstractAs embodied agents operate in increasingly complex environments, the ability to perceive, track, and reason about individual object instances over time becomes essential, especially in tasks requiring sequenced interactions with visually similar objects. In non-Markovian settings, critical decision cues lie in object histories rather than the current scene. Without persistent memory of prior interactions (what was used, where it was placed, or how it changed), visuomotor policies may fail, repeat past actions, or overlook completed ones. To surface this challenge, we introduce LIBERO-Mem, a non-Markovian task suite for stress-testing robotic manipulation under object-level partial observability. It combines short- and long-horizon object tracking with temporally sequenced subgoals, requiring reasoning beyond the current frame. However, vision-language-action (VLA) models often struggle in such settings, with token scaling quickly becoming intractable even for tasks spanning just a few hundred frames. We propose Embodied-SlotSSM, a slot-centric VLA framework built for temporal scalability. It maintains spatio-temporally consistent slot identities and leverages them through two mechanisms: (1) slot-state-space modeling for reconstructing short-term history, and (2) a relational encoder to align the input tokens with action decoding. Together, these components enable temporally grounded, context-aware action prediction. Experiments show Embodied-SlotSSM's baseline performance on LIBERO-Mem and general tasks, offering a scalable solution for non-Markovian reasoning in object-centric policies. Nhat Chung, Taisei Hanyu, Toan Nguyen 0004, Huy Le 0001, Frederick Bumgarner, Duy M. H. Nguyen, Viet-Khoa Vo-Ho, Kashu Yamazaki, Chase Rainwater, Tung Kieu, Anh Nguyen 0003, T. Hoang Ngan Le |
AAAI | 12 |
| 2026 | UNO: Unifying One-stage Video Scene Graph Generation via Object-Centric Visual Representation LearningabstractVideo Scene Graph Generation (VidSGG) aims to represent dynamic visual content by detecting objects and modeling their temporal interactions as structured graphs. Prior studies typically target either coarse-grained box-level or fine-grained panoptic pixel-level VidSGG, often requiring task-specific architectures and multi-stage training pipelines. In this paper, we present UNO (UNified Object-centric VidSGG), a single-stage, unified framework that jointly addresses both tasks within an end-to-end architecture. UNO is designed to minimize task-specific modifications and maximize parameter sharing, enabling generalization across different levels of visual granularity. The core of UNO is an extended slot attention mechanism that decomposes visual features into object and relation slots. To ensure robust temporal modeling, we introduce object temporal consistency learning, which enforces consistent object representations across frames without relying on explicit tracking modules. Additionally, a dynamic triplet prediction module links relation slots to corresponding object pairs, capturing evolving interactions over time. We evaluate UNO on standard box-level and pixel-level VidSGG benchmarks. Results demonstrate that UNO not only achieves competitive performance across both tasks but also offers improved efficiency through a unified, object-centric design. Huy Le 0001, Nhat Chung, Tung Kieu, T. Hoang Ngan Le |
WACV | 5 |
| 2025 | EgoMusic-Driven Human Dance Motion Estimation with Skeleton Mamba
Nhat Le, Baoru Huang, Minh Nhat Vu, Chengcheng Tang, T. Hoang Ngan Le, Thieu Vo, Anh Nguyen 0003 |
ICCV | 7 |
| 2025 | CT-ScanGaze: A Dataset and Baselines for 3D Volumetric Scanpath ModelingabstractUnderstanding radiologists' eye movement during Computed Tomography (CT) reading is crucial for developing effective interpretable computer-aided diagnosis systems. However, CT research in this area has been limited by the lack of publicly available eye-tracking datasets and the three-dimensional complexity of CT volumes. To address these challenges, we present the first publicly available eye gaze dataset on CT, called CT-ScanGaze. Then, we introduce CT-Searcher, a novel 3D scanpath predictor designed specifically to process CT volumes and generate radiologist-like 3D fixation sequences, overcoming the limitations of current scanpath predictors that only handle 2D inputs. Since deep learning models benefit from a pretraining step, we develop a pipeline that converts existing 2D gaze datasets into 3D gaze data to pretrain CT-Searcher. Through both qualitative and quantitative evaluations on CT-ScanGaze, we demonstrate the effectiveness of our approach and provide a comprehensive assessment framework for 3D scanpath prediction in medical imaging. Trong-Thang Pham, Akash Awasthi, Saba Khan, Esteban Duran Marti, Tien-Phat Nguyen, Viet-Khoa Vo-Ho, Cuong Tran 0010, Yuki Ikebe, Anh Totti Nguyen, Anh Nguyen 0003, Zhigang Deng 0001, Carol C. Wu, T. Hoang Ngan Le |
ICCV | 16 |
| 2025 | Robotic-CLIP: Fine-Tuning CLIP on Action Data for Robotic ApplicationsabstractVision language models have played a key role in extracting meaningful features for various robotic applications. Among these, Contrastive Language-Image Pretraining (CLIP) is widely used in robotic tasks that require both vision and natural language understanding. However, CLIP was trained solely on static images paired with text prompts and has not yet been fully adapted for robotic tasks involving dynamic actions. In this paper, we introduce Robotic-CLIP to enhance robotic perception capabilities. We first gather and label large-scale action data, and then build our Robotic-CLIP by fine-tuning CLIP on 309,433 videos (≈ 7.4 million frames) of action data using contrastive learning. By leveraging action data, Robotic-CLIP inherits CLIP's strong image performance while gaining the ability to understand actions in robotic contexts. Intensive experiments show that our Robotic-CLIP outperforms other CLIP-based models across various language-driven robotic tasks. Additionally, we demonstrate the practical effectiveness of Robotic-CLIP in real-world grasping applications. Minh Nhat Vu, Tung D. Ta, Baoru Huang, Thieu Vo, T. Hoang Ngan Le, Anh Nguyen 0003 |
ICRA | 6 |
| 2025 | MAARTA:Multi-agentic Adaptive Radiology Teaching Assistant
Akash Awasthi, Brandon V. Chung, Anh M. Vu, T. Hoang Ngan Le, Rishi Agrawal, Zhigang Deng 0001, Carol C. Wu, Hien Van Nguyen |
MICCAI (5) | 4 |
| 2025 | BiMa: Towards Biases Mitigation for Text-Video Retrieval via Scene Element GuidanceabstractText-video retrieval (TVR) systems often suffer from visual-linguistic biases present in datasets, which cause pre-trained vision-language models to overlook key details. To address this, we propose BiMa, a novel framework designed to mitigate biases in both visual and textual representations. Our approach begins by generating scene elements that characterize each video by identifying relevant entities/objects and activities. For visual debiasing, we integrate these scene elements into the video embeddings, enhancing them to emphasize fine-grained and salient details. For textual debiasing, we introduce a mechanism to disentangle text features into content and bias components, enabling the model to focus on meaningful content while separately handling biased information. Extensive experiments and ablation studies across five major TVR benchmarks (i.e., MSR-VTT, MSVD, LSMDC, ActivityNet, and DiDeMo) demonstrate the competitive performance of BiMa. Additionally, the model's bias mitigation capability is consistently validated by its strong results on out-of-distribution retrieval tasks. Huy Le 0001, Nhat Chung, Tung Kieu, Anh Nguyen 0003, T. Hoang Ngan Le |
ACM Multimedia | 5 |
| 2025 | TolerantECG: A Foundation Model for Imperfect ElectrocardiogramabstractThe electrocardiogram (ECG) is an essential and effective tool for diagnosing heart diseases. However, its effectiveness can be compromised by noise or unavailability of one or more leads of the standard 12-lead recordings, resulting in diagnostic errors or uncertainty. To address these challenges, we propose TolerantECG, a foundation model for ECG signals that is robust to noise and capable of functioning with arbitrary subsets of the standard 12-lead ECG. TolerantECG training combines contrastive and self-supervised learning frameworks to jointly learn ECG signal representations alongside their corresponding knowledge-retrieval-based text report descriptions and corrupted or lead-missing signals. Comprehensive benchmarking results demonstrate that TolerantECG consistently ranks as the best or second-best performer across various ECG signal conditions and class levels in the PTB-XL dataset, and achieves the highest performance on the MIT-BIH Arrhythmia Database. The source is available at this link: https://github.com/Fsoft-AIC/TolerantECG Huynh Dang Nguyen, Trong-Thang Pham, T. Hoang Ngan Le |
ACM Multimedia | 3 |
| 2025 | Interpreting Radiologist's Intention from Eye Movements in Chest X-ray Diagnosis
Trong-Thang Pham, Anh Nguyen 0003, Zhigang Deng 0001, Carol C. Wu, T. Hoang Ngan Le |
ACM Multimedia | 6 |
| 2025 | Learning Human Motion with Temporally Conditional MambaabstractLearning human motion based on a time-dependent input signal presents a challenging yet impactful task with various applications. The goal of this task is to generate or estimate human movement that consistently reflects the temporal patterns of conditioning inputs. Existing methods typically rely on cross-attention mechanisms to fuse the condition with motion. However, this approach primarily captures global interactions and struggles to maintain step-by-step temporal alignment. To address this limitation, we introduce Temporally Conditional Mamba, a new mamba-based model for human motion generation. Our approach integrates conditional information into the recurrent dynamics of the Mamba block, enabling better temporally aligned motion. To validate the effectiveness of our method, we evaluate it on a variety of human motion tasks. Extensive experiments demonstrate that our model significantly improves temporal alignment, motion realism, and condition consistency over state-of-the-art approaches. Our project page is available at https://zquang2202.github.io/TCM. Baoru Huang, Minh Nhat Vu, T. Hoang Ngan Le, Thieu Vo, Anh Nguyen 0003 |
SIGGRAPH Asia | 5 |
| 2025 | GazeSearch: Radiology Findings Search BenchmarkabstractMedical eye-tracking data is an important information source for understanding how radiologists visually inter-pret medical images. This information not only improves the accuracy of deep learning models for X-ray analysis but also their interpretability, enhancing transparency in decision-making. However, the current eye-tracking data is dispersed, unprocessed, and ambiguous, making it difficult to derive meaningful insights. Therefore, there is a need to create a new dataset with more focus and purposeful eye-tracking data, improving its utility for diagnostic applications. In this work, we propose a refinement method inspired by the target-present visual search challenge: there is a specific finding and fixations are guided to locate it. After re-fining the existing eye-tracking datasets, we transform them into a curated visual search dataset, called Gazesearch. specifically for radiology findings, where each fixation sequence is purposefully aligned to the task of locating a particular finding. Subsequently, we introduce a scan path prediction baseline, called ChestSearch, specifically tailored to Gazesearch. Finally, we employ the newly introduced Gazesearch as a benchmark to evaluate the performance of current state-of-the-art methods, offering a comprehensive assessment for visual search in the medical imaging domain. Code is available at https://github.com/UARK-AICV/GazeSearch. Trong-Thang Pham, Tien-Phat Nguyen, Yuki Ikebe, Akash Awasthi, Zhigang Deng 0001, Carol C. Wu, T. Hoang Ngan Le |
WACV | 8 |
| 2025 | ItpCtrl-AI: End-to-end interpretable and controllable artificial intelligence by modeling radiologists' intentions
Trong-Thang Pham, Jacob Brecheisen, Carol C. Wu, Hien Van Nguyen, Zhigang Deng 0001, Donald A. Adjeroh, Gianfranco Doretto, Arabinda Choudhary, T. Hoang Ngan Le |
Artif. Intell. Medicine | 9 |
| 2025 | A2VIS: Amodal-Aware Approach to Video Instance Segmentation
Thang Pham, Winston Bounsavy, Tri Nguyen 0005, T. Hoang Ngan Le |
Image Vis. Comput. | 5 |
| 2025 | Structural chain of thoughts for radiology education
Akash Awasthi, Brandon Chung, Anh M. Vu, Saba Khan, T. Hoang Ngan Le, Zhigang Deng 0001, Rishi Agrawal, Carol C. Wu, Hien Van Nguyen |
Knowl. Based Syst. | 5 |
| 2024 | Guide3D: A Bi-planar X-ray Dataset for 3D Shape Reconstruction
Tudor Jianu, Baoru Huang, Hoan Nguyen, Binod Bhattarai, Tuong KL. Do, Erman Tjiputra, Quang D. Tran, Pierre Berthet-Rayne, T. Hoang Ngan Le, Sebastiano Fichera, Anh Nguyen 0003 |
ACCV (5) | 9 |
| 2024 | FG-CXR: A Radiologist-Aligned Gaze Dataset for Enhancing Interpretability in Chest X-Ray Report Generation
Trong-Thang Pham, Ngoc-Vuong Ho, Nhat-Tan Bui, Thinh Phan, Brijesh Patel 0001, Donald A. Adjeroh, Gianfranco Doretto, Anh Nguyen 0003, Carol C. Wu, T. Hoang Ngan Le |
ACCV (6) | 11 |
| 2024 | Amodal Instance Segmentation with Diffusion Shape Prior Estimation
Viet-Khoa Vo-Ho, Tri Nguyen 0005, T. Hoang Ngan Le |
ACCV (10) | 4 |
| 2024 | PGDS: Pose-Guidance Deep Supervision for Mitigating Clothes-Changing in Person Re-IdentificationabstractPerson Re-Identification (Re-ID) task seeks to enhance the tracking of multiple individuals by surveillance cameras. It supports multimodal tasks, including text-based person retrieval and human matching. One of the most significant challenges faced in Re-ID is clothes-changing, where the same person may appear in different outfits. While previous methods have made notable progress in maintaining clothing data consistency and handling clothing change data, they still rely excessively on clothing information, which can limit performance due to the dynamic nature of human appearances. To mitigate this challenge, we propose the Pose-Guidance Deep Supervision (PGDS), an effective framework for learning pose guidance within the Re-ID task. It consists of three modules: a human encoder, a pose encoder, and a Pose-to-Human Projection module(PHP). Our framework guides the human encoder, i.e., the main re-identification model, with pose information from the pose encoder through multiple layers via the knowledge transfer mechanism from the PHP module, helping the human encoder learn body parts information without increasing computation resources in the inference stage. Through extensive experiments, our method surpasses the performance of current state-of-the-art methods, demonstrating its robustness and effectiveness for real-world applications. Our code is available at https://github.com/huyquoctrinh/PGDS. Quoc-Huy Trinh, Nhat-Tan Bui, Dinh-Hieu Hoang, Phuoc-Thao Vo Thi, Hai-Dang Nguyen, Debesh Jha, Ulas Bagci, T. Hoang Ngan Le, Minh-Triet Tran |
AVSS | 8 |
| 2024 | Language-Driven 6-DoF Grasp Detection Using Negative Prompt Guidance
Toan Nguyen 0004, Minh Nhat Vu, Baoru Huang, An Vuong, T. Hoang Ngan Le, Thieu Vo, Anh Nguyen 0003 |
ECCV (19) | 6 |
| 2024 | WAVER: Writing-Style Agnostic Text-Video Retrieval Via Distilling Vision-Language Models Through Open-Vocabulary KnowledgeabstractText-video retrieval, a prominent sub-field within the domain of multimodal information retrieval, has witnessed remarkable growth in recent years. However, existing methods assume video scenes are consistent with unbiased descriptions. These limitations fail to align with real-world scenarios since descriptions can be influenced by annotator biases, diverse writing styles, and varying textual perspectives. To overcome the aforementioned problems, we introduce WAVER, a cross-domain knowledge distillation framework via vision-language models through open-vocabulary knowledge designed to tackle the challenge of handling different writing styles in video descriptions. WAVER capitalizes on the open-vocabulary properties that lie in pre-trained vision-language models and employs an implicit knowledge distillation approach to transfer text-based knowledge from a teacher model to a vision-based student. Empirical studies conducted across four standard benchmark datasets, encompassing various settings, provide compelling evidence that WAVER can achieve state-of-the-art performance in text-video retrieval task while handling writing-style variations. The code is available at: https://github.com/Fsoft-AIC/WAVER Huy Le 0001, Tung Kieu, Anh Nguyen 0003, T. Hoang Ngan Le |
ICASSP | 4 |
| 2024 | Language-Conditioned Affordance-Pose Detection in 3D Point CloudsabstractAffordance detection and pose estimation are of great importance in many robotic applications. Their combination helps the robot gain an enhanced manipulation capability, in which the generated pose can facilitate the corresponding affordance task. Previous methods for affodance-pose joint learning are limited to a predefined set of affordances, thus limiting the adaptability of robots in real-world environments. In this paper, we propose a new method for language-conditioned affordance-pose joint learning in 3D point clouds. Given a 3D point cloud object, our method detects the affordance region and generates appropriate 6-DoF poses for any unconstrained affordance label. Our method consists of an open-vocabulary affordance detection branch and a language-guided diffusion model that generates 6-DoF poses based on the affordance text. We also introduce a new high-quality dataset for the task of language-driven affordance-pose joint learning. Intensive experimental results demonstrate that our proposed method works effectively on a wide range of open-vocabulary affordances and outperforms other baselines by a large margin. In addition, we illustrate the usefulness of our method in real-world robotic applications. Our code and dataset are publicly available at https://3DAPNet.github.io. Toan Nguyen 0004, Minh Nhat Vu, Baoru Huang, Tuan Van Vo, Vy Truong, T. Hoang Ngan Le, Thieu Vo, Bac Le, Anh Nguyen 0003 |
ICRA | 6 |
| 2024 | Open-Vocabulary Affordance Detection using Knowledge Distillation and Text-Point CorrelationabstractAffordance detection presents intricate challenges and has a wide range of robotic applications. Previous works have faced limitations such as the complexities of 3D object shapes, the wide range of potential affordances on real-world objects, and the lack of open-vocabulary support for affordance understanding. In this paper, we introduce a new open-vocabulary affordance detection method in 3D point clouds, leveraging knowledge distillation and text-point correlation. Our approach employs pre-trained 3D models through knowledge distillation to enhance feature extraction and semantic understanding in 3D point clouds. We further introduce a new text-point correlation method to learn the semantic links between point cloud features and open-vocabulary labels. The intensive experiments show that our approach outperforms previous works and adapts to new affordance labels and unseen objects. Notably, our method achieves the improvement of 7.96% mIOU score compared to the baselines. Furthermore, it offers real-time inference which is well-suitable for robotic manipulation applications. Tuan Van Vo, Minh Nhat Vu, Baoru Huang, Toan Nguyen 0004, T. Hoang Ngan Le, Thieu Vo, Anh Nguyen 0003 |
ICRA | 5 |
| 2024 | Open-Fusion: Real-time Open-Vocabulary 3D Mapping and Queryable Scene RepresentationabstractPrecise 3D environmental mapping with semantics is essential in robotics. Existing methods often rely on pre-defined concepts during training or are time-intensive when generating semantic maps. This paper presents Open-Fusion, an approach for real-time open-vocabulary 3D mapping and queryable scene representation using RGB-D data. Open-Fusion harnesses the power of a pretrained vision-language foundation model (VLFM) for open-set semantic comprehension and employs the Truncated Signed Distance Function (TSDF) for swift 3D scene reconstruction. By leveraging the VLFM, we extract region-based embeddings and their associated confidence maps. These are then integrated with the 3D knowledge from TSDF using an enhanced Hungarian-based feature-matching mechanism. In particular, Open-Fusion delivers outstanding annotation-free 3D segmentation for open vocabulary query without the need for additional 3D training. Benchmark tests on the ScanNet dataset against leading zero-shot methods highlight Open-Fusion’s superiority. Furthermore, it seamlessly combines the strengths of region-based VLFM and TSDF, facilitating real-time 3D scene comprehension that includes object concepts and open-world semantics. We encourage the readers to view the demos on our project page: https://uark-aicv.github.io/OpenFusion Kashu Yamazaki, Taisei Hanyu, Viet-Khoa Vo-Ho, Thang Pham, Gianfranco Doretto, Anh Nguyen 0003, T. Hoang Ngan Le |
ICRA | 8 |
| 2024 | ShapeFormer: Shape Prior Visible-to-Amodal Transformer-based Amodal Instance SegmentationabstractAmodal Instance Segmentation (AIS) presents a challenging task as it involves predicting both visible and occluded parts of objects within images. Existing AIS methods rely on a bidirectional approach, encompassing both the transition from amodal features to visible features (amodal-to-visible) and from visible features to amodal features (visible-to-amodal). Our observation shows that the utilization of amodal features through the amodal-to-visible can confuse the visible features due to the extra information of occluded/hidden segments not presented in visible display. Consequently, this compromised quality of visible features during the subsequent visible-to-amodal transition. To tackle this issue, we introduce ShapeFormer, a decoupled Transformer-based model with a visible-to-amodal transition. It facilitates the explicit relationship between output segmentations and avoids the need for amodal-to-visible transitions. ShapeFormer comprises three key modules: (i) Visible-Occluding Mask Head for predicting visible segmentation with occlusion awareness, (ii) Shape-Prior Amodal Mask Head for predicting amodal and occluded masks, and (iii) Category-Specific Shape Prior Retriever aims to provide shape prior knowledge. Comprehensive experiments and extensive ablation studies across various AIS benchmarks demonstrate the effectiveness of our ShapeFormer. The code is available at: https: //github.com/UARK-AICV/ShapeFormer Winston Bounsavy, Viet-Khoa Vo-Ho, Anh Nguyen 0003, Tri Nguyen 0005, T. Hoang Ngan Le |
IJCNN | 6 |
| 2024 | Unifying Global and Local Scene Entities Modelling for Precise Action SpottingabstractSports videos pose complex challenges, including cluttered backgrounds, camera angle changes, small action-representing objects, and imbalanced action class distribution. Existing methods for detecting actions in sports videos heavily rely on global features, utilizing a backbone network as a black box that encompasses the entire spatial frame. However, these approaches tend to overlook the nuances of the scene and struggle with detecting actions that occupy a small portion of the frame. In particular, they face difficulties when dealing with action classes involving small objects, such as balls or yellow/red cards in soccer, which only occupy a fraction of the screen space. To address these challenges, we introduce a novel approach that analyzes and models scene entities using an adaptive attention mechanism. Particularly, our model disentangles the scene content into the global environment feature and local relevant scene entities feature. To efficiently extract environmental features while considering temporal information with less computational cost, we propose the use of a 2D backbone network with a time-shift mechanism. To accurately capture relevant scene entities, we employ a Vision-Language model in conjunction with the adaptive attention mechanism. Our model has demonstrated outstanding performance, securing the 1st place in the SoccerNet-v2 Action Spotting, FineDiving, and FineGym challenge with a substantial performance improvement of 1.6, 2.0, and 1.3 points in avg-mAP compared to the runner-up methods. Furthermore, our approach offers interpretability capabilities in contrast to other deep learning models, which are often designed as black boxes. Our code and models are released at: https://github.com/Fsoft-AIC/unifying-global-local-feature. Kim Hoang Tran, Phuc Vuong Do, Ngoc Quoc Ly, T. Hoang Ngan Le |
IJCNN | 4 |
| 2024 | Lightweight Language-driven Grasp Detection using Conditional Consistency ModelabstractLanguage-driven grasp detection is a fundamental yet challenging task in robotics with various industrial applications. This work presents a new approach for language-driven grasp detection that leverages lightweight diffusion models to achieve fast inference time. By integrating diffusion processes with grasping prompts in natural language, our method can effectively encode visual and textual information, enabling more accurate and versatile grasp positioning that aligns well with the text query. To overcome the long inference time problem in diffusion models, we leverage the image and text features as the condition in the consistency model to reduce the number of denoising timesteps during inference. The intensive experimental results show that our method outperforms other recent grasp detection methods and lightweight diffusion models by a clear margin. We further validate our method in real-world robotic experiments to demonstrate its fast inference time capability. Minh Nhat Vu, Baoru Huang, An Vuong, T. Hoang Ngan Le, Thieu Vo, Anh Nguyen 0003 |
IROS | 5 |
| 2024 | Language-driven Grasp Detection with Mask-guided AttentionabstractGrasp detection is an essential task in robotics with various industrial applications. However, traditional methods often struggle with occlusions and do not utilize language for grasping. Incorporating natural language into grasp detection remains a challenging task and largely unexplored. To address this gap, we propose a new method for language-driven grasp detection with mask-guided attention by utilizing the transformer attention mechanism with semantic segmentation features. Our approach integrates visual data, segmentation mask features, and natural language instructions, significantly improving grasp detection accuracy. Our work introduces a new framework for language-driven grasp detection, paving the way for language-driven robotic applications. Intensive experiments show that our method outperforms other recent baselines by a clear margin, with a 10.0% success score improvement. We further validate our method in real-world robotic experiments, confirming the effectiveness of our approach. Tuan Van Vo, Minh Nhat Vu, Baoru Huang, An Vuong, T. Hoang Ngan Le, Thieu Vo, Anh Nguyen 0003 |
IROS | 5 |
| 2024 | HENASY: Learning to Assemble Scene-Entities for Interpretable Egocentric Video-Language ModelabstractCurrent video-language models (VLMs) rely extensively on instance-level alignment between video and language modalities, which presents two major limitations: (1) visual reasoning disobeys the natural perception that humans do in first-person perspective, leading to a lack of reasoning interpretation; and (2) learning is limited in capturing inherent fine-grained relationships between two modalities.
In this paper, we take an inspiration from human perception and explore a compositional approach for egocentric video representation. We introduce HENASY (Hierarchical ENtities ASsemblY), which includes a spatiotemporal token grouping mechanism to explicitly assemble dynamically evolving scene entities through time and model their relationship for video representation. By leveraging compositional structure understanding, HENASY possesses strong interpretability via visual grounding with free-form text queries. We further explore a suite of multi-grained contrastive losses to facilitate entity-centric understandings. This comprises three alignment types: video-narration, noun-entity, verb-entities alignments.
Our method demonstrates strong interpretability in both quantitative and qualitative experiments; while maintaining competitive performances on five downstream tasks via zero-shot transfer or as video/text representation, including video/text retrieval, action recognition, multi-choice query, natural language query, and moments query.
Project page: https://uark-aicv.github.io/HENASY Viet-Khoa Vo-Ho, Thinh Phan, Kashu Yamazaki, T. Hoang Ngan Le |
NeurIPS | 5 |
| 2024 | DINTR: Tracking via Diffusion-based InterpolationabstractObject tracking is a fundamental task in computer vision, requiring the localization of objects of interest across video frames. Diffusion models have shown remarkable capabilities in visual generation, making them well-suited for addressing several requirements of the tracking problem. This work proposes a novel diffusion-based methodology to formulate the tracking task. Firstly, their conditional process allows for injecting indications of the target object into the generation process. Secondly, diffusion mechanics can be developed to inherently model temporal correspondences, enabling the reconstruction of actual frames in video. However, existing diffusion models rely on extensive and unnecessary mapping to a Gaussian noise domain, which can be replaced by a more efficient and stable interpolation process. Our proposed interpolation mechanism draws inspiration from classic image-processing techniques, offering a more interpretable, stable, and faster approach tailored specifically for the object tracking task. By leveraging the strengths of diffusion models while circumventing their limitations, our Diffusion-based INterpolation TrackeR (DINTR) presents a promising new paradigm and achieves a superior multiplicity on seven benchmarks across five indicator representations. Pha A. Nguyen, T. Hoang Ngan Le, Jackson David Cothren, Alper Yilmaz 0001, Khoa Luu |
NeurIPS | 2 |
| 2024 | Accelerating Transformers with Spectrum-Preserving Token MergingabstractIncreasing the throughput of the Transformer architecture, a foundational component used in numerous state-of-the-art models for vision and language tasks (e.g., GPT, LLaVa), is an important problem in machine learning. One recent and effective strategy is to merge token representations within Transformer models, aiming to reduce computational and memory requirements while maintaining accuracy. Prior work has proposed algorithms based on Bipartite Soft Matching (BSM), which divides tokens into distinct sets and merges the top $k$ similar tokens. However, these methods have significant drawbacks, such as sensitivity to token-splitting strategies and damage to informative tokens in later layers. This paper presents a novel paradigm called PiToMe, which prioritizes the preservation of informative tokens using an additional metric termed the \textit{energy score}. This score identifies large clusters of similar tokens as high-energy, indicating potential candidates for merging, while smaller (unique and isolated) clusters are considered as low-energy and preserved. Experimental findings demonstrate that PiToMe saved from 40-60\% FLOPs of the base models while exhibiting superior off-the-shelf performance on image classification (0.5\% average performance drop of ViT-MAEH compared to 2.6\% as baselines), image-text retrieval (0.3\% average performance drop of Clip on Flick30k compared to 4.5\% as others), and analogously in visual questions answering with LLaVa-7B. Furthermore, PiToMe is theoretically shown to preserve intrinsic spectral properties to the original token space under mild conditions. Chau Tran, Duy M. H. Nguyen, Duy Nguyen 0003, TrungTin Nguyen, T. Hoang Ngan Le, Pengtao Xie, Daniel Sonntag, James Zou 0001, Mathias Niepert |
NeurIPS | 5 |
| 2024 | MEGANet: Multi-Scale Edge-Guided Attention Network for Weak Boundary Polyp SegmentationabstractEfficient polyp segmentation in healthcare plays a critical role in enabling early diagnosis of colorectal cancer. However, the segmentation of polyps presents numerous challenges, including the intricate distribution of backgrounds, variations in polyp sizes and shapes, and indistinct boundaries. Defining the boundary between the foreground (i.e. polyp itself) and the background (surrounding tissue) is difficult. To mitigate these challenges, we propose Multi-Scale Edge-Guided Attention Network (MEGANet) tailored specifically for polyp segmentation within colonoscopy images. This network draws inspiration from the fusion of a classical edge detection technique with an attention mechanism. By combining these techniques, MEGANet effectively preserves high-frequency information, notably edges and boundaries, which tend to erode as neural networks deepen. MEGANet is designed as an end-to-end framework, encompassing three key modules: an encoder, which is responsible for capturing and abstracting the features from the input image, a decoder, which focuses on salient features, and the Edge-Guided Attention module (EGA) that employs the Laplacian Operator to accentuate polyp boundaries. Extensive experiments, both qualitative and quantitative, on five benchmark datasets, demonstrate that our MEGANet outperforms other existing SOTA methods under six evaluation metrics. Our code is available at https://github.com/UARK-AICV/MEGANet. Nhat-Tan Bui, Dinh-Hieu Hoang, Minh-Triet Tran, T. Hoang Ngan Le |
WACV | 5 |
| 2024 | I-AI: A Controllable & Interpretable AI System for Decoding Radiologists' Intense Focus for Accurate CXR DiagnosesabstractIn the field of chest X-ray (CXR) diagnosis, existing works often focus solely on determining where a radiologist looks, typically through tasks such as detection, segmentation, or classification. However, these approaches are often designed as black-box models, lacking interpretability. In this paper, we introduce Interpretable Artificial Intelligence (I-AI) a novel and unified controllable interpretable pipeline for decoding the intense focus of radiologists in CXR diagnosis. Our I-AI addresses three key questions: where a radiologist looks, how long they focus on specific areas, and what findings they diagnose. By capturing the intensity of the radiologist’s gaze, we provide a unified solution that offers insights into the cognitive process underlying radiological interpretation. Unlike current methods that rely on black-box machine learning models, which can be prone to extracting erroneous information from the entire input image during the diagnosis process, we tackle this issue by effectively masking out irrelevant information. Our proposed I-AI leverages a vision-language model, allowing for precise control over the interpretation process while ensuring the exclusion of irrelevant features.To train our I-AI model, we utilize an eye gaze dataset to extract anatomical gaze information and generate ground truth heatmaps. Through extensive experimentation, we demonstrate the efficacy of our method. We showcase that the attention heatmaps, designed to mimic radiologists’ focus, encode sufficient and relevant information, enabling accurate classification tasks using only a portion of CXR. The code, checkpoints, and data are at https://github.com/UARK-AICV/IAI. Trong-Thang Pham, Jacob Brecheisen, Anh Nguyen 0003, T. Hoang Ngan Le |
WACV | 5 |
| 2024 | ZEETAD: Adapting Pretrained Vision-Language Model for Zero-Shot End-to-End Temporal Action DetectionabstractTemporal action detection (TAD) involves the localization and classification of action instances within untrimmed videos. While standard TAD follows fully supervised learning with closed-set setting on large training data, recent zero-shot TAD methods showcase the promising open-set setting by leveraging large-scale contrastive visual-language (ViL) pretrained models. However, existing zero-shot TAD methods have limitations on how to properly construct the strong relationship between two interdependent tasks of localization and classification and adapt ViL model to video understanding. In this work, we present ZEE-TAD, featuring two modules: dual-localization and zero-shot proposal classification. The former is a Transformer-based module that detects action events while selectively collecting crucial semantic embeddings for later recognition. The latter one, CLIP-based module, generates semantic embeddings from text and frame inputs for each temporal unit. Additionally, we enhance discriminative capability on unseen classes by minimally updating the frozen CLIP encoder with lightweight adapters. Extensive experiments on THUMOS14 and ActivityNet-1.3 datasets demonstrate our approach’s superior performance in zero-shot TAD and effective knowledge transfer from ViL models to unseen action categories. Code is available at https: //github.com/UARK-AICV/ZEETAD. Thinh Phan, Viet-Khoa Vo-Ho, Duy Le 0004, Gianfranco Doretto, Donald A. Adjeroh, T. Hoang Ngan Le |
WACV | 6 |
| 2024 | Multi-camera multi-object tracking on the move via single-stage global association approach
Pha A. Nguyen, Kha Gia Quach, Chi Nhan Duong, Son Lam Phung, T. Hoang Ngan Le, Khoa Luu |
Pattern Recognit. | 5 |
| 2023 | VLTinT: Visual-Linguistic Transformer-in-Transformer for Coherent Video Paragraph CaptioningabstractVideo Paragraph Captioning aims to generate a multi-sentence description of an untrimmed video with multiple temporal event locations in a coherent storytelling. Following the human perception process, where the scene is effectively understood by decomposing it into visual (e.g. human, animal) and non-visual components (e.g. action, relations) under the mutual influence of vision and language, we first propose a visual-linguistic (VL) feature. In the proposed VL feature, the scene is modeled by three modalities including (i) a global visual environment; (ii) local visual main agents; (iii) linguistic scene elements. We then introduce an autoregressive Transformer-in-Transformer (TinT) to simultaneously capture the semantic coherence of intra- and inter-event contents within a video. Finally, we present a new VL contrastive loss function to guarantee the learnt embedding features are consistent with the captions semantics. Comprehensive experiments and extensive ablation studies on the ActivityNet Captions and YouCookII datasets show that the proposed Visual-Linguistic Transformer-in-Transform (VLTinT) outperforms previous state-of-the-art methods in terms of accuracy and diversity. The source code is made publicly available at: https://github.com/UARK-AICV/VLTinT. Kashu Yamazaki, Viet-Khoa Vo-Ho, Quang Sang Truong, Bhiksha Raj, T. Hoang Ngan Le |
AAAI | 5 |
| 2023 | FREDOM: Fairness Domain Adaptation Approach to Semantic Scene UnderstandingabstractAlthough Domain Adaptation in Semantic Scene Segmentation has shown impressive improvement in recent years, the fairness concerns in the domain adaptation have yet to be well defined and addressed. In addition, fairness is one of the most critical aspects when deploying the segmentation models into human-related real-world applications, e.g., autonomous driving, as any unfair predictions could influence human safety. In this paper, we propose a novel Fairness Domain Adaptation (FREDOM) approach to semantic scene segmentation. In particular, from the proposed formulated fairness objective, a new adaptation framework will be introduced based on the fair treatment of class distributions. Moreover, to generally model the context of structural dependency, a new conditional structural constraint is introduced to impose the consistency of predicted segmentation. Thanks to the proposed Conditional Structure Network, the self-attention mechanism has sufficiently modeled the structural information of segmentation. Through the ablation studies, the proposed method has shown the performance improvement of the segmentation models and promoted fairness in the model predictions. The experimental results on the two standard benchmarks, i.e., SYNTHIA$\rightarrow$Cityscapes and GTA5$\rightarrow$Cityscapes, have shown that our method achieved State-of-the-Art (SOTA) performance11The implementation of FREDOM is available at https://github.com/uark-cviu/FREDOM Thanh-Dat Truong, T. Hoang Ngan Le, Bhiksha Raj, Jackson David Cothren, Khoa Luu |
CVPR | 2 |
| 2023 | CLIP-TSA: Clip-Assisted Temporal Self-Attention for Weakly-Supervised Video Anomaly DetectionabstractVideo anomaly detection (VAD) – commonly formulated as a multiple-instance learning problem in a weakly-supervised manner due to its labor-intensive nature – is a challenging problem in video surveillance where the frames of anomaly need to be localized in an untrimmed video. In this paper, we first propose to utilize the ViT-encoded visual features from CLIP, in contrast with the conventional C3D or I3D features in the domain, to efficiently extract discriminative representations in the novel technique. We then model temporal dependencies and nominate the snippets of interest by leveraging our proposed Temporal Self-Attention (TSA). The ablation study confirms the effectiveness of TSA and ViT feature. The extensive experiments show that our proposed CLIP-TSA outperforms the existing state-of-the-art (SOTA) methods by a large margin on three commonly-used benchmark datasets in the VAD problem (UCF-Crime, ShanghaiTech Campus and XD-Violence). Our source code is available at https://github.com/joos2010kj/CLIP-TSA. Hyekang Joo, Viet-Khoa Vo-Ho, Kashu Yamazaki, T. Hoang Ngan Le |
ICIP | 4 |
| 2023 | Open-Vocabulary Affordance Detection in 3D Point CloudsabstractAffordance detection is a challenging problem with a wide variety of robotic applications. Traditional affordance detection methods are limited to a predefined set of affordance labels, hence potentially restricting the adaptability of intelligent robots in complex and dynamic environments. In this paper, we present the Open-Vocabulary Affordance Detection (OpenAD) method, which is capable of detecting an unbounded number of affordances in 3D point clouds. By simultaneously learning the affordance text and the point feature, OpenAD successfully exploits the semantic relationships between affordances. Therefore, our proposed method enables zero-shot detection and can be able to detect previously unseen affordances without a single annotation example. Intensive experimental results show that OpenAD works effectively on a wide range of affordance detection setups and outperforms other baselines by a large margin. Additionally, we demonstrate the practicality of the proposed OpenAD in real-world robotic applications with a fast inference speed. Our project is available at https://openad2023.github.io. Toan Nguyen 0004, Minh Nhat Vu, An Vuong, Dzung Nguyen, Thieu Vo, T. Hoang Ngan Le, Anh Nguyen 0003 |
IROS | 6 |
| 2023 | EmbryosFormer: Deformable Transformer and Collaborative Encoding-Decoding for Embryos Stage Development ClassificationabstractThe timing of cell divisions in early embryos during the In-Vitro Fertilization (IVF) process is a key predictor of embryo viability. However, observing cell divisions in Time-Lapse Monitoring (TLM) is a time-consuming process and highly depends on experts. In this paper, we propose EmbryosFormer, a computational model to automatically detect and classify cell divisions from original time-lapse images. Our proposed network is designed as an encoder-decoder deformable transformer with collaborative heads. The transformer contracting path predicts per-image labels and is optimized by a classification head. The transformer expanding path models the temporal coherency between embryo images to ensure monotonic non-decreasing constraint and is optimized by a segmentation head. Both contracting and expanding paths are synergetically learned by a collaboration head. We have benchmarked our proposed EmbryosFormer on two datasets: a public dataset with mouse embryos with 8-cell stage and an in-house dataset with human embryos with 4-cell stage. Source code: https://github.com/UARK-AICV/Embryos. Tien-Phat Nguyen, Trong-Thang Pham, Tri Nguyen 0005, Hau Lam, Jennifer Fowler, Minh-Triet Tran, T. Hoang Ngan Le |
WACV | 10 |
| 2023 | AOE-Net: Entities Interactions Modeling with Adaptive Attention Mechanism for Temporal Action Proposals Generation
Viet-Khoa Vo-Ho, Sang Truong, Kashu Yamazaki, Bhiksha Raj, Minh-Triet Tran, T. Hoang Ngan Le |
Int. J. Comput. Vis. | 6 |
| 2023 | LIAAD: Lightweight attentive angular distillation for large-scale age-invariant face recognition
Thanh-Dat Truong, Chi Nhan Duong, Kha Gia Quach, T. Hoang Ngan Le, Tien D. Bui, Khoa Luu |
Neurocomputing | 4 |
| 2023 | sCL-ST: Supervised Contrastive Learning With Semantic Transformations for Multiple Lead ECG Arrhythmia ClassificationabstractThe automatic classification of electrocardiogram (ECG) signals has played an important role in cardiovascular diseases diagnosis and prediction. With recent advancements in deep neural networks (DNNs), particularly Convolutional Neural Networks (CNNs), learning deep features automatically from the original data is becoming an effective and widespread approach in a variety of intelligent tasks including biomedical and health informatics. However, most of the existing approaches are trained on either 1D CNNs or 2D CNNs, and they suffer from the limitations of random phenomena (i.e. random initial weights). Furthermore, the ability to train such DNNs in a supervised manner in healthcare is often limited due to the scarcity of labeled training data. To address the problems of weight initialization and limited annotated data, in this work, we leverage recent self-supervised learning technique, namely, contrastive learning, and present supervised contrastive learning (sCL). Different from existing self-supervised contrastive learning approaches, which often generate false negatives because of random selection of negative anchors, our contrastive learning makes use of labeled data to pull the same class closer together and push different classes far apart to avoid potential false negatives. Furthermore, unlike other kinds of signals (e.g. speech, image, video), ECG signal is sensitive to changes, and inappropriate transformation could directly affect diagnosis results. To deal with this issue, we present two semantic transformations, i.e. semantic split-join and semantic weighted peaks noise smoothing. The proposed deep neural network sCL-ST with supervised contrastive learning and semantic transformations is trained as an end-to-end framework for the multi-label classification of 12-lead ECGs. Our sCL-ST network contains two sub-networks i.e. pre-text task and down-stream task. Our experimental results have been evaluated on 12-lead PhysioNet 2020 dataset and shown that our proposed network outperforms the state-of-the-art existing approaches. Sang Truong, Brijesh Patel 0001, Donald A. Adjeroh, T. Hoang Ngan Le |
IEEE J. Biomed. Health Informatics | 5 |
| 2022 | AISFormer: Amodal Instance Segmentation with Transformer
Minh Q. Tran, Viet-Khoa Vo-Ho, Kashu Yamazaki, Arthur A. F. Fernandes, Michael Kidd, T. Hoang Ngan Le |
BMVC | 6 |
| 2022 | Self-Supervised Domain Adaptation in Crowd CountingabstractSelf-training crowd counting has not been attentively explored though it is one of the important challenges in computer vision. In practice, the fully supervised methods usually require an intensive resource of manual annotation. In order to address this challenge, this work introduces a new approach to utilize existing datasets with ground truth to produce more robust predictions on unlabeled datasets, named domain adaptation, in crowd counting. While the network is trained with labeled data, samples without labels from the target domain are also added to the training process. In this process, the entropy map is computed and minimized in addition to the adversarial training process designed in parallel. Experiments on Shanghaitech, UCF_CC_50, and UCF-QNRF datasets prove a more generalized improvement of our method over the other state-of-the-arts in the cross-domain setting. Pha A. Nguyen, Thanh-Dat Truong, Miaoqing Huang, T. Hoang Ngan Le, Khoa Luu |
ICIP | 5 |
| 2022 | VLCAP: Vision-Language with Contrastive Learning for Coherent Video Paragraph CaptioningabstractIn this paper, we leverage the human perceiving process, that involves vision and language interaction, to generate a coherent paragraph description of untrimmed videos. We propose vision-language (VL) features consisting of two modalities, i.e., (i) vision modality to capture global visual content of the entire scene and (ii) language modality to extract scene elements description of both human and non-human objects (e.g. animals, vehicles, etc), visual and non-visual elements (e.g. relations, activities, etc). Furthermore, we propose to train our proposed VLCap under a contrastive learning VL loss. The experiments and ablation studies on ActivityNet Captions and YouCookII datasets show that our VLCap outperforms existing SOTA methods on both accuracy and diversity metrics. Source code: https://github.com/UARK-AICV/VLCAP Kashu Yamazaki, Sang Truong, Viet-Khoa Vo-Ho, Michael Kidd, Chase Rainwater, Khoa Luu, T. Hoang Ngan Le |
ICIP | 7 |
| 2022 | 3DConvCaps: 3DUnet with Convolutional Capsule Encoder for Medical Image SegmentationabstractConvolutional Neural Networks (CNNs) have achieved promising results in medical image segmentation. However, CNNs require lots of training data and are incapable of handling pose and deformation of objects. Furthermore, their pooling layers tend to discard important information such as positions as well as CNNs are sensitive to rotation and affine transformation. Capsule network is a recent new architecture that has achieved better robustness in part-whole representation learning by replacing pooling layers with dynamic routing and convolutional strides, which has shown potential results on popular tasks such as digit classification and object segmentation. In this paper, we propose a 3D encoder-decoder network with Convolutional Capsule Encoder (called 3DConvCaps) to learn lower-level features (short-range attention) with convolutional layers while modeling the higher-level features (long-range dependence) with capsule layers. Our experiments on multiple datasets including iSeg-2017, Hippocampus, and Cardiac demonstrate that our 3D 3DConvCaps network considerably outperforms previous capsule networks and 3D-UNets. We further conduct ablation studies of network efficiency and segmentation performance under various configurations of convolution layers and capsule layers at both contracting and expanding paths. The implementation is available at: https://github.com/UARK-AICV/3DConvCaps Minh Q. Tran, Viet-Khoa Vo-Ho, T. Hoang Ngan Le |
ICPR | 3 |
| 2022 | OTAdapt: Optimal Transport-based Approach For Unsupervised Domain AdaptationabstractUnsupervised domain adaptation is one of the challenging problems in computer vision. This paper presents a novel approach to unsupervised domain adaptations based on the optimal transport-based distance. Our approach allows aligning target and source domains without the requirement of meaningful metrics across domains. In addition, the proposal can associate the correct mapping between source and target domains and guarantee a constraint of topology between source and target domains. The proposed method is evaluated on different datasets in various problems, i.e. (i) digit recognition on MNIST, MNISTM, USPS datasets, (ii) Object recognition on Amazon, Webcam, DSLR, and VisDA datasets, (iii) Insect Recognition on the IP102 dataset. The experimental results show our proposed method consistently improves performance accuracy. Also, our framework can be incorporated with any other CNN frameworks within an end-to-end deep network design for recognition problems to improve their performance. Thanh-Dat Truong, Naga Venkata Sai Raviteja Chappa, Xuan-Bac Nguyen, T. Hoang Ngan Le, Ashley Dowling, Khoa Luu |
ICPR | 4 |
| 2022 | (2+1)D Distilled ShuffleNet: A Lightweight Unsupervised Distillation Network for Human Action RecognitionabstractWhile most existing deep neural networks (DNN) architectures are proposed for increasing performance, they also raise overall model complexity. However, practical applications require lightweight DNN models, that are able to run real-time in edge computing devices. In this work, we present a simple and elegant unsupervised distillation learning paradigm to train a lightweight network to human action recognition called (2+1)D Distilled ShuffleNet. Leveraging the distilling technique, the proposed method allows us to create a lightweight DNN model that achieves high accuracy and real-time speed. Our lightweight (2+1)D Distilled ShuffleNet is designed as an unsupervised paradigm; it does not require labelled data during distilling knowledge from the teacher to the student. Furthermore, to help the student be more "intelligent", we propose to distill the knowledge from two different teachers, i.e., 2D teacher and 3D teacher. The experimental results have shown that our lightweight (2+1)D Distilled ShuffleNet outperforms other state-of-the-art distillation networks with 86.4% and 59.9% top-1 accuracy on UCF101 and HMDB51 datasets, respectively, whereas the inference running time is at 47.16 FPS on CPU with only 17.1M parameters and 12.07 GFLOPs. Duc-Quang Vu, T. Hoang Ngan Le, Jia-Ching Wang |
ICPR | 2 |
| 2022 | Efficient algorithms for mining closed and maximal high utility itemsets
Hai Duong 0001, T. Hoang Ngan Le, Thong Tran, Tin Truong 0001, Bac Le, Philippe Fournier-Viger |
Knowl. Based Syst. | 2 |
| 2022 | Non-volume preserving-based fusion to group-level emotion recognition on crowd videos
Kha Gia Quach, T. Hoang Ngan Le, Chi Nhan Duong, Ibsa Jalata, Kaushik Roy 0003, Khoa Luu |
Pattern Recognit. | 2 |
| 2021 | AEI: Actors-Environment Interaction with Adaptive Attention for Temporal Action Proposals Generation
Viet-Khoa Vo-Ho, Hyekang Joo, Kashu Yamazaki, Sang Truong, Kris Makoto Kitani, Minh-Triet Tran, T. Hoang Ngan Le |
BMVC | 7 |
| 2021 | Agent-Environment Network for Temporal Action Proposal GenerationabstractTemporal action proposal generation is an essential and challenging task that aims at localizing temporal intervals containing human actions in untrimmed videos. Most of existing approaches are unable to follow the human cognitive process of understanding the video context due to lack of attention mechanism to express the concept of an action or an agent who performs the action or the interaction between the agent and the environment. Based on the action definition that a human, known as an agent, interacts with the environment and performs an action that affects the environment, we propose a contextual Agent-Environment Network. Our proposed contextual AEN involves (i) agent pathway, operating at a local level to tell about which humans/agents are acting and (ii) environment pathway operating at a global level to tell about how the agents interact with the environment. Comprehensive evaluations on 20-action THUMOS-14 and 200-action ActivityNet-1.3 datasets with different backbone networks, i.e C3D and SlowFast, show that our method robustly exhibits outperformance against state-of-the-art methods regardless of the employed backbone network. Viet-Khoa Vo-Ho, T. Hoang Ngan Le, Kashu Yamazaki, Akihiro Sugimoto, Minh-Triet Tran |
ICASSP | 2 |
| 2021 | BiMaL: Bijective Maximum Likelihood Approach to Domain Adaptation in Semantic Scene SegmentationabstractSemantic segmentation aims to predict pixel-level labels. It has become a popular task in various computer vision applications. While fully supervised segmentation methods have achieved high accuracy on large-scale vision datasets, they are unable to generalize on a new test environment or a new domain well. In this work, we first introduce a new Unaligned Domain Score to measure the efficiency of a learned model on a new target domain in unsupervised manner. Then, we present the new Bijective Maximum Likelihood1(BiMaL) loss that is a generalized form of the Adversarial Entropy Minimization without any assumption about pixel independence. We have evaluated the proposed BiMaL on two domains. The proposed BiMaL approach consistently outperforms the SOTA methods on empirical experiments on "SYNTHIA to Cityscapes", "GTA5 to Cityscapes", and "SYNTHIA to Vistas". Thanh-Dat Truong, Chi Nhan Duong, T. Hoang Ngan Le, Son Lam Phung, Chase Rainwater, Khoa Luu |
ICCV | 3 |
| 2021 | The Right to Talk: An Audio-Visual Transformer ApproachabstractTurn-taking has played an essential role in structuring the regulation of a conversation. The task of identifying the main speaker (who is properly taking his/her turn of speaking) and the interrupters (who are interrupting or reacting to the main speaker’s utterances) remains a challenging task. Although some prior methods have partially addressed this task, there still remain some limitations. Firstly, a direct association of Audio and Visual features may limit the correlations to be extracted due to different modalities. Secondly, the relationship across temporal segments helping to maintain the consistency of localization, separation and conversation contexts is not effectively exploited. Finally, the interactions between speakers that usually contain the tracking and anticipatory decisions about transition to a new speaker is usually ignored. Therefore, this work introduces a new Audio-Visual Transformer approach to the problem of localization and highlighting the main speaker in both audio and visual channels of a multi-speaker conversation video in the wild. The proposed method exploits different types of correlations presented in both visual and audio signals. The temporal audio-visual relationships across spatial-temporal space are anticipated and optimized via the self-attention mechanism in a Transformer structure. Moreover, a newly collected dataset is introduced for the main speaker detection. To the best of our knowledge, it is one of the first studies that is able to automatically localize and highlight the main speaker in both visual and audio channels in multi-speaker conversation videos. Thanh-Dat Truong, Chi Nhan Duong, The De Vu, Bhiksha Raj, T. Hoang Ngan Le, Khoa Luu |
ICCV | 6 |
| 2021 | PairFlow: Enhancing Portable Chest X-Ray By Flow-Based Deformation For Covid-19 DiagnosingabstractThis work aims to assist physicians improve their speed and diagnostic accuracy when interpreting portable CXR (p_CXR), which are in especially high demand in the setting of the ongoing COVID-19 pandemic. In this paper, we introduce new deep learning frameworks, named Pair-Flow, to align and enhance the quality of p_CXR to be more consistent, and to more closely match higher quality conventional CXR (c_CXR). The contributions of this work are four folds. Firstly, a new database collection of subject-pair CXR is introduced and available to download. Secondly, a new deep learning-based alignment approach is presented to align subject-pairs dataset to obtain pixel-pairs dataset. Thirdly, a new Pair-Flow approach, an end-to-end invertible transfer deep learning method, to enhance the degraded quality of p_CXR. Finally, the performance of the proposed system is evaluated at both image quality and topological properties. T. Hoang Ngan Le, James Sorensen, Toan Duc Bui, Arabinda Choudhary, Khoa Luu, Hien Van Nguyen |
ICIP | 1 |
| 2021 | Point-Unet: A Context-Aware Point-Based Neural Network for Volumetric Segmentation
Ngoc-Vuong Ho, Gia-Han Diep, T. Hoang Ngan Le, Binh-Son Hua |
MICCAI (1) | 4 |
| 2021 | 3D-UCaps: 3D Capsules Unet for Volumetric Image Segmentation
Binh-Son Hua, T. Hoang Ngan Le |
MICCAI (1) | 3 |
| 2021 | Deep reinforcement learning in medical imaging: A literature review
Shaohua Kevin Zhou, T. Hoang Ngan Le, Khoa Luu, Hien Van Nguyen, Nicholas Ayache |
Medical Image Anal. | 2 |
| 2020 | Offset Curves Loss for Imbalanced Problem in Medical SegmentationabstractMedical image segmentation has played an important role in medical analysis and widely developed for many clinical applications. Deep learning-based approaches have achieved high performance in semantic segmentation but they are limited to pixel-wise setting and imbalanced classes data problem. In this paper, we tackle those limitations by developing a new deep learning-based model which takes into account both higher feature level i.e. region inside contour, intermediate feature level i.e. offset curves around the contour and lower feature level i.e. contour. Our proposed Offset Curves (OsC) loss consists of three main fitting terms. The first fitting term focuses on pixel-wise level segmentation whereas the second fitting term acts as attention model which pays attention to the area around the boundaries (offset curves). The third terms plays a role as regularization term which takes the length of boundaries into account. We evaluate our proposed OsC loss on both 2D network and 3D network. Two common medical datasets, i.e. retina DRIVE and brain tumor BRATS 2018 datasets are used to benchmark our proposed loss performance. The experiments have shown that our proposed OsC loss function outperforms other mainstream loss functions such as Cross-Entropy, Dice, Focal on the most common segmentation networks Unet, FCN. T. Hoang Ngan Le, Kashu Yamazaki, Toan Duc Bui, Khoa Luu, Marios Savvides |
ICPR | 1 |
| 2020 | A Multi-task Contextual Atrous Residual Network for Brain Tumor Detection & SegmentationabstractIn recent years, deep neural networks have achieved state-of-the-art performance in a variety of recognition and segmentation tasks in medical imaging including brain tumor segmentation. We investigate that segmenting a brain tumor is facing to the imbalanced data problem where the number of pixels belonging to the background class (non tumor pixel) is much larger than the number of pixels belonging to the foreground class (tumor pixel). To address this problem, we propose a multitask network which is formed as a cascaded structure. Our model consists of two targets, i.e., (i) effectively differentiate the brain tumor regions and (ii) estimate the brain tumor mask. The first objective is performed by our proposed contextual brain tumor detection network, which plays a role of an attention gate and focuses on the region around brain tumor only while ignoring the far neighbor background which is less correlated to the tumor. Different from other existing object detection networks which process every pixel, our contextual brain tumor detection network only processes contextual regions around ground-truth instances and this strategy aims at producing meaningful regions proposals. The second objective is built upon a 3D atrous residual network and under an encode-decode network in order to effectively segment both large and small objects (brain tumor). Our 3D atrous residual network is designed with a skip connection to enables the gradient from the deep layers to be directly propagated to shallow layers, thus, features of different depths are preserved and used for refining each other. In order to incorporate larger contextual information from volume MRI data, our network utilizes the 3D atrous convolution with various kernel sizes, which enlarges the receptive field of filters. Our proposed network has been evaluated on various datasets including BRATS2015, BRATS2017 and BRATS2018 datasets with both validation set and testing set. Our performance has been benchmarked by both region-based metrics and surface-based metrics. We also have conducted comparisons against state-of-the-art approaches.11Code and models will be publicly available. T. Hoang Ngan Le, Kashu Yamazaki, Kha Gia Quach, Dat T. Truong, Marios Savvides |
ICPR | 1 |
| 2020 | Flow-Based Deformation Guidance for Unpaired Multi-contrast MRI Image-to-Image Translation
Toan Duc Bui, Manh Nguyen 0002, T. Hoang Ngan Le, Khoa Luu |
MICCAI (2) | 3 |
| 2019 | Automatic Face Aging in Videos via Deep Reinforcement LearningabstractThis paper presents a novel approach for synthesizing automatically age-progressed facial images in video sequences using Deep Reinforcement Learning. The proposed method models facial structures and the longitudinal face-aging process of given subjects coherently across video frames. The approach is optimized using a long-term reward, Reinforcement Learning function with deep feature extraction from Deep Convolutional Neural Network. Unlike previous age-progression methods that are only able to synthesize an aged likeness of a face from a single input image, the proposed approach is capable of age-progressing facial likenesses in videos with consistently synthesized facial features across frames. In addition, the deep reinforcement learning method guarantees preservation of the visual identity of input faces after age-progression. Results on videos of our new collected aging face AGFW-v2 database demonstrate the advantages of the proposed solution in terms of both quality of age-progressed faces, temporal smoothness, and cross-age face verification. Chi Nhan Duong, Khoa Luu, Kha Gia Quach, Eric Patterson, Tien D. Bui, T. Hoang Ngan Le |
CVPR | 7 |
| 2019 | Learning from Longitudinal Face Demonstration - Where Tractable Deep Modeling Meets Inverse Reinforcement Learning
Chi Nhan Duong, Kha Gia Quach, Khoa Luu, T. Hoang Ngan Le, Marios Savvides, Tien D. Bui |
Int. J. Comput. Vis. | 4 |
| 2018 | Deep Recurrent Level Set for Segmenting Brain Tumors
T. Hoang Ngan Le, Raajitha Gummadi, Marios Savvides |
MICCAI (3) | 1 |
| 2018 | Deep contextual recurrent residual networks for scene labeling
T. Hoang Ngan Le, Chi Nhan Duong, Ligong Han, Khoa Luu, Kha Gia Quach, Marios Savvides |
Pattern Recognit. | 1 |
| 2018 | Reformulating Level Sets as Deep Recurrent Neural Network Approach to Semantic SegmentationabstractVariational Level Set (LS) has been a widely used method in medical segmentation. However, it is limited when dealing with multi-instance objects in the real world. In addition, its segmentation results are quite sensitive to initial settings and highly depend on the number of iterations. To address these issues and boost the classic variational LS methods to a new level of the learnable deep learning approaches, we propose a novel definition of contour evolution named Recurrent Level Set (RLS) 1 to employ Gated Recurrent Unit under the energy minimization of a variational LS functional. The curve deformation process in RLS is formed as a hidden state evolution procedure and updated by minimizing an energy functional composed of fitting forces and contour length. By sharing the convolutional features in a fully end-to-end trainable framework, we extend RLS to Contextual RLS (CRLS) to address semantic segmentation in the wild. The experimental results have shown that our proposed RLS improves both computational time and segmentation accuracy against the classic variational LS-based method whereas the fully end-to-end system CRLS achieves competitive performance compared to the state-of-the-art semantic segmentation approaches. T. Hoang Ngan Le, Kha Gia Quach, Khoa Luu, Chi Nhan Duong, Marios Savvides |
IEEE Trans. Image Process. | 1 |
| 2017 | Temporal Non-volume Preserving Approach to Facial Age-Progression and Age-Invariant Face RecognitionabstractModeling the long-term facial aging process is extremely challenging due to the presence of large and non-linear variations during the face development stages. In order to efficiently address the problem, this work first decomposes the aging process into multiple short-term stages. Then, a novel generative probabilistic model, named Temporal Non-Volume Preserving (TNVP) transformation, is presented to model the facial aging process at each stage. Unlike Generative Adversarial Networks (GANs), which requires an empirical balance threshold, and Restricted Boltzmann Machines (RBM), an intractable model, our proposed TNVP approach guarantees a tractable density function, exact inference and evaluation for embedding the feature transformations between faces in consecutive stages. Our model shows its advantages not only in capturing the non-linear age related variance in each stage but also producing a smooth synthesis in age progression across faces. Our approach can model any face in the wild provided with only four basic landmark points. Moreover, the structure can be transformed into a deep convolutional network while keeping the advantages of probabilistic models with tractable log-likelihood density estimation. Our method is evaluated in both terms of synthesizing age-progressed faces and cross-age face verification and consistently shows the state-of-the-art results in various face aging databases, i.e. FG-NET, MORPH, AginG Faces in the Wild (AGFW), and Cross-Age Celebrity Dataset (CACD). A large-scale face verification on Megaface challenge 1 is also performed to further show the advantages of our proposed approach. Chi Nhan Duong, Kha Gia Quach, Khoa Luu, T. Hoang Ngan Le, Marios Savvides |
ICCV | 4 |
| 2017 | Semi self-training beard/moustache detection and segmentation simultaneously
T. Hoang Ngan Le, Khoa Luu, Chenchen Zhu, Marios Savvides |
Image Vis. Comput. | 1 |
| 2017 | DeepSafeDrive: A grammar-aware driver parsing approach to Driver Behavioral Situational Awareness (DB-SAW)
T. Hoang Ngan Le, Chenchen Zhu, Yutong Zheng, Khoa Luu, Marios Savvides |
Pattern Recognit. | 1 |
| 2016 | Robust hand detection in VehiclesabstractThe problems of hand detection have been widely addressed in many areas, e.g. human computer interaction environment, driver behaviors monitoring, etc. However, the detection accuracy in recent hand detection systems are still far away from the demands in practice due to a number of challenges, e.g. hand variations, highly occlusions, low-resolution and strong lighting conditions. This paper presents the Multiple Scale Faster Region-based Convolutional Neural Network (MS-FRCNN) to handle the problems of hand detection in given digital images collected under challenging conditions. Our proposed method introduces a multiple scale deep feature extraction approach in order to handle the challenging factors to provide a robust hand detection algorithm. The method is evaluated on the challenging hand database, i.e. the Vision for Intelligent Vehicles and Applications (VIVA) Challenge, and compared against various recent hand detection methods. Our proposed method achieves the state-of-the-art results with 20% of the detection accuracy higher than the second best one in the VIVA challenge. T. Hoang Ngan Le, Chenchen Zhu, Yutong Zheng, Khoa Luu, Marios Savvides |
ICPR | 1 |
| 2016 | A novel Shape Constrained Feature-based Active Contour model for lips/mouth segmentation in the wild
T. Hoang Ngan Le, Marios Savvides |
Pattern Recognit. | 1 |
| 2015 | Hierarchical segmentation and tracking of coronary arteries in 2D X-ray Angiography sequencesabstractCoronary arteries (CA) segmentation from an angiographic sequence is essential to guide the cardiologists during percutaneous interventions for the treatment and diagnosis of pathologies. Segmentation of the CA from X-ray angiograms is a very challenging problem due to the changes in contrast in the sequence in addition to the CA's complex topology. In this paper, we propose a hierarchical segmentation method that extends the Vessel Walker model using a temporal prior and multiscale information to extract CA with a higher level of accuracy. In this method, the vessel located in frame Itat time t is extracted by utilizing the segmentation result at frame It−1together with Histogram of Oriented Gradient (HOG) features and a shape matching technique. Our experiments conducted on five paediatric angiograms have shown promising qualitative and quantitative results with a mean Dice coefficient of 64% and 53% Recall and 85% in Precision. Faten M'hiri, T. Hoang Ngan Le, Luc Duong, Christian Desrosiers, Mohamed Cheriet |
ICIP | 2 |
| 2015 | A robust contour sampling and tensor-based approach to facial beard and mustache shape segmentation and matchingabstractIn this paper, we propose a novel system for beard and mustache segmentation and matching in facial images. We first segment out facial hair contours from the image by utilizing a sparse dictionary on self-quotient images to classify regions as either skin or facial hair. We then landmark the shape contour to obtain points around the contour of the image using a combination of two algorithms, a novel non-uniform sampling algorithm, and points obtained from SIFT. We utilize these landmark points to extract inner distance-based shape context features. Finally, these features are used as inputs for a tensor product graph-based matching system. We run experiments on the Multiple Biometric Grand Challenge (MBGC) and the PINELLAS mugshot databases. Our pipeline achieves 90.3% matching accuracy on a subset of the PINELLAS database when divided into four types of facial hair. Karanhaar Singh, Khoa Luu, T. Hoang Ngan Le, Marios Savvides |
ICIP | 3 |
| 2015 | Facial aging and asymmetry decomposition based approaches to identification of twins
T. Hoang Ngan Le, Keshav Seshadri, Khoa Luu, Marios Savvides |
Pattern Recognit. | 1 |
| 2014 | A novel eyebrow segmentation and eyebrow shape-based identificationabstractRecent studies in biometrics have shown that the periocular region of the face is sufficiently discriminative for robust recognition, and particularly effective in certain scenarios such as extreme occlusions, and illumination variations where traditional face recognition systems are unreliable. In this paper, we first propose a fully automatic, robust and fast graph-cut based eyebrow segmentation technique to extract the eyebrow shape from a given face image. We then propose an eyebrow shape-based identification system for periocular face recognition. Our experiments have been conducted over large datasets from the MBGC and AR databases and the resilience of the proposed approach has been evaluated under varying data conditions. The experimental results show that the proposed eyebrow segmentation achieves high accuracy with an F-Measure of 99.4% and the identification system achieves rates of 76.0% on the AR database and 85.0% on the MBGC database. T. Hoang Ngan Le, Utsav Prabhu, Marios Savvides |
IJCB | 1 |
| 2013 | SparCLeS: Dynamic 퓁 1 Sparse Classifiers With Level Sets for Robust Beard/Moustache Detection and SegmentationabstractRobust facial hair detection and segmentation is a highly valued soft biometric attribute for carrying out forensic facial analysis. In this paper, we propose a novel and fully automatic system, called SparCLeS, for beard/moustache detection and segmentation in challenging facial images. SparCLeS uses the multiscale self-quotient (MSQ) algorithm to preprocess facial images and deal with illumination variation. Histogram of oriented gradients (HOG) features are extracted from the preprocessed images and a dynamic sparse classifier is built using these features to classify a facial region as either containing skin or facial hair. A level set based approach, which makes use of the advantages of both global and local information, is then used to segment the regions of a face containing facial hair. Experimental results demonstrate the effectiveness of our proposed system in detecting and segmenting facial hair regions in images drawn from three databases, i.e., the NIST Multiple Biometric Grand Challenge (MBGC) still face database, the NIST Color Facial Recognition Technology FERET database, and the Labeled Faces in the Wild (LFW) database. T. Hoang Ngan Le, Khoa Luu, Marios Savvides |
IEEE Trans. Image Process. | 1 |
| 2012 | A novel energy based filter for cross-blink eye detectionabstractBased on the fact that eye regions may be considered as a sequence of consecutive low and high spatial frequency regions, we propose a novel and efficient filter based on Isotropic Gaussian Energy which can be used for eye detection. When applied to facial images, the designed filter highlights the eye region with a prominent pattern, which can then be localized by template matching technique such as MACE correlation filter. The proposed filter is proved useful to detect both open and closed eyes. We demonstrate the effectiveness of the proposed filter by conducting experiments on both FERET and MBGC face databases. T. Hoang Ngan Le, Khoa Luu, Utsav Prabhu, Marios Savvides |
ICIP | 1 |
| 2012 | Beard and mustache segmentation using sparse classifiers on self-quotient imagesabstractIn this paper, we propose a novel system for beard and mustache detection and segmentation in challenging facial images. Our system first eliminates illumination artifacts using the self-quotient algorithm. A sparse classifier is then used on these self-quotient images to classify a region as either containing skin or facial hair. We conduct experiments on the MBGC and color FERET databases to demonstrate the effectiveness of our proposed system. T. Hoang Ngan Le, Khoa Luu, Keshav Seshadri, Marios Savvides |
ICIP | 1 |
| 2012 | Facecut - a robust approach for facial feature segmentationabstractSegmentation of facial features is a key pre-processing step in enabling facial recognition, building of 3D facial models, expression analysis, and pose estimation. Recently, graph cuts based algorithms have been adapted to carry out this task but many of these methods require manual initialization of points in the foreground and background. In this paper, we propose a novel and fully automatic approach, named Face-Cut, to perform accurate facial feature segmentation. FaceCut combines the positive features of the Modified Active Shape Model (MASM) and GrowCut algorithms to ensure highly accurate and completely automatic segmentation of facial features. We demonstrate the effectiveness of FaceCut on images from two challenging databases. Khoa Luu, T. Hoang Ngan Le, Keshav Seshadri, Marios Savvides |
ICIP | 2 |
| 2011 | Ternary Entropy-Based Binarization of Degraded Document Images Using Morphological OperatorsabstractA vast number of historical and badly degraded document images can be found in libraries, public, and national archives. Due to the complex nature of different artifacts, such poor quality documents are hard to read and to process. In this paper, a novel adaptive binarization algorithm using ternary entropy-based approach is proposed. Given an input image, the contrast of intensity is first estimated by a grayscale morphological closing operator. A double-threshold is generated by our Shannon entropy-based ternarizing method to classify pixels into text, near-text, and non-text regions. The pixels in the second region are relabeled by the local mean and the standard deviation. Our proposed method classifies noise into two categories which are processed by binary morphological operators, shrink and swell filters, and graph searching strategy. The method is tested with three databases that have been used in the Document Image Binarization Contest 2009 (DIBCO 2009), the Handwriting Document Image Binarization Contest 2010 (H-DBCIO 2010), and the International Conference on Frontier in Handwriting Recognition 2010 (ICFHR 2010). The evaluation is based upon nine distinct measures. Experimental results show that our proposed algorithm outperforms other state-of-the-art methods. T. Hoang Ngan Le, Tien D. Bui, Ching Y. Suen |
ICDAR | 1 |
| 2010 | High payload steganography mechanism using hybrid edge detector
Wen-Jan Chen, Chin-Chen Chang 0001, T. Hoang Ngan Le |
Expert Syst. Appl. | 3 |
| 2009 | Sharing a verifiable secret image using two shadows
Chin-Chen Chang 0001, Chia-Chen Lin 0001, T. Hoang Ngan Le, Bac Le |
Pattern Recognit. | 3 |
| 2009 | Self-verifying visual secret sharing using error diffusion and interpolation techniquesabstractIn this paper, we propose a novel scheme called a self-verifying visual secret sharing scheme, which can be applied to both grayscale and color images. This scheme uses two halftone images. The first, considered to be the host image, is created by directly applying a halftoning technique to the original secret image. The other, regarded as the logo, is generated from the host image by exploiting the interpolation and error diffusion techniques. Because the set of shadows and the reconstructed secret image are generated by simple Boolean operations, no computational complexity and no pixel expansion occur in our scheme. Experimental results confirm that each shadow generated by our scheme is a noise-like image and eight times smaller than the secret image. Moreover, the peak signal-to-noise ratio value of the reconstructed secret image is larger than 33 dB. Based on the extracted halftone logo, the proposed scheme provides an effective solution for verifying the reliability of the set of collected shadows as well as the reconstructed secret image. Furthermore, the reconstructed secret image can be established completely if and only ifkout ofnvalid shadows have been collected. To achieve our objectives, four techniques were adopted: error diffusion, image clustering, interpolation, and inverse halftoning-based edge detection. Chin-Chen Chang 0001, Chia-Chen Lin 0001, T. Hoang Ngan Le, Bac Le |
IEEE Trans. Inf. Forensics Secur. | 3 |