VLDB 2026 Research / reviewers in the wild / expert
Anh Nguyen 0003
dblp:52/5285-3
· DBLP profile ↗
86ranked-venue papers
8as first author
77since 2021 · last 2026
0000-0002-1449-211XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 66 · 8 first-author · 57 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 36 since 2021Systems, architecture and hardware · 28 · 6 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 10 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric PerspectiveabstractAs embodied agents operate in increasingly complex environments, the ability to perceive, track, and reason about individual object instances over time becomes essential, especially in tasks requiring sequenced interactions with visually similar objects. In non-Markovian settings, critical decision cues lie in object histories rather than the current scene. Without persistent memory of prior interactions (what was used, where it was placed, or how it changed), visuomotor policies may fail, repeat past actions, or overlook completed ones. To surface this challenge, we introduce LIBERO-Mem, a non-Markovian task suite for stress-testing robotic manipulation under object-level partial observability. It combines short- and long-horizon object tracking with temporally sequenced subgoals, requiring reasoning beyond the current frame. However, vision-language-action (VLA) models often struggle in such settings, with token scaling quickly becoming intractable even for tasks spanning just a few hundred frames. We propose Embodied-SlotSSM, a slot-centric VLA framework built for temporal scalability. It maintains spatio-temporally consistent slot identities and leverages them through two mechanisms: (1) slot-state-space modeling for reconstructing short-term history, and (2) a relational encoder to align the input tokens with action decoding. Together, these components enable temporally grounded, context-aware action prediction. Experiments show Embodied-SlotSSM's baseline performance on LIBERO-Mem and general tasks, offering a scalable solution for non-Markovian reasoning in object-centric policies. Nhat Chung, Taisei Hanyu, Toan Nguyen 0004, Huy Le 0001, Frederick Bumgarner, Duy M. H. Nguyen, Viet-Khoa Vo-Ho, Kashu Yamazaki, Chase Rainwater, Tung Kieu, Anh Nguyen 0003, T. Hoang Ngan Le |
AAAI | 11 |
| 2026 | Physics-Guided 3D Convolutional Learning for Accurate Springback Error Prediction in Single Point Incremental Forming
Frans Coenen, Mariluz Penalva Oscoz, Ander Martin Rebe, Yang Hai, Anh Nguyen 0003 |
ICPR (1) | 6 |
| 2026 | Fine-Grained Alignment in Vision-and-Language Navigation Through Bayesian Optimization
Yuhang Song 0008, Mario Gianni, Chenguang Yang 0001, Kunyang Lin, Te-Chuan Chiu, Anh Nguyen 0003, Chun-Yi Lee |
ICPR (4) | 6 |
| 2026 | MEUR: A Benchmark for Evaluating Vision-Language Models on Multimodal Event Understanding and Reasoning
Tong Chen 0005, Changyu Zeng, Hongbin Na, Nijia Han, Fuyu Xing, Qi Chen 0026, Qiufeng Wang 0001, Anh Nguyen 0003, Shuihua Wang, Ling Chen 0006, Jionglong Su, Haiyang Zhang 0004, Wei Wang 0042 |
LREC | 10 |
| 2026 | SAGE: Sustainable Agent-Guided Expert-tuning for Culturally Attuned Translation in Low-Resource Southeast Asia
Zhixiang Lu, Chong Zhang 0006, Yulong Li 0002, Angelos Stefanidis, Anh Nguyen 0003, Muhammad Imran Razzak, Jionglong Su, Zhengyong Jiang |
WWW | 5 |
| 2026 | CoRe: Contrast and reconstruction combination self-supervised point cloud representation learning
Changyu Zeng, Jimin Xiao, Anh Nguyen 0003, Xuming Hu, Wei Wang 0042, Yutao Yue |
Expert Syst. Appl. | 3 |
| 2026 | Directly extracting maximal multi-level high-utility patterns
Trinh D. D. Nguyen, N. T. Tung, Loan T. T. Nguyen, An Mai, Anh Nguyen 0003 |
Knowl. Based Syst. | 5 |
| 2025 | Fractal Calibration for Long-tailed Object DetectionabstractReal-world datasets follow an imbalanced distribution, which poses significant challenges in rare-category object detection. Recent studies tackle this problem by developing re-weighting and re-sampling methods, that utilise the class frequencies of the dataset. However, these techniques focus solely on the frequency statistics and ignore the distribution of the classes in image space, missing important information. In contrast to them, we propose FRActal CALibration (FRACAL): a novel post-calibration method for long-tailed object detection. FRACAL devises a logit adjustment method that utilises the fractal dimension to estimate how uniformly classes are distributed in image space. During inference, it uses the fractal dimension to inversely down-weight the probabilities of uniformly spaced class predictions achieving balance in two axes: between frequent and rare categories, and between uniformly spaced and sparsely spaced classes. FRACAL is a post-processing method and it does not require any training, also it can be combined with many off-the-shelf models such as one-stage sigmoid detectors and two-stage instance segmentation models. FRACAL boosts the rare class performance by up to 8.6% and surpasses all previous methods on LVIS dataset, while showing good generalisation to other datasets such as COCO, V3Det and OpenImages. We provide the code at https://github.com/kostas1515/FRACAL. Konstantinos Panagiotis Alexandridis, Ismail Elezi, Jiankang Deng, Anh Nguyen 0003, Shan Luo 0001 |
CVPR | 4 |
| 2025 | LP-Diff: Towards Improved Restoration of Real-World Degraded License PlateabstractLicense plate (LP) recognition is crucial in intelligent traffic management systems. However, factors such as long distances and poor camera quality often lead to severe degradation of captured LP images, posing challenges to accurate recognition. The design of License Plate Image Restoration (LPIR) methods frequently relies on synthetic degraded data, which limits their effectiveness on real-world severely degraded LP images. To address this issue, we introduce the first paired LPIR dataset collected in real-world scenarios, named MDLP, including 10,245 pairs of multi-frame severely degraded LP images and their corresponding clear images. To better restore severely degraded LP, we propose a novel Diffusion-based network, called LP-Diff, to tackle real-world LPIR tasks. Our approach incorporates (1) an Inter-frame Cross Attention Module to fuse temporal information across multiple frames, (2) a Texture Enhancement Module to restore texture information in degraded images, and (3) a Dual-Pathway Fusion Module to select effective features from both channel and spatial dimensions. Extensive experiments demonstrate the reliability of our dataset for model training and evaluation. Our proposed LP-Diff consistently outperforms other state-of-the-art image restoration methods on real-world LPIR tasks. The dataset and code are available at https://github.com/haoyGONG/LP-Diff. Haoyan Gong, Yuzheng Feng, Anh Nguyen 0003, Hongbin Liu 0007 |
CVPR | 4 |
| 2025 | Are Spatial-Temporal Graph Convolution Networks for Human Action Recognition Over-Parameterized?abstractSpatial-temporal graph convolutional networks (ST-GCNs) showcase impressive performance in skeleton-based human action recognition (HAR). However, despite the development of numerous models, their recognition performance does not differ significantly after aligning the input settings. With this observation, we hypothesize that ST-GCNs are over-parameterized for HAR, a conjecture subsequently confirmed through experiments employing the lottery ticket hypothesis. Additionally, a novel sparse ST-GCNs generator is proposed, which trains a sparse architecture from a randomly initialized dense network while maintaining comparable performance levels to the dense components. Moreover, we generate multi-level sparsity ST-GCNs by integrating sparse structures at various sparsity levels and demonstrate that the assembled model yields a significant enhancement in HAR performance. Thorough experiments on four datasets, including NTU-RGB+D 60(120), Kinetics-400, and FineGYM, demonstrate that the proposed sparse ST-GCNs can achieve comparable performance to their dense components. Even with 95% fewer parameters, the sparse ST-GCNs exhibit a degradation of1% in top-1 accuracy. The code is available at https://github.com/davelailai/Sparse-ST-GCN. Jianyang Xie, Yitian Zhao, Yanda Meng, He Zhao 0002, Anh Nguyen 0003, Yalin Zheng |
CVPR | 5 |
| 2025 | EgoMusic-Driven Human Dance Motion Estimation with Skeleton Mamba
Nhat Le, Baoru Huang, Minh Nhat Vu, Chengcheng Tang, T. Hoang Ngan Le, Thieu Vo, Anh Nguyen 0003 |
ICCV | 9 |
| 2025 | CT-ScanGaze: A Dataset and Baselines for 3D Volumetric Scanpath ModelingabstractUnderstanding radiologists' eye movement during Computed Tomography (CT) reading is crucial for developing effective interpretable computer-aided diagnosis systems. However, CT research in this area has been limited by the lack of publicly available eye-tracking datasets and the three-dimensional complexity of CT volumes. To address these challenges, we present the first publicly available eye gaze dataset on CT, called CT-ScanGaze. Then, we introduce CT-Searcher, a novel 3D scanpath predictor designed specifically to process CT volumes and generate radiologist-like 3D fixation sequences, overcoming the limitations of current scanpath predictors that only handle 2D inputs. Since deep learning models benefit from a pretraining step, we develop a pipeline that converts existing 2D gaze datasets into 3D gaze data to pretrain CT-Searcher. Through both qualitative and quantitative evaluations on CT-ScanGaze, we demonstrate the effectiveness of our approach and provide a comprehensive assessment framework for 3D scanpath prediction in medical imaging. Trong-Thang Pham, Akash Awasthi, Saba Khan, Esteban Duran Marti, Tien-Phat Nguyen, Viet-Khoa Vo-Ho, Cuong Tran 0010, Yuki Ikebe, Anh Totti Nguyen, Anh Nguyen 0003, Zhigang Deng 0001, Carol C. Wu, T. Hoang Ngan Le |
ICCV | 12 |
| 2025 | Towards a Universal 3D Medical Multi-Modality Generalization via Learning Personalized Invariant Representation
Zhaorui Tan, Xi Yang 0008, Tan Pan, Chen Jiang 0006, Xin Guo 0010, Qiufeng Wang 0001, Anh Nguyen 0003, Yuan Qi 0001, Kaizhu Huang |
ICCV | 8 |
| 2025 | More Reliable Pseudo-Labels, Better Performance: A Generalized Approach to Single Positive Multi-Label LearningabstractMulti-label learning is a challenging computer vision task that requires assigning multiple categories to each image. However, fully annotating large-scale datasets is often impractical due to high costs and effort, motivating the study of learning from partially annotated data. In the extreme case of Single Positive Multi-Label Learning (SPML), each image is provided with only one positive label, while all other labels remain unannotated. Traditional SPML methods that treat missing labels as unknown or negative tend to yield inaccuracies and false negatives, and integrating various pseudo-labeling strategies can introduce additional noise. To address these challenges, we propose the Generalized Pseudo-Label Robust Loss (GPR Loss), a novel loss function that effectively learns from diverse pseudo-labels while mitigating noise. Complementing this, we introduce a simple yet effective Dynamic Augmented Multi-focus Pseudo-labeling (DAMP) technique. Together, these contributions form the Adaptive and Efficient Vision-Language Pseudo-Labeling (AEVLP) framework. Extensive experiments on four benchmark datasets demonstrate that our framework significantly advances multi-label classification, achieving state-of-the-art results. Luong Tran, Thieu Vo, Anh Nguyen 0003, Sang Dinh |
ICCV | 3 |
| 2025 | FedEFM: Federated Endovascular Foundation Model with Unseen DataabstractIn endovascular surgery, the precise identification of catheters and guidewires in X-ray images is essential for reducing intervention risks. However, accurately segmenting catheter and guidewire structures is challenging due to the limited availability of labeled data. Foundation models offer a promising solution by enabling the collection of similar-domain data to train models whose weights can be fine-tuned for downstream tasks. Nonetheless, large-scale data collection for training is constrained by the necessity of maintaining patient privacy. This paper proposes a new method to train a foundation model in a decentralized federated learning setting for endovascular intervention. To ensure the feasibility of the training, we tackle the unseen data issue using differentiable Earth Mover's Distance within a knowledge distillation frame-work. Once trained, our foundation model's weights provide valuable initialization for downstream tasks, thereby enhancing task-specific performance. Intensive experiments show that our approach achieves new state-of-the-art results, contributing to advancements in endovascular intervention and robotic-assisted endovascular surgery, while addressing the critical issue of data sharing in the medical domain. Tuong KL. Do, Nghia Vu, Tudor Jianu, Baoru Huang, Minh Nhat Vu, Jionglong Su, Erman Tjiputra, Quang D. Tran, Te-Chuan Chiu, Anh Nguyen 0003 |
ICRA | 10 |
| 2025 | Robotic-CLIP: Fine-Tuning CLIP on Action Data for Robotic ApplicationsabstractVision language models have played a key role in extracting meaningful features for various robotic applications. Among these, Contrastive Language-Image Pretraining (CLIP) is widely used in robotic tasks that require both vision and natural language understanding. However, CLIP was trained solely on static images paired with text prompts and has not yet been fully adapted for robotic tasks involving dynamic actions. In this paper, we introduce Robotic-CLIP to enhance robotic perception capabilities. We first gather and label large-scale action data, and then build our Robotic-CLIP by fine-tuning CLIP on 309,433 videos (≈ 7.4 million frames) of action data using contrastive learning. By leveraging action data, Robotic-CLIP inherits CLIP's strong image performance while gaining the ability to understand actions in robotic contexts. Intensive experiments show that our Robotic-CLIP outperforms other CLIP-based models across various language-driven robotic tasks. Additionally, we demonstrate the practical effectiveness of Robotic-CLIP in real-world grasping applications. Minh Nhat Vu, Tung D. Ta, Baoru Huang, Thieu Vo, T. Hoang Ngan Le, Anh Nguyen 0003 |
ICRA | 7 |
| 2025 | Weakly-Supervised Learning via Multi-Lateral Decoder Branching for Tool Segmentation in Robot-Assisted Cardiovascular CatheterizationabstractRobot-assisted catheterization has garnered a good attention for its potentials in treating cardiovascular diseases. However, advancing surgeon-robot collaboration still requires further research, particularly on task-specific automation. For instance, automated tool segmentation can assist surgeons in visualizing and tracking endovascular tools during procedures. While learning-based models have demonstrated state-of-the-art segmentation performances, generating ground-truth labels for fully-supervised methods is laborintensive, time consuming, and costly. In this study, we developed a weakly-supervised learning method that is based on multi-lateral pseudo labeling for tool segmentation in cardiovascular angiogram datasets. The method utilizes a modified U-Net architecture featuring one encoder and multiple laterally branched decoders. The decoders generate diverse pseudo labels under different perturbations to augment the available partial annotation for model training. A mixed loss function with shared consistency was adapted for this purpose. The weakly-supervised model was trained end-to-end and validated using partially annotated angiogram data from three cardiovascular catheterization procedures. Validation results show that the weakly-supervised model could perform closer to fully-supervised models. Furthermore, the proposed multi-lateral approach outperforms three well known weakly-supervised learning methods, offering the highest segmentation performance across the three angiogram datasets. Numerous ablation studies confirmed the model's consistent performance under different settings. Finally, the model was applied for tool segmentation in a robot-assisted catheterization experiments. The model enhanced visualization with high connectivity indices for guidewire and catheter, and a mean segmentation time of 35.26±11.29 ms per frame. This study provides a fast, stable, and less expensive method for segmentation and visualization of endovascular tools in robot-assisted cardiac catheterization. Olatunji Mumini Omisore, Toluwanimi Oluwadara Akinyemi, Anh Nguyen 0003, Lei Wang 0029 |
ICRA | 3 |
| 2025 | Hybrid Gripper with Passive Pneumatic Soft Joints for Grasping Deformable Thin ObjectsabstractGrasping a variety of objects remains a key challenge in the development of versatile robotic systems. The human hand is remarkably dexterous, capable of grasping and manipulating objects with diverse shapes, mechanical properties, and textures. Inspired by how humans use two fingers to pick up thin and large objects such as fabric or sheets of paper, we aim to develop a gripper optimized for grasping such deformable objects. Observing how the soft and flexible fingertip joints of the hand approach and grasp thin materials, a hybrid gripper design that incorporates both soft and rigid components was proposed. The gripper utilizes a soft pneumatic ring wrapped around a rigid revolute joint to create a flexible two-fingered gripper. Experiments were conducted to characterize and evaluate the gripper's performance in handling sheets of paper and other objects. Compared to rigid grippers, the proposed design improves grasping efficiency and reduces the gripping distance by up to eightfold. Ngoc-Duy Tran, Hoang-Hiep Ly, Xuan-Thuan Nguyen, Thi Thoa Mac, Anh Nguyen 0003, Tung D. Ta |
ICRA | 5 |
| 2025 | Online Trajectory Replanner for Dynamically Grasping Irregular ObjectsabstractThis paper presents a new trajectory replanner for grasping irregular objects. Unlike conventional grasping tasks where the object's geometry is assumed simple, we aim to achieve a “dynamic grasp” of the irregular objects, which requires continuous adjustment during the grasping process. To effectively handle irregular objects, we propose a trajectory optimization framework that comprises two phases. Firstly, in a specified time limit of 10 s, initial offline trajectories are computed for a seamless motion from an initial configuration of the robot to grasp the object and deliver it to a predefined target location. Secondly, fast online trajectory optimization is implemented to update robot trajectories in real-time within 100 ms. This helps to mitigate pose estimation errors from the vision system. To account for model inaccuracies, disturbances, and other non-modeled effects, trajectory tracking controllers for both the robot and the gripper are implemented to execute the optimal trajectories from the proposed framework. The intensive experimental results effectively demonstrate the performance of our trajectory planning framework in both simulation and real-world scenarios. Minh Nhat Vu, Florian Grander, Anh Nguyen 0003, Christoph Unger |
ICRA | 3 |
| 2025 | Towards Autonomous Wood-Log Grasping with a Forestry Crane: Simulator and BenchmarkingabstractForestry machines operated in forest production environments face challenges when performing manipulation tasks, especially regarding the complicated dynamics of underactuated crane systems and the heavy weight of logs to be grasped. This study investigates the feasibility of using reinforcement learning for forestry crane manipulators in grasping and lifting heavy wood logs autonomously. We first build a simulator using Mujoco physics engine to create realistic scenarios, including modeling a forestry crane with 8 degrees of freedom from CAD data and wood logs of different sizes. We further implement a velocity controller for autonomous log grasping with deep reinforcement learning using a curriculum strategy. Utilizing our new simulator, the proposed control strategy exhibits a success rate of 96% when grasping logs of different diameters and under random initial configurations of the forestry crane. In addition, reward functions and reinforcement learning baselines are implemented to provide an open-source benchmark for the community in large-scale manipulation tasks. A video with several demonstrations can be seen at https://www.acin.tuwien.ac.at/en/d18a/. Minh Nhat Vu, Alexander Wachter, Gerald Ebmer, Marc-Philip Ecker, Tobias Glück, Anh Nguyen 0003, Wolfgang Kemmetmüller, Andreas Kugi |
ICRA | 6 |
| 2025 | Attention-SSM Network for Predicting Springback Error in Single Point Incremental FormingabstractPredicting springback when employing a Single Point Incremental Forming (SPIF) process is a complex industrial manufacturing task. Current methodologies predominantly depend on Recurrent Neural Networks (RNNs), such as LSTM and GRU, for the processing of sequential 3D point cloud data representations. However, these systems exhibit constrained scalability, inefficiency stemming from sequential processing, and a deficiency in interpretability, rendering them less appropriate for practical manufacturing contexts. We introduce the use of Attention-State-Space Modeling(SSM) Network in this study. This innovative approach utilizes Transformer-based self-attention mechanisms in conjunction with SSM to improve sprinback prediction efficiency, scalability, and explainability. Our model encodes the 3D point sequence through self-attention layers and processes the encoded sequence with the Mamba module, effectively capturing contextual dependencies and essential spatial information essential for forecasting springback failures. Extensive evaluation indicates that our proposed approach attains state-of-the-art performance across diverse grid sizes, providing insights into essential sequence characteristics through visualizations. Our contributions are threefold: (1) the introduction of the Attention-SSM Network for SPIF springback prediction, (2) state-of-the-art performance tested on benchmark datasets, and (3) a comprehensive feature analysis emphasizing model explainability. Our code and models are available at https://github.com/DarrenChen0923/Mamba-back. Mariluz Penalva Oscoz, Ander Martin Rebe, Yang Hai, Frans Coenen, Anh Nguyen 0003 |
IJCNN | 6 |
| 2025 | Modeling The States of Liquid Phase Change Pouch Actuators by Reservoir ComputingabstractLiquid phase change pouch actuators (liquid pouch motors) hold great promise for a wide range of robotic applications, from artificial organs to pneumatic manipulators for dexterous manipulation. However, the usability of liquid pouch motors remains challenging due to the nonlinear intrinsic properties of liquids and their highly dynamic implications for liquid-gas phase changes, which complicate state modeling and estimation. To address these issues, we propose a reservoir computing-based method for modeling the inflation states of a customized liquid pouch motor, which serves as an actuator, featuring four Peltier heating junctions. We use a motion capture system to track the landmark movements on the pouch as a proxy for its volumetric profile. These movements represent the internal liquid-gas phase changes of the pouch at stable room temperature, atmospheric pressure, and in the presence of electrical noise. The motion coordinates are thus learned by our reservoir computing framework, PhysRes, to model the states based on prior observations. Through training, our model achieves excellent results on the test set, with a normalized root mean squared error of 0.0041 in estimating the states and a corresponding volumetric error of 0.0160%. To further demonstrate how such actuators could be implemented in the future, we also design a dual-pouch actuator-based robotic gripper to control the grasping of soft objects. Our design and source code are available at: https://github.com/tatung/liquidpouch_reservoir. Cedric Caremel, Anh Nguyen 0003, Manfred Huber, Yoshihiro Kawahara, Tung D. Ta |
IROS | 3 |
| 2025 | Lightweight Temporal Transformer Decomposition for Federated Autonomous DrivingabstractTraditional vision-based autonomous driving systems often face difficulties in navigating complex environments when relying solely on single-image inputs. To overcome this limitation, incorporating temporal data such as past image frames or steering sequences, has proven effective in enhancing robustness and adaptability in challenging scenarios. While previous high-performance methods exist, they often rely on resource-intensive fusion networks, making them impractical for training and unsuitable for federated learning. To address these challenges, we propose lightweight temporal transformer decomposition, a method that processes sequential image frames and temporal steering data by breaking down large attention maps into smaller matrices. This approach reduces model complexity, enabling efficient weight updates for convergence and real-time predictions while leveraging temporal information to enhance autonomous driving performance. Intensive experiments on three datasets demonstrate that our method outperforms recent approaches by a clear margin while achieving real-time performance. Additionally, real robot experiments further confirm the effectiveness of our method. Our source code can be found at: https://github.com/aioz-ai/LTFed. Tuong KL. Do, Binh X. Nguyen, Quang D. Tran, Erman Tjiputra, Te-Chuan Chiu, Anh Nguyen 0003 |
IROS | 6 |
| 2025 | SplineFormer: An Explainable Transformer Network for Autonomous Endovascular NavigationabstractRobot-assisted endovascular navigation provides significant advantages, including reduced radiation exposure for surgeons and improved patient safety. However, a major challenge is to control curvilinear instruments like guidewires precisely for smooth and accurate navigation while adapting to anatomical variations and external forces. Traditional segmentation-based approaches struggle with real-time prediction of the guidewire’s evolving shape, limiting their effectiveness in navigation tasks. In this paper, we propose SplineFormer, an explainable transformer network that predicts the continuous, structured representation of the guidewire as a B-spline. This formulation enables a compact, smooth, and explainable state representation that facilitates downstream navigation. By leveraging SplineFormer’s predictions within an imitation learning framework, our system successfully performs autonomous endovascular navigation. Experimental results show that SplineFormer achieves a 50% success rate when fully autonomously cannulating the Brachiocephalic Artery in a real robotic setup, demonstrating its potential for improved autonomous navigation in endovascular interventions. Tudor Jianu, Shayan Doust, Mengyun Li, Baoru Huang, Tuong KL. Do, Hoan Nguyen, Karl Bates, Tung D. Ta, Sebastiano Fichera, Pierre Berthet-Rayne, Anh Nguyen 0003 |
IROS | 11 |
| 2025 | GraspMAS: Zero-Shot Language-driven Grasp Detection with Multi-Agent SystemabstractLanguage-driven grasp detection has the potential to revolutionize human-robot interaction by allowing robots to understand and execute grasping tasks based on natural language commands. However, existing approaches face two key challenges. First, they often struggle to interpret complex text instructions or operate ineffectively in densely cluttered environments. Second, most methods require a training or fine-tuning step to adapt to new domains, limiting their generation in real-world applications. In this paper, we introduce GraspMAS, a new multi-agent system framework for language-driven grasp detection. GraspMAS is designed to reason through ambiguities and improve decision-making in real-world scenarios. Our framework consists of three specialized agents: Planner, responsible for strategizing complex queries; Coder, which generates and executes source code; and Observer, which evaluates the outcomes and provides feedback. Intensive experiments on two large-scale datasets demonstrate that our GraspMAS significantly outperforms existing baselines. Additionally, robot experiments conducted in both simulation and real-world settings further validate the effectiveness of our approach. Our project page is available at https://zquang2202.github.io/GraspMAS. Thieu Vo, Tung D. Ta, Baoru Huang, Minh Nhat Vu, Anh Nguyen 0003 |
IROS | 8 |
| 2025 | GraspMamba: A Mamba-based Language-driven Grasp Detection Framework with Hierarchical Feature LearningabstractGrasp detection is a fundamental robotic task critical to the success of many industrial applications. However, current language-driven models for this task often struggle with cluttered images, lengthy textual descriptions, or slow inference speed. We introduce GraspMamba, a new language-driven grasp detection method that employs hierarchical feature fusion with Mamba vision to tackle these challenges. By leveraging rich visual features of the Mamba-based backbone alongside textual information, our approach effectively enhances the fusion of multimodal features. GraspMamba represents the first Mamba-based grasp detection model to extract vision and language features at multiple scales, delivering robust performance and rapid inference time. Intensive experiments show that GraspMamba outperforms recent methods by a clear margin. We validate our approach through real-world robotic experiments, highlighting its fast inference speed. An Vuong, Anh Nguyen 0003, Ian D. Reid 0001, Minh Nhat Vu |
IROS | 3 |
| 2025 | BiMa: Towards Biases Mitigation for Text-Video Retrieval via Scene Element GuidanceabstractText-video retrieval (TVR) systems often suffer from visual-linguistic biases present in datasets, which cause pre-trained vision-language models to overlook key details. To address this, we propose BiMa, a novel framework designed to mitigate biases in both visual and textual representations. Our approach begins by generating scene elements that characterize each video by identifying relevant entities/objects and activities. For visual debiasing, we integrate these scene elements into the video embeddings, enhancing them to emphasize fine-grained and salient details. For textual debiasing, we introduce a mechanism to disentangle text features into content and bias components, enabling the model to focus on meaningful content while separately handling biased information. Extensive experiments and ablation studies across five major TVR benchmarks (i.e., MSR-VTT, MSVD, LSMDC, ActivityNet, and DiDeMo) demonstrate the competitive performance of BiMa. Additionally, the model's bias mitigation capability is consistently validated by its strong results on out-of-distribution retrieval tasks. Huy Le 0001, Nhat Chung, Tung Kieu, Anh Nguyen 0003, T. Hoang Ngan Le |
ACM Multimedia | 4 |
| 2025 | Interpreting Radiologist's Intention from Eye Movements in Chest X-ray Diagnosis
Trong-Thang Pham, Anh Nguyen 0003, Zhigang Deng 0001, Carol C. Wu, T. Hoang Ngan Le |
ACM Multimedia | 2 |
| 2025 | Learning Human Motion with Temporally Conditional MambaabstractLearning human motion based on a time-dependent input signal presents a challenging yet impactful task with various applications. The goal of this task is to generate or estimate human movement that consistently reflects the temporal patterns of conditioning inputs. Existing methods typically rely on cross-attention mechanisms to fuse the condition with motion. However, this approach primarily captures global interactions and struggles to maintain step-by-step temporal alignment. To address this limitation, we introduce Temporally Conditional Mamba, a new mamba-based model for human motion generation. Our approach integrates conditional information into the recurrent dynamics of the Mamba block, enabling better temporally aligned motion. To validate the effectiveness of our method, we evaluate it on a variety of human motion tasks. Extensive experiments demonstrate that our model significantly improves temporal alignment, motion realism, and condition consistency over state-of-the-art approaches. Our project page is available at https://zquang2202.github.io/TCM. Baoru Huang, Minh Nhat Vu, T. Hoang Ngan Le, Thieu Vo, Anh Nguyen 0003 |
SIGGRAPH Asia | 7 |
| 2025 | A Novel Approach to Differential Expression Analysis of Co-Occurrence Networks for Small-Sampled Microbiome DataabstractGraph-based machine learning methods are valuable tools for identifying and predicting variation in genetic data. In particular, understanding phenotypic effects at the cellular level is an accelerating area in pharmacogenomics. Insight into how drugs or disease affect bio-networks could aid drug development and precision medicine. This article proposes a novel graph-theoretic approach to infer a co-occurrence network from 16S microbiome data, designed specifically for smallsample datasets. Such datasets pose challenges due to sparsity, compositionality, and complex interactions. The methodology includes steps to enrich and statistically filter the inferred networks. The approach extracts informative, feature-rich, biologically meaningful, and statistically significant networks from limited data. While tailored for small datasets, it is broadly applicable and can be extended to multi-omics integration. The method is tested on data from chickens vaccinated and challenged with Eimeria tenella. Genetic reads are processed, and networks inferred to characterize intestinal ecosystems at three disease progression stages. Analysis of network features yields biologically intuitive conclusions using statistical methods. Notably, the distribution of node features evolves with disease progression, and distributions reveal mutualistic and parasitic species clusters. A sub-network consistently appears across all conditions, suggesting a 'persistent microbiome'. A clustering algorithm is also applied to demonstrate the methods utility for downstream analysis. Nandini Amit Gadhia, Michalis Smyrnakis, Po-Yu Liu, Damer Blake, Melanie Hay, Anh Nguyen 0003, Dominic Richards, Dong Xia, Ritesh Krishna |
IEEE Trans. Comput. Biol. Bioinform. | 6 |
| 2024 | Dynamic Semantic-Based Spatial Graph Convolution Network for Skeleton-Based Human Action RecognitionabstractGraph convolutional networks (GCNs) have attracted great attention and achieved remarkable performance in skeleton-based action recognition. However, most of the previous works are designed to refine skeleton topology without considering the types of different joints and edges, making them infeasible to represent the semantic information. In this paper, we proposed a dynamic semantic-based graph convolution network (DS-GCN) for skeleton-based human action recognition, where the joints and edge types were encoded in the skeleton topology in an implicit way. Specifically, two semantic modules, the joints type-aware adaptive topology and the edge type-aware adaptive topology, were proposed. Combining proposed semantics modules with temporal convolution, a powerful framework named DS-GCN was developed for skeleton-based action recognition. Extensive experiments in two datasets, NTU-RGB+D and Kinetics-400 show that the proposed semantic modules were generalized enough to be utilized in various backbones for boosting recognition accuracy. Meanwhile, the proposed DS-GCN notably outperformed state-of-the-art methods. The code is released here https://github.com/davelailai/DS-GCN Jianyang Xie, Yanda Meng, Yitian Zhao, Anh Nguyen 0003, Xiaoyun Yang, Yalin Zheng |
AAAI | 4 |
| 2024 | Guide3D: A Bi-planar X-ray Dataset for 3D Shape Reconstruction
Tudor Jianu, Baoru Huang, Hoan Nguyen, Binod Bhattarai, Tuong KL. Do, Erman Tjiputra, Quang D. Tran, Pierre Berthet-Rayne, T. Hoang Ngan Le, Sebastiano Fichera, Anh Nguyen 0003 |
ACCV (5) | 11 |
| 2024 | FG-CXR: A Radiologist-Aligned Gaze Dataset for Enhancing Interpretability in Chest X-Ray Report Generation
Trong-Thang Pham, Ngoc-Vuong Ho, Nhat-Tan Bui, Thinh Phan, Brijesh Patel 0001, Donald A. Adjeroh, Gianfranco Doretto, Anh Nguyen 0003, Carol C. Wu, T. Hoang Ngan Le |
ACCV (6) | 8 |
| 2024 | Generating Valid and Natural Adversarial Examples with Large Language ModelsabstractDeep learning-based natural language processing (NLP) models, particularly pre-trained language models (PLMs), have been revealed to be vulnerable to adversarial attacks. However, the adversarial examples generated by many mainstream word-level adversarial attack models are neither valid nor natural, leading to the loss of semantic maintenance, grammaticality, and human imperceptibility. Based on the exceptional capacity of language understanding and generation of large language models (LLMs), we propose LLM-Attack, which aims at generating both valid and natural adversarial examples with LLMs. The method consists of two stages: word importance ranking (which searches for the most vulnerable words) and word synonym replacement (which substitutes them with their synonyms obtained from LLMs). Experimental results on the Movie Review (MR), IMDB, and Yelp Review Polarity datasets against the baseline adversarial attack models illustrate the effectiveness of LLM-Attack, and it outperforms the baselines in human and GPT-4 evaluation by a significant margin. The model can generate adversarial examples that are typically valid and natural, with the preservation of semantic meaning, grammaticality, and human imperceptibility. Wei Wang 0042, Qi Chen 0026, Qiufeng Wang 0001, Anh Nguyen 0003 |
CSCWD | 5 |
| 2024 | Language-driven Grasp DetectionabstractGrasp detection is a persistent and intricate challenge with various industrial applications. Recently, many meth-ods and datasets have been proposed to tackle the grasp detection problem. However, most of them do not consider using natural language as a condition to detect the grasp poses. In this paper, we introduce Grasp-Anything++, a new language-driven grasp detection dataset featuring 1M samples, over 3M objects, and upwards of 10M grasping in-structions. We utilize foundation models to create a large-scale scene corpus with corresponding images and grasp prompts. We approach the language-driven grasp detection task as a conditional generation problem. Drawing on the success of diffusion models in generative tasks and given that language plays a vital role in this task, we propose a new language-driven grasp detection method based on dif-fusion models. Our key contribution is the contrastive training objective, which explicitly contributes to the denoising process to detect the grasp pose given the language instructions. We illustrate that our approach is theoretically sup-portive. The intensive experiments show that our method outperforms state-of-the-art approaches and allows real-world robotic grasping. Finally, we demonstrate our large-scale dataset enables zero-short grasp detection and is a challenging benchmark for future work. Vuong Dinh An, Minh Nhat Vu, Baoru Huang, Thieu Vo, Anh Nguyen 0003 |
CVPR | 7 |
| 2024 | Adaptive Parametric Activation
Konstantinos Panagiotis Alexandridis, Jiankang Deng, Anh Nguyen 0003, Shan Luo 0001 |
ECCV (54) | 3 |
| 2024 | Scalable Group Choreography via Variational Phase Manifold Learning
Nhat Le, Khoa Do, Xuan Bui, Tuong KL. Do, Erman Tjiputra, Quang D. Tran, Anh Nguyen 0003 |
ECCV (18) | 7 |
| 2024 | Language-Driven 6-DoF Grasp Detection Using Negative Prompt Guidance
Toan Nguyen 0004, Minh Nhat Vu, Baoru Huang, An Vuong, T. Hoang Ngan Le, Thieu Vo, Anh Nguyen 0003 |
ECCV (19) | 8 |
| 2024 | WAVER: Writing-Style Agnostic Text-Video Retrieval Via Distilling Vision-Language Models Through Open-Vocabulary KnowledgeabstractText-video retrieval, a prominent sub-field within the domain of multimodal information retrieval, has witnessed remarkable growth in recent years. However, existing methods assume video scenes are consistent with unbiased descriptions. These limitations fail to align with real-world scenarios since descriptions can be influenced by annotator biases, diverse writing styles, and varying textual perspectives. To overcome the aforementioned problems, we introduce WAVER, a cross-domain knowledge distillation framework via vision-language models through open-vocabulary knowledge designed to tackle the challenge of handling different writing styles in video descriptions. WAVER capitalizes on the open-vocabulary properties that lie in pre-trained vision-language models and employs an implicit knowledge distillation approach to transfer text-based knowledge from a teacher model to a vision-based student. Empirical studies conducted across four standard benchmark datasets, encompassing various settings, provide compelling evidence that WAVER can achieve state-of-the-art performance in text-video retrieval task while handling writing-style variations. The code is available at: https://github.com/Fsoft-AIC/WAVER Huy Le 0001, Tung Kieu, Anh Nguyen 0003, T. Hoang Ngan Le |
ICASSP | 3 |
| 2024 | Grasp-Anything: Large-scale Grasp Dataset from Foundation ModelsabstractFoundation models such as ChatGPT have made significant strides in robotic tasks due to their universal representation of real-world domains. In this paper, we leverage foundation models to tackle grasp detection, a persistent challenge in robotics with broad industrial applications. Despite numerous grasp datasets, their object diversity remains limited compared to real-world figures. Fortunately, foundation models possess an extensive repository of real-world knowledge, including objects we encounter in our daily lives. As a consequence, a promising solution to the limited representation in previous grasp datasets is to harness the universal knowledge embedded in these foundation models. We present Grasp-Anything, a new large-scale grasp dataset synthesized from foundation models to implement this solution. Grasp-Anything excels in diversity and magnitude, boasting 1M samples with text descriptions and more than 3M objects, surpassing prior datasets. Empirically, we show that Grasp-Anything successfully facilitates zero-shot grasp detection on vision-based tasks and real-world robotic experiments. Our dataset and code are available at https://airvlab.github.io/grasp-anything/. Vuong Dinh An, Minh Nhat Vu, Baoru Huang, Huynh Thi Thanh Binh, Thieu Vo, Andreas Kugi, Anh Nguyen 0003 |
ICRA | 8 |
| 2024 | Reducing Non-IID Effects in Federated Autonomous Driving with Contrastive Divergence LossabstractFederated learning has been widely applied in autonomous driving since it enables training a learning model among vehicles without sharing users’ data. However, data from autonomous vehicles usually suffer from the non-independent-and-identically-distributed (non-IID) problem, which may cause negative effects on the convergence of the learning process. In this paper, we propose a new contrastive divergence loss to address the non-IID problem in autonomous driving by reducing the impact of divergence factors from transmitted models during the local learning process of each silo. We also analyze the effects of contrastive divergence in various autonomous driving scenarios, under multiple network infrastructures, and with different centralized/distributed learning schemes. Our intensive experiments on three datasets demonstrate that our proposed contrastive divergence loss significantly improves the performance over current state-of-the-art approaches. Our source code is available at https://github.com/aioz-ai/CDL. Tuong KL. Do, Binh X. Nguyen, Quang D. Tran, Erman Tjiputra, Te-Chuan Chiu, Anh Nguyen 0003 |
ICRA | 7 |
| 2024 | Language-Conditioned Affordance-Pose Detection in 3D Point CloudsabstractAffordance detection and pose estimation are of great importance in many robotic applications. Their combination helps the robot gain an enhanced manipulation capability, in which the generated pose can facilitate the corresponding affordance task. Previous methods for affodance-pose joint learning are limited to a predefined set of affordances, thus limiting the adaptability of robots in real-world environments. In this paper, we propose a new method for language-conditioned affordance-pose joint learning in 3D point clouds. Given a 3D point cloud object, our method detects the affordance region and generates appropriate 6-DoF poses for any unconstrained affordance label. Our method consists of an open-vocabulary affordance detection branch and a language-guided diffusion model that generates 6-DoF poses based on the affordance text. We also introduce a new high-quality dataset for the task of language-driven affordance-pose joint learning. Intensive experimental results demonstrate that our proposed method works effectively on a wide range of open-vocabulary affordances and outperforms other baselines by a large margin. In addition, we illustrate the usefulness of our method in real-world robotic applications. Our code and dataset are publicly available at https://3DAPNet.github.io. Toan Nguyen 0004, Minh Nhat Vu, Baoru Huang, Tuan Van Vo, Vy Truong, T. Hoang Ngan Le, Thieu Vo, Bac Le, Anh Nguyen 0003 |
ICRA | 9 |
| 2024 | Open-Vocabulary Affordance Detection using Knowledge Distillation and Text-Point CorrelationabstractAffordance detection presents intricate challenges and has a wide range of robotic applications. Previous works have faced limitations such as the complexities of 3D object shapes, the wide range of potential affordances on real-world objects, and the lack of open-vocabulary support for affordance understanding. In this paper, we introduce a new open-vocabulary affordance detection method in 3D point clouds, leveraging knowledge distillation and text-point correlation. Our approach employs pre-trained 3D models through knowledge distillation to enhance feature extraction and semantic understanding in 3D point clouds. We further introduce a new text-point correlation method to learn the semantic links between point cloud features and open-vocabulary labels. The intensive experiments show that our approach outperforms previous works and adapts to new affordance labels and unseen objects. Notably, our method achieves the improvement of 7.96% mIOU score compared to the baselines. Furthermore, it offers real-time inference which is well-suitable for robotic manipulation applications. Tuan Van Vo, Minh Nhat Vu, Baoru Huang, Toan Nguyen 0004, T. Hoang Ngan Le, Thieu Vo, Anh Nguyen 0003 |
ICRA | 7 |
| 2024 | Open-Fusion: Real-time Open-Vocabulary 3D Mapping and Queryable Scene RepresentationabstractPrecise 3D environmental mapping with semantics is essential in robotics. Existing methods often rely on pre-defined concepts during training or are time-intensive when generating semantic maps. This paper presents Open-Fusion, an approach for real-time open-vocabulary 3D mapping and queryable scene representation using RGB-D data. Open-Fusion harnesses the power of a pretrained vision-language foundation model (VLFM) for open-set semantic comprehension and employs the Truncated Signed Distance Function (TSDF) for swift 3D scene reconstruction. By leveraging the VLFM, we extract region-based embeddings and their associated confidence maps. These are then integrated with the 3D knowledge from TSDF using an enhanced Hungarian-based feature-matching mechanism. In particular, Open-Fusion delivers outstanding annotation-free 3D segmentation for open vocabulary query without the need for additional 3D training. Benchmark tests on the ScanNet dataset against leading zero-shot methods highlight Open-Fusion’s superiority. Furthermore, it seamlessly combines the strengths of region-based VLFM and TSDF, facilitating real-time 3D scene comprehension that includes object concepts and open-world semantics. We encourage the readers to view the demos on our project page: https://uark-aicv.github.io/OpenFusion Kashu Yamazaki, Taisei Hanyu, Viet-Khoa Vo-Ho, Thang Pham, Gianfranco Doretto, Anh Nguyen 0003, T. Hoang Ngan Le |
ICRA | 7 |
| 2024 | ShapeFormer: Shape Prior Visible-to-Amodal Transformer-based Amodal Instance SegmentationabstractAmodal Instance Segmentation (AIS) presents a challenging task as it involves predicting both visible and occluded parts of objects within images. Existing AIS methods rely on a bidirectional approach, encompassing both the transition from amodal features to visible features (amodal-to-visible) and from visible features to amodal features (visible-to-amodal). Our observation shows that the utilization of amodal features through the amodal-to-visible can confuse the visible features due to the extra information of occluded/hidden segments not presented in visible display. Consequently, this compromised quality of visible features during the subsequent visible-to-amodal transition. To tackle this issue, we introduce ShapeFormer, a decoupled Transformer-based model with a visible-to-amodal transition. It facilitates the explicit relationship between output segmentations and avoids the need for amodal-to-visible transitions. ShapeFormer comprises three key modules: (i) Visible-Occluding Mask Head for predicting visible segmentation with occlusion awareness, (ii) Shape-Prior Amodal Mask Head for predicting amodal and occluded masks, and (iii) Category-Specific Shape Prior Retriever aims to provide shape prior knowledge. Comprehensive experiments and extensive ablation studies across various AIS benchmarks demonstrate the effectiveness of our ShapeFormer. The code is available at: https: //github.com/UARK-AICV/ShapeFormer Winston Bounsavy, Viet-Khoa Vo-Ho, Anh Nguyen 0003, Tri Nguyen 0005, T. Hoang Ngan Le |
IJCNN | 4 |
| 2024 | Lightweight Language-driven Grasp Detection using Conditional Consistency ModelabstractLanguage-driven grasp detection is a fundamental yet challenging task in robotics with various industrial applications. This work presents a new approach for language-driven grasp detection that leverages lightweight diffusion models to achieve fast inference time. By integrating diffusion processes with grasping prompts in natural language, our method can effectively encode visual and textual information, enabling more accurate and versatile grasp positioning that aligns well with the text query. To overcome the long inference time problem in diffusion models, we leverage the image and text features as the condition in the consistency model to reduce the number of denoising timesteps during inference. The intensive experimental results show that our method outperforms other recent grasp detection methods and lightweight diffusion models by a clear margin. We further validate our method in real-world robotic experiments to demonstrate its fast inference time capability. Minh Nhat Vu, Baoru Huang, An Vuong, T. Hoang Ngan Le, Thieu Vo, Anh Nguyen 0003 |
IROS | 7 |
| 2024 | Language-driven Grasp Detection with Mask-guided AttentionabstractGrasp detection is an essential task in robotics with various industrial applications. However, traditional methods often struggle with occlusions and do not utilize language for grasping. Incorporating natural language into grasp detection remains a challenging task and largely unexplored. To address this gap, we propose a new method for language-driven grasp detection with mask-guided attention by utilizing the transformer attention mechanism with semantic segmentation features. Our approach integrates visual data, segmentation mask features, and natural language instructions, significantly improving grasp detection accuracy. Our work introduces a new framework for language-driven grasp detection, paving the way for language-driven robotic applications. Intensive experiments show that our method outperforms other recent baselines by a clear margin, with a 10.0% success score improvement. We further validate our method in real-world robotic experiments, confirming the effectiveness of our approach. Tuan Van Vo, Minh Nhat Vu, Baoru Huang, An Vuong, T. Hoang Ngan Le, Thieu Vo, Anh Nguyen 0003 |
IROS | 7 |
| 2024 | HabiCrowd: A High Performance Simulator for Crowd-Aware Visual NavigationabstractVisual navigation, a foundational aspect of Embodied AI (E-AI) and robotics has been extensively studied in the past few years. While many 3D simulators have been introduced for the visual navigation tasks, scarcely works have combined human dynamics, creating the gap between simulation and real-world applications. Furthermore, current 3D simulators incorporating human dynamics have several limitations, particularly in terms of computational efficiency, which is a promise of modern simulators. To overcome these issues, we introduce HabiCrowd, the new standard benchmark for crowd-aware visual navigation that includes a crowd dynamics model with diverse human settings into photorealistic environments. Empirical evaluations demonstrate that our proposed human dynamics model achieves state-of-the-art performance in collision avoidance while exhibiting superior computational efficiency compared to its counterparts. We leverage HabiCrowd to conduct several comprehensive studies on crowd-aware visual navigation tasks and human-robot interactions. The source code and data can be found at https://habicrowd.github.io/. An Vuong, Toan Nguyen 0004, Minh Nhat Vu, Baoru Huang, Huynh Thi Thanh Binh, Thieu Vo, Anh Nguyen 0003 |
IROS | 7 |
| 2024 | Multi-disease Detection in Retinal Images Guided by Disease Causal Estimation
Jianyang Xie, Xiuju Chen, Yitian Zhao, Yanda Meng, He Zhao 0002, Anh Nguyen 0003, Yalin Zheng |
MICCAI (1) | 6 |
| 2024 | CLIP-DR: Textual Knowledge-Guided Diabetic Retinopathy Grading with Ranking-Aware Prompting
Qinkai Yu, Jianyang Xie, Anh Nguyen 0003, He Zhao 0002, Jiong Zhang 0004, Huazhu Fu, Yitian Zhao, Yalin Zheng, Yanda Meng |
MICCAI (1) | 3 |
| 2024 | Interpret Your Decision: Logical Reasoning Regularization for Generalization in Visual ClassificationabstractVision models excel in image classification but struggle to generalize to unseen data, such as classifying images from unseen domains or discovering novel categories. In this paper, we explore the relationship between logical reasoning and deep learning generalization in visual classification. A logical regularization termed L-Reg is derived which bridges a logical analysis framework to image classification. Our work reveals that L-Reg reduces the complexity of the model in terms of the feature distribution and classifier weights. Specifically, we unveil the interpretability brought by L-Reg, as it enables the model to extract the salient features, such as faces to persons, for classification. Theoretical analysis and experiments demonstrate that L-Reg enhances generalization across various scenarios, including multi-domain generalization and generalized category discovery. In complex real-world scenarios where images span unknown classes and unseen domains, L-Reg consistently improves generalization, highlighting its practical efficacy. Zhaorui Tan, Xi Yang 0008, Qiufeng Wang 0001, Anh Nguyen 0003, Kaizhu Huang |
NeurIPS | 4 |
| 2024 | Domain-specific Guided Summarization for Mental Health Posts
Lu Qian, Haiyang Zhang 0004, Wei Wang 0042, Anh Nguyen 0003 |
PACLIC | 7 |
| 2024 | I-AI: A Controllable & Interpretable AI System for Decoding Radiologists' Intense Focus for Accurate CXR DiagnosesabstractIn the field of chest X-ray (CXR) diagnosis, existing works often focus solely on determining where a radiologist looks, typically through tasks such as detection, segmentation, or classification. However, these approaches are often designed as black-box models, lacking interpretability. In this paper, we introduce Interpretable Artificial Intelligence (I-AI) a novel and unified controllable interpretable pipeline for decoding the intense focus of radiologists in CXR diagnosis. Our I-AI addresses three key questions: where a radiologist looks, how long they focus on specific areas, and what findings they diagnose. By capturing the intensity of the radiologist’s gaze, we provide a unified solution that offers insights into the cognitive process underlying radiological interpretation. Unlike current methods that rely on black-box machine learning models, which can be prone to extracting erroneous information from the entire input image during the diagnosis process, we tackle this issue by effectively masking out irrelevant information. Our proposed I-AI leverages a vision-language model, allowing for precise control over the interpretation process while ensuring the exclusion of irrelevant features.To train our I-AI model, we utilize an eye gaze dataset to extract anatomical gaze information and generate ground truth heatmaps. Through extensive experimentation, we demonstrate the efficacy of our method. We showcase that the attention heatmaps, designed to mimic radiologists’ focus, encode sufficient and relevant information, enabling accurate classification tasks using only a portion of CXR. The code, checkpoints, and data are at https://github.com/UARK-AICV/IAI. Trong-Thang Pham, Jacob Brecheisen, Anh Nguyen 0003, T. Hoang Ngan Le |
WACV | 3 |
| 2024 | Self-supervised learning for point cloud data: A surveyabstract3D point clouds are a crucial type of data collected by LiDAR sensors and widely used in transportation applications due to its concise descriptions and accurate localization. Deep neural networks (DNNs) have achieved remarkable success in processing large amount of disordered and sparse 3D point clouds, especially in various computer vision tasks, such as pedestrian detection and vehicle recognition. Among all the learning paradigms, Self-Supervised Learning (SSL), an unsupervised training paradigm that mines effective information from the data itself, is considered as an essential solution to solve the time-consuming and labor-intensive data labelling problems via smart pre-training task design. This paper provides a comprehensive survey of recent advances on SSL for point clouds. We first present an innovative taxonomy, categorizing the existing SSL methods into four broad categories based on the pretexts’ characteristics. Under each category, we then further categorize the methods into more fine-grained groups and summarize the strength and limitations of the representative methods. We also compare the performance of the notable SSL methods in literature on multiple downstream tasks on benchmark datasets both quantitatively and qualitatively. Finally, we propose a number of future research directions based on the identified limitations of existing SSL research on point clouds. Changyu Zeng, Wei Wang 0042, Anh Nguyen 0003, Jimin Xiao, Yutao Yue |
Expert Syst. Appl. | 3 |
| 2024 | Zero-shot text classification with knowledge resources under label-fully-unseen setting
Wei Wang 0042, Qi Chen 0026, Kaizhu Huang, Anh Nguyen 0003, Suparna De |
Neurocomputing | 5 |
| 2024 | Film-GAN: towards realistic analog film photo generation
Haoyan Gong, Jionglong Su, Kah Phooi Seng, Anh Nguyen 0003, Hongbin Liu 0007 |
Neural Comput. Appl. | 4 |
| 2024 | Dynamic Semantic-Based Spatial-Temporal Graph Convolution Network for Skeleton-Based Human Action RecognitionabstractHuman action recognition is an essential topic in computer vision and image processing. Graph convolutional networks (GCNs) have attracted significant attention and achieved noteworthy performance in skeleton-based human action recognition tasks. However, most of the previous graph-based works are designed to refine skeleton topology without considering the types of different joints and edges and the occurrence order of the frames. Such a limitation makes them insufficient to represent intrinsic semantic information. Differently, we proposed a dynamic semantic-based spatial-temporal graph convolution network (DS-STGCN) to address the challenge. DS-STGCN has two dynamic semantic modules for spatial and temporal contexts respectively. Specifically, the joints and edge types were encoded in the spatial module implicitly, and the occurrence order of frames was encoded in the temporal module implicitly. Extensive experiments on four datasets including NTU-RGB+D 60(120), Kinetics-400, and FineGYM show that our proposed two semantic modules can bring consistent recognition performance improvement with various backbones. Meanwhile, the proposed DS-STGCN notably surpassed state-of-the-art methods on these datasets. Notably, in the more challenging dataset, such as Kinetics-400, our model significantly outperformed other state-of-the-art GCN-based methods by a large margin. The code has been released at https://github.com/davelailai/DS-STGCN. Jianyang Xie, Yanda Meng, Yitian Zhao, Anh Nguyen 0003, Xiaoyun Yang, Yalin Zheng |
IEEE Trans. Image Process. | 4 |
| 2023 | Learning to Terminate in Object Navigation
Yuhang Song 0008, Anh Nguyen 0003, Chun-Yi Lee |
ACML | 2 |
| 2023 | Music-Driven Group ChoreographyabstractMusic-driven choreography is a challenging problem with a wide variety of industrial applications. Recently, many methods have been proposed to synthesize dance motions from music for a single dancer. However, generating dance motion for a group remains an open problem. In this paper, we present AIOZ - GDANCE, a new large-scale dataset for music-driven group dance generation. Unlike existing datasets that only support single dance, our new dataset contains group dance videos, hence supporting the study of group choreography. We propose a semi-autonomous labeling method with humans in the loop to obtain the 3D ground truth for our dataset. The proposed dataset consists of 16.7 hours of paired music and 3D motion from in-the-wild videos, covering 7 dance styles and 16 music genres. We show that naively applying single dance generation technique to creating group dance motion may lead to unsatisfactory results, such as inconsistent movements and collisions between dancers. Based on our new dataset, we propose a new method that takes an input music sequence and a set of 3D positions of dancers to efficiently produce multiple group-coherent choreographies. We propose new evaluation metrics for measuring group dance quality and perform intensive experiments to demonstrate the effectiveness of our method. Our project facilitates future research on group dance generation and is available at https://aioz-ai.github.io/AIOZ-GDANCE/. Nhat Le, Trong-Thang Pham, Tuong KL. Do, Erman Tjiputra, Quang D. Tran, Anh Nguyen 0003 |
CVPR | 6 |
| 2023 | Reducing Training Time in Cross-Silo Federated Learning using Multigraph TopologyabstractFederated learning is an active research topic since it enables several participants to jointly train a model without sharing local data. Currently, cross-silo federated learning is a popular training setting that utilizes a few hundred reliable data silos with high-speed access links to training a model. While this approach has been widely applied in real-world scenarios, designing a robust topology to reduce the training time remains an open problem. In this paper, we present a new multigraph topology for cross-silo federated learning. We first construct the multigraph using the overlay graph. We then parse this multigraph into different simple graphs with isolated nodes. The existence of isolated nodes allows us to perform model aggregation without waiting for other nodes, hence effectively reducing the training time. Intensive experiments on three public datasets show that our proposed method significantly reduces the training time compared with recent state-of-the-art topologies while maintaining the accuracy of the learned model. Our code can be found at: https://github.com/aioz-ai/MultigraphFL Tuong KL. Do, Binh X. Nguyen, Vuong Pham, Erman Tjiputra, Quang D. Tran, Anh Nguyen 0003 |
ICCV | 7 |
| 2023 | Open-Vocabulary Affordance Detection in 3D Point CloudsabstractAffordance detection is a challenging problem with a wide variety of robotic applications. Traditional affordance detection methods are limited to a predefined set of affordance labels, hence potentially restricting the adaptability of intelligent robots in complex and dynamic environments. In this paper, we present the Open-Vocabulary Affordance Detection (OpenAD) method, which is capable of detecting an unbounded number of affordances in 3D point clouds. By simultaneously learning the affordance text and the point feature, OpenAD successfully exploits the semantic relationships between affordances. Therefore, our proposed method enables zero-shot detection and can be able to detect previously unseen affordances without a single annotation example. Intensive experimental results show that OpenAD works effectively on a wide range of affordance detection setups and outperforms other baselines by a large margin. Additionally, we demonstrate the practicality of the proposed OpenAD in real-world robotic applications with a fast inference speed. Our project is available at https://openad2023.github.io. Toan Nguyen 0004, Minh Nhat Vu, An Vuong, Dzung Nguyen, Thieu Vo, T. Hoang Ngan Le, Anh Nguyen 0003 |
IROS | 7 |
| 2023 | Detecting the Sensing Area of a Laparoscopic Probe in Minimally Invasive Cancer Surgery
Baoru Huang, Anh Nguyen 0003, Stamatia Giannarou, Daniel S. Elson |
MICCAI (9) | 3 |
| 2023 | Language-driven Scene Synthesis using Multi-conditional Diffusion ModelabstractScene synthesis is a challenging problem with several industrial applications. Recently, substantial efforts have been directed to synthesize the scene using human motions, room layouts, or spatial graphs as the input. However, few studies have addressed this problem from multiple modalities, especially combining text prompts. In this paper, we propose a language-driven scene synthesis task, which is a new task that integrates text prompts, human motion, and existing objects for scene synthesis. Unlike other single-condition synthesis tasks, our problem involves multiple conditions and requires a strategy for processing and encoding them into a unified space. To address the challenge, we present a multi-conditional diffusion model, which differs from the implicit unification approach of other diffusion literature by explicitly predicting the guiding points for the original data distribution. We demonstrate that our approach is theoretically supportive. The intensive experiment results illustrate that our method outperforms state-of-the-art benchmarks and enables natural scene editing applications. The source code and dataset can be accessed at https://lang-scene-synth.github.io/. Vuong Dinh An, Minh Nhat Vu, Toan Nguyen 0004, Baoru Huang, Dzung Nguyen, Thieu Vo, Anh Nguyen 0003 |
NeurIPS | 7 |
| 2023 | Uncertainty-aware Label Distribution Learning for Facial Expression RecognitionabstractDespite significant progress over the past few years, ambiguity is still a key challenge in Facial Expression Recognition (FER). It can lead to noisy and inconsistent annotation, which hinders the performance of deep learning models in real-world scenarios. In this paper, we propose a new uncertainty-aware label distribution learning method to improve the robustness of deep models against uncertainty and ambiguity. We leverage neighborhood information in the valence-arousal space to adaptively construct emotiona distributions for training samples. We also consider the uncertainty of provided labels when incorporating them into the label distributions. Our method can be easily integrated into a deep network to obtain more training supervision and improve recognition accuracy. Intensive experiments on several datasets under various noisy and ambiguous settings show that our method achieves competitive results and outperforms recent state-of-the-art approaches. Our code and models are available at https://github.com/minhnhatvt/label-distribution-learning-fer-tf. Nhat Le, Quang D. Tran, Erman Tjiputra, Bac Le, Anh Nguyen 0003 |
WACV | 6 |
| 2023 | Semi-supervised adversarial discriminative domain adaptation
Thai-Vu Nguyen, Anh Nguyen 0003, Nghia Le, Bac Le |
Appl. Intell. | 2 |
| 2023 | Semantic Similarity Distance: Towards better text-image consistency metric in text-to-image generation
Zhaorui Tan, Xi Yang 0008, Zihan Ye, Qiufeng Wang 0001, Yuyao Yan, Anh Nguyen 0003, Kaizhu Huang |
Pattern Recognit. | 6 |
| 2023 | Inverse Image Frequency for Long-Tailed Image RecognitionabstractThe long-tailed distribution is a common phenomenon in the real world. Extracted large scale image datasets inevitably demonstrate the long-tailed property and models trained with imbalanced data can obtain high performance for the over-represented categories, but struggle for the under-represented categories, leading to biased predictions and performance degradation. To address this challenge, we propose a novel de-biasing method named Inverse Image Frequency (IIF). IIF is a multiplicative margin adjustment transformation of the logits in the classification layer of a convolutional neural network. Our method achieves stronger performance than similar works and it is especially useful for downstream tasks such as long-tailed instance segmentation as it produces fewer false positive detections. Our extensive experiments show that IIF surpasses the state of the art on many long-tailed benchmarks such as ImageNet-LT, CIFAR-LT, Places-LT and LVIS, reaching 55.8% top-1 accuracy with ResNet50 on ImageNet-LT and 26.3% segmentation AP with MaskRCNN ResNet50 on LVIS. Code available at https://github.com/kostas1515/iif. Konstantinos Panagiotis Alexandridis, Shan Luo 0001, Anh Nguyen 0003, Jiankang Deng, Stefanos Zafeiriou |
IEEE Trans. Image Process. | 3 |
| 2023 | Controllable Group Choreography Using Contrastive DiffusionabstractMusic-driven group choreography poses a considerable challenge but holds significant potential for a wide range of industrial applications. The ability to generate synchronized and visually appealing group dance motions that are aligned with music opens up opportunities in many fields such as entertainment, advertising, and virtual performances. However, most of the recent works are not able to generate high-fidelity long-term motions, or fail to enable controllable experience. In this work, we aim to address the demand for high-quality and customizable group dance generation by effectively governing the consistency and diversity of group choreographies. In particular, we utilize a diffusion-based generative approach to enable the synthesis of flexible number of dancers and long-term group dances, while ensuring coherence to the input music. Ultimately, we introduce a Group Contrastive Diffusion (GCD) strategy to enhance the connection between dancers and their group, presenting the ability to control the consistency or diversity level of the synthesized group animation via the classifier-guidance sampling technique. Through intensive experiments and evaluation, we demonstrate the effectiveness of our approach in producing visually captivating and consistent group dance motions. The experimental results show the capability of our method to achieve the desired levels of consistency and diversity, while maintaining the overall quality of the generated group choreography. Nhat Le, Tuong KL. Do, Khoa Do, Erman Tjiputra, Quang D. Tran, Anh Nguyen 0003 |
ACM Trans. Graph. | 7 |
| 2022 | Long-Tailed Instance Segmentation Using Gumbel Optimized Loss
Konstantinos Panagiotis Alexandridis, Jiankang Deng, Anh Nguyen 0003, Shan Luo 0001 |
ECCV (10) | 3 |
| 2022 | Deep Federated Learning for Autonomous DrivingabstractAutonomous driving is an active research topic in both academia and industry. However, most of the existing solutions focus on improving the accuracy by training learnable models with centralized large-scale data. Therefore, these methods do not take into account the user’s privacy. In this paper, we present a new approach to learn autonomous driving policy while respecting privacy concerns. We propose a peer-to-peer Deep Federated Learning (DFL) approach to train deep architectures in a fully decentralized manner and remove the need for central orchestration. We design a new Federated Autonomous Driving network (FADNet) that can improve the model stability, ensure convergence, and handle imbalanced data distribution problems while is being trained with federated learning methods. Intensively experimental results on three datasets show that our approach with FADNet and DFL achieves superior accuracy compared with other recent methods. Furthermore, our approach can maintain privacy by not collecting user data to a central server. Our source code can be found at: https://github.com/aioz-ai/FADNet Anh Nguyen 0003, Tuong KL. Do, Binh X. Nguyen, Chien Duong, Tu Phan, Erman Tjiputra, Quang D. Tran |
IV | 1 |
| 2022 | Self-supervised Depth Estimation in Laparoscopic Image Using 3D Geometric Consistency
Baoru Huang, Jian-Qing Zheng, Anh Nguyen 0003, Ioannis Gkouzionis, Kunal Vyas, David Tuch, Stamatia Giannarou, Daniel S. Elson |
MICCAI (8) | 3 |
| 2022 | Generalised Zero-shot Learning for Entailment-based Text Classification with External KnowledgeabstractText classification techniques have been substantially important to many smart computing applications, e.g. topic extraction and event detection. However, classification is always challenging when only insufficient amount of labelled data for model training is available. To mitigate this issue, zero-shot learning (ZSL) has been introduced for models to recognise new classes that have not been observed during the training stage. We propose an entailment-based zero-shot text classification model, named as S-BERT-CAM, to better capture the relationship between the premise and hypothesis in the BERT embedding space. Two widely used textual datasets are utilised to conduct the experiments. We fine-tune our model using 50% of the labels for each dataset and evaluate it on the label space containing all labels (including both seen and unseen labels). The experimental results demonstrate that our model is more robust to the generalised ZSL and significantly improves the overall performance against baselines. Wei Wang 0042, Qi Chen 0026, Kaizhu Huang, Anh Nguyen 0003, Suparna De |
SMARTCOMP | 5 |
| 2022 | Global-local attention for emotion recognitionabstractAbstract Human emotion recognition is an active research area in artificial intelligence and has made substantial progress over the past few years. Many recent works mainly focus on facial regions to infer human affection, while the surrounding context information is not effectively utilized. In this paper, we proposed a new deep network to effectively recognize human emotions using a novel global-local attention mechanism. Our network is designed to extract features from both facial and context regions independently, then learn them together using the attention module. In this way, both the facial and contextual information is used to infer human emotions, therefore enhancing the discrimination of the classifier. The intensive experiments show that our method surpasses the current state-of-the-art methods on recent emotion datasets by a fair margin. Qualitatively, our global-local attention module can extract more meaningful attention maps than previous methods. The source code and trained model of our network are available at https://github.com/minhnhatvt/glamor-net . Nhat Le, Anh Nguyen 0003, Bac Le |
Neural Comput. Appl. | 3 |
| 2022 | Light-Weight Deformable Registration Using Adversarial Learning With Distilling KnowledgeabstractDeformable registration is a crucial step in many medical procedures such as image-guided surgery and radiation therapy. Most recent learning-based methods focus on improving the accuracy by optimizing the non-linear spatial correspondence between the input images. Therefore, these methods are computationally expensive and require modern graphic cards for real-time deployment. In this paper, we introduce a new Light-weight Deformable Registration network that significantly reduces the computational cost while achieving competitive accuracy. In particular, we propose a new adversarial learning with distilling knowledge algorithm that successfully leverages meaningful information from the effective but expensive teacher network to the student network. We design the student network such as it is light-weight and well suitable for deployment on a typical CPU. The extensively experimental results on different public datasets show that our proposed method achieves state-of-the-art accuracy while significantly faster than recent methods. We further show that the use of our adversarial learning algorithm is essential for a time-efficiency deformable registration method. Finally, our source code and trained models are available at https://github.com/aioz-ai/LDR_ALDK. Minh Q. Tran, Tuong KL. Do, Erman Tjiputra, Quang D. Tran, Anh Nguyen 0003 |
IEEE Trans. Medical Imaging | 6 |
| 2021 | Speech Emotion Recognition Using Semantic InformationabstractSpeech emotion recognition is a crucial problem manifesting in a multitude of applications such as human computer interaction and education. Although several advancements have been made in the recent years, especially with the advent of Deep Neural Networks (DNN), most of the studies in the literature fail to consider the semantic information in the speech signal. In this paper, we propose a novel framework that can capture both the semantic and the paralinguistic information in the signal. In particular, our framework is comprised of a semantic feature extractor, that captures the semantic information, and a paralinguistic feature extractor, that captures the paralinguistic information. Both semantic and paraliguistic features are then combined to a unified representation using a novel attention mechanism. The unified feature vector is passed through a LSTM to capture the temporal dynamics in the signal, before the final prediction. To validate the effectiveness of our framework, we use the popular SEWA dataset of the AVEC challenge series and compare with the three winning papers. Our model provides state-of-the-art results in the valence and liking dimensions.1 Panagiotis Tzirakis, Anh Nguyen 0003, Stefanos Zafeiriou, Björn W. Schuller |
ICASSP | 2 |
| 2021 | Multiple Meta-model Quantifying for Medical Visual Question Answering
Tuong KL. Do, Binh X. Nguyen, Erman Tjiputra, Quang D. Tran, Anh Nguyen 0003 |
MICCAI (5) | 6 |
| 2021 | Self-supervised Generative Adversarial Network for Depth Estimation in Laparoscopic Images
Baoru Huang, Jian-Qing Zheng, Anh Nguyen 0003, David Tuch, Kunal Vyas, Stamatia Giannarou, Daniel S. Elson |
MICCAI (4) | 3 |
| 2020 | Collaborative Robot-Assisted Endovascular Catheterization with Generative Adversarial Imitation LearningabstractMaster-slave systems for endovascular catheterization have brought major clinical benefits including reduced radiation doses to the operators, improved precision and stability of the instruments, as well as reduced procedural duration. Emerging deep reinforcement learning (RL) technologies could potentially automate more complex endovascular tasks with enhanced success rates, more consistent motion and reduced fatigue and cognitive workload of the operators. However, the complexity of the pulsatile flows within the vasculature and non-linear behavior of the instruments hinder the use of model-based approaches for RL. This paper describes model-free generative adversarial imitation learning to automate a standard arterial catherization task. The automation policies have been trained in a pre-clinical setting. Detailed validation results show high success rates after skill transfer to a different vascular anatomical model. The quality of the catheter motions also shows less mean and maximum contact forces compared to manual-based approaches. Wenqiang Chi, Giulio Dagnino, Trevor M. Y. Kwok, Anh Nguyen 0003, Dennis Kundrat, Mohamed E. M. K. Abdelaziz, Celia V. Riga, Colin D. Bicknell, Guang-Zhong Yang |
ICRA | 4 |
| 2020 | End-to-End Real-time Catheter Segmentation with Optical Flow-Guided Warping during Endovascular InterventionabstractAccurate real-time catheter segmentation is an important pre-requisite for robot-assisted endovascular intervention. Most of the existing learning-based methods for catheter segmentation and tracking are only trained on smallscale datasets or synthetic data due to the difficulties of ground-truth annotation. Furthermore, the temporal continuity in intraoperative imaging sequences is not fully utilised. In this paper, we present FW-Net, an end-to-end and real-time deep learning framework for endovascular intervention. The proposed FW-Net has three modules: a segmentation network with encoder-decoder architecture, a flow network to extract optical flow information, and a novel flow-guided warping function to learn the frame-to-frame temporal continuity. We show that by effectively learning temporal continuity, the network can successfully segment and track the catheters in real-time sequences using only raw ground-truth for training. Detailed validation results confirm that our FW-Net outperforms stateof-the-art techniques while achieving real-time performance. Anh Nguyen 0003, Dennis Kundrat, Giulio Dagnino, Wenqiang Chi, Mohamed E. M. K. Abdelaziz, Yao Guo 0002, YingLiang Ma, Trevor M. Y. Kwok, Celia V. Riga, Guang-Zhong Yang |
ICRA | 1 |
| 2020 | Autonomous Navigation in Complex Environments with Deep Multimodal Fusion NetworkabstractAutonomous navigation in complex environments is a crucial task in time-sensitive scenarios such as disaster response or search and rescue. However, complex environments pose significant challenges for autonomous platforms to navigate due to their challenging properties: constrained narrow passages, unstable pathway with debris and obstacles, or irregular geological structures and poor lighting conditions. In this work, we propose a multimodal fusion approach to address the problem of autonomous navigation in complex environments such as collapsed cites, or natural caves. We first simulate the complex environments in a physics-based simulation engine and collect a large-scale dataset for training. We then propose a Navigation Multimodal Fusion Network (NMFNet) which has three branches to effectively handle three visual modalities: laser, RGB images, and point cloud data. The extensively experimental results show that our NMFNet outperforms recent state of the art by a fair margin while achieving real-time performance. We further show that the use of multiple modalities is essential for autonomous navigation in complex environments. Finally, we successfully deploy our network to both simulated and real mobile robots. Anh Nguyen 0003, Ngoc Nguyen, Kim Tran, Erman Tjiputra, Quang D. Tran |
IROS | 1 |
| 2018 | AffordanceNet: An End-to-End Deep Learning Approach for Object Affordance DetectionabstractWe propose AffordanceNet, a new deep learning approach to simultaneously detect multiple objects and their affordances from RGB images. Our AffordanceNet has two branches: an object detection branch to localize and classify the object, and an affordance detection branch to assign each pixel in the object to its most probable affordance label. The proposed framework employs three key components for effectively handling the multiclass problem in the affordance mask: a sequence of deconvolutional layers, a robust resizing strategy, and a multi-task loss function. The experimental results on the public datasets show that our AffordanceNet outperforms recent state-of-the-art methods by a fair margin, while its end-to-end architecture allows the inference at the speed of 150ms per image. This makes our AffordanceNet well suitable for real-time robotic applications. Furthermore, we demonstrate the effectiveness of AffordanceNet in different testing environments and in real robotic applications. The source code is available at https://github.com/nqanh/affordance-net. Thanh-Toan Do, Anh Nguyen 0003, Ian D. Reid 0001 |
ICRA | 2 |
| 2018 | Translating Videos to Commands for Robotic Manipulation with Deep Recurrent Neural NetworksabstractWe present a new method to translate videos to commands for robotic manipulation using Deep Recurrent Neural Networks (RNN). Our framework first extracts deep features from the input video frames with a deep Convolutional Neural Networks (CNN). Two RNN layers with an encoder-decoder architecture are then used to encode the visual features and sequentially generate the output words as the command. We demonstrate that the translation accuracy can be improved by allowing a smooth transaction between two RNN layers and using the state-of-the-art feature extractor. The experimental results on our new challenging dataset show that our approach outperforms recent methods by a fair margin. Furthermore, we combine the proposed translation module with the vision and planning system to let a robot perform various manipulation tasks. Finally, we demonstrate the effectiveness of our framework on a full-size humanoid robot WALK-MAN. Anh Nguyen 0003, Dimitrios Kanoulas, Luca Muratore, Darwin G. Caldwell, Nikolaos G. Tsagarakis |
ICRA | 1 |
| 2017 | Object-based affordances detection with Convolutional Neural Networks and dense Conditional Random FieldsabstractWe present a new method to detect object affordances in real-world scenes using deep Convolutional Neural Networks (CNN), an object detector and dense Conditional Random Fields (CRF). Our system first trains an object detector to generate bounding box candidates from the images. A deep CNN is then used to learn the depth features from these bounding boxes. Finally, these feature maps are post-processed with dense CRF to improve the prediction along class boundaries. The experimental results on our new challenging dataset show that the proposed approach outperforms recent state-of-the-art methods by a substantial margin. Furthermore, from the detected affordances we introduce a grasping method that is robust to noisy data. We demonstrate the effectiveness of our framework on the full-size humanoid robot WALK-MAN using different objects in real-world scenarios. Anh Nguyen 0003, Dimitrios Kanoulas, Darwin G. Caldwell, Nikolaos G. Tsagarakis |
IROS | 1 |
| 2016 | Preparatory object reorientation for task-oriented graspingabstractThis paper describes a new task-oriented grasping method to reorient a rigid object to its nominal pose, which is defined as the configuration that it needs to be grasped from, in order to successfully execute a particular manipulation task. Our method combines two key insights: (1) a visual 6 Degree-of-Freedom (DoF) pose estimation technique based on 2D-3D point correspondences is used to estimate the object pose in real-time and (2) the rigid transformation from the current to the nominal pose is computed online and the object is reoriented over a sequence of steps. The outcome of this work is a novel method that can be effectively used in the preparatory phase of a manipulation task, to permit a robot to start from arbitrary object placements and configure the manipulated objects to the nominal pose, as required for the execution of a subsequent task. We experimentally demonstrate the effectiveness of our approach on a full-size humanoid robot (WALK-MAN) using different objects with various pose settings under real-time constraints. Anh Nguyen 0003, Dimitrios Kanoulas, Darwin G. Caldwell, Nikolaos G. Tsagarakis |
IROS | 1 |
| 2016 | Detecting object affordances with Convolutional Neural NetworksabstractWe present a novel and real-time method to detect object affordances from RGB-D images. Our method trains a deep Convolutional Neural Network (CNN) to learn deep features from the input data in an end-to-end manner. The CNN has an encoder-decoder architecture in order to obtain smooth label predictions. The input data are represented as multiple modalities to let the network learn the features more effectively. Our method sets a new benchmark on detecting object affordances, improving the accuracy by 20% in comparison with the state-of-the-art methods that use hand-designed geometric features. Furthermore, we apply our detection method on a full-size humanoid robot (WALK-MAN) to demonstrate that the robot is able to perform grasps after efficiently detecting the object affordances. Anh Nguyen 0003, Dimitrios Kanoulas, Darwin G. Caldwell, Nikolaos G. Tsagarakis |
IROS | 1 |
| 2014 | Contextual Labeling 3D Point Clouds with Conditional Random Fields
Anh Nguyen 0003, Bac Le |
ACIIDS (1) | 1 |