EDBT 2026 Demo / reviewers in the wild / expert
Viet-Khoa Vo-Ho
dblp:224/1953 · also Khoa Vo 0001
· DBLP profile ↗
16ranked-venue papers
4as first author
16since 2021 · last 2026
0000-0003-0277-7094ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 12 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric PerspectiveabstractAs embodied agents operate in increasingly complex environments, the ability to perceive, track, and reason about individual object instances over time becomes essential, especially in tasks requiring sequenced interactions with visually similar objects. In non-Markovian settings, critical decision cues lie in object histories rather than the current scene. Without persistent memory of prior interactions (what was used, where it was placed, or how it changed), visuomotor policies may fail, repeat past actions, or overlook completed ones. To surface this challenge, we introduce LIBERO-Mem, a non-Markovian task suite for stress-testing robotic manipulation under object-level partial observability. It combines short- and long-horizon object tracking with temporally sequenced subgoals, requiring reasoning beyond the current frame. However, vision-language-action (VLA) models often struggle in such settings, with token scaling quickly becoming intractable even for tasks spanning just a few hundred frames. We propose Embodied-SlotSSM, a slot-centric VLA framework built for temporal scalability. It maintains spatio-temporally consistent slot identities and leverages them through two mechanisms: (1) slot-state-space modeling for reconstructing short-term history, and (2) a relational encoder to align the input tokens with action decoding. Together, these components enable temporally grounded, context-aware action prediction. Experiments show Embodied-SlotSSM's baseline performance on LIBERO-Mem and general tasks, offering a scalable solution for non-Markovian reasoning in object-centric policies. Nhat Chung, Taisei Hanyu, Toan Nguyen 0004, Huy Le 0001, Frederick Bumgarner, Duy M. H. Nguyen, Viet-Khoa Vo-Ho, Kashu Yamazaki, Chase Rainwater, Tung Kieu, Anh Nguyen 0003, T. Hoang Ngan Le |
AAAI | 7 |
| 2025 | CT-ScanGaze: A Dataset and Baselines for 3D Volumetric Scanpath ModelingabstractUnderstanding radiologists' eye movement during Computed Tomography (CT) reading is crucial for developing effective interpretable computer-aided diagnosis systems. However, CT research in this area has been limited by the lack of publicly available eye-tracking datasets and the three-dimensional complexity of CT volumes. To address these challenges, we present the first publicly available eye gaze dataset on CT, called CT-ScanGaze. Then, we introduce CT-Searcher, a novel 3D scanpath predictor designed specifically to process CT volumes and generate radiologist-like 3D fixation sequences, overcoming the limitations of current scanpath predictors that only handle 2D inputs. Since deep learning models benefit from a pretraining step, we develop a pipeline that converts existing 2D gaze datasets into 3D gaze data to pretrain CT-Searcher. Through both qualitative and quantitative evaluations on CT-ScanGaze, we demonstrate the effectiveness of our approach and provide a comprehensive assessment framework for 3D scanpath prediction in medical imaging. Trong-Thang Pham, Akash Awasthi, Saba Khan, Esteban Duran Marti, Tien-Phat Nguyen, Viet-Khoa Vo-Ho, Cuong Tran 0010, Yuki Ikebe, Anh Totti Nguyen, Anh Nguyen 0003, Zhigang Deng 0001, Carol C. Wu, T. Hoang Ngan Le |
ICCV | 6 |
| 2024 | Amodal Instance Segmentation with Diffusion Shape Prior Estimation
Viet-Khoa Vo-Ho, Tri Nguyen 0005, T. Hoang Ngan Le |
ACCV (10) | 2 |
| 2024 | Open-Fusion: Real-time Open-Vocabulary 3D Mapping and Queryable Scene RepresentationabstractPrecise 3D environmental mapping with semantics is essential in robotics. Existing methods often rely on pre-defined concepts during training or are time-intensive when generating semantic maps. This paper presents Open-Fusion, an approach for real-time open-vocabulary 3D mapping and queryable scene representation using RGB-D data. Open-Fusion harnesses the power of a pretrained vision-language foundation model (VLFM) for open-set semantic comprehension and employs the Truncated Signed Distance Function (TSDF) for swift 3D scene reconstruction. By leveraging the VLFM, we extract region-based embeddings and their associated confidence maps. These are then integrated with the 3D knowledge from TSDF using an enhanced Hungarian-based feature-matching mechanism. In particular, Open-Fusion delivers outstanding annotation-free 3D segmentation for open vocabulary query without the need for additional 3D training. Benchmark tests on the ScanNet dataset against leading zero-shot methods highlight Open-Fusion’s superiority. Furthermore, it seamlessly combines the strengths of region-based VLFM and TSDF, facilitating real-time 3D scene comprehension that includes object concepts and open-world semantics. We encourage the readers to view the demos on our project page: https://uark-aicv.github.io/OpenFusion Kashu Yamazaki, Taisei Hanyu, Viet-Khoa Vo-Ho, Thang Pham, Gianfranco Doretto, Anh Nguyen 0003, T. Hoang Ngan Le |
ICRA | 3 |
| 2024 | ShapeFormer: Shape Prior Visible-to-Amodal Transformer-based Amodal Instance SegmentationabstractAmodal Instance Segmentation (AIS) presents a challenging task as it involves predicting both visible and occluded parts of objects within images. Existing AIS methods rely on a bidirectional approach, encompassing both the transition from amodal features to visible features (amodal-to-visible) and from visible features to amodal features (visible-to-amodal). Our observation shows that the utilization of amodal features through the amodal-to-visible can confuse the visible features due to the extra information of occluded/hidden segments not presented in visible display. Consequently, this compromised quality of visible features during the subsequent visible-to-amodal transition. To tackle this issue, we introduce ShapeFormer, a decoupled Transformer-based model with a visible-to-amodal transition. It facilitates the explicit relationship between output segmentations and avoids the need for amodal-to-visible transitions. ShapeFormer comprises three key modules: (i) Visible-Occluding Mask Head for predicting visible segmentation with occlusion awareness, (ii) Shape-Prior Amodal Mask Head for predicting amodal and occluded masks, and (iii) Category-Specific Shape Prior Retriever aims to provide shape prior knowledge. Comprehensive experiments and extensive ablation studies across various AIS benchmarks demonstrate the effectiveness of our ShapeFormer. The code is available at: https: //github.com/UARK-AICV/ShapeFormer Winston Bounsavy, Viet-Khoa Vo-Ho, Anh Nguyen 0003, Tri Nguyen 0005, T. Hoang Ngan Le |
IJCNN | 3 |
| 2024 | HENASY: Learning to Assemble Scene-Entities for Interpretable Egocentric Video-Language ModelabstractCurrent video-language models (VLMs) rely extensively on instance-level alignment between video and language modalities, which presents two major limitations: (1) visual reasoning disobeys the natural perception that humans do in first-person perspective, leading to a lack of reasoning interpretation; and (2) learning is limited in capturing inherent fine-grained relationships between two modalities.
In this paper, we take an inspiration from human perception and explore a compositional approach for egocentric video representation. We introduce HENASY (Hierarchical ENtities ASsemblY), which includes a spatiotemporal token grouping mechanism to explicitly assemble dynamically evolving scene entities through time and model their relationship for video representation. By leveraging compositional structure understanding, HENASY possesses strong interpretability via visual grounding with free-form text queries. We further explore a suite of multi-grained contrastive losses to facilitate entity-centric understandings. This comprises three alignment types: video-narration, noun-entity, verb-entities alignments.
Our method demonstrates strong interpretability in both quantitative and qualitative experiments; while maintaining competitive performances on five downstream tasks via zero-shot transfer or as video/text representation, including video/text retrieval, action recognition, multi-choice query, natural language query, and moments query.
Project page: https://uark-aicv.github.io/HENASY Viet-Khoa Vo-Ho, Thinh Phan, Kashu Yamazaki, T. Hoang Ngan Le |
NeurIPS | 1 |
| 2024 | ZEETAD: Adapting Pretrained Vision-Language Model for Zero-Shot End-to-End Temporal Action DetectionabstractTemporal action detection (TAD) involves the localization and classification of action instances within untrimmed videos. While standard TAD follows fully supervised learning with closed-set setting on large training data, recent zero-shot TAD methods showcase the promising open-set setting by leveraging large-scale contrastive visual-language (ViL) pretrained models. However, existing zero-shot TAD methods have limitations on how to properly construct the strong relationship between two interdependent tasks of localization and classification and adapt ViL model to video understanding. In this work, we present ZEE-TAD, featuring two modules: dual-localization and zero-shot proposal classification. The former is a Transformer-based module that detects action events while selectively collecting crucial semantic embeddings for later recognition. The latter one, CLIP-based module, generates semantic embeddings from text and frame inputs for each temporal unit. Additionally, we enhance discriminative capability on unseen classes by minimally updating the frozen CLIP encoder with lightweight adapters. Extensive experiments on THUMOS14 and ActivityNet-1.3 datasets demonstrate our approach’s superior performance in zero-shot TAD and effective knowledge transfer from ViL models to unseen action categories. Code is available at https: //github.com/UARK-AICV/ZEETAD. Thinh Phan, Viet-Khoa Vo-Ho, Duy Le 0004, Gianfranco Doretto, Donald A. Adjeroh, T. Hoang Ngan Le |
WACV | 2 |
| 2023 | VLTinT: Visual-Linguistic Transformer-in-Transformer for Coherent Video Paragraph CaptioningabstractVideo Paragraph Captioning aims to generate a multi-sentence description of an untrimmed video with multiple temporal event locations in a coherent storytelling. Following the human perception process, where the scene is effectively understood by decomposing it into visual (e.g. human, animal) and non-visual components (e.g. action, relations) under the mutual influence of vision and language, we first propose a visual-linguistic (VL) feature. In the proposed VL feature, the scene is modeled by three modalities including (i) a global visual environment; (ii) local visual main agents; (iii) linguistic scene elements. We then introduce an autoregressive Transformer-in-Transformer (TinT) to simultaneously capture the semantic coherence of intra- and inter-event contents within a video. Finally, we present a new VL contrastive loss function to guarantee the learnt embedding features are consistent with the captions semantics. Comprehensive experiments and extensive ablation studies on the ActivityNet Captions and YouCookII datasets show that the proposed Visual-Linguistic Transformer-in-Transform (VLTinT) outperforms previous state-of-the-art methods in terms of accuracy and diversity. The source code is made publicly available at: https://github.com/UARK-AICV/VLTinT. Kashu Yamazaki, Viet-Khoa Vo-Ho, Quang Sang Truong, Bhiksha Raj, T. Hoang Ngan Le |
AAAI | 2 |
| 2023 | CLIP-TSA: Clip-Assisted Temporal Self-Attention for Weakly-Supervised Video Anomaly DetectionabstractVideo anomaly detection (VAD) – commonly formulated as a multiple-instance learning problem in a weakly-supervised manner due to its labor-intensive nature – is a challenging problem in video surveillance where the frames of anomaly need to be localized in an untrimmed video. In this paper, we first propose to utilize the ViT-encoded visual features from CLIP, in contrast with the conventional C3D or I3D features in the domain, to efficiently extract discriminative representations in the novel technique. We then model temporal dependencies and nominate the snippets of interest by leveraging our proposed Temporal Self-Attention (TSA). The ablation study confirms the effectiveness of TSA and ViT feature. The extensive experiments show that our proposed CLIP-TSA outperforms the existing state-of-the-art (SOTA) methods by a large margin on three commonly-used benchmark datasets in the VAD problem (UCF-Crime, ShanghaiTech Campus and XD-Violence). Our source code is available at https://github.com/joos2010kj/CLIP-TSA. Hyekang Joo, Viet-Khoa Vo-Ho, Kashu Yamazaki, T. Hoang Ngan Le |
ICIP | 2 |
| 2023 | AOE-Net: Entities Interactions Modeling with Adaptive Attention Mechanism for Temporal Action Proposals Generation
Viet-Khoa Vo-Ho, Sang Truong, Kashu Yamazaki, Bhiksha Raj, Minh-Triet Tran, T. Hoang Ngan Le |
Int. J. Comput. Vis. | 1 |
| 2022 | AISFormer: Amodal Instance Segmentation with Transformer
Minh Q. Tran, Viet-Khoa Vo-Ho, Kashu Yamazaki, Arthur A. F. Fernandes, Michael Kidd, T. Hoang Ngan Le |
BMVC | 2 |
| 2022 | VLCAP: Vision-Language with Contrastive Learning for Coherent Video Paragraph CaptioningabstractIn this paper, we leverage the human perceiving process, that involves vision and language interaction, to generate a coherent paragraph description of untrimmed videos. We propose vision-language (VL) features consisting of two modalities, i.e., (i) vision modality to capture global visual content of the entire scene and (ii) language modality to extract scene elements description of both human and non-human objects (e.g. animals, vehicles, etc), visual and non-visual elements (e.g. relations, activities, etc). Furthermore, we propose to train our proposed VLCap under a contrastive learning VL loss. The experiments and ablation studies on ActivityNet Captions and YouCookII datasets show that our VLCap outperforms existing SOTA methods on both accuracy and diversity metrics. Source code: https://github.com/UARK-AICV/VLCAP Kashu Yamazaki, Sang Truong, Viet-Khoa Vo-Ho, Michael Kidd, Chase Rainwater, Khoa Luu, T. Hoang Ngan Le |
ICIP | 3 |
| 2022 | 3DConvCaps: 3DUnet with Convolutional Capsule Encoder for Medical Image SegmentationabstractConvolutional Neural Networks (CNNs) have achieved promising results in medical image segmentation. However, CNNs require lots of training data and are incapable of handling pose and deformation of objects. Furthermore, their pooling layers tend to discard important information such as positions as well as CNNs are sensitive to rotation and affine transformation. Capsule network is a recent new architecture that has achieved better robustness in part-whole representation learning by replacing pooling layers with dynamic routing and convolutional strides, which has shown potential results on popular tasks such as digit classification and object segmentation. In this paper, we propose a 3D encoder-decoder network with Convolutional Capsule Encoder (called 3DConvCaps) to learn lower-level features (short-range attention) with convolutional layers while modeling the higher-level features (long-range dependence) with capsule layers. Our experiments on multiple datasets including iSeg-2017, Hippocampus, and Cardiac demonstrate that our 3D 3DConvCaps network considerably outperforms previous capsule networks and 3D-UNets. We further conduct ablation studies of network efficiency and segmentation performance under various configurations of convolution layers and capsule layers at both contracting and expanding paths. The implementation is available at: https://github.com/UARK-AICV/3DConvCaps Minh Q. Tran, Viet-Khoa Vo-Ho, T. Hoang Ngan Le |
ICPR | 2 |
| 2021 | AEI: Actors-Environment Interaction with Adaptive Attention for Temporal Action Proposals Generation
Viet-Khoa Vo-Ho, Hyekang Joo, Kashu Yamazaki, Sang Truong, Kris Makoto Kitani, Minh-Triet Tran, T. Hoang Ngan Le |
BMVC | 1 |
| 2021 | Offboard 3D Object Detection From Point Cloud SequencesabstractWhile current 3D object recognition research mostly focuses on the real-time, onboard scenario, there are many offboard use cases of perception that are largely underexplored, such as using machines to automatically generate high-quality 3D labels. Existing 3D object detectors fail to satisfy the high-quality requirement for offboard uses due to the limited input and speed constraints. In this paper, we propose a novel offboard 3D object detection pipeline using point cloud sequence data. Observing that different frames capture complementary views of objects, we design the offboard detector to make use of the temporal points through both multi-frame object detection and novel objectcentric refinement models. Evaluated on the Waymo Open Dataset, our pipeline named 3D Auto Labeling shows significant gains compared to the state-of-the-art onboard detectors and our offboard baselines. Its performance is even on par with human labels verified through a human label study. Further experiments demonstrate the application of auto labels for semi-supervised learning and provide extensive analysis to validate various design choices. Charles R. Qi, Mahyar Najibi, Viet-Khoa Vo-Ho, Boyang Deng, Dragomir Anguelov |
CVPR | 5 |
| 2021 | Agent-Environment Network for Temporal Action Proposal GenerationabstractTemporal action proposal generation is an essential and challenging task that aims at localizing temporal intervals containing human actions in untrimmed videos. Most of existing approaches are unable to follow the human cognitive process of understanding the video context due to lack of attention mechanism to express the concept of an action or an agent who performs the action or the interaction between the agent and the environment. Based on the action definition that a human, known as an agent, interacts with the environment and performs an action that affects the environment, we propose a contextual Agent-Environment Network. Our proposed contextual AEN involves (i) agent pathway, operating at a local level to tell about which humans/agents are acting and (ii) environment pathway operating at a global level to tell about how the agents interact with the environment. Comprehensive evaluations on 20-action THUMOS-14 and 200-action ActivityNet-1.3 datasets with different backbone networks, i.e C3D and SlowFast, show that our method robustly exhibits outperformance against state-of-the-art methods regardless of the employed backbone network. Viet-Khoa Vo-Ho, T. Hoang Ngan Le, Kashu Yamazaki, Akihiro Sugimoto, Minh-Triet Tran |
ICASSP | 1 |