VLDB 2026 Research / reviewers in the wild / expert
Hoang Bao Le
dblp:86/888 · also Hoang-Bao Le
· DBLP profile ↗
7ranked-venue papers
3as first author
6since 2021 · last 2026
0009-0000-2496-4347ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FIGROTD: A Friendly-to-Handle Dataset for Image Guided Retrieval with Optional Text
Hoang Bao Le, Ly-Duyen Tran, Binh T. Nguyen 0001, Liting Zhou, Cathal Gurrin |
MMM (1) | 1 |
| 2025 | Vision Projector: Improving Zero-Shot Composed Image Retrieval at InferenceabstractComposed Image Retrieval (CIR) involves retrieving a target image based on a query composed of a reference image and a textual modification. Zero-Shot CIR extends this task by removing the need for labeled triplets during training. Most state-of-the-art (SOTA) methods share a common structure: a vision-language encoder followed by a matching module using Transformers or contrastive learning. Instead of increasing data or model complexity, we wonder that: Can we improve retrieval performance at inference time? To answer this, we propose the Vision Projector (VP)-a lightweight, plug-and-play module that enhances visual representations without retraining. Integrated directly into MagicLens, VP consistently improves performance across CIRR, FashionIQ, and CIRCO. Notably, it boosts MagicLens by 18% on CIRCO, despite not using its strongest variant. Code is available at: https://github.com/baohl00/VisionProjector_ZSCIR. Hoang Bao Le, Ly-Duyen Tran, Binh T. Nguyen 0001, Liting Zhou, Cathal Gurrin |
CBMI | 1 |
| 2025 | The CASTLE 2024 Dataset: Advancing the Art of Multimodal UnderstandingabstractEgocentric video has seen increased interest in recent years, as it is used in a range of areas. However, most existing datasets are limited to a single perspective. In this paper, we present the CASTLE 2024 dataset, a multimodal collection containing ego- and exo-centric (i.e., first- and third-person perspective) video and audio from 15 time-aligned sources, as well as other sensor streams and auxiliary data. The dataset was recorded by volunteer participants over four days in a common location and includes the point of view of 10 participants, with an additional 5 fixed cameras providing an exocentric perspective. The entire dataset contains over 600 hours of UHD video recorded at 50 frames per second. In contrast to other datasets, CASTLE 2024 does not contain any partial censoring, such as blurred faces or distorted audio. The dataset is available via https://castle-dataset.github.io/. Luca Rossetto, Werner Bailer, Duc-Tien Dang-Nguyen, Graham Healy, Björn Þór Jónsson 0001, Onanong Kongmeesub, Hoang Bao Le, Stevan Rudinac, Klaus Schöffmann, Florian Spiess 0001, Ly-Duyen Tran, Minh-Triet Tran, Quang-Linh Tran, Cathal Gurrin |
ACM Multimedia | 7 |
| 2025 | Extending Lifelog Retrieval to Multi-stream Video Retrieval at the CASTLE Challenge 2025abstractWe present the DCU team's system for the CASTLE Challenge at ACM Multimedia 2025, which explores video retrieval and question answering in egocentric, multi-user environments. Our system adapts techniques developed for lifelogging, particularly event-based semantic retrieval and QA pipelines, to the CASTLE dataset with minimal architectural changes. It combines vision-language embeddings, transcript-based retrieval, and person tracking to support both automatic and interactive search workflows. In the interactive track, we introduce a modular interface for narrative reconstruction and exploratory search. Qualitative results show that the system can generate plausible, evidence-based answers to complex multimodal queries. These findings suggest that lifelog retrieval systems offer a viable foundation for broader egocentric video analysis. Quang-Linh Tran, Hoang Bao Le, Thang-Long Nguyen-Ho, Graham Healy, Liting Zhou, Ly-Duyen Tran |
ACM Multimedia | 2 |
| 2025 | ExGra-Med: Extended Context Graph Alignment for Medical Vision-Language ModelsabstractState-of-the-art medical multi-modal LLMs (med-MLLMs), such as LLaVA-Med and BioMedGPT, primarily depend on scaling model size and data volume, with training driven largely by autoregressive objectives. However, we reveal that this approach can lead to weak vision-language alignment, making these models overly dependent on costly instruction-following data. To address this, we introduce ExGra-Med, a novel multi-graph alignment framework that jointly aligns images, instruction responses, and extended captions in the latent space, advancing semantic grounding and cross-modal coherence. To scale to large LLMs (e.g., LLaMa-7B), we develop an efficient end-to-end training scheme using black-box gradient estimation, enabling fast and scalable optimization. Empirically, ExGra-Med matches LLaVA-Med’s performance using just 10\% of pre-training data, achieving a 20.13\% gain on VQA-RAD and approaching full-data performance. It also outperforms strong baselines like BioMedGPT and RadFM on visual chatbot and zero-shot classification tasks, demonstrating its promise for efficient, high-quality vision-language integration in medical AI. Duy M. H. Nguyen, Nghiem Tuong Diep, Hoang Bao Le, Tai D. Nguyen, Anh-Tien Nguyen, TrungTin Nguyen, Nhat Ho, Pengtao Xie, Roger Wattenhofer, Daniel Sonntag, James Zou 0001, Mathias Niepert |
NeurIPS | 4 |
| 2024 | VidBasys: A User-Friendly Interactive Video Retrieval System for Novice Users in IVR4BabstractIn this paper, we present the VidBasys interactive video retrieval system for novice users, an upgraded version of VideoCLIP 2.0 that participated in the Video Browser Showdown 2024. While the novel user interface is designed in a more user-friendly way for newbies to accommodate the target of the Interactive Video Retrieval for Beginner (IVR4B), the core search engine is enhanced with the advance of the recent CLIP model to bridge the gap in semantics between image and text. This version is designed to focus on novice users with a simple, easy-to-use but effective user interface. The system supports freetext search to enhance the user experience and minimise the number of actions required for filtering. The new user interface supports simple search and filters with clearly designed freetext search boxes. In addition, the retrieved results are displayed in an optimised layout to maximise image display space and minimise user interactions. The improvements are expected to support novice users in accurately retrieving the desired videos. Thao-Nhu Nguyen, Quang-Linh Tran, Hoang Bao Le, Binh T. Nguyen 0001, Liting Zhou, Gareth J. F. Jones, Cathal Gurrin |
CBMI | 3 |
| 2008 | Optimal control of first order linear systems with fixed proportional-integral structure controllerabstractThis paper proposes a new semi-analytic robust mixed H2/H-infinity design method for fixed structure controllers (i.e. PID, the most widely used structure in industry). Precisely, the method consists in determining the parameters of a given structure controller that minimizes the influence of a step load disturbance to the process output with the respect of robustness constraints, i.e. constraints on maximum amplification of measurement noise, minimum module margin and minimum phase margin. The design objective and the robustness constraints are expressed as H2 and H-infinity norms in function of unknown controller parameters. The controller design problem is then reformulated into a nonlinear optimization problem with a set of inequality constraints that can be efficiently solved numerically. Finally, we obtain a controller design tool which provides, when it exists, the unique optimal controller that fulfills the design specifications. The method is based on generic models that can represent common industrial plants. Further, the proposed method enables a graphical representation of the different design tradeoffs. To demonstrate the results, we apply this method for first order processes controlled by PI controller. Hoang Bao Le, Eduardo Mendes |
ICARCV | 1 |