Thao-Nhu Nguyen

dblp:283/7047 · DBLP profile ↗
← Back
9ranked-venue papers
7as first author
9since 2021 · last 2024
0000-0003-1356-9434ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 7 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 VidBasys: A User-Friendly Interactive Video Retrieval System for Novice Users in IVR4B
abstract
In this paper, we present the VidBasys interactive video retrieval system for novice users, an upgraded version of VideoCLIP 2.0 that participated in the Video Browser Showdown 2024. While the novel user interface is designed in a more user-friendly way for newbies to accommodate the target of the Interactive Video Retrieval for Beginner (IVR4B), the core search engine is enhanced with the advance of the recent CLIP model to bridge the gap in semantics between image and text. This version is designed to focus on novice users with a simple, easy-to-use but effective user interface. The system supports freetext search to enhance the user experience and minimise the number of actions required for filtering. The new user interface supports simple search and filters with clearly designed freetext search boxes. In addition, the retrieved results are displayed in an optimised layout to maximise image display space and minimise user interactions. The improvements are expected to support novice users in accurately retrieving the desired videos.
Thao-Nhu Nguyen, Quang-Linh Tran, Hoang Bao Le, Binh T. Nguyen 0001, Liting Zhou, Gareth J. F. Jones, Cathal Gurrin
CBMI1
2024 A Parallel Transformer Framework for Video Moment Retrieval
abstract
In the realm of video understanding, Video Moment Retrieval (VMR) is an important yet challenging task that aims to locate the boundary of a moment of interest within a long untrimmed video. Existing VMR methods often focus on the visual content extracted from the video only (or frame sequences), however, the rich semantic information at the object level that describes the image's content has not been explored yet. To overcome those limitations, we propose PaTF, an attention-based Parallel Transformer Framework that enriches the feature representations by exploring both low-level visual cues and high-level relational contexts of video-query pairs. Our framework consists of two parallel transformers: one for the visual-textual stream and the other for the semantic-textual stream. The visual-textual stream extracts the links between global visual features and textual information, while the semantic-textual stream emphasises the relations between objects via scene graph representations. Furthermore, our comprehensive experiment conducted on the Charades-STA dataset demonstrates that the proposed framework outperforms the state-of-the-art methods by a large margin, roughly 5% and 7% at Recall@1 with IoU = 0.5 and IoU = 0.7, respectively.
Thao-Nhu Nguyen, Zongyao Li 0004, Satoshi Yamazaki, Jianquan Liu, Cathal Gurrin
ICMR1
2024 VideoCLIP 2.0: An Interactive CLIP-Based Video Retrieval System for Novice Users at VBS2024
Thao-Nhu Nguyen, Le Minh Quang, Graham Healy, Binh T. Nguyen 0001, Cathal Gurrin
MMM (4)1
2023 Efficient Search with an Interactive Video Retrieval System for Novice Users in IVR4B
abstract
In this paper, we present the second release of VideoCLIP, an interactive CLIP-based video retrieval system that participated in the Video Browser Showdown 2023. While we continue to use the underlying architecture to map the content between image and text, we concentrate on improving the user experience for novice users. Specifically, we have implemented three different query modalities and redesigned the user interface in order to adapt to the context of the Interactive Video Retrieval for Beginners (IVR4B) workshop. These modifications ultimately aim to provide newcomers with a simple and efficient user experience to locate the desired videos.
Thao-Nhu Nguyen, Bunyarit Puangthamawathanakun, Chonlameth Arpnikanondt, Cathal Gurrin, Annalina Caputo, Graham Healy
CBMI1
2023 VideoCLIP: An Interactive CLIP-based Video Retrieval System at VBS2023
Thao-Nhu Nguyen, Bunyarit Puangthamawathanakun, Annalina Caputo, Graham Healy, Binh T. Nguyen 0001, Chonlameth Arpnikanondt, Cathal Gurrin
MMM (1)1
2023 Interactive video retrieval in the age of effective joint embedding deep models: lessons from the 11th VBS
Jakub Lokoc, Stelios Andreadis, Werner Bailer, Aaron Duane, Cathal Gurrin, Zhixin Ma 0001, Nicola Messina, Thao-Nhu Nguyen, Ladislav Peska, Luca Rossetto, Loris Sauter, Konstantin Schall, Klaus Schöffmann, Omar Shahbaz Khan, Florian Spiess 0001, Lucia Vadicamo, Stefanos Vrochidis
Multim. Syst.8
2023 LifeSeeker: an interactive concept-based retrieval system for lifelog data
abstract
Lifelogging was introduced as the process of passively capturing personal daily events via wearable devices. It ultimately creates a visual diary encoding every aspect of one's life with the aim of future sharing or recollecting. In this paper, we present LifeSeeker, a lifelog image retrieval system participating in the Lifelog Search Challenge (LSC) for 3 years, since 2019. Our objective is to support users to seek specific life moments using a combination of textual descriptions, spatial relationships, location information, and image similarities. In addition to the LSC challenge results, a further experiment was conducted in order to evaluate the power retrieval of our system on both expert and novice users. This experiment informed us about the effectiveness of the user's interaction with the system when involving non-experts.
Thao-Nhu Nguyen, Tu-Khiem Le, Van-Tu Ninh, Annalina Caputo, Graham Healy, Sinéad Smyth, Minh-Triet Tran, Binh T. Nguyen 0001
Multim. Tools Appl.1
2022 Videofall - A Hierarchical Search Engine for VBS2022
Thao-Nhu Nguyen, Bunyarit Puangthamawathanakun, Graham Healy, Binh T. Nguyen 0001, Cathal Gurrin, Annalina Caputo
MMM (2)1
2021 A VR Interface for Browsing Visual Spaces at VBS2021
Ly-Duyen Tran, Duy Nguyen 0003, Thao-Nhu Nguyen, Graham Healy, Annalina Caputo, Binh T. Nguyen 0001, Cathal Gurrin
MMM (2)3