Phuong Anh Nguyen 0002

dblp:91/9506-2 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0003-1289-3785ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2026 A Pruning-based Question-Answering for Interactive Video Search: A Simple Baseline
abstract
There are various factors affecting the performance of video search. An imprecise query will enlarge search space and reduce the discriminative power of ranking functions. This problem is further exacerbated by the presence of numerous visually or semantically similar videos in large datasets. Consequently, users need to painstakingly browse through many highly similar candidates to locate the search target, leading to increased cognitive load and inefficient searching. Ideally, engaging users through interactive questioning to resolve uncertainties in the search process is an effective strategy for progressively narrowing down the search space. However, despite rapid advances in deep learning, generating informative questions conditioned on the user query and search result remains a highly difficult problem.
Yu-Tong Cheng, Phuong Anh Nguyen 0002, Chong-Wah Ngo
ICMR2
2021 Chinese White Dolphin Detection in the Wild
abstract
For ecological protection of the ocean, biologists usually conduct line-transect vessel surveys to measure sea species’ population density within their habitat (such as dolphins). However, sea species observation via vessel surveys consumes a lot of manpower resources and is more challenging compared to observing common objects, due to the scarcity of the object in the wild, tiny-size of the objects, and similar-sized distracter objects (e.g., floating trash). To reduce the human experts’ workload and improve the observation accuracy, in this paper, we develop a practical system to detect Chinese White Dolphins in the wild automatically. First, we construct a dataset named Dolphin-14k with more than 2.6k dolphin instances. To improve the dataset annotation efficiency caused by the rarity of dolphins, we design an interactive dolphin box annotation strategy to annotate sparse dolphin instances in long videos efficiently. Second, we compare the performance and efficiency of three off-the-shelf object detection algorithms, including Faster-RCNN, FCOS, and YoloV5, on the Dolphin-14k dataset and pick YoloV5 as the detector, where a new category (Distracter) is added to the model training to reject the false positives. Finally, we incorporate the dolphin detector into a system prototype, which detects dolphins in video frames at 100.99 FPS per GPU with high accuracy (i.e., 90.95 [email protected]).
Hao Zhang 0047, Qi Zhang 0041, Phuong Anh Nguyen 0002, Victor C. S. Lee, Antoni B. Chan
MMAsia3
2021 SQL-Like Interpretable Interactive Video Search
Jiaxin Wu 0001, Phuong Anh Nguyen 0002, Zhixin Ma 0001, Chong-Wah Ngo
MMM (2)2
2021 Interactive Video Retrieval in the Age of Deep Learning - Detailed Evaluation of VBS 2019
abstract
Despite the fact that automatic content analysis has made remarkable progress over the last decade - mainly due to significant advances in machine learning - interactive video retrieval is still a very challenging problem, with an increasing relevance in practical applications. The Video Browser Showdown (VBS) is an annual evaluation competition that pushes the limits of interactive video retrieval with state-of-the-art tools, tasks, data, and evaluation metrics. In this paper, we analyse the results and outcome of the 8th iteration of the VBS in detail. We first give an overview of the novel and considerably larger V3C1 dataset and the tasks that were performed during VBS 2019. We then go on to describe the search systems of the six international teams in terms of features and performance. And finally, we perform an in-depth analysis of the per-team success ratio and relate this to the search strategies that were applied, the most popular features, and problems that were experienced. A large part of this analysis was conducted based on logs that were collected during the competition itself. This analysis gives further insights into the typical search behavior and differences between expert and novice users. Our evaluation shows that textual search and content browsing are the most important aspects in terms of logged user interactions. Furthermore, we observe a trend towards deep learning based features, especially in the form of labels generated by artificial neural networks. But nevertheless, for some tasks, very specific content-based search features are still being used. We expect these findings to contribute to future improvements of interactive video search systems.
Luca Rossetto, Ralph Gasser, Jakub Lokoc, Werner Bailer, Klaus Schöffmann, Bernd Münzer, Tomás Soucek, Phuong Anh Nguyen 0002, Paolo Bolettieri, Andreas Leibetseder, Stefanos Vrochidis
IEEE Trans. Multim.8
2021 Interactive Search vs. Automatic Search: An Extensive Study on Video Retrieval
abstract
This article conducts user evaluation to study the performance difference between interactive and automatic search. Particularly, the study aims to provide empirical insights of how the performance landscape of video search changes, with tens of thousands of concept detectors freely available to exploit for query formulation. We compare three types of search modes: free-to-play (i.e., search from scratch), non-free-to-play (i.e., search by inspecting results provided by automatic search), and automatic search including concept-free and concept-based retrieval paradigms. The study involves a total of 40 participants; each performs interactive search over 15 queries of various difficulty levels using two search modes on the IACC.3 dataset provided by TRECVid organizers. The study suggests that the performance of automatic search is still far behind interactive search. Furthermore, providing users with the result of automatic search for exploration does not show obvious advantage over asking users to search from scratch. The study also analyzes user behavior to reveal insights of how users compose queries, browse results, and discover new query terms for search, which can serve as guideline for future research of both interactive and automatic search.
Phuong Anh Nguyen 0002, Chong-Wah Ngo
ACM Trans. Multim. Comput. Commun. Appl.1
2020 VIREO @ Video Browser Showdown 2020
Phuong Anh Nguyen 0002, Jiaxin Wu 0001, Chong-Wah Ngo, Danny Francis, Benoit Huet
MMM (2)1
2019 VIREO @ Video Browser Showdown 2019
Phuong Anh Nguyen 0002, Chong-Wah Ngo, Danny Francis, Benoit Huet
MMM (2)1
2019 Interactive Search or Sequential Browsing? A Detailed Analysis of the Video Browser Showdown 2018
abstract
This work summarizes the findings of the 7th iteration of the Video Browser Showdown (VBS) competition organized as a workshop at the 24th International Conference on Multimedia Modeling in Bangkok. The competition focuses on video retrieval scenarios in which the searched scenes were either previously observed or described by another person (i.e., an example shot is not available). During the event, nine teams competed with their video retrieval tools in providing access to a shared video collection with 600 hours of video content. Evaluation objectives, rules, scoring, tasks, and all participating tools are described in the article. In addition, we provide some insights into how the different teams interacted with their video browsers, which was made possible by a novel interaction logging mechanism introduced for this iteration of the VBS. The results collected at the VBS evaluation server confirm that searching for one particular scene in the collection when given a limited time is still a challenging task for many of the approaches that were showcased during the event. Given only a short textual description, finding the correct scene is even harder. In ad hoc search with multiple relevant scenes, the tools were mostly able to find at least one scene, whereas recall was the issue for many teams. The logs also reveal that even though recent exciting advances in machine learning narrow the classical semantic gap problem, user-centric interfaces are still required to mediate access to specific content. Finally, open challenges and lessons learned are presented for future VBS events.
Jakub Lokoc, Gregor Kovalcík, Bernd Münzer, Klaus Schöffmann, Werner Bailer, Ralph Gasser, Stefanos Vrochidis, Phuong Anh Nguyen 0002, Sitapa Watcharapinchai, Kai Uwe Barthel
ACM Trans. Multim. Comput. Commun. Appl.8
2018 Enhanced VIREO KIS at VBS 2018
Phuong Anh Nguyen 0002, Yi-Jie Lu, Hao Zhang 0047, Chong-Wah Ngo
MMM (2)1
2017 Color-Sketch Simulator: A Guide for Color-Based Visual Known-Item Search
Jakub Lokoc, Phuong Anh Nguyen 0002, Marta Vomlelová, Chong-Wah Ngo
ADMA2
2017 Concept-Based Interactive Search System
Yi-Jie Lu, Phuong Anh Nguyen 0002, Hao Zhang 0047, Chong-Wah Ngo
MMM (2)2
2015 Temporal Matching Kernel with Explicit Feature Maps
abstract
This paper proposes a framework for content-based video retrieval that addresses various tasks as particular event retrieval, copy detection or video synchronization. Given a video query, the method is able to efficiently retrieve, from a large collection, similar video events or near-duplicates with temporarily consistent excerpts. As a byproduct of the representation, it provides a precise temporal alignment of the query and the detected video excerpts.
Sébastien Poullot, Shunsuke Tsukatani, Phuong Anh Nguyen 0002, Hervé Jégou, Shin'ichi Satoh 0001
ACM Multimedia3