Quang-Linh Tran

dblp:348/4770 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
10since 2021 · last 2025
0000-0002-5409-0916ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Multi-modal Context Reranking for Lifelog Question Answering
abstract
Lifelog question-answering (QA) involves seeking answers to users' questions from within their personal lifelog. A lifelog consists of passively collected multimodal personal information from the owner's life experiences, including images, biometrics, geolocation data, and textual descriptions. Data in a lifelog can become vast, spanning the user's lifetime. QA for such large collections presents significant challenges: finding the most relevant lifelog events (contexts) that may contain the answer before generating a response. We propose a reranking model designed to improve retrieval accuracy by effectively ranking the relevant contexts. Our model integrates multimodal information from images and text, employing a combination of visual extractors and language models. We experiment with three visual extractor models, Vision Transformer, BLIP2, and CLIP, as well as three language models, namely BERT, MiniLM, and ModernBERT. Compared with a retrieval baseline using cosine similarity ranking from Stella-1.5B embeddings, our experimental results on two lifelog QA datasets demonstrate a substantial improvement. Recall@1 is observed to increase from 37.65 % to 65.57% and Precision@1 from 52.65% to 85.88% when using the ModernBERT +ViT reranking model for the OpenLifelogQA task. These findings show the robustness of multimodal reranking in context selection for lifelog QA and provide a mechanism for accurate and efficient retrieval in lifelog applications.
Quang-Linh Tran, Ly-Duyen Tran, Binh T. Nguyen 0001, Gareth J. F. Jones, Cathal Gurrin
CBMI1
2025 A RAG Approach for Multi-Modal Open-ended Lifelog Question-Answering
abstract
Lifelogging is the passive collection, storage and analysis of daily data through wearable sensors. Question Answering (QA) for lifelog data enables natural language interactions with personal daily life records, providing insights into individual routines and behaviours. While this task has great potential for personal analytics and memory augmentation, progress has been limited due to the challenges of lifelog management, since they can comprise of enormous multi-modal data sets spanning a lifetime. We introduce a Retrieval-Augmented Generation (RAG) approach for addressing the lifelog QA task. A RAG approach first includes a retrieval model finding the correct lifelog events containing answers and then a large language model (LLM) generating answers from the questions. In addition, we construct an open-ended lifelog QA benchmark with 14,187 QA pairs to examine the RAG approach to lifelog QA. Using an embedding-based retrieval approach, our lifelog context retriever achieves a performance of 77.67% Recall@5 and 94.35% Recall@20 using an embedding-based retrieval approach with the Stella 1.5B model. Combined with the Mistral 7B model, the model achieves scores of 39.54% ROUGE-L and 3.475 Accuracy on a scale of 5 scored by GPT-4o. This approach potentially provides an effective approach to lifelog QA with high performance that does not require fine-tuning.
Quang-Linh Tran, Ngo Ngoc Diep Pham, Quoc Trung Truong, Minh Hung Nguyen, Hong Cat Le, Dang Khoi Vu, Van Minh Thien Nguyen, Van Kinh Nguyen, Luu Phuong Ngoc Lam Nguyen, Tan Le, Minh Phuc Dang, Binh T. Nguyen 0001, Gareth J. F. Jones, Cathal Gurrin
ICMR1
2025 AIQAM'25: The 2nd ACM Workshop on AI-powered Question Answering Systems for Multimedia
abstract
The Second ACM Workshop on AI-Powered Question & Answering Systems for Multimedia (AIQAM'25) was held on 27 October 2025 in Dublin, Ireland, co-located with ACM Multimedia 2025. The workshop's main objective is to create a collaborative and inclusive space for researchers at the intersection of Artificial Inteligence (AI), large language models (LLMs), multimodal information retrieval, and question answering systems. Building on the success of its first edition (AIQAM'24) at ICMR 2024, AIQAM'25 provided a forum for presenting novel methods, applications, and surveys, which address the challenges of integrating text, image, audio, and video data into QA systems. The programme featured contributions ranging from methodological advances in reasoning and evaluation frameworks, domain-specific applications in education and sustainability, to a survey work reviewing the state of multimedia retrieval-augmented QA. This summary paper outlines the objectives and scope of the workshop, describes its format, highlights the keynote and accepted contributions, and acknowledges the efforts of the organising committee.
Tai Tan Mai, Ly-Duyen Tran, Quang-Linh Tran, Duc-Tien Dang-Nguyen, Cathal Gurrin
ACM Multimedia3
2025 The CASTLE 2024 Dataset: Advancing the Art of Multimodal Understanding
abstract
Egocentric video has seen increased interest in recent years, as it is used in a range of areas. However, most existing datasets are limited to a single perspective. In this paper, we present the CASTLE 2024 dataset, a multimodal collection containing ego- and exo-centric (i.e., first- and third-person perspective) video and audio from 15 time-aligned sources, as well as other sensor streams and auxiliary data. The dataset was recorded by volunteer participants over four days in a common location and includes the point of view of 10 participants, with an additional 5 fixed cameras providing an exocentric perspective. The entire dataset contains over 600 hours of UHD video recorded at 50 frames per second. In contrast to other datasets, CASTLE 2024 does not contain any partial censoring, such as blurred faces or distorted audio. The dataset is available via https://castle-dataset.github.io/.
Luca Rossetto, Werner Bailer, Duc-Tien Dang-Nguyen, Graham Healy, Björn Þór Jónsson 0001, Onanong Kongmeesub, Hoang Bao Le, Stevan Rudinac, Klaus Schöffmann, Florian Spiess 0001, Ly-Duyen Tran, Minh-Triet Tran, Quang-Linh Tran, Cathal Gurrin
ACM Multimedia13
2025 Extending Lifelog Retrieval to Multi-stream Video Retrieval at the CASTLE Challenge 2025
abstract
We present the DCU team's system for the CASTLE Challenge at ACM Multimedia 2025, which explores video retrieval and question answering in egocentric, multi-user environments. Our system adapts techniques developed for lifelogging, particularly event-based semantic retrieval and QA pipelines, to the CASTLE dataset with minimal architectural changes. It combines vision-language embeddings, transcript-based retrieval, and person tracking to support both automatic and interactive search workflows. In the interactive track, we introduce a modular interface for narrative reconstruction and exploratory search. Qualitative results show that the system can generate plausible, evidence-based answers to complex multimodal queries. These findings suggest that lifelog retrieval systems offer a viable foundation for broader egocentric video analysis.
Quang-Linh Tran, Hoang Bao Le, Thang-Long Nguyen-Ho, Graham Healy, Liting Zhou, Ly-Duyen Tran
ACM Multimedia1
2025 VideoEase at VBS2025: An Interactive Video Retrieval System
abstract
We present the VideoEase interactive video retrieval system, which we used to participate in VBS2025. This is the first time that VideoEase has taken part in the VBS challenge. VideoEase is built on the Milvus vector database, which supports vector searches on massive datasets with high-dimensional vectors. The CLIP, BLIP2, and OpenCLIP models play a crucial role in encoding keyframe images and queries into embeddings. A new user interface with simple yet effective components is also introduced. We experiment with VBS24 queries to evaluate the performance of the VideoEase system. Our results show that answer rank is improved by merging outputs from several models with appropriate weights.
Quang-Linh Tran, Binh T. Nguyen 0001, Gareth J. F. Jones, Cathal Gurrin
MMM (5)1
2024 VidBasys: A User-Friendly Interactive Video Retrieval System for Novice Users in IVR4B
abstract
In this paper, we present the VidBasys interactive video retrieval system for novice users, an upgraded version of VideoCLIP 2.0 that participated in the Video Browser Showdown 2024. While the novel user interface is designed in a more user-friendly way for newbies to accommodate the target of the Interactive Video Retrieval for Beginner (IVR4B), the core search engine is enhanced with the advance of the recent CLIP model to bridge the gap in semantics between image and text. This version is designed to focus on novice users with a simple, easy-to-use but effective user interface. The system supports freetext search to enhance the user experience and minimise the number of actions required for filtering. The new user interface supports simple search and filters with clearly designed freetext search boxes. In addition, the retrieved results are displayed in an optimised layout to maximise image display space and minimise user interactions. The improvements are expected to support novice users in accurately retrieving the desired videos.
Thao-Nhu Nguyen, Quang-Linh Tran, Hoang Bao Le, Binh T. Nguyen 0001, Liting Zhou, Gareth J. F. Jones, Cathal Gurrin
CBMI2
2024 The First ACM Workshop on AI-Powered Question Answering Systems for Multimedia
abstract
The advent of large language models (LLMs) has energised research in Question-Answering (QA) tasks, enabling responses across varied domains like economics and mathematics. Despite their capabilities, LLMs often lack explainability due to their complex parameter embeddings. Additionally, integrating multimedia data into QA systems introduces challenges in processing and interpreting diverse data types such as text, images, audio, and video. This necessitates sophisticated algorithms for accurate information retrieval across media while ensuring the reliability of the data and responses remains a significant challenge. The AIQAM workshop aims to bring together researchers and practitioners to address these challenges and enhance QA systems with multimedia data. The focus is on promoting innovations that improve the accuracy, explainability, and trustworthiness of QA systems, contributing to the development of the field.
Tai Tan Mai, Quang-Linh Tran, Ly-Duyen Tran, Van-Tu Ninh, Duc-Tien Dang-Nguyen, Cathal Gurrin
ICMR2
2024 MemoriLens: a Low-cost Lifelog Camera Using Raspberry Pi Zero
abstract
Lifelogging is the process of automatically logging data about an individual's daily life, which can then be used in various domains, such as behavior analysis and health monitoring. Various technological devices, including wearable cameras and smartwatches, can help record lifelog data, but getting access to lifelog cameras has proven difficult in recent years, due to a lack of such devices on the market. Creating a lifelog camera that is not only easy to use and cost-efficient but also provides comprehensive functions to log all images about life is challenging due to the lack of hardware and software. This paper introduces MemoriLens, a low-cost camera that efficiently collects, organizes, and stores lifelog data using a readily available custom-designed Raspberry Pi Zero board. The camera is designed to capture images automatically and send them to a private account in cloud services for storage. We open-source the implementing of the camera at: https://github.com/linh222/raspberry_lifelog_camera and we encourage lifelog researchers to use our designs and software as required.
Quang-Linh Tran, Binh T. Nguyen 0001, Gareth J. F. Jones, Cathal Gurrin
ICMR1
2022 Aspect-based Sentiment Analysis for Vietnamese Reviews about Beauty Product on E-commerce Websites
Quang-Linh Tran, Phan Thanh Dat Le, Trong-Hop Do
PACLIC1