Ly-Duyen Tran

dblp:267/2642 · also Allie Tran · DBLP profile ↗
← Back
22ranked-venue papers
10as first author
22since 2021 · last 2026
0000-0002-9597-1832ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 10 first-author · 22 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Keep Your Memories, Lose the Evidence: Privacy-Aware Lifelogging
abstract
The continuous, passive capture of everyday experience, known as Lifelogging, holds significant promise for memory augmentation, health monitoring, and behavioural self-understanding, yet remains far from widespread adoption. A key barrier is architectural: existing systems treat privacy as a post hoc concern, retaining raw sensory data and offering users little meaningful control. We present SelfHealth, a wearable lifelogging prototype that addresses these challenges through a continuous analytics paradigm, in which semantic meaning is extracted at the point of capture and raw data exposure is minimised by design. The system combines real-time anonymisation of faces, licence plates, and sensitive text; privacy-preserving semantic retrieval via user-specific embedding transformations; and customisable behavioural dashboards for tracking binary state, burst actions, and episodic activity. Initial real-world pilot suggests that privacy-aware continuous lifelogging is practically feasible, and surfaces important insights around the trade-off between anonymisation and recall utility, and the role of visible user control in bystander social negotiation.
Ly-Duyen Tran
ICMR1
2026 Introduction to the 9th Annual Lifelog Search Challenge, LSC'26
abstract
The ACM Lifelog Search Challenge (LSC) is an annual comparative benchmarking exercise that brings together researchers in the field of multimedia retrieval to evaluate interactive search systems using a large-scale multimodal lifelog dataset. This paper presents an overview of the ninth edition of the challenge (LSC’26), held as a workshop during the ACM International Conference on Multimedia Retrieval (ICMR ’26) in Amsterdam. To broaden its scope, the workshop now features three submission tracks: the traditional Challenge Track for real-time search performance, a new General Lifelog Research Track for theoretical and architectural advancements, and an additional Open Source Track aimed at enhancing reproducibility and reducing barriers to entry for new participants.
Ly-Duyen Tran, Werner Bailer, Duc-Tien Dang-Nguyen, Graham Healy, Steve Hodges 0001, Björn Þór Jónsson 0001, Wolfgang Hürst, Luca Rossetto, Klaus Schöffmann, Minh-Triet Tran, Liting Zhou, Cathal Gurrin
ICMR1
2026 FIGROTD: A Friendly-to-Handle Dataset for Image Guided Retrieval with Optional Text
Hoang Bao Le, Ly-Duyen Tran, Binh T. Nguyen 0001, Liting Zhou, Cathal Gurrin
MMM (1)2
2026 H-EAGLE: Hierarchical Extension of EAGLE for Multi-level Semantic Video Retrieval
Thang-Long Nguyen-Ho, Viet-Tham Huynh, Ly-Duyen Tran, Minh-Triet Tran, Cathal Gurrin, Graham Healy
MMM (4)3
2026 On the Brittleness of CLIP Text Encoders
Ly-Duyen Tran, Luca Rossetto
MMM (2)1
2025 Vision Projector: Improving Zero-Shot Composed Image Retrieval at Inference
abstract
Composed Image Retrieval (CIR) involves retrieving a target image based on a query composed of a reference image and a textual modification. Zero-Shot CIR extends this task by removing the need for labeled triplets during training. Most state-of-the-art (SOTA) methods share a common structure: a vision-language encoder followed by a matching module using Transformers or contrastive learning. Instead of increasing data or model complexity, we wonder that: Can we improve retrieval performance at inference time? To answer this, we propose the Vision Projector (VP)-a lightweight, plug-and-play module that enhances visual representations without retraining. Integrated directly into MagicLens, VP consistently improves performance across CIRR, FashionIQ, and CIRCO. Notably, it boosts MagicLens by 18% on CIRCO, despite not using its strongest variant. Code is available at: https://github.com/baohl00/VisionProjector_ZSCIR.
Hoang Bao Le, Ly-Duyen Tran, Binh T. Nguyen 0001, Liting Zhou, Cathal Gurrin
CBMI2
2025 Multi-modal Context Reranking for Lifelog Question Answering
abstract
Lifelog question-answering (QA) involves seeking answers to users' questions from within their personal lifelog. A lifelog consists of passively collected multimodal personal information from the owner's life experiences, including images, biometrics, geolocation data, and textual descriptions. Data in a lifelog can become vast, spanning the user's lifetime. QA for such large collections presents significant challenges: finding the most relevant lifelog events (contexts) that may contain the answer before generating a response. We propose a reranking model designed to improve retrieval accuracy by effectively ranking the relevant contexts. Our model integrates multimodal information from images and text, employing a combination of visual extractors and language models. We experiment with three visual extractor models, Vision Transformer, BLIP2, and CLIP, as well as three language models, namely BERT, MiniLM, and ModernBERT. Compared with a retrieval baseline using cosine similarity ranking from Stella-1.5B embeddings, our experimental results on two lifelog QA datasets demonstrate a substantial improvement. Recall@1 is observed to increase from 37.65 % to 65.57% and Precision@1 from 52.65% to 85.88% when using the ModernBERT +ViT reranking model for the OpenLifelogQA task. These findings show the robustness of multimodal reranking in context selection for lifelog QA and provide a mechanism for accurate and efficient retrieval in lifelog applications.
Quang-Linh Tran, Ly-Duyen Tran, Binh T. Nguyen 0001, Gareth J. F. Jones, Cathal Gurrin
CBMI2
2025 Introduction to the 8th Annual Lifelog Search Challenge, LSC'25
abstract
For the eighth time since 2018, the ACM Lifelog Search Challenge (LSC) was run to compare interactive lifelog search systems in a live metrics-based challenge. The goal of the LSC workshop is to comparatively evaluate the capabilities of systems accessing a large multimodal lifelog. LSC'25 attracted eleven participating teams, each of which had developed an innovative interactive lifelog retrieval system. The benchmark was organised in a hybrid manner (due to political issues in 2025) at the LSC workshop at ACM ICMR'25 in Chicago, USA. This short paper summarises the LSC workshop setting and presents the participating lifelog search systems.
Cathal Gurrin, Liting Zhou, Graham Healy, Ly-Duyen Tran, Luca Rossetto, Werner Bailer, Duc-Tien Dang-Nguyen, Steve Hodges 0001, Björn Þór Jónsson 0001, Minh-Triet Tran, Klaus Schöffmann
ICMR4
2025 AIQAM'25: The 2nd ACM Workshop on AI-powered Question Answering Systems for Multimedia
abstract
The Second ACM Workshop on AI-Powered Question & Answering Systems for Multimedia (AIQAM'25) was held on 27 October 2025 in Dublin, Ireland, co-located with ACM Multimedia 2025. The workshop's main objective is to create a collaborative and inclusive space for researchers at the intersection of Artificial Inteligence (AI), large language models (LLMs), multimodal information retrieval, and question answering systems. Building on the success of its first edition (AIQAM'24) at ICMR 2024, AIQAM'25 provided a forum for presenting novel methods, applications, and surveys, which address the challenges of integrating text, image, audio, and video data into QA systems. The programme featured contributions ranging from methodological advances in reasoning and evaluation frameworks, domain-specific applications in education and sustainability, to a survey work reviewing the state of multimedia retrieval-augmented QA. This summary paper outlines the objectives and scope of the workshop, describes its format, highlights the keynote and accepted contributions, and acknowledges the efforts of the organising committee.
Tai Tan Mai, Ly-Duyen Tran, Quang-Linh Tran, Duc-Tien Dang-Nguyen, Cathal Gurrin
ACM Multimedia2
2025 The CASTLE 2024 Dataset: Advancing the Art of Multimodal Understanding
abstract
Egocentric video has seen increased interest in recent years, as it is used in a range of areas. However, most existing datasets are limited to a single perspective. In this paper, we present the CASTLE 2024 dataset, a multimodal collection containing ego- and exo-centric (i.e., first- and third-person perspective) video and audio from 15 time-aligned sources, as well as other sensor streams and auxiliary data. The dataset was recorded by volunteer participants over four days in a common location and includes the point of view of 10 participants, with an additional 5 fixed cameras providing an exocentric perspective. The entire dataset contains over 600 hours of UHD video recorded at 50 frames per second. In contrast to other datasets, CASTLE 2024 does not contain any partial censoring, such as blurred faces or distorted audio. The dataset is available via https://castle-dataset.github.io/.
Luca Rossetto, Werner Bailer, Duc-Tien Dang-Nguyen, Graham Healy, Björn Þór Jónsson 0001, Onanong Kongmeesub, Hoang Bao Le, Stevan Rudinac, Klaus Schöffmann, Florian Spiess 0001, Ly-Duyen Tran, Minh-Triet Tran, Quang-Linh Tran, Cathal Gurrin
ACM Multimedia11
2025 Overview of the First CASTLE Grand Challenge at ACM Multimedia 2025
abstract
The inaugural edition of the CASTLE grand challenge was held at ACM Multimedia 2025. The focus of the CASTLE challenge is to advance the state-of-the-art in analysis and understanding of multimodal data, especially centered around multistream ego- and exo-centric video. In this first instance of the challenge, participants had to solve three types of tasks - event instance search, object instance search, and question answering - in a fully automatic or interactive setting.
Luca Rossetto, Werner Bailer, Cathal Gurrin, Duc-Tien Dang-Nguyen, Klaus Schöffmann, Ly-Duyen Tran
ACM Multimedia6
2025 Extending Lifelog Retrieval to Multi-stream Video Retrieval at the CASTLE Challenge 2025
abstract
We present the DCU team's system for the CASTLE Challenge at ACM Multimedia 2025, which explores video retrieval and question answering in egocentric, multi-user environments. Our system adapts techniques developed for lifelogging, particularly event-based semantic retrieval and QA pipelines, to the CASTLE dataset with minimal architectural changes. It combines vision-language embeddings, transcript-based retrieval, and person tracking to support both automatic and interactive search workflows. In the interactive track, we introduce a modular interface for narrative reconstruction and exploratory search. Qualitative results show that the system can generate plausible, evidence-based answers to complex multimodal queries. These findings suggest that lifelog retrieval systems offer a viable foundation for broader egocentric video analysis.
Quang-Linh Tran, Hoang Bao Le, Thang-Long Nguyen-Ho, Graham Healy, Liting Zhou, Ly-Duyen Tran
ACM Multimedia6
2025 SLIVeR: A Narrative VR Experience for Immersive Lifelog Exploration
abstract
We present SLIVeR (Someone else's Lifelog in Virtual Reality), an interactive system that reimagines lifelog data through narrative-based Virtual Reality (VR). Instead of passively viewing chronological data, users navigate personal history through cinematic scenes and existential prompts embedded in a memory-reconstruction storyline. Guided by life-oriented questions, users progress from disorientation to recollection using curated lifelog clips. Built in Unity and deployed on Meta Quest 3, SLIVeR transforms lifelogging into a reflective, game-like journey that blends storytelling, gamification, and immersive visuals to enhance user engagement.
Songkai Jia, Cathal Gurrin, Ly-Duyen Tran
ACM Multimedia4
2025 Through Someone Else's Eyes: Lifelogging Meets Narrative Virtual Reality
Songkai Jia, Cathal Gurrin, Monica Ward, Ly-Duyen Tran
ACM Multimedia5
2024 The First ACM Workshop on AI-Powered Question Answering Systems for Multimedia
abstract
The advent of large language models (LLMs) has energised research in Question-Answering (QA) tasks, enabling responses across varied domains like economics and mathematics. Despite their capabilities, LLMs often lack explainability due to their complex parameter embeddings. Additionally, integrating multimedia data into QA systems introduces challenges in processing and interpreting diverse data types such as text, images, audio, and video. This necessitates sophisticated algorithms for accurate information retrieval across media while ensuring the reliability of the data and responses remains a significant challenge. The AIQAM workshop aims to bring together researchers and practitioners to address these challenges and enhance QA systems with multimedia data. The focus is on promoting innovations that improve the accuracy, explainability, and trustworthiness of QA systems, contributing to the development of the field.
Tai Tan Mai, Quang-Linh Tran, Ly-Duyen Tran, Van-Tu Ninh, Duc-Tien Dang-Nguyen, Cathal Gurrin
ICMR3
2024 Interactive Question Answering for Multimodal Lifelog Retrieval
abstract
Supporting Question Answering (QA) tasks is the next step for lifelog retrieval systems, similar to the progression of the parent field of information retrieval. In this paper, we propose a new pipeline to tackle the QA task in the context of lifelogging, which is based on the open-domain QA pipeline. We incorporate this pipeline into a multimodal lifelog retrieval system, which allows users to submit questions prevalent to a lifelog and then suggests possible text answers based on multimodal data. A test collection is developed to facilitate the user study, the aim of which is to evaluate the effectiveness of the proposed system compared to a conventional lifelog retrieval system. The results show that the proposed system is more effective than the conventional system, in terms of both effectiveness and user satisfaction. The results also suggest that the proposed system is more valuable for novice users, while both systems are equally effective for experienced users.
Ly-Duyen Tran, Liting Zhou, Binh T. Nguyen 0001, Cathal Gurrin
MMM (5)1
2023 VAISL: Visual-Aware Identification of Semantic Locations in Lifelog
Ly-Duyen Tran, Dongyun Nie, Liting Zhou, Binh T. Nguyen 0001, Cathal Gurrin
MMM (2)1
2023 Myscéal: a deeper analysis of an interactive lifelog search engine
Ly-Duyen Tran, Duy Nguyen 0003, Binh T. Nguyen 0001, Liting Zhou
Multim. Tools Appl.1
2022 An Exploration into the Benefits of the CLIP model for Lifelog Retrieval
abstract
In this paper, we attempt to fine-tune the CLIP (Contrastive Language-Image Pre-Training) model on the Lifelog Question Answering dataset (LLQA) to investigate retrieval performance of the fine-tuned model over the zero-shot baseline model. We train the model adopting a weight space ensembling approach using a modified loss function to take into account the differences in our dataset (LLQA) when compared with the dataset the CLIP model was originally pretrained on. We further evaluate our fine-tuned model using visual as well as multimodal queries on multiple retrieval tasks, demonstrating improved performance over the zero-shot baseline model.
Ly-Duyen Tran, Naushad Alam, Yvette Graham, Linh Khanh Vo, Nghiem Tuong Diep, Binh T. Nguyen 0001, Liting Zhou, Cathal Gurrin
CBMI1
2022 LLQA - Lifelog Question Answering Dataset
abstract
Recollecting details from lifelog data involves a higher level of granularity and reasoning than a conventional lifelog retrieval task. Investigating the task of Question Answering (QA) in lifelog data could help in human memory recollection, as well as improve traditional lifelog retrieval systems. However, there has not yet been a standardised benchmark dataset for the lifelog-based QA. In order to provide a first dataset and baseline benchmark for QA on lifelog data, we present a novel dataset, LLQA, which is an augmented 85-day lifelog collection and includes over 15,000 multiple-choice questions. We also provide different baselines for the evaluation of future works. The results showed that lifelog QA is a challenging task that requires more exploration. The dataset is publicly available at https://github.com/allie-tran/LLQA .
Ly-Duyen Tran, Thanh Cong Ho, Lan Anh Pham, Binh T. Nguyen 0001, Cathal Gurrin, Liting Zhou
MMM (1)1
2022 A Virtual Reality Reminiscence Interface for Personal Lifelogs
abstract
Letters, diaries, postcards, photo albums, home videos, and lifelogs! These are artefacts of our personal history, they represent how we cherish and preserve memories, re-engage with our past and share our experiences with others. In this demonstration paper, we explore an Virtual Reality (VR) approach to help people reminisce about the past through lifelogs. Our user study found that most participants enjoyed the experience, although for some, the VR environment was overwhelming.
Ly-Duyen Tran, Diarmuid Kennedy, Liting Zhou, Binh T. Nguyen 0001, Cathal Gurrin
MMM (2)1
2021 A VR Interface for Browsing Visual Spaces at VBS2021
Ly-Duyen Tran, Duy Nguyen 0003, Thao-Nhu Nguyen, Graham Healy, Annalina Caputo, Binh T. Nguyen 0001, Cathal Gurrin
MMM (2)1