Khanh-Duy Le

dblp:188/2629 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0002-8297-5666ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 10 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Enhancing Pointing Gestures of Non-HMD Users in Asymmetric Collocated Mixed Reality Collaboration
abstract
A common collocated group setting in mixed-reality (MR) collaboration is a person wearing a MR headset (HMD user) and presenting MR contents to audiences who are not provided with such specialized devices (Non-HMD users). In this setting, while Non-HMD users can view the MR environment shown on a large physical display, it still remains challenging for the HMD user to interpret their pointing gesture when they spatially refer to objects in the MR environment. To address this, we designed and evaluated two pointing techniques—SCREEN and SCREEN+SPACE—that support Non-HMD users in referring to MR content. Screen pointing allows users to refer to MR objects by pointing at their representation on the display, while Screen+Space pointing additionally enables direct spatial references toward the HMD user’s environment. In a study involving 36 participants (18 pairs) performing an expert-novice collaboration task, both techniques improved communication fluency and reduced cognitive workload. Notably, Screen+Space pointing was preferred over Screen pointing, where Non-HMD users often pointed relative to the HMD user’s physical space rather than the screen. This suggests that people naturally treat MR content as embedded in shared space, and that space pointing enables more natural and effective communication in asymmetric MR collaborations.
Nam-Dang Vo, Van-Vinh Thai, Anthony Tang 0001, Khanh-Duy Le
AVI4
2025 RedirectedStepper: Exploring Walking-In-Place Locomotion in VR Using a Mini Stepper for Ascents
Quang-Tri Le, Nham Huynh-Duc, Tanh Quang Tran, Morten Fjeld, Wolfgang Stuerzlinger, Markus Zank, Minh-Triet Tran, Khanh-Duy Le
CHI8
2025 Designing Interactive Multimodal Information Retrieval and Access for Heads Up Computing (DIMIRA-HUC)
abstract
The advancement of wearable intelligent systems presents a unique opportunity to transform how humans interact with digital content. This workshop explores the design of Interactive Multimodal Information Retrieval and Access systems specifically tailored for Heads-Up Computing environments. By leveraging multimodal inputs, such as voice, gaze, and gesture, these systems enable real-time, hands-free access to digital information, facilitating seamless and efficient interaction. The goal is to support tasks requiring rapid information access in dynamic environments while ensuring users remain "heads-up" and engaged with the real world. This half-day workshop will share research outcomes and best practices, foster community building, and facilitate discussions on key challenges. By bringing together researchers and practitioners, it aims to drive further advancements in both research and practical applications within this rapidly evolving field.
Haiming Liu 0002, Shengdong Zhao 0001, Silang Wang, Preben Hansen, Ian Oakley, Khanh-Duy Le
CHIIR6
2025 Semi-Supervised Semantic Segmentation using Redesigned Self-Training for White Blood Cells
abstract
Artificial Intelligence (AI) in the medical field, especially in the diagnosis of cancer relating to white blood cells, is held back by two primary obstacles: the scarcity of extensively labeled datasets for white blood cell (WBC) segmentation and outdated segmentation techniques. These challenges stall the development of more accurate and modern techniques to diagnose cancer relating to white blood cells. To overcome the first challenge, there is a need to devise a semi-supervised learning framework to efficiently make use of the huge dataset that is not annotated. Our study tackles this by introducing an innovative self-training approach that integrates FixMatch. Self-training is a semi-supervised framework that leverages the model trained on labeled data to generate pseudo-labels for the unlabeled data and then re-train on both of them. FixMatch is used to regularize the model against weak and strong perturbations in the input image. We discover that by incorporating FixMatch in the self-training pipeline, the performance improves in many cases. When tested, our method achieved superior results on the UNet and DeepLab-V3 networks with ResNet-50 architecture, attaining accuracies of 90.87%, 89.05%, and 75.13% on the Zheng 1, Zheng 2, and LISC datasets, respectively.
Quoc-Vinh Luu, Khanh-Duy Le, Thanh-Huy Nguyen, Thanh-Minh Nguyen, Tien-Thinh Nguyen, Quang-Vinh Dinh
IPAS2
2025 Event-Enriched Image Analysis Grand Challenge At ACM Multimedia 2025
abstract
The Event-Enriched Image Analysis (EVENTA) Grand Challenge, hosted at ACM Multimedia 2025, introduces the first large-scale benchmark for event-level multimodal understanding. Traditional captioning and retrieval tasks largely focus on surface-level recognition of people, objects, and scenes, often overlooking the contextual and semantic dimensions that define real-world events. EVENTA addresses this gap by integrating contextual, temporal, and semantic information to capture the who, when, where, what, and why behind an image. Built upon the OpenEvents V1 dataset, the challenge features two tracks: Event-Enriched Image Retrieval and Captioning, and Event-Based Image Retrieval. A total of 45 teams from six countries participated, with evaluation conducted through Public and Private Test phases to ensure fairness and reproducibility. The top three teams were invited to present their solutions at ACM Multimedia 2025. EVENTA establishes a foundation for context-aware, narrative-driven multimedia AI, with applications in journalism, media analysis, cultural archiving, and accessibility. Further details about the challenge are available at the official homepage: https://ltnghia.github.io/eventa/eventa-2025.
Thien-Phuc Tran, Minh-Quang Nguyen, Minh-Triet Tran, Tam V. Nguyen 0002, Trong-Le Do, Duy-Nam Ly, Viet-Tham Huynh, Khanh-Duy Le, Mai-Khiem Tran, Trung-Nghia Le
ACM Multimedia8
2025 Enhancing Spatial Understanding in Mixed-Reality Presentations
abstract
Mixed reality (MR) presentations often involve a presenter wearing a head-mounted display (HMD) and an audience watching via a large display, making it difficult for audiences to perceive spatial relationships between the presenter and virtual objects. We report two experiments testing three design variations: (1) scene camera placement (audience-aligned vs. opposite), (2) overlaying the presenter’s first-person view, and (3) highlighting objects in the presenter’s view. Results show that audience-aligned cameras and object highlighting improve spatial understanding, while combining third- and first-person views can further aid perception. We derive design guidelines for configuring MR presentations to better support audience comprehension.
Nam-Dang Vo, Van-Vinh Thai, Nam H. Do, Viet-Tham Huynh, Anthony Tang 0001, Khanh-Duy Le
VRST6
2024 DataDive: Supporting Readers' Contextualization of Statistical Statements with Data Exploration
abstract
Statistical statements that refer to data to support narratives or claims are commonly used to inform readers about the magnitude of social issues. While contextualizing statistical statements with relevant data supports readers in building their own interpretation of statements, the complexity of finding contextual information on the web and linking statistical statements with it impedes readers’ efforts to do so. We present DataDive, an interactive tool for contextualizing statistical statements for the readers of online texts. Based on users’ selections of statistical statements, our tool uses an LLM-powered pipeline to generate candidates of relevant contexts and poses them as guiding questions to the user as potential contexts for exploration. When the user selects a question, DataDive employs visualizations to further help the user compare and explore contextually relevant data. A technical evaluation shows that DataDive generates important and diverse questions that facilitate exploration around statistical statements and retrieves relevant data for comparison. Moreover, a user study with 21 participants suggests that DataDive facilitates users to explore diverse contexts and to be more aware of how statistical data could relate to the text.
Khanh-Duy Le, Gionnieve Lim, Daehyun Kim 0005, Yoo Jin Hong, Juho Kim 0001
IUI2
2023 TextANIMAR: Text-based 3D animal fine-grained retrieval
Trung-Nghia Le, Tam V. Nguyen 0002, Minh-Quan Le, Viet-Tham Huynh, Trong-Le Do, Khanh-Duy Le, Mai-Khiem Tran, Nhat Hoang-Xuan, Thang-Long Nguyen-Ho, Vinh-Tiep Nguyen, Tuong-Nghiem Diep, Khanh-Duy Ho, Xuan-Hieu Nguyen, Thien-Phuc Tran, Tuan-Anh Yang, Kim-Phat Tran, Nhu-Vinh Hoang, Minh-Quang Nguyen, E-Ro Nguyen, Minh-Khoi Nguyen-Nhat, Tuan-An To, Trung-Truc Huynh-Le, Nham-Tan Nguyen, Hoang-Chau Luong, Truong Hoai Phong, Nhat-Quynh Le-Pham, Huu-Phuc Pham, Trong-Vu Hoang, Quang-Binh Nguyen, Hai-Dang Nguyen, Akihiro Sugimoto, Minh-Triet Tran
Comput. Graph.7
2023 SketchANIMAR: Sketch-based 3D animal fine-grained retrieval
Trung-Nghia Le, Tam V. Nguyen 0002, Minh-Quan Le, Viet-Tham Huynh, Trong-Le Do, Khanh-Duy Le, Mai-Khiem Tran, Nhat Hoang-Xuan, Thang-Long Nguyen-Ho, Vinh-Tiep Nguyen, Nhat-Quynh Le-Pham, Huu-Phuc Pham, Trong-Vu Hoang, Quang-Binh Nguyen, Trong-Hieu Nguyen Mau, Tuan-Luc Huynh, Thanh-Danh Le, Ngoc-Linh Nguyen-Ha, Tuong-Vy Truong-Thuy, Truong Hoai Phong, Tuong-Nghiem Diep, Khanh-Duy Ho, Xuan-Hieu Nguyen, Thien-Phuc Tran, Tuan-Anh Yang, Kim-Phat Tran, Nhu-Vinh Hoang, Minh-Quang Nguyen, Hoai-Danh Vo, Minh-Hoa Doan, Hai-Dang Nguyen, Akihiro Sugimoto, Minh-Triet Tran
Comput. Graph.7
2021 VXSlate: Exploring Combination of Head Movements and Mobile Touch for Large Virtual Display Interaction
abstract
Virtual Reality (VR) headsets can open opportunities for users to accomplish complex tasks on large virtual displays using compact and portable devices. However, interacting with such large virtual displays using existing interaction techniques might cause fatigue, especially for precise manipulation tasks, due to the lack of physical surfaces. To deal with this issue, we explored the design of VXSlate, an interaction technique that uses a large virtual display as an expansion of a tablet. We combined a user’s head movements as tracked by the VR headset, and touch interaction on the tablet. Using VXSlate, a user head movements positions a virtual representation of the tablet together with the user’s hand, on the large virtual display. This allows the user to perform fine-tuned multi-touch content manipulations. In a user study with seventeen participants, we investigated the effects of VXSlate on users in problem-solving tasks involving content manipulations at different levels of difficulty, such as translation, rotation, scaling, and sketching. As a baseline for comparison, off-the-shelf touch-controller interactions were used. Overall, VXSlate allowed participants to complete the task with completion times and accuracy that are comparable to touch-controller interactions. After an interval of use, VXSlate significantly reduced users’ time to perform scaling tasks in content manipulations, as well as reducing perceived effort. We reflected on the advantages and disadvantages of VXSlate in content manipulation on large virtual displays and explored further applications within the VXSlate design space.
Khanh-Duy Le, Tanh Quang Tran, Karol Chlasta, Krzysztof Krejtz, Morten Fjeld, Andreas M. Kunz
Conference on Designing Interactive Systems1
2021 VRQUEST: Designing and Evaluating a Virtual Reality System for Factory Training
Khanh-Duy Le, Saad Azhar 0001, David Lindh, Dawid Ziobro
INTERACT (5)1
2019 GazeLens: Guiding Attention to Improve Gaze Interpretation in Hub-Satellite Collaboration
Khanh-Duy Le, Ignacio Avellino, Cédric Fleury, Morten Fjeld, Andreas M. Kunz
INTERACT (2)1
2019 DigiMetaplan: supporting facilitated brainstorming for distributed business teams
abstract
While facilitated brainstorming is a proven ideation method for professional teams, distributed teams are currently not able to enjoy its benefits. As workers are shifting towards collaborating in distributed settings, understanding how interactive systems can support facilitated brainstorming is becoming necessary to ensure that distributed teams remain creative. To address this challenge we designed, implemented, and evaluated DigiMetaplan---an interactive surface-based system for distributed facilitated brainstorming for co-located and remote users. The design of DigiMetaplan was inspired by a widely-used facilitated brainstorming method called Metaplan, where the brainstorming process of a group is coordinated by a facilitator. We evaluated the usability of DigiMetaplan with five hybrid teams consisting of a co-located facilitator and two team members connected with one remote participant. Results showed that the features used in DigiMetaplan on interactive surfaces effectively supported teams in performing facilitated collaborative brainstorming in partially distributed settings. We contribute knowledge on how a brainstorming environment translated into a multi-surface distributed system affects facilitated collaboration.
Khanh-Duy Le, Pawel W. Wozniak, Ali Alavi, Morten Fjeld, Andreas M. Kunz
MUM1
2018 OctoBubbles: A Multi-view interactive environment for concurrent visualization and synchronization of UML models and code
abstract
The process of software understanding often requires developers to consult both high- and low-level software artifacts (i.e. models and code). The creation and persistence of such artifacts often take place in different environments, as well as seldom in one single environment. In both cases, software models and code fragments are viewable separately making the workspace overcrowded with many opened interfaces and tabs. In such a situation, developers might lose the big picture and spend unnecessary effort on navigation and locating the artifact of interest. To assist program comprehension and tackle the problem of software navigation, we present OctoBubbles, a multi-view interactive environment for concurrent visualization and synchronization of software models and code. A preliminary evaluation of OctoBubbles with 15 professional developers shows a high level of interest, and points out to potential benefits. Furthermore, we present a future plan to quantitatively investigate the effectiveness of the environment.
Rodi Jolak, Khanh-Duy Le, Kaan Burak Sener, Michel R. V. Chaudron
SANER2
2017 Mirrortablet: exploring a low-cost mobile system for capturing unmediated hand gestures in remote collaboration
abstract
Direct and natural images of hand gestures have been shown to benefit remote collaboration on physical tasks in several settings, including ad-hoc ones. However, to capture such unmediated hand gestures within collaborative tasks, existing approaches require stationary hardware systems or heavily instrumented mobile devices, making them unfeasible for use at ad-hoc workplaces. We present MirrorTablet, a low-cost mirror-based system taking advantage of the built-in front-facing camera of a tablet to capture the user's unmediated hand interactions on and above the screen. This system requires minimal instrumentation of the tablet and can be easily (un-)mounted, making it suitable for mobile usage. A user study with ten pairs of participants on a helper-worker setup working on construction tasks showed that MirrorTablet improved task completion time and had positive effects on participants' perceived workload when working with unfamiliar tasks compared to using a common sketch-only interface. In addition, qualitative feedback yielded design considerations on hand visualization for mobile device remote collaboration on physical tasks.
Khanh-Duy Le, Kening Zhu, Morten Fjeld
MUM1
2017 Immersive environment for distributed creative collaboration
abstract
While videoconferencing has been available for years, tools for distributed creative collaboration such as brainstorming have not yet reached professional use. Unlike regular conferencing, brainstorming relies on information exchange across multiple channels in shared task and communication spaces. If a participant joins a facilitated brainstorming session remotely through his/her mobile device, perceiving these spaces is challenging. By proposing an immersive environment for the remote participant, this tech note addresses how to provide him/her a stronger engagement. Offered as a client application for the remote participant's tablet device, it enables him/her to see all channels and select which one to interact with. Our proof-of-concept client allows us to examine how this immersive environment for distributed creative collaboration can provide the remote participant with an increased engagement.
Khanh-Duy Le, Morten Fjeld, Ali Alavi, Andreas M. Kunz
VRST1
2016 NowAndThen: a social network-based photo recommendation tool supporting reminiscence
abstract
People frequently post their photos on social network sites (e.g. Facebook, Instagram) as a way to share memorable moments, emotions, or locations visited. While sharing photos with the same subjects as past photos could lead to user reminiscence and potentially create valuable benefits, existing products still cannot support this adequately. Based on a survey on user habits in sharing and revisiting photos on social network sites, we propose NowAndThen, a photo recommendation concept and tool that assists reminiscence when sharing photos on social network sites. By combining visual features of the photos and associated tags, a prototype we developed can recommend old photos that have common subjects with the user's current photos of interest. Our study with the prototype shows that this approach can help users positively revive past memories and connections with their friends. In addition, our results include various design insights and implications for future reminiscence-support systems.
Vinh-Tiep Nguyen, Khanh-Duy Le, Minh-Triet Tran, Morten Fjeld
MUM2