Hui-Yin Wu

dblp:99/6187 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0001-7315-210XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 12 since 2021Human-computer interaction and ubiquitous computing · 10 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Empathetic Storytelling in Interactive Extended Reality
abstract
Extended reality (XR) technologies have emerged as a powerful medium for emotional and empathetic storytelling, yet its affective design remains an open challenge. Designers currently lack reliable methods to confirm whether the emotions they intend to elicit are experienced by players. Another problem is the lack of methods to characterize and understand emotional responses in such interactive environments. Additionally, little is known about how interaction shapes emotional engagement in narrative contexts. This doctoral research addresses these gaps through a systematic literature review and a series of user studies. The literature review surveys target emotions, the methods used for evoking emotions, and measurement methods across XR applications. The studies involve purpose-built interactive virtual reality (VR) narratives. Each scene is designed to elicit specific emotions and assessed using questionnaires and multimodal physiological, attentional and behavioral measurement, such as electrodermal activity (EDA), heart rate (HR), facial expressions, eye tracking, and voice. This thesis investigates the gap between designer intent and player experience, examines how interactivity modulates emotional state within narrative contexts, and works toward a reusable characterization framework for affective XR scenes. We are currently conducting our first study, An Unusual Day, a six-scene interactive VR narrative, each with one target emotion, to establish the foundation for this framework.
Pauline Devictor, Hui-Yin Wu, Marco Winckler
IMX2
2026 Quantifying the Cost of Manual Navigation: A Comparison of Gesture-Based Magnification versus Direct Access Reading in Digital Layout-based Documents
abstract
Understanding how diverse audiences engage with structured media is critical to ensure a consistent quality of experience. In this context, we quantify the behavioral and performance cost of manual navigation (e.g., pinch and zoom) versus direct structural access in layout-based digital documents. We specifically investigate newspaper reading when visual access to structural cues (headlines as entry points) is constrained. Participants completed two tasks—reading all headlines aloud and locating target articles—under two conditions: (1) original edition with gesture-based magnification (pan and zoom), which is the industry standard for digital documents, and (2) large-print edition supporting direct-access reading. We collected performance measures (success ratio and completion time), behavioral integrity through reading path analysis, alongside perceived workload and preferences (NASA-TLX). Results from linear mixed-effects models show that the large-print condition yielded not only better performance than gesture-based magnification (18% improvement in reading speed, 30% improvement in speed to locate a target), but more importantly, restored the natural reading strategy that gesture-based magnification interaction disrupts. Readers also reported lower workload and higher preference. These findings highlight the importance of developing automated methods for generating large-print editions, where layout adaptation complements font scaling to support accessibility and quality of experience.
Sebastián Gallardo Díaz, Aurélie Calabrèse, Hui-Yin Wu, Monica Di Meo, Stéphanie Baillif, Dorian Mazauric, Pierre Kornprobst
IMX3
2025 The Triangle of Misunderstanding in Interactive Virtual Narratives: Gulfs Between System, Designers and Players
abstract
Designers of storytelling experiences in virtual reality (VR) can take advantage of the medium's realism and immersion to communicate their intentions.However, interaction freedom comes with unpredictability, raising the risk of miscommunication between the experience sought by the designer and the player's interpretation.To better understand such miscommunications, we revisit Don Norman's work on stages of action to propose a model of designerplayer gulfs in VR that incorporates eight classes of communication gulfs.We designed a two-phase study where 10 participants designed VR scenarios and then played scenarios created by previous participants.Through coupled structured interviews, we identified 127 issues in VR-mediated communication that were mapped to our model to understand their impact on the player's interpretation of the narrative experience.Our work provides a roadmap to identifying sources of miscommunication in VR, a first step to conceiving principles and guidelines for achieving effective communication in storytelling experiences.
Florent Robert, Hui-Yin Wu, Lucile Sassatelli, Marco Winckler
CHI2
2025 WristNotes: Detachable Menu for Annotations in Immersive Environments
Clément Quere, Aline Menin, Eliezer E. Bernart, Carla M. D. S. Freitas, Luciana Porcher Nedel, Paulo Moura, Hui-Yin Wu, Marco Winckler
INTERACT (3)7
2025 Re-examining Concept-based Explainable Models for Multimodal Interpretative Tasks
abstract
Concept-based models have been proposed as a new line of research for explainable by-design deep learning models. However, those models show their whole power when applied to benchmarks where the concepts are well defined and the concepts' attributes easily extractable from the raw data. In this paper, we challenge the most recent concept-based model initially developed for image classification, on more complex interpretative tasks from a recently proposed video benchmark where they perform poorly. We conduct a root cause analysis of the poor performances of state-of-the-art explainable concept-based models for these multimodal interpretative tasks, and propose adaptations to design robust explainable models for detecting character objectification in this novel challenging video benchmark. We show that the optimal architectural choice may vary depending on the modality setting, thereby showing that designing multimodal concept-based approaches remains an open challenge and calls for further investigation.
Julie Tores, Elisa Ancarani, Rémy Sun, Lucile Sassatelli, Hui-Yin Wu, Frédéric Precioso
ACM Multimedia5
2024 Visual Objectification in Films: Towards a New AI Task for Video Interpretation
abstract
In film gender studies, the concept of “male gaze” refers to the way the characters are portrayed on-screen as objects of desire rather than subjects. In this article, we introduce a novel video-interpretation task, to detect character objectification in films. The purpose is to reveal and quantify the usage of complex temporal patterns operated in cinema to produce the cognitive perception of objectification. We introduce the ObyGaze12 dataset, made of 1914 movie clips densely annotated by experts for objectification concepts identified in film studies and psychology. We evaluate recent vision models, show the feasibility of the task and where the challenges remain with concept bottleneck models. Our new dataset and code are made available to the community.
Julie Tores, Lucile Sassatelli, Hui-Yin Wu, Clement Bergman, Lea Andolfi, Victor Ecrement, Frédéric Precioso, Thierry Devars, Magali Guaresi, Virginie Julliard, Sarah Lecossais
CVPR3
2024 AMD Journee: A Patient Co-designed VR Experience to Raise Awareness Towards the Impact of AMD on Social Interactions
abstract
We present a virtual reality (VR) experience designed to raise awareness towards the impact of low-vision conditions on social interactions for patients. Specifically, we look at age-related macular degeneration (AMD) that results in the loss of central visual field acuity (a.k.a. a scotoma), which hinders AMD patients from perceiving facial expressions and gestures, and can bring about awkward interactions, misunderstandings, and feelings of isolation. Using VR, we co-designed an experience composed of four scenes from the life of AMD patients through structured interviews with the patients and orthoptists. The experience takes the perspective of a patient, and throughout the scenarios, provides voiceovers on their feelings, the challenges they face, how they adapt to their situation, and also bits of advice on how their quality of life was improved through considerate actions from people in their social circles. A virtual scotoma is designed to follow the gaze of the user using the HTC Vive Focus 3 headset with an eye-tracking module. Setting out from a formal definition of awareness, we evaluate our experience on three components of awareness – knowledge, engagement, and empathy – through established questionnaires, continuous measures of gaze and skin conductance, and qualitative feedback. Carrying out a experiment with 29 participants, we found not only that our experience had a positive and strong impact on the awareness of participants towards AMD, but also that the scotoma and events had observable influences on gaze activity and emotions. We believe this work outlines the advantages of immersive technologies for public awareness towards conditions such as AMD, and opens avenues to conducting studies with fine-grained, multimodal analysis of user behaviour for designing more engaging experiences.
Johanna Delachambre, Hui-Yin Wu, Sebastian Vizcay, Monica Di Meo, Frédérique Lagniez, Christine Morfin-Bourlat, Stéphanie Baillif, Pierre Kornprobst
IMX2
2024 HandyNotes: using the hands to create semantic representations of contextually aware real-world objects
abstract
This paper uses Mixed Reality (MR) technologies to provide a seamless integration of digital information in physical environments through human-made annotations. Creating digital annotations of physical objects evokes many challenges for performing (simple) tasks such as adding digital notes and connecting them to real-world objects. For that, we have developed an MR system using the Microsoft HoloLens2 to create semantic representations of contextually-aware real-world objects while interacting with holographic virtual objects. User interaction is enhanced with use of fingers as placeholders for menu items. We demonstrate our approach through two real-world scenarios. We also discuss the challenges for using MR technologies.
Clément Quere, Aline Menin, Raphaël Julien, Hui-Yin Wu, Marco Winckler
VR4
2024 Task-based methodology to characterise immersive user experience with multivariate data
abstract
Virtual Reality (VR) technologies enable strong emotions compared to traditional media, stimulating the brain in ways comparable to real-life interactions. This makes VR systems promising for research and applications in training or rehabilitation, to imitate realistic situations. Nonetheless, the evaluation of the user experience in immersive environments is daunting, the richness of the media presents challenges to synchronise context with behavioural metrics in order to provide fine-grained personalised feedback or performance evaluation. The variety of scenarios and interaction modalities multiplies this difficulty of user understanding in face of lifelike training scenarios, complex interactions, and rich context.We propose a task-based methodology that provides fine-grained descriptions and analyses of the experiential user experience (UX) in VR that (1) aligns low-level tasks (i.e. take an object, go somewhere) with multivariate behaviour metrics: gaze, motion, skin conductance, (2) defines performance components (i.e., attention, decision, and efficiency) with baseline values to evaluate task performance, and (3) characterises task performance with multivariate user behaviour data. To illustrate our approach, we apply the task-based methodology to an existing dataset from a road crossing study in VR. We find that the task-based methodology allows us to better observe the experiential UX by highlighting fine-grained relations between behaviour profiles and task performance, opening pathways to personalised feedback and experiences in future VR applications.
Florent Robert, Hui-Yin Wu, Lucile Sassatelli, Marco Winckler
VR2
2023 An Integrated Framework for Understanding Multimodal Embodied Experiences in Interactive Virtual Reality
abstract
Virtual Reality (VR) technology enables “embodied interactions” in realistic environments where users can freely move and interact, with deep physical and emotional states. However, a comprehensive understanding of the embodied user experience is currently limited by the extent to which one can make relevant observations, and the accuracy at which observations can be interpreted.
Florent Robert, Hui-Yin Wu, Lucile Sassatelli, Stephen Ramanoël, Auriane Gros, Marco Winckler
IMX2
2022 On The Link Between Emotion, Attention And Content In Virtual Immersive Environments
abstract
While immersive media have been shown to generate more intense emotions, saliency information has been shown to be a key component for the assessment of their quality, owing to the various portions of the sphere (viewports) a user can attend. In this article, we investigate the tri-partite connection between user attention, user emotion and visual content in immersive environments. To do so, we present a new dataset enabling the analysis of different types of saliency, both low-level and high-level, in connection with the user’s state in 360◦videos. Head and gaze movements are recorded along with self-reports and continuous physiological measurements of emotions. We then study how the accuracy of saliency estimators in predicting user attention depends on user-reported and physiologically-sensed emotional perceptions. Our results show that high-level saliency better predicts user attention for higher levels of arousal. We discuss how this work serves as a first step to understand and predict user attention and intents in immersive interactive environments.
Quentin Guimard, Florent Robert, Camille Bauce, Aldric Ducreux, Lucile Sassatelli, Hui-Yin Wu, Marco Winckler, Auriane Gros
ICIP6
2022 PEM360: a dataset of 360° videos with continuous physiological measurements, subjective emotional ratings and motion traces
abstract
From a user perspective, immersive content can elicit more intense emotions than flat-screen presentations. From a system perspective, efficient storage and distribution remain challenging, and must consider user attention. Understanding the connection between user attention, user emotions and immersive content is therefore key. In this article, we present a new dataset, PEM360 of user head movements and gaze recordings in 360° videos, along with self-reported emotional ratings of valence and arousal, and continuous physiological measurement of electrodermal activity and heart rate. The stimuli are selected to enable the spatiotemporal analysis of the connection between content, user motion and emotion. We describe and provide a set of software tools to process the various data modalities, and introduce a joint instantaneous visualization of user attention and emotion we name Emotional maps. We exemplify new types of analyses the PEM360 dataset can enable. The entire data and code are made available in a reproducible framework.
Quentin Guimard, Florent Robert, Camille Bauce, Aldric Ducreux, Lucile Sassatelli, Hui-Yin Wu, Marco Winckler, Auriane Gros
MMSys6
2022 Analyzing and understanding embodied interactions in virtual reality systems: research proposal
abstract
Virtual reality (VR) offers opportunities in human-computer interaction research, to embody users in immersive environments and observe how they interact with 3D scenarios under well-controlled environments. VR content has stronger influences on users physical and emotional states as compared to traditional 2D media, however, a fuller understanding of this kind of embodied interaction is currently limited by the extent to which attention and behavior can be observed in a VR environment, and the accuracy at which these observations can be interpreted as, and mapped to, real-world interactions and intentions. This thesis aims at the creation of a system to help designers in the analysis of the entire user experience in VR environment: how they feel, what is their intentions when interacting with a certain object, provide them guidance based on their needs and attention. A controlled environment in which the user is guided will help to establish a better intersubjectivity between designer intention who created the experience and users who lived it and will lead to a more efficient analysis of the user behavior in VR systems for the design of better experiences.
Florent Robert, Marco Winckler, Hui-Yin Wu, Lucile Sassatelli
MMSys3
2022 Designing Guided User Tasks in VR Embodied Experiences
abstract
Virtual reality (VR) offers extraordinary opportunities in user behavior research to study and observe how people interact in immersive 3D environments. A major challenge of designing these 3D experiences and user tasks, however, lies in bridging the inter-relational gaps of perception between the designer, the user, and the 3D scene. Paul Dourish identified three gaps of perception: ontology between the scene representation and the user and designer interpretation, intersubjectivity of task communication between designer and user, and intentionality between the user's intentions and designer's interpretations. We present the GUsT-3D framework for designing Guided User Tasks in embodied VR experiences, i.e., tasks that require the user to carry out a series of interactions guided by the constraints of the 3D scene. GUsT-3D is implemented as a set of tools that support a 4-step workflow to (1) annotate entities in the scene with navigation and interaction possibilities, (2) define user tasks with interactive and timing constraints, (3) manage interactions, task validation, and user logging in real-time, and (4) conduct post-scenario analysis through spatio-temporal queries using ontology definitions. To illustrate the diverse possibilities enabled by our framework, we present two case studies with an indoor scene and an outdoor scene, and conducted a formative evaluation involving six expert interviews to assess the framework and the implemented workflow. Analysis of the responses show that the GUsT-3D framework fits well into a designer's creative process, providing a necessary workflow to create, manage, and understand VR embodied experiences.
Hui-Yin Wu, Florent Robert, Théo Fafet, Brice Graulier, Barthelemy Passin-Cauneau, Lucile Sassatelli, Marco Winckler
Proc. ACM Hum. Comput. Interact.1
2021 Towards accessible news reading design in virtual reality for low vision
abstract
Low-vision conditions resulting in partial loss of the central visual field strongly affect patients’ daily tasks and routines, and none more prominently than the ability to access text. Though vision aids such as magnifiers, digital screens, and text-to-speech devices can improve overall accessibility to text, news media, which is non-linear and has complex and volatile formatting, is still inaccessible, barring low-vision patients from easy access to essential news content. This position paper proposes virtual reality as a promising solution towards accessible and enjoyable news reading for low vision. We first provide an extensive review into existing research on low-vision reading technologies and visual accessibility solutions for modern news media. From previous research and studies, we then conduct an analysis into the advantages of virtual reality for low-vision reading and propose comprehensive guidelines for visual accessibility design in virtual reality, with a focus on reading. This is coupled with a hands-on survey of eight reading applications in virtual reality to evaluate how accessibility design is currently implemented in existing products. Finally, we present an open toolbox using browser-based graphics (WebGL) that implements the design principles from our study. A proof-of-concept is created using this toolbox to demonstrate the feasibility of our proposal with modern virtual reality technology.
Hui-Yin Wu, Aurélie Calabrèse, Pierre Kornprobst
Multim. Tools Appl.1
2020 Joint Attention for Automated Video Editing
abstract
Joint attention refers to the shared focal points of attention for occupants in a space. In this work, we introduce a computational definition of joint attention for the automated editing of meetings in multi-camera environments from the AMI corpus. Using extracted head pose and individual headset amplitude as features, we developed three editing methods: (1) a naive audio-based method that selects the camera using only the headset input, (2) a rule-based edit that selects cameras at a fixed pacing using pose data, and (3) an editing algorithm using LSTM (Long-short term memory) learned joint-attention from both pose and audio data, trained on expert edits. The methods are evaluated qualitatively against the human edit, and quantitatively in a user study with 22 participants. Results indicate that LSTM-trained joint attention produces edits that are comparable to the expert edit, offering a wider range of camera views than audio, while being more generalizable as compared to rule-based methods.
Hui-Yin Wu, Trevor Santarra, Michael Leece, Rolando Vargas, Arnav Jhala
IMX1
2018 Thinking Like a Director: Film Editing Patterns for Virtual Cinematographic Storytelling
abstract
This article introduces Film Editing Patterns (FEP) , a language to formalize film editing practices and stylistic choices found in movies. FEP constructs are constraints, expressed over one or more shots from a movie sequence, that characterize changes in cinematographic visual properties, such as shot sizes, camera angles, or layout of actors on the screen. We present the vocabulary of the FEP language, introduce its usage in analyzing styles from annotated film data, and describe how it can support users in the creative design of film sequences in 3D. More specifically, (i) we define the FEP language, (ii) we present an application to craft filmic sequences from 3D animated scenes that uses FEPs as a high level mean to select cameras and perform cuts between cameras that follow best practices in cinema, and (iii) we evaluate the benefits of FEPs by performing user experiments in which professional filmmakers and amateurs had to create cinematographic sequences. The evaluation suggests that users generally appreciate the idea of FEPs, and that it can effectively help novice and medium experienced users in crafting film sequences with little training.
Hui-Yin Wu, Francesca Palù, Roberto Ranon, Marc Christie
ACM Trans. Multim. Comput. Commun. Appl.1
2012 Structure-conforming XML document transformation based on graph homomorphism
abstract
We propose a principled method to specify XML document transformation so that the outcome of a transformation can be ensured to conform to certain structural constraints as required by the target XML document type. We view XML document types as graphs, and model transformations as relations between the two graphs. Starting from this abstraction, we use and extend graph homomorphism as a formalism for the specifications of transformations between XML document types. A specification can then be checked to ensure whether results from the transformation will always be structure-conforming.
Tyng-Ruey Chuang, Hui-Yin Wu
ACM Symposium on Document Engineering2