Yue Jiang 0002

dblp:89/6600-2 · DBLP profile ↗
← Back
20ranked-venue papers
9as first author
16since 2021 · last 2026
0000-0003-0022-6512ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 18 · 8 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 SeekUI: Predicting Visual Search Behavior on Graphical User Interfaces with a Reward-Augmented Vision Language Model
abstract
Visual search is key to understanding and improving interaction with graphical user interfaces (GUIs), yet predicting scanpaths on real GUIs remains an open challenge. Unlike free-viewing, visual search is goal-driven and shaped by both linguistic and visual features of the GUI. State-of-the-art models of visual search, trained on natural images, fail with GUIs because they cannot capture the effects of grouping and semantics on search strategies. We present SeekUI, a reward-augmented Vision Language Model (VLM) that predicts scanpaths directly from a GUI screenshot and a text cue describing the desired target. Our model extends the capability of VLMs to reproduce human-like visual search behavior on GUIs and outperforms baseline models across different types of GUIs. Importantly, it reproduces key empirical phenomena established in eye-tracking studies of visual search, including the Guess–Scan–Confirm strategy. In sum, SeekUI provides a foundation for predicting visual search behavior and has potential for informing GUI evaluation and optimization.
Zixin Guo, Yue Jiang 0002, Luis A. Leiva, Antti Oulasvirta
CHI2
2026 Is this it? Benchmarking Scanpath Metrics for Information Display
abstract
Scanpath prediction is a fundamental task in human visual attention research, aiming to simulate user viewing behaviour for given stimuli. While scanpath prediction methods have matured in natural scenes, recent research has expanded to information displays, such as graphical user interfaces and data visualisations. However, there is currently no consensus on which scanpath metrics to use for evaluation, raising concerns regarding the validity and comparability of the proposed methods. This paper benchmarks ten commonly used scanpath metrics across the MASSVIS and UEyes datasets by comparing model predictions with empirical gaze data. We evaluate these metrics with subjective expert ratings of scanpath similarity. Our analysis reveals that vector-based and region-based metrics align more closely with expert ratings than pixel-based and recurrence-based metrics. Based on these findings, we provide best practices for evaluating visual scanpaths in information displays, emphasising the urgent need for appropriate metrics to ensure the validity of future research.
Yao Wang 0018, Junichi Nagasawa, Danqing Shi, Chuhan Jiao, Yue Jiang 0002, Andreas Bulling
ETRA5
2025 ILuvUI: Instruction-tuned LangUage-Vision modeling of UIs from Machine Conversations
abstract
Publisher Copyright: © 2025 Copyright held by the owner/author(s).
Yue Jiang 0002, Eldon Schoop, Amanda Swearngin, Jeffrey Nichols 0001
IUI1
2025 Understanding visual search in graphical user interfaces
abstract
How do we find items within graphical user interfaces (GUIs)? Current understanding of this issue relies on studies using symbol matrices, natural scenes, and other non-GUI stimuli. To understand whether the effects discovered in those environments extend to mobile, desktop, and web interfaces, this paper reports on visual search performance and eye movements with 900 real-world GUIs. In an eye-tracking study, participants (N=84) were given a cue (textual or image) describing a target to find within a GUI. The study found that the type of GUI, the absence/presence of the target, and cue type affected search time more than visual complexity did. We also compared visual search to free-viewing in GUIs, concluding that these two tasks are distinctly different. Synthesis of the results points to a Guess-Scan-Confirm pattern in visual search: in the first few fixations, gaze is frequently directed toward the top-left corner of the screen, a pattern possibly related to the top left being a statistically likely location of the target or of information that could aid in finding it; attention then gets more selectively guided, in line with the GUI’s structure and the features of the target; and, finally, the user must confirm whether the target has been identified or, instead, that no target is visible. The VSGUI10K eye-tracking dataset (10,282 trials) is released for study and modeling of visual search.
Aini Putkonen, Yue Jiang 0002, Jingchun Zeng, Olli Tammilehto, Jussi P. P. Jokinen, Antti Oulasvirta
Int. J. Hum. Comput. Stud.2
2024 Graph4GUI: Graph Neural Networks for Representing Graphical User Interfaces
abstract
Present-day graphical user interfaces (GUIs) exhibit diverse arrangements of text, graphics, and interactive elements such as buttons and menus, but representations of GUIs have not kept up. They do not encapsulate both semantic and visuo-spatial relationships among elements. To seize machine learning’s potential for GUIs more efficiently, Graph4GUI exploits graph neural networks to capture individual elements’ properties and their semantic—visuo-spatial constraints in a layout. The learned representation demonstrated its effectiveness in multiple tasks, especially generating designs in a challenging GUI autocompletion task, which involved predicting the positions of remaining unplaced elements in a partially completed GUI. The new model’s suggestions showed alignment and visual appeal superior to the baseline method and received higher subjective ratings for preference. Furthermore, we demonstrate the practical benefits and efficiency advantages designers perceive when utilizing our model as an autocompletion plug-in.
Yue Jiang 0002, Changkong Zhou, Vikas Garg 0001, Antti Oulasvirta
CHI1
2024 AXNav: Replaying Accessibility Tests from Natural Language
abstract
Developers and quality assurance testers often rely on manual testing to test accessibility features throughout the product lifecycle. Unfortunately, manual testing can be tedious, often has an overwhelming scope, and can be difficult to schedule amongst other development milestones. Recently, Large Language Models (LLMs) have been used for a variety of tasks including automation of UIs. However, to our knowledge, no one has yet explored the use of LLMs in controlling assistive technologies for the purposes of supporting accessibility testing. In this paper, we explore the requirements of a natural language based accessibility testing workflow, starting with a formative study. From this we build a system that takes a manual accessibility test instruction in natural language (e.g., “Search for a show in VoiceOver”) as input and uses an LLM combined with pixel-based UI Understanding models to execute the test and produce a chaptered, navigable video. In each video, to help QA testers, we apply heuristics to detect and flag accessibility issues (e.g., Text size not increasing with Large Text enabled, VoiceOver navigation loops). We evaluate this system through a 10-participant user study with accessibility QA professionals who indicated that the tool would be very useful in their current work and performed tests similarly to how they would manually test the features. The study also reveals insights for future work on using LLMs for accessibility testing.
Maryam Taeb, Amanda Swearngin, Eldon Schoop, Ruijia Cheng, Yue Jiang 0002, Jeffrey Nichols 0001
CHI5
2024 DreamStruct: Understanding Slides and User Interfaces via Synthetic Data Generation
Yi-Hao Peng, Faria Huq, Yue Jiang 0002, Jason Wu 0001, Xin Yue Li, Jeffrey P. Bigham, Amy Pavel
ECCV (24)3
2024 EyeFormer: Predicting Personalized Scanpaths with Transformer-Guided Reinforcement Learning
abstract
From a visual-perception perspective, modern graphical user interfaces (GUIs) comprise a complex graphics-rich two-dimensional visuospatial arrangement of text, images, and interactive objects such as buttons and menus. While existing models can accurately predict regions and objects that are likely to attract attention “on average”, no scanpath model has been capable of predicting scanpaths for an individual. To close this gap, we introduce EyeFormer, which utilizes a Transformer architecture as a policy network to guide a deep reinforcement learning algorithm that predicts gaze locations. Our model offers the unique capability of producing personalized predictions when given a few user scanpath samples. It can predict full scanpath information, including fixation positions and durations, across individuals and various stimulus types. Additionally, we demonstrate applications in GUI layout optimization driven by our model.
Yue Jiang 0002, Zixin Guo, Hamed Rezazadegan Tavakoli, Luis A. Leiva, Antti Oulasvirta
UIST1
2024 FlexDoc: Flexible Document Adaptation through Optimizing both Content and Layout
abstract
Designing adaptive documents that are visually appealing across various devices and for diverse viewers is a challenging task. This is due to the wide variety of devices and different viewer requirements and preferences. Alterations to a document’s content, style, or layout often necessitate numerous adjustments, potentially leading to a complete layout redesign. We introduce FlexDoc, a framework for creating and consuming documents that seamlessly adapt to different devices, author, and viewer preferences and interactions. It eliminates the need to manually create multiple document layouts, as FlexDoc enables authors to define desired document properties using templates and employs both discrete and continuous optimization in a novel comprehensive optimization process, which leverages automatic text summarization and image carving techniques to adapt both layout and content during consumption dynamically. Further, we demonstrate FlexDoc in real-world scenarios.
Yue Jiang 0002, Christof Lutteroth, Rajiv Jain, Chris Tensmeyer, Varun Manjunatha, Wolfgang Stuerzlinger, Vlad I. Morariu
VL/HCC1
2024 Impact of Design Decisions in Scanpath Modeling
abstract
Modeling visual saliency in graphical user interfaces (GUIs) allows to understand how people perceive GUI designs and what elements attract their attention. One aspect that is often overlooked is the fact that computational models depend on a series of design parameters that are not straightforward to decide. We systematically analyze how different design parameters affect scanpath evaluation metrics using a state-of-the-art computational model (DeepGaze++). We particularly focus on three design parameters: input image size, inhibition-of-return decay, and masking radius. We show that even small variations of these design parameters have a noticeable impact on standard evaluation metrics such as DTW or Eyenalysis. These effects also occur in other scanpath models, such as UMSS and ScanGAN, and in other datasets such as MASSVIS. Taken together, our results put forward the impact of design decisions for predicting users' viewing behavior on GUIs.
Parvin Emami, Yue Jiang 0002, Zixin Guo, Luis A. Leiva
Proc. ACM Hum. Comput. Interact.2
2024 VisRecall++: Analysing and Predicting Visualisation Recallability from Gaze Behaviour
abstract
Question answering has recently been proposed as a promising means to assess the recallability of information visualisations. However, prior works are yet to study the link between visually encoding a visualisation in memory and recall performance. To fill this gap, we propose VisRecall++ -- a novel 40-participant recallability dataset that contains gaze data on 200 visualisations and 1,000 questions, including identifying the title and retrieving values. We measured recallability by asking participants questions after they observed the visualisation for 10 seconds. Our analyses reveal several insights, such as saccade amplitude, number of fixations, and fixation duration significantly differ between high and low recallability groups. Finally, we propose GazeRecallNet -- a novel computational method to predict recallability from gaze behaviour that outperforms the state-of-the-art model RecallNet and three other baselines on this task. Taken together, our results shed light on assessing recallability from gaze behaviour and inform future work on recallability-based visualisation optimisation.
Yao Wang 0018, Yue Jiang 0002, Zhiming Hu 0003, Constantin Ruhdorfer, Mihai Bâce, Andreas Bulling
Proc. ACM Hum. Comput. Interact.2
2023 UEyes: Understanding Visual Saliency across User Interface Types
abstract
While user interfaces (UIs) display elements such as images and text in a grid-based layout, UI types differ significantly in the number of elements and how they are displayed. For example, webpage designs rely heavily on images and text, whereas desktop UIs tend to feature numerous small images. To examine how such differences affect the way users look at UIs, we collected and analyzed a large eye-tracking-based dataset, UEyes (62 participants and 1,980 UI screenshots), covering four major UI types: webpage, desktop UI, mobile UI, and poster. We analyze its differences in biases related to such factors as color, location, and gaze direction. We also compare state-of-the-art predictive models and propose improvements for better capturing typical tendencies across UI types. Both the dataset and the models are publicly available.
Yue Jiang 0002, Luis A. Leiva, Hamed Rezazadegan Tavakoli, Paul R. B. Houssel, Julia Kylmälä, Antti Oulasvirta
CHI1
2022 Pretty Princess vs. Successful Leader: Gender Roles in Greeting Card Messages
abstract
People write personalized greeting cards on various occasions. While prior work has studied gender roles in greeting card messages, systematic analysis at scale and tools for raising the awareness of gender stereotyping remain under-investigated. To this end, we collect a large greeting card message corpus covering three different occasions (birthday, Valentine’s Day and wedding) from three sources (exemplars from greeting message websites, real-life greetings from social media and language model generated ones). We uncover a wide range of gender stereotypes in this corpus via topic modeling, odds ratio and Word Embedding Association Test (WEAT). We further conduct a survey to understand people’s perception of gender roles in messages from this corpus and if gender stereotyping is a concern. The results show that people want to be aware of gender roles in the messages, but remain unconcerned unless the perceived gender roles conflict with the recipient’s true personality. In response, we developed GreetA, an interactive visualization and writing assistant tool to visualize fine-grained topics in greeting card messages drafted by the users and the associated gender perception scores, but without suggesting text changes as an intervention.
Jiao Sun, Sherry Tongshuang Wu, Yue Jiang 0002, Ronil Awalegaonkar, Xi Victoria Lin, Diyi Yang
CHI3
2021 ReverseORC: Reverse Engineering of Resizable User Interface Layouts with OR-Constraints
abstract
Reverse engineering (RE) of user interfaces (UIs) plays an important role in software evolution. However, the large diversity of UI technologies and the need for UIs to be resizable make this challenging. We propose ReverseORC, a novel RE approach able to discover diverse layout types and their dynamic resizing behaviours independently of their implementation, and to specify them by using OR constraints. Unlike previous RE approaches, ReverseORC infers flexible layout constraint specifications by sampling UIs at different sizes and analyzing the differences between them. It can create specifications that replicate even some non-standard layout managers with complex dynamic layout behaviours. We demonstrate that ReverseORC works across different platforms with very different layout approaches, e.g., for GUIs as well as for the Web. Furthermore, it can be used to detect and fix problems in legacy UIs, extend UIs with enhanced layout behaviours, and support the creation of flexible UI layouts.
Yue Jiang 0002, Wolfgang Stuerzlinger, Christof Lutteroth
CHI1
2021 Vid2Doppler: Synthesizing Doppler Radar Data from Videos for Training Privacy-Preserving Activity Recognition
abstract
Millimeter wave (mmWave) Doppler radar is a new and promising sensing approach for human activity recognition, offering signal richness approaching that of microphones and cameras, but without many of the privacy-invading downsides. However, unlike audio and computer vision approaches that can draw from huge libraries of videos for training deep learning models, Doppler radar has no existing large datasets, holding back this otherwise promising sensing modality. In response, we set out to create a software pipeline that converts videos of human activities into realistic, synthetic Doppler radar data. We show how this cross-domain translation can be successful through a series of experimental results. Overall, we believe our approach is an important stepping stone towards significantly reducing the burden of training such as human sensing systems, and could help bootstrap uses in human-computer interaction.
Karan Ahuja, Yue Jiang 0002, Mayank Goel, Chris Harrison 0001
CHI2
2021 "Positive Energy": Perceptions and Attitudes Towards COVID-19 Information on Social Media in China
abstract
The COVID-19 outbreak has resulted in a worldwide public health crisis. In such times of crisis, access to relevant and accurate information is critical. For many people in China, domestic social media platforms such as WeChat and Weibo have become dominant sources of COVID-19-related information and news. People have to evaluate the trustworthiness of COVID-19-related information and make sharing decisions using platforms that have to contend with government censorship policies, astroturfers, and other government interventions. We interviewed 33 Chinese WeChat users to understand how individuals were seeking COVID-19-related information and how they identified and evaluated specific COVID-19-related misinformation. This work exposes how COVID-19-related content with "positive energy" was prevalent on social media in China. A significant number of interviewees exhibited a willingness to prioritize information valence over veracity when evaluating and sharing content with others. Further, the work revealed how Chinese citizens' understanding of information ecosystems played an important role in their attitudes towards censorship and official media, and also influenced their evaluation of domestic and international information during a global crisis.
Zhicong Lu, Yue Jiang 0002, Chenxinran Shen, Margaret C. Jack, Daniel J. Wigdor, Mor Naaman
Proc. ACM Hum. Comput. Interact.2
2020 ORCSolver: An Efficient Solver for Adaptive GUI Layout with OR-Constraints
abstract
OR-constrained (ORC) graphical user interface layouts unify conventional constraint-based layouts with flow layouts, which enables the definition of flexible layouts that adapt to screens with different sizes, orientations, or aspect ratios with only a single layout specification. Unfortunately, solving ORC layouts with current solvers is time-consuming and the needed time increases exponentially with the number of widgets and constraints. To address this challenge, we propose ORCSolver, a novel solving technique for adaptive ORC layouts, based on a branch-and-bound approach with heuristic preprocessing. We demonstrate that ORCSolver simplifies ORC specifications at runtime and our approach can solve ORC layout specifications efficiently at near-interactive rates.
Yue Jiang 0002, Wolfgang Stuerzlinger, Matthias Zwicker, Christof Lutteroth
CHI1
2020 The Government's Dividend: Complex Perceptions of Social Media Misinformation in China
abstract
The social media environment in China has become the dominant source of information and news over the past decade. This news environment has naturally suffered from challenges related to mis- and dis-information, encumbered by an increasingly complex landscape of factors and players including social media services, fact-checkers, censorship policies, and astroturfing. Interviews with 44 Chinese WeChat users were conducted to understand how individuals perceive misinformation and how it impacts their news consumption practices. Overall, this work exposes the diverse attitudes and coping strategies that Chinese users employ in complex social media environments. Due to the complex nature of censorship in China and participants' lack of understanding of censor-ship, they expressed varied opinions about its influence on the credibility of online information sources. Further, although most participants claimed that their opinions would not be easily swayed by astroturfers, many admitted that they could not effectively distinguish astroturfers from ordinary Internet users. Participants' inability to make sense of comments found online lead many participants to hold pro-censorship attitudes: the Government's Dividend.
Zhicong Lu, Yue Jiang 0002, Mor Naaman, Daniel J. Wigdor
CHI2
2020 SDFDiff: Differentiable Rendering of Signed Distance Fields for 3D Shape Optimization
abstract
We propose SDFDiff, a novel approach for image-based shape optimization using differentiable rendering of 3D shapes represented by signed distance functions (SDFs). Compared to other representations, SDFs have the advantage that they can represent shapes with arbitrary topology, and that they guarantee watertight surfaces. We apply our approach to the problem of multi-view 3D reconstruction, where we achieve high reconstruction quality and can capture complex topology of 3D objects. In addition, we employ a multi-resolution strategy to obtain a robust optimization algorithm. We further demonstrate that our SDF-based differentiable renderer can be integrated with deep learning models, which opens up options for learning approaches on 3D objects without 3D supervision. In particular, we apply our method to single-view 3D reconstruction and achieve state-of-the-art results.
Yue Jiang 0002, Dantong Ji, Zhizhong Han, Matthias Zwicker
CVPR1
2019 ORC Layout: Adaptive GUI Layout with OR-Constraints
abstract
We propose a novel approach for constraint-based graphical user interface (GUI) layout based on OR-constraints (ORC) in standard soft/hard linear constraint systems. ORC layout unifies grid layout and flow layout, supporting both their features as well as cases where grid and flow layouts individually fail. We describe ORC design patterns that enable designers to safely create flexible layouts that work across different screen sizes and orientations. We also present the ORC Editor, a GUI editor that enables designers to apply ORC in a safe and effective manner, mixing grid, flow and new ORC layout features as appropriate. We demonstrate that our prototype can adapt layouts to screens with different aspect ratios with only a single layout specification, easing the burden of GUI maintenance. Finally, we show that ORC specifications can be modified interactively and solved efficiently at runtime.
Yue Jiang 0002, Ruofei Du, Christof Lutteroth, Wolfgang Stuerzlinger
CHI1