Yao Wang 0018

dblp:72/628-18 · DBLP profile ↗
← Back
17ranked-venue papers
7as first author
15since 2021 · last 2026
0000-0002-3633-8623ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 8 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Is this it? Benchmarking Scanpath Metrics for Information Display
abstract
Scanpath prediction is a fundamental task in human visual attention research, aiming to simulate user viewing behaviour for given stimuli. While scanpath prediction methods have matured in natural scenes, recent research has expanded to information displays, such as graphical user interfaces and data visualisations. However, there is currently no consensus on which scanpath metrics to use for evaluation, raising concerns regarding the validity and comparability of the proposed methods. This paper benchmarks ten commonly used scanpath metrics across the MASSVIS and UEyes datasets by comparing model predictions with empirical gaze data. We evaluate these metrics with subjective expert ratings of scanpath similarity. Our analysis reveals that vector-based and region-based metrics align more closely with expert ratings than pixel-based and recurrence-based metrics. Based on these findings, we provide best practices for evaluating visual scanpaths in information displays, emphasising the urgent need for appropriate metrics to ensure the validity of future research.
Yao Wang 0018, Junichi Nagasawa, Danqing Shi, Chuhan Jiao, Yue Jiang 0002, Andreas Bulling
ETRA1
2026 DiffGaze: A Diffusion Model for Modelling Fine-grained Human Gaze Behaviour on 360\({}^{\circ}\) Images
abstract
Modelling human gaze behaviour on 360 \({}^{\circ}\) images is important for various human–computer interaction applications. However, existing methods are limited to predicting discrete fixation sequences or aggregated saliency maps, thereby neglecting fine-grained gaze behaviour such as saccadic eye movements that can be captured by commercial eye-trackers. We introduce a more challenging task— fine-grained gaze sequence generation . This task aims to generate eye-tracker-like gaze data for given stimuli. We propose DiffGaze , a diffusion-based method for generating realistic and diverse fine-grained human gaze sequences conditioned on 360 \({}^{\circ}\) images. We evaluate DiffGaze on two 360 \({}^{\circ}\) image benchmarks for fine-grained gaze sequence generation as well as two downstream tasks, scanpath prediction and saliency prediction. Our evaluations show that DiffGaze outperforms the fine-grained gaze generation baselines in all tasks on both benchmarks. We also report a 21-participant survey study showing that our method generates gaze sequences that are indistinguishable from real human sequences. Taken together, our evaluations not only demonstrate the effectiveness of DiffGaze but also point towards a new generation of methods that faithfully model the rich spatial and temporal nature of natural human gaze behaviour.
Chuhan Jiao, Yao Wang 0018, Mihai Bâce, Zhiming Hu 0003, Andreas Bulling
ACM Trans. Interact. Intell. Syst.2
2026 Tell Me Without Telling Me: Two-Way Prediction of Visualization Literacy and Visual Attention
abstract
Accounting for individual differences can improve the effectiveness of visualization design. While the role of visual attention in visualization interpretation is well recognized, existing work often overlooks how this behavior varies based on visual literacy levels. Based on data from a 235-participant user study covering three visualization tests (mini-VLAT, CALVI, and SGL), we show that distinct attention patterns in visual data exploration can correlate with participants' literacy levels: While experts (high-scorers) generally show a strong attentional focus, novices (low-scorers) focus less and explore more. We then propose two computational models leveraging these insights: Lit2Sal - a novel visual saliency model that predicts observer attention given their visualization literacy level, and Sal2Lit - a model to predict visual literacy from human visual attention data. Our quantitative and qualitative evaluation demonstrates that Lit2Sal outperforms state-of-the-art saliency models with literacy-aware considerations. Sal2Lit predicts literacy with 86% accuracy using a single attention map, providing a time-efficient supplement to literacy assessment that only takes less than a minute. Taken together, our unique approach to consider individual differences in salience models and visual attention in literacy assessments paves the way for new directions in personalized visual data communication to enhance understanding.
Minsuk Chang, Yao Wang 0018, Huichen Will Wang, Yuanhong Zhou, Andreas Bulling, Cindy Xiong Bearfield
IEEE Trans. Vis. Comput. Graph.2
2025 Chartist: Task-driven Eye Movement Control for Chart Reading
abstract
| openaire: EC/HE/101141916/EU//Artificial User
Danqing Shi, Yao Wang 0018, Yunpeng Bai, Andreas Bulling, Antti Oulasvirta
CHI2
2025 ChartQC: Question Classification from Human Attention Data on Charts
Takumi Nishiyasu, Tobias Kostorz, Yao Wang 0018, Yoichi Sato 0001, Andreas Bulling
ETRA3
2025 Towards a Better Understanding of Graph Perception in Immersive Environments
abstract
As Immersive Analytics (IA) increasingly uses Virtual Reality (VR) for stereoscopic 3D (S3D) graph visualisation, it is crucial to understand how users perceive network structures in these immersive environments. However, little is known about how humans read S3D graphs during task solving, and how gaze behaviour indicates task performance. To address this gap, we report a user study with 18 participants asked to perform three analytical tasks on S3D graph visualisations in a VR environment. Our findings reveal systematic relationships between network structural properties and gaze behaviour. Based on these insights, we contribute a comprehensive eye tracking methodology for analysing human perception in immersive environments and establish eye tracking as a valuable tool for objectively evaluating cognitive load in S3D graph visualisation.
Lin Zhang 0042, Yao Wang 0018, Wilhelm Kerle-Malcharek, Karsten Klein 0001, Falk Schreiber, Andreas Bulling
GD2
2025 AttentionLeak: What Does Human Attention Reveal About Information Visualisation?
Malte Sönnichsen, Mayar Elfares, Yao Wang 0018, Ralf Küsters, Alina Roitberg, Andreas Bulling
ICDAR (4)3
2025 RelEYEance: Gaze-based Assessment of Users' AI-reliance at Run-time
abstract
In time-critical detection tasks, such as drone monitoring, a key condition for users to effectively leverage AI assistance is to find an appropriate trade-off between making fast decisions and verifying AI suggestions, which we refer to as appropriate user reliance. However, assessing such reliance is often oversimplified by focusing solely on task outcomes, potentially overlooking whether users properly verify AI messages. We collected eye-tracking data from an AI-assisted monitoring task and developed a gaze-based reliance model: RelEYEance, to assess the extent of user reliance on AI-suggested alarms. We found that gaze patterns related to verification behaviors distinguish between appropriate reliance, over-reliance, and under-reliance, influencing task performance. We validated our model in a second user study, showing it can reliably detect users' over- and under-reliance at run-time, which could be used e.g. for issuing intervention messages. The results demonstrate the potential for real-time human-AI reliance assessment, facilitating adaptive reliance calibration.
Zekun Wu 0001, Yao Wang 0018, Markus Langer, Anna Maria Feit
Proc. ACM Hum. Comput. Interact.2
2024 SalChartQA: Question-driven Saliency on Information Visualisations
abstract
Understanding the link between visual attention and users’ information needs when visually exploring information visualisations is under-explored due to a lack of large and diverse datasets to facilitate these analyses. To fill this gap we introduce SalChartQA – a novel crowd-sourced dataset that uses the BubbleView interface to track user attention and a question-answering (QA) paradigm to induce different information needs in users. SalChartQA contains 74,340 answers to 6,000 questions on 3,000 visualisations. Informed by our analyses demonstrating the close correlation between information needs and visual saliency, we propose the first computational method to predict question-driven saliency on visualisations. Our method outperforms state-of-the-art saliency models for several metrics, such as the correlation coefficient and the Kullback-Leibler divergence. These results show the importance of information needs for shaping attentive behaviour and pave the way for new applications, such as task-driven optimisation of visualisations or explainable AI in chart question-answering.
Yao Wang 0018, Weitian Wang, Abdullah Abdelhafez, Mayar Elfares, Zhiming Hu 0003, Mihai Bâce, Andreas Bulling
CHI1
2024 Saliency3D: A 3D Saliency Dataset Collected on Screen
abstract
While visual saliency has recently been studied in 3D, the experimental setup for collecting 3D saliency data can be expensive and cumbersome. To address this challenge, we propose a novel experimental design that utilises an eye tracker on a screen to collect 3D saliency data, which could reduce the cost and complexity of data collection. We first collected gaze data on a computer screen and then mapped the 2D points to 3D saliency data through perspective transformation. Using this method, we propose Saliency3D, a 3D saliency dataset (49,276 fixations) comprising 10 participants looking at sixteen objects. We examined the viewing preferences for objects and our results indicate potential preferred viewing directions and a correlation between salient features and the variation in viewing directions.
Yao Wang 0018, Mihai Bâce, Karsten Klein 0001, Andreas Bulling
ETRA1
2024 VisRecall++: Analysing and Predicting Visualisation Recallability from Gaze Behaviour
abstract
Question answering has recently been proposed as a promising means to assess the recallability of information visualisations. However, prior works are yet to study the link between visually encoding a visualisation in memory and recall performance. To fill this gap, we propose VisRecall++ -- a novel 40-participant recallability dataset that contains gaze data on 200 visualisations and 1,000 questions, including identifying the title and retrieving values. We measured recallability by asking participants questions after they observed the visualisation for 10 seconds. Our analyses reveal several insights, such as saccade amplitude, number of fixations, and fixation duration significantly differ between high and low recallability groups. Finally, we propose GazeRecallNet -- a novel computational method to predict recallability from gaze behaviour that outperforms the state-of-the-art model RecallNet and three other baselines on this task. Taken together, our results shed light on assessing recallability from gaze behaviour and inform future work on recallability-based visualisation optimisation.
Yao Wang 0018, Yue Jiang 0002, Zhiming Hu 0003, Constantin Ruhdorfer, Mihai Bâce, Andreas Bulling
Proc. ACM Hum. Comput. Interact.1
2024 Scanpath Prediction on Information Visualisations
abstract
We propose Unified Model of Saliency and Scanpaths (UMSS)- a model that learns to predict multi-duration saliency and scanpaths (i.e. sequences of eye fixations) on information visualisations. Although scanpaths provide rich information about the importance of different visualisation elements during the visual exploration process, prior work has been limited to predicting aggregated attention statistics, such as visual saliency. We present in-depth analyses of gaze behaviour for different information visualisation elements (e.g. Title, Label, Data) on the popular MASSVIS dataset. We show that while, overall, gaze patterns are surprisingly consistent across visualisations and viewers, there are also structural differences in gaze dynamics for different elements. Informed by our analyses, UMSS first predicts multi-duration element-level saliency maps, then probabilistically samples scanpaths from them. Extensive experiments on MASSVIS show that our method consistently outperforms state-of-the-art methods with respect to several, widely used scanpath and saliency evaluation metrics. Our method achieves a relative improvement in sequence score of 11.5% for scanpath prediction, and a relative improvement in Pearson correlation coefficient of up to 23.6% for saliency prediction. These results are auspicious and point towards richer user models and simulations of visual attention on visualisations without the need for any eye tracking equipment.
Yao Wang 0018, Mihai Bâce, Andreas Bulling
IEEE Trans. Vis. Comput. Graph.1
2023 Improving neural saliency prediction with a cognitive model of human visual attention
Ekta Sood, Lei Shi 0032, Matteo Bortoletto, Yao Wang 0018, Philipp Müller 0001, Andreas Bulling
CogSci4
2022 Impact of Gaze Uncertainty on AOIs in Information Visualisations
abstract
Gaze-based analysis of areas of interest (AOIs) is widely used in information visualisation research to understand how people explore visualisations or assess the quality of visualisations concerning key characteristics such as memorability. However, nearby AOIs in visualisations amplify the uncertainty caused by the gaze estimation error, which strongly influences the mapping between gaze samples or fixations and different AOIs. We contribute a novel investigation into gaze uncertainty and quantify its impact on AOI-based analysis on visualisations using two novel metrics: the Flipping Candidate Rate (FCR) and Hit Any AOI Rate (HAAR). Our analysis of 40 real-world visualisations, including human gaze and AOI annotations, shows that gaze uncertainty frequently and significantly impacts the analysis conducted in AOI-based studies. Moreover, we analysed four visualisation types and found that bar and scatter plots are usually designed in a way that causes more uncertainty than line and pie plots in gaze-based analysis.
Yao Wang 0018, Maurice Koch, Mihai Bâce, Daniel Weiskopf, Andreas Bulling
ETRA1
2022 VisRecall: Quantifying Information Visualisation Recallability via Question Answering
abstract
Despite its importance for assessing the effectiveness of communicating information visually, fine-grained recallability of information visualisations has not been studied quantitatively so far. In this work, we propose a question-answering paradigm to study visualisation recallability and present VisRecall - a novel dataset consisting of 200 visualisations that are annotated with crowd-sourced human (N = 305) recallability scores obtained from 1,000 questions of five question types. Furthermore, we present the first computational method to predict recallability of different visualisation elements, such as the title or specific data values. We report detailed analyses of our method on VisRecall and demonstrate that it outperforms several baselines in overall recallability and FE-, F-, RV-, and U-question recallability. Our work makes fundamental contributions towards a new generation of methods to assist designers in optimising visualisations.
Yao Wang 0018, Chuhan Jiao, Mihai Bâce, Andreas Bulling
IEEE Trans. Vis. Comput. Graph.1
2018 Sobel Heuristic Kernel for Aerial Semantic Segmentation
abstract
Misclassification in semantic segmentation mostly occurs in the pixels around the semantic contour. In this work, we address the task of aerial image segmentation by borrowing the kernel prior from classical edge detecting operator. We propose a module called Sobel Heuristic Kernel(SHK). Our work makes several main contributions and experimentally shows good performance. To the best of our knowledge, we are the first to combine traditional edge detection method and deep learning method in semantic segmentation. Our SHK module reaches state of the art in the Inria Aerial Image Labeling dataset.
Yao Wang 0018, Yisong Chen, Peng Lu 0007
ICIP2
2018 Large-Scale Structure from Motion with Semantic Constraints of Aerial Images
Yao Wang 0018, Peng Lu 0007, Yisong Chen
PRCV (1)2