Huichen Will Wang

dblp:385/7549 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2026
0009-0007-5941-4047ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Human-computer interaction and pervasive computing
5 papers
Human-AI interaction · 46% Usability and user experience research · 38% User interface design and tools · 15%
Computer graphics and multimedia
5 papers
Visualization and visual analytics · 100%
Artificial intelligence
1 paper
Vision and language · 87% Trustworthy machine learning · 13%

Topics — the 11 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visualization and visual analytics › visualization evaluation
trust in visualization
1.012026
Do You "Trust" This Visualization? An Inventory to Measure Trust in Visualizations · IEEE Trans. Vis. Comput. Graph. 2026
Visualization and visual analytics › visual attention
visual attention modeling
1.012026
Tell Me Without Telling Me: Two-Way Prediction of Visualization Literacy and Visual Attention · IEEE Trans. Vis. Comput. Graph. 2026
Visualization and visual analytics
visualization evaluation
1.012026
Do You "Trust" This Visualization? An Inventory to Measure Trust in Visualizations · IEEE Trans. Vis. Comput. Graph. 2026
Visualization and visual analytics
visualization literacy
1.012026
Tell Me Without Telling Me: Two-Way Prediction of Visualization Literacy and Visual Attention · IEEE Trans. Vis. Comput. Graph. 2026
Computer vision › Vision and language › vision-language model › multimodal large language model
multimodal large language model reasoning
0.912025
Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark · ICML 2025
Computer vision › Vision and language › multimodal reasoning
multimodal reasoning benchmark
0.912025
Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark · ICML 2025
User interface design and tools › visualization
visualization design
0.912025
DracoGPT: Extracting Visualization Design Preferences from Large Language Models · IEEE Trans. Vis. Comput. Graph. 2025
Usability and user experience research › visual perception
visualization perception
0.912025
How Aligned are Human Chart Takeaways and LLM Predictions? A Case Study on Bar Charts with Varying Layouts · IEEE Trans. Vis. Comput. Graph. 2025
Visualization and visual analytics
data storytelling
0.312025
Jupybara: Operationalizing a Design Space for Actionable Data Analysis and Storytelling with LLMs · CHI 2025
Visualization and visual analytics
graphical perception
0.312025
How Aligned are Human Chart Takeaways and LLM Predictions? A Case Study on Bar Charts with Varying Layouts · IEEE Trans. Vis. Comput. Graph. 2025
Visualization and visual analytics
visualization recommendation
0.312025
DracoGPT: Extracting Visualization Design Preferences from Large Language Models · IEEE Trans. Vis. Comput. Graph. 2025

Methods — techniques the papers use, named apart from their topics

large language model prompting · 3.5user study · 2.0trust game · 2.0saliency modeling · 2.0psychometric validation · 2.0exploratory factor analysis · 2.0computational modeling · 2.0factor analysis · 1.7draco · 1.7design space · 1.7test-time compute scaling · 0.9chain-of-thought prompting · 0.9
YearPublicationVenuePosition
2026 Tell Me Without Telling Me: Two-Way Prediction of Visualization Literacy and Visual Attention
abstract
Accounting for individual differences can improve the effectiveness of visualization design. While the role of visual attention in visualization interpretation is well recognized, existing work often overlooks how this behavior varies based on visual literacy levels. Based on data from a 235-participant user study covering three visualization tests (mini-VLAT, CALVI, and SGL), we show that distinct attention patterns in visual data exploration can correlate with participants' literacy levels: While experts (high-scorers) generally show a strong attentional focus, novices (low-scorers) focus less and explore more. We then propose two computational models leveraging these insights: Lit2Sal - a novel visual saliency model that predicts observer attention given their visualization literacy level, and Sal2Lit - a model to predict visual literacy from human visual attention data. Our quantitative and qualitative evaluation demonstrates that Lit2Sal outperforms state-of-the-art saliency models with literacy-aware considerations. Sal2Lit predicts literacy with 86% accuracy using a single attention map, providing a time-efficient supplement to literacy assessment that only takes less than a minute. Taken together, our unique approach to consider individual differences in salience models and visual attention in literacy assessments paves the way for new directions in personalized visual data communication to enhance understanding.
Minsuk Chang, Yao Wang 0018, Huichen Will Wang, Yuanhong Zhou, Andreas Bulling, Cindy Xiong Bearfield
IEEE Trans. Vis. Comput. Graph.3
2026 Do You "Trust" This Visualization? An Inventory to Measure Trust in Visualizations
abstract
Trust plays a critical role in visual data communication and decision-making, yet existing visualization research employs varied trust measures, making it challenging to compare and synthesize findings across studies. In this work, we first took a bottom-up, data-driven approach to understand what visualization readers mean when they say they "trust" a visualization. We compiled and adapted a broad set of trust-related statements from existing inventories and collected responses to visualizations with varying degrees of trustworthiness. Through exploratory factor analysis, we derived an operational definition of trust in visualizations. Our findings indicate that people perceive a trustworthy visualization as one that presents credible information and is comprehensible and usable. Building on this insight, we developed an eight-item inventory: four core items measuring trust in visualizations and four optional items controlling for individual differences in baseline trust tendency. We established the inventory's internal consistency reliability using McDonald's omega, confirmed its content validity by demonstrating alignment with theoretically-grounded trust dimensions, and validated its criterion validity through two trust games with real-world stakes. Finally, we illustrate how this standardized inventory can be applied across diverse visualization research contexts. Utilizing our inventory, future research can examine how design choices, tasks, and domains influence trust, and how to foster appropriate trusting behavior in human-data interactions.
Huichen Will Wang, Kylie R. Lin, Andrew Cohen, Ryan Kennedy, Zach Zwald, Carolina Nobre, Cindy Xiong Bearfield
IEEE Trans. Vis. Comput. Graph.1
2025 Jupybara: Operationalizing a Design Space for Actionable Data Analysis and Storytelling with LLMs
abstract
Mining and conveying actionable insights from complex data is a key challenge of exploratory data analysis (EDA) and storytelling.To address this challenge, we present a design space for actionable EDA and storytelling.Synthesizing theory and expert interviews, we highlight how semantic precision, rhetorical persuasion, and pragmatic relevance underpin effective EDA and storytelling.We also show how this design space subsumes common challenges in actionable EDA and storytelling, such as identifying appropriate analytical strategies and leveraging relevant domain knowledge.Building on the potential of LLMs to generate coherent narratives with commonsense reasoning, we contribute Jupybara, an AI-enabled assistant for actionable EDA and storytelling implemented as a Jupyter Notebook extension.Jupybara employs two strategiesdesign-space-aware prompting and multi-agent architectures-to operationalize our design space.An expert evaluation confirms Jupybara's usability, steerability, explainability, and reparability, as well as the effectiveness of our strategies in operationalizing the design space framework with LLMs.
Huichen Will Wang, Lawrence Birnbaum, Vidya Setlur
CHI1
2025 Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
abstract
The ability to organically reason over and with both text and images is a pillar of human intelligence, yet the ability of Multimodal Large Language Models (MLLMs) to perform such multimodal reasoning remains under-explored. Existing benchmarks often emphasize text-dominant reasoning or rely on shallow visual cues, failing to adequately assess integrated visual and textual reasoning. We introduce EMMA (Enhanced MultiModal reAsoning), a benchmark targeting organic multimodal reasoning across mathematics, physics, chemistry, and coding. EMMA tasks demand advanced cross-modal reasoning that cannot be addressed by reasoning independently in each modality, offering an enhanced test suite for MLLMs' reasoning capabilities. Our evaluation of state-of-the-art MLLMs on EMMA reveals significant limitations in handling complex multimodal and multi-step reasoning tasks, even with advanced techniques like Chain-of-Thought prompting and test-time compute scaling underperforming. These findings underscore the need for improved multimodal architectures and training paradigms to close the gap between human and model reasoning in multimodality.
Yunzhuo Hao, Jiawei Gu, Huichen Will Wang, Zhengyuan Yang, Yu Cheng 0001
ICML3
2025 DracoGPT: Extracting Visualization Design Preferences from Large Language Models
abstract
Trained on vast corpora, Large Language Models (LLMs) have the potential to encode visualization design knowledge and best practices. However, if they fail to do so, they might provide unreliable visualization recommendations. What visualization design preferences, then, have LLMs learned? We contribute DracoGPT, a method for extracting, modeling, and assessing visualization design preferences from LLMs. To assess varied tasks, we develop two pipelines-DracoGPT-Rank and DracoGPT-Recommend-to model LLMs prompted to either rank or recommend visual encoding specifications. We use Draco as a shared knowledge base in which to represent LLM design preferences and compare them to best practices from empirical research. We demonstrate that DracoGPT can accurately model the preferences expressed by LLMs, enabling analysis in terms of Draco design constraints. Across a suite of backing LLMs, we find that DracoGPT-Rank and DracoGPT-Recommend moderately agree with each other, but both substantially diverge from guidelines drawn from human subjects experiments. Future work can build on our approach to expand Draco's knowledge base to model a richer set of preferences and to provide a robust and cost-effective stand-in for LLMs.
Huichen Will Wang, Mitchell Gordon, Leilani Battle, Jeffrey Heer
IEEE Trans. Vis. Comput. Graph.1
2025 How Aligned are Human Chart Takeaways and LLM Predictions? A Case Study on Bar Charts with Varying Layouts
abstract
Large Language Models (LLMs) have been adopted for a variety of visualizations tasks, but how far are we from perceptually aware LLMs that can predict human takeaways? Graphical perception literature has shown that human chart takeaways are sensitive to visualization design choices, such as spatial layouts. In this work, we examine the extent to which LLMs exhibit such sensitivity when generating takeaways, using bar charts with varying spatial layouts as a case study. We conducted three experiments and tested four common bar chart layouts: vertically juxtaposed, horizontally juxtaposed, overlaid, and stacked. In Experiment 1, we identified the optimal configurations to generate meaningful chart takeaways by testing four LLMs, two temperature settings, nine chart specifications, and two prompting strategies. We found that even state-of-the-art LLMs struggled to generate semantically diverse and factually accurate takeaways. In Experiment 2, we used the optimal configurations to generate 30 chart takeaways each for eight visualizations across four layouts and two datasets in both zero-shot and one-shot settings. Compared to human takeaways, we found that the takeaways LLMs generated often did not match the types of comparisons made by humans. In Experiment 3, we examined the effect of chart context and data on LLM takeaways. We found that LLMs, unlike humans, exhibited variation in takeaway comparison types for different bar charts using the same bar layout. Overall, our case study evaluates the ability of LLMs to emulate human interpretations of data and points to challenges and opportunities in using LLMs to predict human chart takeaways.
Huichen Will Wang, Jane Hoffswell, Sao Myat Thazin Thane, Victor S. Bursztyn, Cindy Xiong Bearfield
IEEE Trans. Vis. Comput. Graph.1